Color map layer for labeling
By generating a six-dimensional shading point cloud and transforming it into a five-dimensional map tiles, the problem of point color inaccuracy caused by misalignment in the existing technology is solved, and the accuracy and efficiency of map annotations are improved, and semantic mapping and sensor fusion are supported.
Patent Information
- Application Number
- CN202380085098.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-11-21
- Filing Date
- 2023-09-29
- Publication Date
- 2025-07-18
AI Technical Summary
The existing color map layers are based on misaligned point clouds, resulting in inaccurate point colors, affecting the accuracy and efficiency of map annotations.
By generating a six-dimensional shading point cloud and transforming it into a five-dimensional map tile, fusion is used to use posture maps and camera images to remove dynamic object points, enhance road marking colors, and accurate alignment and fusion of point clouds and images.
Improve the accuracy and efficiency of map annotations, support semantic mapping processing and downstream tasks, provide multimodal fusion between point clouds and images, and enhance sensor fusion capabilities.
Smart Images

Figure CN120345006A_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 416,454, filed on October 14, 2022, and U.S. Patent Application No. 17 / 991,794, filed on November 21, 2022, the entire contents of which are incorporated herein by reference. Background Art
[0003] Map annotation provides a way to highlight specific regions on a map and provide additional information related to those specific regions. Color map layers can be used for map annotation. Existing color map layers are based on misaligned points in a point cloud and include inaccurate point colors. Brief Description of the Drawings
[0004] Figure 1 is an example environment of a vehicle that can implement one or more components of an autonomous system;
[0005] Figure 2 is a diagram of one or more systems of a vehicle including an autonomous system;
[0006] Figure 3 is Figure 1 and Figure 2 a diagram of one or more devices and / or components of one or more systems;
[0007] Figure 4 is a diagram of certain components of an autonomous system;
[0008] Figure 5 is a diagram of an implementation of a process for generating a color map layer for annotation;
[0009] Figure 6 is an example flowchart of a process for generating a color map layer for map annotation;
[0010] Figure 7 is a diagram of an example architecture for generating a 6D (six - dimensional) colored point cloud;
[0011] Figure 8 is an example flowchart of a process for image processing;
[0012] Figure 9 is an example flowchart of a process for point cloud projection;
[0013] Figure 10 is an example flowchart of a process for result refinement; and
[0014] Figure 11 is an example flowchart of a process for generating a color map layer for map annotation. Detailed implementation manners
[0015] In the following description, for the purpose of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it will be apparent that the embodiments described in the present disclosure may be implemented without these specific details. In some instances, well-known structures and devices are illustrated in block diagram form to avoid unnecessarily obscuring aspects of the present disclosure.
[0016] In the drawings, for ease of description, a specific arrangement or order of schematic elements (such as those representing systems, devices, modules, instruction blocks, and / or data elements, etc.) is illustrated. However, those skilled in the art will understand that unless explicitly described, the specific order or arrangement of the schematic elements in the drawings is not intended to imply a required processing order or sequence, or a separation of processes. In addition, unless explicitly described, the inclusion of schematic elements in the drawings is not intended to mean that such elements are required in all embodiments, nor does it mean that the features represented by such elements cannot be included in some embodiments or cannot be combined with other elements in some embodiments.
[0017] In addition, in the drawings, connecting elements (such as solid lines, dashed lines, or arrows, etc.) are used to illustrate the connection, relationship, or association between two or more other schematic elements or among them. The absence of any such connecting element is not intended to mean that there cannot be a connection, relationship, or association. In other words, some connections, relationships, or associations between elements are not illustrated in the drawings so as not to obscure the present disclosure. In addition, for ease of illustration, a single connecting element may be used to represent multiple connections, relationships, or associations between elements. For example, if the connecting element represents the communication of a signal, data, or instruction (such as a "software instruction"), those skilled in the art should understand that such an element may represent one or more signal paths (such as a bus) that may be required to affect the communication.
[0018] Although terms such as "first", "second", and / or "third" are used to describe various elements, these elements should not be limited by these terms. The terms "first", "second", and / or "third" are only used to distinguish one element from another. For example, without departing from the scope of the described embodiments, a first contact may be referred to as a second contact, and similarly, a second contact may be referred to as a first contact. Both the first contact and the second contact are contacts, but they are not the same contact.
[0019] The terms used in the description of the various embodiments described herein are included only for the purpose of describing particular embodiments and are not intended to be limiting. As used in the description of the various embodiments and the appended claims, the singular forms "a", "an", and "the" are also intended to include the plural forms and may be used interchangeably with "one or more than one" or "at least one", unless the context clearly dictates otherwise. It will also be understood that the term "and / or" as used herein refers to and includes any and all possible combinations of one or more of the associated listed items. It will also be understood that when the terms "comprise", "include", "have", and / or "with" are used in this specification, it specifies the presence of the stated features, integers, steps, operations, elements, and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0020] As used herein, the terms "communicate" and "communicating" refer to at least one of receiving, receiving, transmitting, conveying, and / or providing information (or information represented by, for example, data, signals, messages, instructions, and / or commands, etc.). For a unit (e.g., a device, a system, a component of a device or system, and / or a combination thereof, etc.) that is to communicate with another unit, this means that the unit can directly or indirectly receive information from the other unit and / or send (e.g., transmit) information to the other unit. This may refer to a direct or indirect connection that is inherently wired and / or wireless. Additionally, two units can communicate with each other even if the information transmitted between the first unit and the second unit is modified, processed, relayed, and / or routed. For example, even if the first unit receives information passively and does not actively transmit information to the second unit, the first unit can communicate with the second unit. As another example, if at least one intermediate unit (e.g., a third unit located between the first unit and the second unit) processes the information received from the first unit and transmits the processed information to the second unit, the first unit can communicate with the second unit. In some embodiments, a message may refer to a network packet (e.g., a data packet, etc.) that includes data.
[0021] As used herein, depending on the context, the term "if" is optionally interpreted to mean "when", "at the time of", "in response to determining that", and / or "in response to detecting", etc. Similarly, depending on the context, the phrase "if it has been determined" or "if [the stated condition or event] is detected" is optionally interpreted to mean "when determining...", "in response to determining that", or "at the time of detecting [the stated condition or event]" and / or "in response to detecting [the stated condition or event]", etc. Further, as used herein, terms such as "has", "have", or "possess" are intended to be open-ended terms. Additionally, unless otherwise expressly stated, the phrase "based on" is intended to mean "at least partially based on".
[0022] Reference will now be made in detail to the embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the various described embodiments. However, it will be apparent to one of ordinary skill in the art that the various described embodiments may be practiced without these specific details. In other instances, well-known methods, procedures, components, circuits, and networks have not been described in detail so as not to unnecessarily obscure aspects of the embodiments.
[0023] General Overview
[0024] In some aspects and / or embodiments, the systems, methods, and computer program products described herein include and / or implement generating a color map layer composed of five-dimensional (5D) tiles (x, y, r, g, b). A six-dimensional (x, y, z, r, g, b) colored point cloud is transformed into 5D colored map tiles. In some embodiments, the point cloud from the pose map and the camera images from the driving log are used to generate a 6D colored point cloud. In some embodiments, generating a 6D colored point cloud includes operations such as image processing, point cloud projection, result refinement, and pose map update.
[0025] In some embodiments, image processing further includes: for a camera image lacking poses at some timestamps, performing pose interpolation based on the poses of two adjacent images having corresponding poses. Extracting the poses corresponding to the two adjacent images from the point cloud at the same or substantially similar timestamps. Additionally, overlapping images (identical images, substantially identical images, or images having a similarity greater than a specific similarity value) indicating that the driving distance is less than a specific distance value (due to, for example, heavy traffic or waiting for a green light) are filtered. Obtaining image pixel labels (such as vehicles, buildings, trees, etc.) from a pre-trained image segmentation network. In some embodiments, point cloud projection includes projecting the point cloud in the point cloud coordinate system onto the image coordinate system, and combining the image including coloring / RGB data with the point cloud to generate a six-dimensional colored point cloud. In some embodiments, result refinement includes: removing points from the colored point cloud (e.g., removing points representing dynamic objects (such as pedestrians, vehicles, bicycles) and points having different labels between the image and the point cloud); when multiple image pixels (multiple pixel labels) correspond to a single point, selecting the point color of the point as the pixel color of the closest image pixel; and enhancing the road marking color by adding the point cloud intensity value (obtained, for example, by a LiDAR sensor) to each point in the point cloud.
[0026] By virtue of the implementation of the systems, methods, and computer program products described herein, some of the advantages of these techniques include: providing a more accurate alignment between pixel colors and points to generate a colored map layer (e.g., 5D colored map tiles) to accelerate the map annotation process. In an example, the colored map layer is used in, for example, semantic mapping processes implemented by map annotation tools. Additionally, these techniques facilitate downstream semantic tasks such as lane extractor networks. In particular, these techniques provide a baseline multi-modality (e.g., camera and LiDAR sensors) fusion between the point cloud and the image. These techniques incorporate multiple surrounding environment images for sensor fusion.
[0027] Now refer to Figure 1, an exemplary environment 100 is illustrated, in which vehicles including autonomous systems and vehicles not including autonomous systems operate. As illustrated, the environment 100 includes vehicles 102a - 102n, objects 104a - 104n, routes 106a - 106n, area 108, vehicle - to - infrastructure (V2I) devices 110, network 112, remote autonomous vehicle (AV) system 114, queue management system 116, and V2I system 118. The vehicles 102a - 102n, vehicle - to - infrastructure (V2I) devices 110, network 112, autonomous vehicle (AV) system 114, queue management system 116, and V2I system 118 are interconnected via a wired connection, a wireless connection, or a combination of wired and wireless connections (e.g., establishing a connection for communication, etc.). In some embodiments, the objects 104a - 104n are interconnected with at least one of the vehicles 102a - 102n, vehicle - to - infrastructure (V2I) devices 110, network 112, autonomous vehicle (AV) system 114, queue management system 116, and V2I system 118 via a wired connection, a wireless connection, or a combination of wired and wireless connections.
[0028] The vehicles 102a - 102n (individually referred to as vehicle 102 and collectively referred to as vehicles 102) include at least one device configured to transport goods and / or passengers. In some embodiments, the vehicle 102 is configured to communicate with the V2I device 110, remote AV system 114, queue management system 116, and / or V2I system 118 via the network 112. In some embodiments, the vehicle 102 includes a car, a bus, a truck, and / or a train, etc. In some embodiments, the vehicle 102 is the same as or similar to the vehicle 200 described herein (see Figure 2 ). In some embodiments, the vehicles 200 in the set of vehicles 200 are associated with an autonomous queue manager. In some embodiments, as described herein, the vehicle 102 travels along the corresponding routes 106a - 106n (individually referred to as route 106 and collectively referred to as routes 106). In some embodiments, one or more than one vehicle 102 includes an autonomous system (e.g., an autonomous system that is the same as or similar to the autonomous system 202).
[0029] Objects 104a - 104n (individually referred to as object 104 and collectively as objects 104) include, for example, at least one vehicle, at least one pedestrian, at least one cyclist, and / or at least one structure (e.g., building, sign, fire hydrant, etc.). Each object 104 (e.g., located at a fixed location and over a period of time) is stationary or (e.g., having a speed and associated with at least one trajectory) moving. In some embodiments, object 104 is associated with a corresponding location in region 108.
[0030] Routes 106a - 106n (individually referred to as route 106 and collectively as routes 106) are each associated with (e.g., define) a sequence of actions (also referred to as a trajectory) that a connected AV can navigate along. Each route 106 begins at an initial state (e.g., a state corresponding to a first spatio - temporal location and / or speed, etc.) and ends at a final target state (e.g., a state corresponding to a second spatio - temporal location different from the first spatio - temporal location) or a target zone (e.g., a subspace of acceptable states (e.g., termination states)). In some embodiments, the first state includes a location where one or more individuals will board the AV, and the second state or zone includes one or more locations where one or more individuals boarding the AV will disembark. In some embodiments, route 106 includes multiple acceptable sequences of states (e.g., multiple sequences of spatio - temporal locations) that are associated with (e.g., define) multiple trajectories. In an example, route 106 includes only high - level actions or imprecise state locations, such as a series of connected roads indicating a direction change at a roadway intersection. Additionally or alternatively, route 106 can include more precise actions or states, such as, for example, a specific target lane or precise location within a lane region and a target rate at those locations. In an example, route 106 includes multiple precise state sequences along at least one high - level action with a finite look - ahead horizon to reach an intermediate target, where the combination of successive iterations of the finite - horizon state sequences cumulatively corresponds to multiple trajectories that together form a high - level route terminating at the final target state or zone.
[0031] Region 108 includes a physical region (e.g., a geographical region) that the vehicle 102 can navigate. In an example, region 108 includes at least one state (e.g., a country, a province, an individual state among multiple states included in a country, etc.), at least a portion of a state, at least one city, at least a portion of a city, etc. In some embodiments, region 108 includes at least one named arterial road (referred to herein as a "road"), such as a highway, an interstate highway, a parkway, an urban street, etc. Additionally or alternatively, in some examples, region 108 includes at least one unnamed road, such as a lane, a section of a parking lot, a section of a vacant and / or undeveloped area, a dirt road, etc. In some embodiments, a road includes at least one lane (e.g., the portion of the road that the vehicle 102 can traverse). In an example, a road includes at least one lane associated with (e.g., identified based on) at least one lane marking line.
[0032] A vehicle-to-infrastructure (V2I) device 110 (sometimes referred to as a vehicle-to-infrastructure or vehicle-to-everything (V2X) device) includes at least one device configured to communicate with the vehicle 102 and / or the V2I system 118. In some embodiments, the V2I device 110 is configured to communicate with the vehicle 102, the remote AV system 114, the queue management system 116, and / or the V2I system 118 via the network 112. In some embodiments, the V2I device 110 includes a radio frequency identification (RFID) device, a sign, a camera (e.g., a two-dimensional (2D) and / or three-dimensional (3D) camera), a lane marking, a streetlight, a parking meter, etc. In some embodiments, the V2I device 110 is configured to communicate directly with the vehicle 102. Additionally or alternatively, in some embodiments, the V2I device 110 is configured to communicate with the vehicle 102, the remote AV system 114, and / or the queue management system 116 via the V2I system 118. In some embodiments, the V2I device 110 is configured to communicate with the V2I system 118 via the network 112.
[0033] The network 112 includes one or more wired and / or wireless networks. In an example, the network 112 includes a cellular network (e.g., a Long-Term Evolution (LTE) network, a third-generation (3G) network, a fourth-generation (4G) network, a fifth-generation (5G) network, a Code Division Multiple Access (CDMA) network, etc.), a Public Land Mobile Network (PLMN), a Local Area Network (LAN), a Wide Area Network (WAN), a Metropolitan Area Network (MAN), a telephone network (e.g., a Public Switched Telephone Network (PSTN)), a private network, an ad hoc network, an intranet, the Internet, a fiber-optic-based network, a cloud computing network, etc., and / or a combination of some or all of these networks.
[0034] The remote AV system 114 includes at least one device configured to communicate with the vehicle 102, the V2I device 110, the network 112, the queue management system 116, and / or the V2I system 118 via the network 112. In an example, the remote AV system 114 includes a server, a server group, and / or other similar devices. In some embodiments, the remote AV system 114 is co-located with the queue management system 116. In some embodiments, the remote AV system 114 participates in the installation of some or all of the components of the vehicle (including autonomous systems, autonomous vehicle computing, and / or software implemented by autonomous vehicle computing, etc.). In some embodiments, the remote AV system 114 maintains (e.g., updates and / or replaces) these components and / or software during the life of the vehicle.
[0035] The queue management system 116 includes at least one device configured to communicate with the vehicle 102, the V2I device 110, the remote AV system 114, and / or the V2I system 118. In an example, the queue management system 116 includes a server, a server group, and / or other similar devices. In some embodiments, the queue management system 116 is associated with a ridesharing company (e.g., an organization for controlling the operation of multiple vehicles (e.g., vehicles including autonomous systems and / or vehicles not including autonomous systems), etc.).
[0036] In some embodiments, the V2I system 118 includes at least one device configured to communicate with the vehicle 102, the V2I device 110, the remote AV system 114, and / or the queue management system 116 via the network 112. In some examples, the V2I system 118 is configured to communicate with the V2I device 110 via a connection different from the network 112. In some embodiments, the V2I system 118 includes a server, a server group, and / or other similar devices. In some embodiments, the V2I system 118 is associated with a municipal authority or a private institution (e.g., a private institution for maintaining the V2I device 110, etc.).
[0037] Provide Figure 1 The number and arrangement of the illustrated elements are provided as examples. Compared with Figure 1 the illustrated elements, there may be additional elements, fewer elements, different elements, and / or elements with different arrangements. Additionally or alternatively, at least one element of the environment 100 may perform one or more functions described as being performed by Figure 1 at least one different element. Additionally or alternatively, at least one set of elements of the environment 100 may perform one or more functions described as being performed by at least one different set of elements of the environment 100.
[0038] Now refer to Figure 2 , the vehicle 200 (which may be the same as or similar to the vehicle 102 of Figure 1 ) includes an autonomous system 202, a powertrain control system 204, a steering control system 206, and a braking system 208, or is associated with the autonomous system 202, the powertrain control system 204, the steering control system 206, and the braking system 208. In some embodiments, the vehicle 200 is the same as or similar to the vehicle 102 (see Figure 1 ). In some embodiments, the autonomous system 202 is configured to endow the vehicle 200 with autonomous driving capabilities (e.g., implement at least one of the following driving functions, features, and / or devices that are automatic or based on maneuvers, and the at least one driving function, feature, and / or device that is automatic or based on maneuvers enables the vehicle 200 to operate partially or completely without human intervention, including but not limited to fully autonomous vehicles (e.g., vehicles that abandon reliance on human intervention, such as level 5 ADS-operated vehicles), highly autonomous vehicles (e.g., vehicles that abandon reliance on human intervention in certain situations, such as level 4 ADS-operated vehicles), and / or conditionally autonomous vehicles (e.g., vehicles that abandon reliance on human intervention in limited situations, such as level 3 ADS-operated vehicles), etc.). In one embodiment, the autonomous system 202 includes the operational or tactical functionality required to operate the vehicle 200 in road traffic and continuously perform part or all of the dynamic driving task (DDT). In another embodiment, the autonomous system 202 includes an advanced driver assistance system (ADAS) that includes driver support features. The autonomous system 202 supports various levels of driving automation ranging from no driving automation (e.g., level 0) to full driving automation (e.g., level 5). For a detailed description of fully autonomous vehicles and highly autonomous vehicles, reference can be made to SAE International standard J3016: Taxonomy and Definitions for Terms Related to On-Road Motor Vehicle Automated Driving Systems, the entire content of which is incorporated by reference. In some embodiments, the vehicle 200 is associated with an autonomous queue manager and / or a ridesharing company.
[0039] The autonomous system 202 includes a sensor suite that includes one or more devices such as a camera 202a, a LiDAR sensor 202b, a Radar sensor 202c, and a microphone 202d. In some embodiments, the autonomous system 202 may include more or fewer devices and / or different devices (e.g., ultrasonic sensors, inertial sensors, GPS receivers (discussed below), and / or odometer sensors for generating data associated with an indication of the distance the vehicle 200 has traveled, etc.). In some embodiments, the autonomous system 202 uses one or more devices included in the autonomous system 202 to generate data associated with the environment 100 described herein. The data generated by one or more devices of the autonomous system 202 can be used by one or more of the systems described herein to observe the environment (e.g., environment 100) in which the vehicle 200 is located. In some embodiments, the autonomous system 202 includes a communication device 202e, an autonomous vehicle compute 202f, a drive-by-wire (DBW) system 202h, and a safety controller 202g.
[0040] The camera 202a includes at least one device configured to communicate with the communication device 202e, the autonomous vehicle compute 202f, and / or the safety controller 202g via a bus (e.g., a bus 302 that is the same or similar to Figure 3 the bus). The camera 202a includes at least one camera (e.g., a digital camera using an optical sensor such as a charge-coupled device (CCD), a thermal camera, an infrared (IR) camera, and / or an event camera, etc.) configured to capture images including physical objects (e.g., cars, buses, curbs, and / or people, etc.). In some embodiments, the camera 202a generates camera data as output. In some examples, the camera 202a generates camera data that includes image data associated with the image. In such an example, the image data may specify at least one parameter corresponding to the image (e.g., image characteristics such as exposure, brightness, etc., and / or an image timestamp, etc.). In such an example, the image may be in a format (e.g., RAW, JPEG, and / or PNG, etc.). In some embodiments, the camera 202a includes a plurality of independent cameras configured (e.g., positioned) on the vehicle for the purpose of capturing images for stereovision (stereo vision). In some examples, the camera 202a includes generating image data and transmitting the image data to the autonomous vehicle compute 202f and / or a queue management system (e.g., to Figure 1a plurality of cameras of the same or similar queuing management system as queuing management system 116. In such an example, the autonomous vehicle computing 202f determines the depth to one or more objects in the field of view of at least two of the plurality of cameras based on image data from at least two cameras. In some embodiments, the camera 202a is configured to capture images of objects within a distance relative to the camera 202a (e.g., up to 100 meters and / or up to 1 kilometer, etc.). Thus, the camera 202a includes features such as sensors and lenses optimized for sensing objects at one or more distances relative to the camera 202a.
[0041] In an embodiment, the camera 202a includes at least one camera configured to capture one or more images associated with one or more traffic lights, street signs, and / or other physical objects that provide visual navigation information. In some embodiments, the camera 202a generates traffic light data associated with one or more images. In some examples, the camera 202a generates TLD (Traffic Light Detection) data associated with one or more images including a format (e.g., RAW, JPEG, and / or PNG, etc.). In some embodiments, the camera 202a that generates TLD data is different from other systems incorporating cameras described herein in that the camera 202a may include one or more cameras having a wide field of view (e.g., a wide-angle lens, a fish-eye lens, and / or a lens having a viewing angle of about 120 degrees or greater, etc.) to generate images related to as many physical objects as possible.
[0042] The Light Detection and Ranging (LiDAR) sensor 202b includes being configured to communicate with the communication device 202e, the autonomous vehicle computing 202f, and / or the safety controller 202g via a bus (e.g., with Figure 3at least one device that communicates via a bus (e.g., a bus identical or similar to bus 302). The LiDAR sensor 202b includes a system configured to emit light from a light emitter (e.g., a laser emitter). The light emitted by the LiDAR sensor 202b includes light outside the visible spectrum (e.g., infrared light, etc.). In some embodiments, during operation, the light emitted by the LiDAR sensor 202b encounters a physical object (e.g., a vehicle) and is reflected back to the LiDAR sensor 202b. In some embodiments, the light emitted by the LiDAR sensor 202b does not penetrate the physical object it encounters. The LiDAR sensor 202b also includes at least one light detector that detects the light after the light emitted from the light emitter encounters a physical object. In some embodiments, at least one data processing system associated with the LiDAR sensor 202b generates an image (e.g., a point cloud and / or a combined point cloud, etc.) representing the objects included in the field of view of the LiDAR sensor 202b. In some examples, at least one data processing system associated with the LiDAR sensor 202b generates an image representing the boundary of the physical object and / or the surface of the physical object (e.g., the topology of the surface), etc. In such examples, the image is used to determine the boundary of the physical object in the field of view of the LiDAR sensor 202b.
[0043] A Radio Detection and Ranging (Radar) sensor 202c includes at least one device configured to communicate with a communication device 202e, an autonomous vehicle computer 202f, and / or a safety controller 202g via a bus (e.g., a bus identical or similar to Figure 3 bus 302). The Radar sensor 202c includes a system configured to emit (pulsed or continuous) radio waves. The radio waves emitted by the Radar sensor 202c include radio waves within a specific spectrum. In some embodiments, during operation, the radio waves emitted by the Radar sensor 202c encounter a physical object and are reflected back to the Radar sensor 202c. In some embodiments, the radio waves emitted by the Radar sensor 202c are not reflected by some objects. In some embodiments, at least one data processing system associated with the Radar sensor 202c generates a signal representing the objects included in the field of view of the Radar sensor 202c. For example, at least one data processing system associated with the Radar sensor 202c generates an image representing the boundary of the physical object and / or the surface of the physical object (e.g., the topology of the surface), etc. In some examples, the image is used to determine the boundary of the physical object in the field of view of the Radar sensor 202c.
[0044] The microphone 202d includes at least one device configured to communicate with the communication device 202e, the autonomous vehicle computing 202f, and / or the safety controller 202g via a bus (e.g., a bus identical or similar to the Figure 3 bus 302). The microphone 202d includes one or more microphones (e.g., an array microphone and / or an external microphone, etc.) that capture an audio signal and generate data associated with (e.g., representing) the audio signal. In some examples, the microphone 202d includes a transducer device and / or a similar device. In some embodiments, one or more of the systems described herein can receive the data generated by the microphone 202d and determine the position (e.g., distance, etc.) of an object relative to the vehicle 200 based on the audio signal associated with the data.
[0045] The communication device 202e includes at least one device configured to communicate with the camera 202a, the LiDAR sensor 202b, the Radar sensor 202c, the microphone 202d, the autonomous vehicle computing 202f, the safety controller 202g, and / or the DBW (drive-by-wire) system 202h. For example, the communication device 202e can include a device identical or similar to the Figure 3 communication interface 314. In some embodiments, the communication device 202e includes a vehicle-to-vehicle (V2V) communication device (e.g., a device for enabling wireless communication of data between vehicles).
[0046] The autonomous vehicle computing 202f includes at least one device configured to communicate with the camera 202a, the LiDAR sensor 202b, the Radar sensor 202c, the microphone 202d, the communication device 202e, the safety controller 202g, and / or the DBW system 202h. In some examples, the autonomous vehicle computing 202f includes devices such as a client device, a mobile device (e.g., a cellular phone and / or a tablet, etc.), and / or a server (e.g., a computing device including one or more central processing units and / or graphics processing units, etc.). In some embodiments, the autonomous vehicle computing 202f is identical or similar to the autonomous vehicle (AV) computing 400 described herein. Additionally or alternatively, in some embodiments, the autonomous vehicle computing 202f is configured to communicate with an autonomous vehicle system (e.g., an autonomous vehicle system identical or similar to the Figure 1 remote AV system 114), a queue management system (e.g., a queue management system identical or similar to the Figure 1 queue management system 116), a V2I device (e.g., a V2I device identical or similar to the Figure 1 V2I device 110), and / or a V2I system (e.g., a V2I system identical or similar to the Figure 1communicate with a V2I system 118 that is the same as or similar to the V2I system).
[0047] The safety controller 202g includes at least one device configured to communicate with the camera 202a, the LiDAR sensor 202b, the Radar sensor 202c, the microphone 202d, the communication device 202e, the autonomous vehicle computing 202f, and / or the DBW system 202h. In some examples, the safety controller 202g includes one or more controllers (such as an electrical controller and / or an electromechanical controller, etc.) configured to generate and / or transmit control signals to operate one or more devices of the vehicle 200 (such as the powertrain control system 204, the steering control system 206, and / or the braking system 208, etc.). In some embodiments, the safety controller 202g is configured to generate control signals that are prior to (e.g., override) the control signals generated and / or transmitted by the autonomous vehicle computing 202f.
[0048] The DBW system 202h includes at least one device configured to communicate with the communication device 202e and / or the autonomous vehicle computing 202f. In some examples, the DBW system 202h includes one or more controllers (such as an electrical controller and / or an electromechanical controller, etc.) configured to generate and / or transmit control signals to operate one or more devices of the vehicle 200 (such as the powertrain control system 204, the steering control system 206, and / or the braking system 208, etc.). Additionally or alternatively, one or more controllers of the DBW system 202h are configured to generate and / or transmit control signals to operate at least one different device of the vehicle 200 (such as turn signals, headlights, door locks, and / or windshield wipers, etc.).
[0049] The powertrain control system 204 includes at least one device configured to communicate with the DBW system 202h. In some examples, the powertrain control system 204 includes at least one controller and / or actuator, etc. In some embodiments, the powertrain control system 204 receives control signals from the DBW system 202h, and the powertrain control system 204 causes the vehicle 200 to perform longitudinal vehicle movements (such as starting to move forward, stopping moving forward, starting to move backward, stopping moving backward, accelerating in a certain direction, decelerating in a certain direction, etc.), or perform lateral vehicle movements (such as making a left turn and / or making a right turn, etc.). In an example, the powertrain control system 204 increases, maintains the same, or decreases the energy (such as fuel and / or electricity, etc.) provided to the motor of the vehicle, thereby causing at least one wheel of the vehicle 200 to rotate or not rotate.
[0050] The steering control system 206 includes at least one device configured to rotate one or more wheels of the vehicle 200. In some examples, the steering control system 206 includes at least one controller and / or actuator, etc. In some embodiments, the steering control system 206 rotates two front wheels and / or two rear wheels of the vehicle 200 to the left or right to turn the vehicle 200 left or right. In other words, the steering control system 206 causes the activities required to regulate the y-axis component of the vehicle's movement.
[0051] The braking system 208 includes at least one device configured to actuate one or more brakes to decelerate the vehicle 200 and / or keep it stationary. In some examples, the braking system 208 includes at least one controller and / or actuator configured to close one or more calipers associated with one or more wheels of the vehicle 200 on the corresponding rotors of the vehicle 200. Additionally or alternatively, in some examples, the braking system 208 includes an automatic emergency braking (AEB) system and / or a regenerative braking system, etc.
[0052] In some embodiments, the vehicle 200 includes at least one platform sensor (not explicitly illustrated) for measuring or inferring the nature of the state or condition of the vehicle 200. In some examples, the vehicle 200 includes platform sensors such as a global positioning system (GPS) receiver, an inertial measurement unit (IMU), a wheel speed sensor, a wheel brake pressure sensor, a wheel torque sensor, an engine torque sensor, and / or a steering angle sensor. Although the braking system 208 is illustrated as being proximal to the vehicle 200 in Figure 2 the vehicle 200, the braking system 208 can be located anywhere in the vehicle 200.
[0053] Now refer to Figure 3 , a schematic diagram illustrating the device 300. As illustrated, the device 300 includes a processor 304, a memory 306, a storage component 308, an input interface 310, an output interface 312, a communication interface 314, and a bus 302. In some embodiments, the device 300 corresponds to: at least one device of the vehicle 102 (e.g., at least one device of the system of the vehicle 102); and / or one or more devices of the network 112 (e.g., one or more devices of the system of the network 112). In some embodiments, one or more devices of the vehicle 102 (e.g., one or more devices of the system of the vehicle 102), and / or one or more devices of the network 112 (e.g., one or more devices of the system of the network 112) include at least one device 300 and / or at least one component of the device 300. As Figure 3As shown, device 300 includes bus 302, processor 304, memory 306, storage component 308, input interface 310, output interface 312, and communication interface 314.
[0054] Bus 302 includes components that permit communication among the components of device 300. In some cases, processor 304 includes a processor (e.g., a central processing unit (CPU), a graphics processing unit (GPU), and / or an accelerated processing unit (APU), etc.), a microphone, a digital signal processor (DSP), and / or any processing component that can be programmed to perform at least one function (e.g., a field programmable gate array (FPGA) and / or an application specific integrated circuit (ASIC), etc.). Memory 306 includes random access memory (RAM), read only memory (ROM), and / or another type of dynamic and / or static storage device that stores data and / or instructions for use by processor 304 (e.g., flash memory, magnetic memory, and / or optical memory, etc.).
[0055] Storage component 308 stores data and / or software related to the operation and use of device 300. In some examples, storage component 308 includes a hard disk (e.g., a magnetic disk, an optical disk, a magneto-optical disk, and / or a solid state disk, etc.), a compact disk (CD), a digital versatile disk (DVD), a floppy disk, a cassette tape, a magnetic tape, a CD-ROM, a RAM, a PROM, an EPROM, a FLASH-EPROM, an NV-RAM, and / or another type of computer-readable medium, and corresponding drives.
[0056] Input interface 310 includes components that permit device 300 to receive information such as via a user input (e.g., a touch screen display, a keyboard, a keypad, a mouse, a button, a switch, a microphone, and / or a camera, etc.). Additionally or alternatively, in some embodiments, input interface 310 includes sensors for sensing information (e.g., a global positioning system (GPS) receiver, an accelerometer, a gyroscope, and / or an actuator, etc.). Output interface 312 includes components for providing output information from device 300 (e.g., a display, a speaker, and / or one or more light emitting diodes (LEDs), etc.).
[0057] In some embodiments, communication interface 314 includes transceiver-like components (e.g., a transceiver and / or separate receiver and transmitter, etc.) that permit device 300 to communicate with other devices via a wired connection, a wireless connection, or a combination of a wired connection and a wireless connection. In some examples, communication interface 314 permits device 300 to receive information from and / or provide information to another device. In some examples, communication interface 314 includes an Ethernet interface, an optical interface, a coaxial interface, an infrared interface, a radio frequency (RF) interface, a universal serial bus (USB) interface, an interface and / or a cellular network interface, etc.
[0058] In some embodiments, the device 300 performs one or more processes described herein. The device 300 performs these processes based on software instructions executed by a processor 304 that are stored in a computer-readable medium such as a memory 306 and / or a storage component 308. A computer-readable medium (e.g., a non-transitory computer-readable medium) is defined herein as a non-transitory memory device. A non-transitory memory device includes a storage space located within a single physical storage device or a storage space distributed across multiple physical storage devices.
[0059] In some embodiments, software instructions are read into the memory 306 and / or the storage component 308 from another computer-readable medium or from another device via a communication interface 314. When the software instructions stored in the memory 306 and / or the storage component 308 are executed, they cause the processor 304 to perform one or more processes described herein. Additionally or alternatively, instead of software instructions or in combination with software instructions, hardwired circuitry is used to perform one or more processes described herein. Thus, unless otherwise explicitly stated, the embodiments described herein are not limited to any particular combination of hardware circuitry and software.
[0060] The memory 306 and / or the storage component 308 includes a data storage portion or at least one data structure (e.g., a database, etc.). The device 300 is capable of receiving information from the data storage portion or at least one data structure in the memory 306 or the storage component 308, storing the information in the data storage portion or at least one data structure, communicating the information to the data storage portion or at least one data structure, or searching for information stored in the data storage portion or at least one data structure. In some examples, the information includes network data, input data, output data, or any combination thereof.
[0061] In some embodiments, the device 300 is configured to execute software instructions stored in the memory 306 and / or the memory of another device (e.g., another device that is the same as or similar to the device 300). As used herein, the term "module" refers to at least one instruction stored in the memory 306 and / or the memory of another device, which, when executed by the processor 304 and / or the processor of another device (e.g., another device that is the same as or similar to the device 300), causes the device 300 (e.g., at least one component of the device 300) to perform one or more processes described herein. In some embodiments, the module is implemented in software, firmware, and / or hardware, etc.
[0062] Provide Figure 3 The number and arrangement of the illustrated components are provided as an example. In some embodiments, compared with Figure 3Compared with the illustrated components, apparatus 300 may include additional components, fewer components, different components, or components arranged differently. Additionally or alternatively, a set of components of apparatus 300 (e.g., one or more than one component) may perform one or more than one function described as being performed by another component or another set of components of apparatus 300.
[0063] Now referring to Figure 4 , an example block diagram of an autonomous vehicle computing 400 (sometimes referred to as an "AV stack") is illustrated. As illustrated, autonomous vehicle computing 400 includes a perception system 402 (sometimes referred to as a perception module), a planning system 404 (sometimes referred to as a planning module), a localization system 406 (sometimes referred to as a localization module), a control system 408 (sometimes referred to as a control module), and a database 410. In some embodiments, perception system 402, planning system 404, localization system 406, control system 408, and database 410 are included in and / or implemented in an automatic navigation system of a vehicle (e.g., autonomous vehicle computing 202f of vehicle 200). Additionally or alternatively, in some embodiments, perception system 402, planning system 404, localization system 406, control system 408, and database 410 are included in one or more than one independent system (e.g., one or more than one system the same or similar to autonomous vehicle computing 400, etc.). In some examples, perception system 402, planning system 404, localization system 406, control system 408, and database 410 are included in one or more than one independent system located in a vehicle and / or at least one remote system as described herein. In some embodiments, any and / or all of the systems included in autonomous vehicle computing 400 are implemented in software (e.g., software instructions stored in a memory), computer hardware (e.g., via a microprocessor, a microcontroller, an application specific integrated circuit (ASIC), and / or a field programmable gate array (FPGA), etc.), or a combination of computer software and computer hardware. It will also be understood that, in some embodiments, autonomous vehicle computing 400 is configured to communicate with remote systems (e.g., an autonomous vehicle system the same or similar to remote AV system 114, a queue management system 116 the same or similar to queue management system 116, and / or a V2I system 118 the same or similar to V2I system 118, etc.).
[0064] In some embodiments, the perception system 402 receives data associated with at least one physical object in the environment (e.g., data used by the perception system 402 to detect at least one physical object), and classifies the at least one physical object. In some examples, the perception system 402 receives image data captured by at least one camera (e.g., camera 202a), the image being associated with one or more physical objects within the field of view of the at least one camera (e.g., representing the one or more physical objects). In such examples, the perception system 402 classifies the at least one physical object based on one or more groupings of physical objects (e.g., bicycles, vehicles, traffic signs, and / or pedestrians, etc.). In some embodiments, based on the classification of the physical objects by the perception system 402, the perception system 402 transmits data associated with the classification of the physical objects to the planning system 404.
[0065] In some embodiments, the planning system 404 receives data associated with a destination and generates data associated with at least one route (e.g., route 106) along which a vehicle (e.g., vehicle 102) can travel towards the destination. In some embodiments, the planning system 404 periodically or continuously receives data from the perception system 402 (e.g., the data associated with the classification of the physical objects described above), and the planning system 404 updates at least one trajectory or generates at least one different trajectory based on the data generated by the perception system 402. In other words, the planning system 404 can perform tasks related to the tactical functions required to operate the vehicle 102 in road traffic. Tactical efforts involve maneuvering the vehicle in traffic during the journey, which includes but is not limited to deciding whether and when to overtake another vehicle, change lanes, or select an appropriate speed, acceleration, deceleration, etc. In some embodiments, the planning system 404 receives data associated with the updated position of the vehicle (e.g., vehicle 102) from the positioning system 406, and the planning system 404 updates at least one trajectory or generates at least one different trajectory based on the data generated by the positioning system 406.
[0066] In some embodiments, the positioning system 406 receives data associated with (e.g., representing) the location of a vehicle (e.g., vehicle 102) in an area. In some examples, the positioning system 406 receives LiDAR data associated with at least one point cloud generated by at least one LiDAR sensor (e.g., LiDAR sensor 202b). In certain examples, the positioning system 406 receives data associated with at least one point cloud from multiple LiDAR sensors, and the positioning system 406 generates a combined point cloud based on the respective point clouds. In these examples, the positioning system 406 compares the at least one point cloud or the combined point cloud with two-dimensional (2D) and / or three-dimensional (3D) maps of the area stored in the database 410. Then, based on the positioning system 406 comparing the at least one point cloud or the combined point cloud with the maps, the positioning system 406 determines the position of the vehicle in the area. In some embodiments, the map includes a combined point cloud of the area generated prior to the navigation of the vehicle. In some embodiments, the map includes, but is not limited to, a high-precision map of the roadway geometry, a map describing the connectivity of the road network, a map describing the physical properties of the roadways (such as traffic speed, traffic flow, the number of vehicle and bicycle traffic lanes, lane width, lane traffic direction, or the type and location of lane markings, or a combination thereof, etc.), and a map describing the spatial location of road features (such as crosswalks, traffic signs, or various other driving signals, etc.). In some embodiments, the map is generated in real time based on the data received by the perception system.
[0067] In another example, the positioning system 406 receives Global Navigation Satellite System (GNSS) data generated by a Global Positioning System (GPS) receiver. In some examples, the positioning system 406 receives GNSS data associated with the location of a vehicle in an area, and the positioning system 406 determines the latitude and longitude of the vehicle in the area. In such examples, the positioning system 406 determines the position of the vehicle in the area based on the latitude and longitude of the vehicle. In some embodiments, the positioning system 406 generates data associated with the position of the vehicle. In some examples, based on the positioning system 406 determining the position of the vehicle, the positioning system 406 generates data associated with the position of the vehicle. In such examples, the data associated with the position of the vehicle includes data associated with one or more semantic properties corresponding to the position of the vehicle.
[0068] In some embodiments, the control system 408 receives data associated with at least one trajectory from the planning system 404, and the control system 408 controls the operation of the vehicle. In some examples, the control system 408 receives data associated with at least one trajectory from the planning system 404, and the control system 408 controls the operation of the vehicle by generating and transmitting control signals to cause the powertrain control system (e.g., the DBW system 202h and / or the powertrain control system 204, etc.), the steering control system (e.g., the steering control system 206), and / or the braking system (e.g., the braking system 208) to operate. For example, the control system 408 is configured to perform operational functions such as lateral vehicle motion control or longitudinal vehicle motion control. Lateral vehicle motion control causes activities required to regulate the y-axis component of the vehicle motion. Longitudinal vehicle motion control causes activities required to regulate the x-axis component of the vehicle motion. In an example, in the case where the trajectory includes a left turn, the control system 408 transmits a control signal to cause the steering control system 206 to adjust the steering angle of the vehicle 200, thereby causing the vehicle 200 to turn left. Additionally or alternatively, the control system 408 generates and transmits control signals to cause other devices of the vehicle 200 (e.g., headlights, turn signals, door locks, and / or windshield wipers, etc.) to change states.
[0069] In some embodiments, the perception system 402, the planning system 404, the positioning system 406, and / or the control system 408 implement at least one machine learning model (e.g., at least one multi-layer perceptron (MLP), at least one convolutional neural network (CNN), at least one recurrent neural network (RNN), at least one autoencoder, and / or at least one transformer, etc.). In some examples, the perception system 402, the planning system 404, the positioning system 406, and / or the control system 408 implement at least one machine learning model alone or in combination with one or more of the above systems. In some examples, the perception system 402, the planning system 404, the positioning system 406, and / or the control system 408 implement at least one machine learning model as part of a pipeline (e.g., a pipeline for identifying one or more objects located in the environment, etc.).
[0070] The database 410 stores data transmitted to, received from, and / or updated by the perception system 402, the planning system 404, the positioning system 406, and / or the control system 408. In some examples, the database 410 includes a storage component (e.g., associated with Figure 3the same or similar storage components as the storage component 308). In some embodiments, the database 410 stores data associated with 2D and / or 3D maps of at least one area. In some examples, the database 410 stores data associated with 2D and / or 3D maps of a part of a city, multiple parts of multiple cities, multiple cities, counties, states, and / or countries (e.g., nations), etc. In such examples, a vehicle (e.g., a vehicle the same or similar to the vehicle 102 and / or the vehicle 200) can drive along one or more drivable areas (e.g., single-lane roads, multi-lane roads, highways, back roads, and / or off-road paths, etc.), and cause at least one LiDAR sensor (e.g., a LiDAR sensor the same or similar to the LiDAR sensor 202b) to generate data associated with an image representing the objects included in the field of view of the at least one LiDAR sensor.
[0071] In some embodiments, the database 410 can be implemented across multiple devices. In some examples, the database 410 is included in a vehicle (e.g., a vehicle the same or similar to the vehicle 102 and / or the vehicle 200), an autonomous vehicle system (e.g., an autonomous vehicle system the same or similar to the remote AV system 114), a queue management system (e.g., a queue management system the same or similar to Figure 1 the queue management system 116) and / or a V2I system (e.g., a V2I system the same or similar to Figure 1 the V2I system 118), etc.
[0072] Now referring to Figure 5 FIG. 500 which illustrates an implementation of a process for generating a colored map layer for annotation. In some embodiments, the implementation 500 includes Figure 2 the autonomous system 202, and the autonomous system 202 includes a LiDAR sensor 202b and an AV computing 202f. In some embodiments, data generated by the Figure 3 LiDAR sensor 202b is obtained by the device 300 to generate a high-definition (HD) map with a colored map layer.
[0073] In the implementation 500, the AV computing 504 (e.g., Figure 2 the AV computing 202f) includes a planning system 506 (e.g., Figure 4 the planning system 404) and a control system 508 (e.g., Figure 4 the control system 408). The planning system 506 determines a trajectory (514) for AV navigation. For example, the planning system 506 periodically or continuously receives data from a perception system (e.g., Figure 4 the perception system 402), and the data includes information about the environment (e.g.,Figure 1 raw sensor data associated with an object in an environment 100). The planning system 506 determines at least one trajectory based on the sensor data generated by the perception system 402. The trajectory is transmitted (516) to a control system for controlling the operation of the vehicle.
[0074] In some embodiments, the raw sensor data includes camera images obtained as the vehicle navigates along a trajectory through the environment (e.g., data associated with at least one image generated by at least one camera 202a). In some embodiments, the raw sensor data is LiDAR data obtained as the vehicle navigates along a trajectory through the environment (e.g., data associated with at least one point cloud generated by at least one LiDAR sensor 202b). In an example, the raw sensor data includes camera images and corresponding LiDAR data recorded in a driving log. LiDAR data is a collection of 2D or 3D points (also known as a point cloud) used to construct a representation of the environment. In an example, as the vehicle traverses the environment according to the trajectory, the LiDAR device repeatedly scans the environment with a 360-degree sweep. The rotational scan of the environment by LiDAR is commonly referred to as a full sweep. The sweeps typically overlap such that LiDAR data (e.g., point clouds) at the same location at different timestamps are represented in the sweeps. In an example, a pose graph is generated that includes nodes representing the pose of the vehicle in the environment and edges representing the transformations between consecutive poses. In an example, features (e.g., feature maps) are used to compute pose-to-pose constraints in the pose graph.
[0075] In some embodiments, the pose graph and camera images are merged into an HD map by generating a color map layer. In an example, an HD map is a high-precision map that enables a computer-based navigation system to determine precise trajectories and other information for navigating in the environment. The HD map is comprehensive and is constructed to support safe and efficient decision-making. The HD map includes several layers, such as a standard base map layer, a geometric layer that describes the geometric properties of the roadways and the connectivity of the road network, and a semantic layer that describes the physical properties of the roadways (e.g., the number of vehicle and bicycle traffic lanes, lane widths, lane traffic directions, or the type and location of lane markings, or any combination thereof, etc.) and the spatial locations of road features such as crosswalks, traffic signs, or various other driving signals, etc.). In operation, a positioning system (e.g., positioning system 406) compares the captured sensor data with the stored map to determine the location of the vehicle, including a computer-based navigation system, in the area. Creating and updating the HD map includes visualizing the map such that a human annotator can do so at a user interface (e.g., Figure 3Verify and further annotate the HD map at the input interface 310). In the example, the pose map derived from the LiDAR data captured along the determined trajectory (514) is interpolated to generate poses at the timestamps corresponding to the camera image data. A color map layer is generated by fusing the poses and the camera image information to generate 5D map tiles.
[0076] In the example, the base map layer is a less detailed map containing general feature information associated with the environment. For example, the base map in 2D form is a standard map obtained from a third party (such as a map provider, etc.). The base map is a standardized map without customization. In the example, the base map is a standard definition map and does not include the spatial locations of road geometries, physical properties, and road features (such as crosswalks, traffic signs, or various other driving signals of various types, etc.) with connectivity properties. The geometry layer of the HD map describes the lane geometry properties and the road network connectivity properties. The semantic layer describes the spatial locations of the physical properties of the lanes (e.g., the number of vehicle and bicycle traffic lanes, lane widths, lane traffic directions, or the types and locations of lane markings, or any combination thereof, etc.) and road features (such as crosswalks, traffic signs, or various other driving signals of various types, etc.). In operation, the positioning system (e.g., positioning system 406) compares the captured sensor data with the stored map to determine the position of the vehicle, including the computer-based navigation system, in the area. Creating and updating the HD map includes visualizing the map via a map annotation tool so that a human annotator can verify and further annotate the HD map at the user interface (e.g., Figure 3 the input interface 310). In the example, at the output interface (e.g., Figure 3 the output interface 312), output the visualization.
[0077] Figure 6 is an example flowchart of a process 600 for generating a color map layer for map annotation. In some embodiments, one or more steps of the steps of process 600 are performed by a device or system (or a group of devices and / or a group of systems) that is separate from or includes the autonomous system (e.g., fully and / or partially, etc.). For example, it can be performed by Figure 1 the remote AV system 114, Figure 1 the vehicle 102 or Figure 2 the vehicle 200 (e.g., the autonomous system 202 of the vehicle 102 or 200), Figure 3 the device 300 and / or Figure 4The AV calculation 400 (e.g., one or more systems of the AV calculation 400) performs one or more steps of the processing 600 (e.g., fully and / or partially). In some embodiments, the steps of the processing 600 can be performed among any of the above systems in a collaborative manner with each other.
[0078] At block 602, the device (e.g., Figure 3 the device 300) generates a six-dimensional (6D: X, Y, Z, R, G, B) point cloud based on the three-dimensional (3D: X, Y, Z) point cloud from the pose map and the camera image from the driving log. The six dimensions include the X coordinate, Y coordinate, Z coordinate, red value, green value, and blue value. The three dimensions include the X coordinate, Y coordinate, and Z coordinate. The camera image includes RGB color information, and the 3D point cloud from the pose map is painted with the RGB colors from the camera image to generate a 6D colored point cloud. In some embodiments, the camera image is registered with the 3D point cloud in the pose map. In an example, image registration refers to transforming the camera image and the 3D point cloud into the same coordinate system.
[0079] At block 604, the device transforms the 6D colored point cloud into a 5D map tile. In some embodiments, the device can transform the 6D colored point cloud into a 5D map tile by ignoring or discarding the Z coordinate value.
[0080] In some embodiments, the device first removes points with a height (e.g., Z coordinate) greater than a specific height value (e.g., 1.8 meters) from the 6D colored point cloud, and then transforms the truncated 6D colored point cloud (the 6D colored point cloud including the remaining points) into a 5D map tile by ignoring or discarding the Z coordinate values. For example, some points representing green plants (e.g., tree branches) are removed from the 6D colored point cloud because the green plants may obscure more important objects on the ground (e.g., vehicles, pedestrians, bicycles, lane markings, crosswalks, sidewalks, traffic lights, etc.). In the example, the points representing green plants correspond to heights greater than 1.8 meters in the point cloud. The device then transforms the truncated 6D colored point cloud (the 6D colored point cloud including the remaining points) into a 5D map tile. In the example, the device can transform the truncated 6D colored point cloud into a 5D map tile by ignoring or discarding the Z coordinate values. In some embodiments, the 5D map tile represents a top view (or "bird's-eye" view) of a part of the environment (or an object in the environment) represented by the pose map and the camera image. The 5D map tile (X, Y, R, G, B) is provided to a map rasterizer for map annotation. In some embodiments, the map rasterizer is software for converting an image described in the format of a vector map (e.g., a 5D map tile and / or a set of 5D map tiles) into a raster image (a series of pixels). In the example, the raster image is an array of cells or pixels organized in rows and columns (e.g., a grid), where each cell or pixel contains a value representing information.
[0081] Figure 7 FIG. is a diagram of an exemplary architecture 700 for generating a 6D colored point cloud. In some embodiments, one or more steps described as being performed by the 6D colored point cloud generator 712 may be performed by Figures 1 to 4 one or more devices (e.g., by the remote AV system 114, etc.). As Figure 7 shown, the pose map 702 includes sensors (e.g., Figure 2The 3D point cloud 704 generated by the optical detection and ranging (LiDAR) sensor 202b) (the 3D point cloud 704 is an element of the pose graph 702). The combination of the position (translation) and orientation (rotation) relative to a point (or a set of points representing the environment) is called a pose. In some embodiments, a pose can be associated with an object represented by a portion of the point cloud. The position is a 2D translation vector representing the X and Y coordinates. The orientation is represented by a 2×2 matrix, which is also called a rotation matrix. The translation vector and the rotation matrix can be used to represent the position and orientation. The pose graph 702 includes nodes (each node represents a pose associated with a point cloud) connected by edges (each edge represents the relative pose of one point cloud with respect to another point cloud). In an example, the edges represent spatial information. For example, the pose of each point cloud can include the position or orientation of the vehicle relative to other features (such as trees, buildings, etc.) and the environment in which the vehicle and other features are located (the position or orientation at the time when each camera image 708 is captured).
[0082] The driving log 706 includes the camera images 708 captured by the Figure 2 camera 202a and the calibration data 710. The calibration data 710 includes the transformation matrix (K, R, t) between the point cloud generated by the LiDAR sensor (e.g., Figure 2 the LiDAR sensor 202b) and the camera images captured by the camera (e.g., Figure 2 the camera 202a). The LiDAR sensor and the camera are mounted at different positions in the AV. The transformation matrix provides information indicating the relative positions of the LiDAR sensor and the camera.
[0083] The 3D point cloud 704, the camera images 708, and the calibration data 710 are input into a 6D colored point cloud generator 712 (implemented as Figures 1 to 4 one or more devices (e.g., implemented by the remote AV system 114, etc.)). The camera images 708 include RGB color information, and the 3D point cloud 704 from the pose graph 702 is painted with the RGB color from the camera images 708 to generate a 6D colored point cloud 714. The 6D colored point cloud generator 712 outputs the 6D colored point cloud 714. Then the 6D colored point cloud 714 is stored in the pose graph 716. In some embodiments, the 6D colored point cloud generator 712 performs Figure 8 image processing of Figure 9 point cloud projection of Figure 10 result refinement of Figure 10 and pose graph update of
[0084] In some embodiments, pose map 716 and pose map 702 can be the same pose map. The 6D colored point cloud 714 can replace the 3D point cloud 704 (or a part of the 3D point cloud 704) in the pose map 702. In some embodiments, pose map 716 and pose map 702 can be different pose maps.
[0085] Figure 8 is an example flowchart of a process 800 for image processing. In some embodiments, one or more steps of the steps of process 800 are performed by another device or system or another group of devices and / or systems that are separate from or include the autonomous system (e.g., fully and / or partially). For example, it can be performed by Figure 1 the remote AV system 114, Figure 1 the vehicle 102 or Figure 2 the vehicle 200 (e.g., the autonomous system 202 of vehicle 102 or 200), Figure 3 the device 300, Figure 4 the AV computing 400 (e.g., one or more systems of the AV computing 400) and / or Figure 7 the 6D colored point cloud generator 712 to perform one or more steps of process 800 (e.g., fully and / or partially).
[0086] In some embodiments, the steps of process 800 can be performed among any of the above systems in a collaborative manner.
[0087] In some embodiments, at block 802, the device obtains a camera image of the AV from the driving log 706 (e.g., Figure 2 captured by one or more cameras 202a of Figure 7 the camera image 708).
[0088] At block 804, the first pose of the first camera image at the first timestamp is estimated. The first camera image is associated with the first timestamp, and the pose map does not include the pose at the first timestamp. In an example, the first pose is estimated based on the poses corresponding to two adjacent camera images (the two adjacent camera images having two corresponding 3D point clouds from the pose map 702). In an example, the first pose is estimated based on the poses corresponding to the two closest camera images (the two adjacent camera images). A pose refers to the position and orientation of sensor data. In some embodiments, not all camera images captured by a camera (e.g., Figure 2 the camera 202a) are in the pose map (e.g., Figure 7has a corresponding pose in the pose map 702), which is due to the different frequencies of the camera and the LiDAR sensor. The key camera image has a corresponding point cloud (the point cloud is generated when the key camera image is captured, i.e., the key camera image and the corresponding point cloud are captured synchronously). The corresponding point cloud is used to extract the poses to be included in the pose map. For non-key camera images that do not have corresponding point clouds (no point cloud is generated when the non-key camera image is captured, and thus no matching pose is found in the pose map), the pose for each non-key camera image can be estimated based on the poses corresponding to two adjacent key camera images (the two adjacent key camera images have two corresponding poses in the pose map).
[0089] The adjacent poses are obtained from the pose map at the timestamps corresponding to two adjacent camera images. The two adjacent camera images are the two camera images closest to the first camera image that has no corresponding pose in the pose map. There is a timestamp for each camera image or each point cloud. The two adjacent poses include a second pose corresponding to a second key image with a second timestamp that is immediately earlier than the timestamp of a specific non-key camera image (e.g., the first camera image), and a third pose corresponding to a third key image with a timestamp that is immediately later than the timestamp of the specific non-key camera image (e.g., the first camera image), and these two adjacent poses are used to estimate the first pose of the specific non-key camera image (e.g., the first camera image). The two adjacent poses are poses extracted from two different point clouds (the second point cloud corresponding to the second key image and the third point cloud corresponding to the third key image). In the example, the second point cloud and the third point cloud have sequential timestamps and are obtained from the same LiDAR device.
[0090] The pose estimation is completed by linear interpolation. An intermediate timestamp of a specific non-key camera image (e.g., the first camera image) is set between the two timestamps of two adjacent key camera images (e.g., the second camera image and the third camera image). A linear polynomial within the pose range of the two adjacent key camera images is used to calculate the corresponding intermediate pose (e.g., the first pose) of the specific non-key camera image.
[0091] At block 806, the device filters out camera images with a high degree of overlap based on the specific driving distance indicated in the overlapping camera images. In some embodiments, the same camera images, substantially the same camera images, or camera images with a similarity greater than a specific similarity value are filtered out or removed. For example, camera images captured within a specific driving distance (e.g., two meters) are filtered out or removed because these camera images may be the same, substantially the same, or similar to each other. If all of the acquired camera images 708 are within the specific driving distance, the device returns to block 802 to obtain new camera images 708.
[0092] At block 808, the device obtains image pixel labels (such as cars, buildings, trees, etc.) from a pre-trained image segmentation network. The device identifies objects (such as cars, pedestrians, buildings, trees, lane markings, traffic lights, etc.) and provides image pixel labels for each object (e.g., pedestrians, roads, buildings, cars, trees, etc.). Image segmentation is a process of dividing a camera image into multiple segments. In this process, each pixel in the camera image is associated with an object type.
[0093] Figure 9 is an example flowchart of a process 900 for point cloud projection. In some embodiments, one or more steps of the steps of process 900 are performed by another device or system or another group of devices and / or systems that are separate from or include the autonomous system (e.g., fully and / or partially). For example, it can be performed by Figure 1 the remote AV system 114, Figure 1 the vehicle 102 or Figure 2 the vehicle 200 (e.g., the autonomous system 202 of vehicle 102 or 200), Figure 3 the device 300, Figure 4 the AV computing 400 (e.g., one or more systems of AV computing 400) and / or Figure 7 the 6D colored point cloud generator 712 to perform one or more steps of process 900 (e.g., fully and / or partially). In some embodiments, the steps of process 900 can be performed between any of the above systems in a collaborative manner.
[0094] In some embodiments, at block 902, the device obtains from a pose graph (e.g., Figure 7Obtain a point cloud patch (3D point cloud with corrected pose) from the pose graph 702. Query the 3D point cloud with the corresponding corrected pose from the pose graph through Structured Query Language (SQL). For example, a query is a command or instruction in a domain-specific language for managing data held in a relational database management system. The query enables communication with the database to utilize the data for tasks, functions, and queries. In an example, the query can be used to search the database and for other functions such as creating tables, adding data to tables, modifying data, and dropping tables. The pose graph includes "corrected pose". The poses of the raw point clouds obtained from the LiDAR sensors are not accurate when they are projected onto the world coordinate system because the position of the AV may shift after a long drive. A Simultaneous Localization and Mapping (SLAM) algorithm is introduced to globally optimize all the point cloud poses to obtain point clouds with corrected poses. SLAM is an algorithm that enables the simultaneous real-time construction of a map of the surrounding environment and the localization on that map. The point cloud patch includes a 3D point cloud with the corresponding corrected pose.
[0095] At block 904, the apparatus obtains calibration data (which may be referred to as "transformation data") for transforming the point cloud patch from a driving log (e.g., Figure 7 the driving log 706). The calibration data includes a transformation matrix (K, R, t) between the point cloud generated by a LiDAR sensor (e.g., Figure 2 the LiDAR sensor 202b) and the camera image captured by a camera (e.g., Figure 2 the camera 202a). The transformation matrix (K, R, t) is used to project the point cloud from 3D coordinates to 2D image coordinates. K is the intrinsic matrix of the camera (e.g., Figure 2 the camera 202a); R is the rotation matrix; and t is the translation vector. The intrinsic matrix (K) of the camera includes fixed parameters when the camera is manufactured, and K is provided by the camera manufacturer. The rotation matrix (R) and the translation vector (t) represent the relative positions between one or more LiDARs and one or more cameras in the AV. The distances and angles between one or more LiDARs and one or more cameras can be measured by a measuring device (e.g., a ruler or the like).
[0096] At block 906, the device projects points in the 3D point cloud into image coordinates. The 3D point cloud is in a point cloud coordinate system different from the image coordinate system. In some embodiments, points in the 3D point cloud are projected into the 2D image coordinate system via a pinhole camera model to determine the pixel coordinates of each point.
[0097] The pinhole camera model is represented by Equation 1 below. The pinhole camera model is used to map from a 3D scene to a 2D image. The calibration data obtained at block 904 includes the values of K, R, and t in Equation 1.
[0098] x = K [R t] X (Equation 1)
[0099] where X is the 3D point cloud coordinate; x is the 2D image coordinate; K is the intrinsic matrix of the camera (e.g., Figure 2 camera 202a); R is the rotation matrix; and t is the translation vector. K represents parameters of the camera, such as the focal length of the camera, etc. K is used to project 3D points from the camera coordinates to 2D positions on the same image plane. R represents the orientation of the 3D point cloud relative to the image coordinates. t represents the position of the 3D point cloud relative to the image coordinates. Multiply the 3D point cloud patch obtained at block 902 by the calibration data obtained at block 904 (using Equation 1) to output 2D image coordinates.
[0100] Figure 10 FIG. 1000 is an example flowchart of process 1000 for result refinement. In some embodiments, one or more steps of process 1000 are performed by another device or system or another group of devices and / or systems that are separate from or include the autonomous system (e.g., fully and / or partially). For example, it can be performed by Figure 1 the remote AV system 114, Figure 1 the vehicle 102, or Figure 2 the vehicle 200 (e.g., the autonomous system 202 of vehicle 102 or 200), Figure 3 the device 300, Figure 4 the AV computing 400 (e.g., one or more systems of AV computing 400), and / or Figure 7 the 6D colored point cloud generator 712 (e.g., fully and / or partially). In some embodiments, the steps of process 1000 can be performed among any of the above systems in a collaborative manner.
[0101] In some embodiments, at block 1002, the device applies a first filter to filter out or remove points corresponding to dynamic objects (e.g., pedestrians, vehicles) that may be unnecessary for annotating the map. The device applies the first filter to further filter out or remove points that have different labels (such as semantic labels of pedestrians, roads, buildings, cars, trees, etc.) between the camera image and the point cloud. Semantic labels for the image pixels and the point cloud are obtained from an object detection network. In an example, the object detection network detects instances of certain classified objects (such as pedestrians, buildings, or vehicles, etc.). The semantic labels are derived from semantic segmentation for assigning classification labels to each pixel. Semantic labels are determined using image-based object detection, LiDAR-based object detection, or any combination thereof. In an example, there is a certain misalignment for the aligned image pixels and the point cloud.
[0102] The device applies a second filter to identify points that have at least two different pixel labels (points with at least two aligned image pixels). For each point that has at least two different pixel labels, the device selects the pixel label (image pixel) that is closest in distance to the point. For example, the pose of each image pixel can be obtained from a driving log (e.g., Figure 7 the calibration data 710 in the driving log 706). The distance between an image pixel and a point is the square root of the positional difference between the image pixel and the point. The device assigns the pixel color (R, G, B) from the closest image pixel to each point.
[0103] At block 1004, the device enhances the road marking color by adding an intensity value to the points corresponding to the road markings. In an example, a road marking is a device or indicator applied to a road for specifying traffic flow and / or control. In an example, a road marking is a raised device or sign along the road. Additionally, in an example, a road marking is a painted line, indicator, or other graphic on the road surface. Pixel colors (R, G, B) are assigned to each point in the 3D (X, Y, Z) point cloud (e.g., Figure 7 the 3D point cloud 704) to generate a 6D (X, Y, Z, R, G, B) colored point cloud (e.g., Figure 7 the 6D colored point cloud 714).
[0104] At block 1006, the 6D colored point cloud (updated colored point cloud) is stored in a pose graph (e.g., Figure 7 the pose graph 702) to replace the 3D point cloud.
[0105] Figure 11is an example flowchart of process 1100 for generating a color map layer for map annotation. In some embodiments, one or more steps of process 1100 are performed by another device or system or another group of devices and / or systems that are separate from or include the autonomous system (e.g., fully and / or partially). For example, it can be performed by Figure 1 the remote AV system 114, Figure 1 the vehicle 102, or Figure 2 the vehicle 200 (e.g., the autonomous system 202 of vehicle 102 or 200), Figure 3 the device 300, Figure 4 the AV computing 400 (e.g., one or more systems of AV computing 400) and / or Figure 7 the 6D coloring point cloud generator 712 (e.g., fully and / or partially) to perform one or more steps of process 1100. In some embodiments, the steps of process 1100 can be performed among any of the above systems in a collaborative manner.
[0106] In some embodiments, at block 1102, a processor (e.g., Figure 3 the processor 304) receives a point cloud (e.g., Figure 7 the 3D point cloud 704) from a pose map (e.g., Figure 7 the pose map 702).
[0107] At block 1102, the processor receives a 3D point cloud from a pose map of the vehicle (e.g., Figure 7 the pose map 702). The pose map includes a 3D point cloud with corresponding poses. The 3D point cloud is generated by one or more LiDAR sensors on the vehicle (e.g., LiDAR sensor 202b). One or more LiDAR sensors detect objects while the vehicle is moving.
[0108] At block 1104, the processor receives camera images from the vehicle's driving log (e.g., Figure 7 the driving log 706). Some of the camera images correspond to the 3D point cloud from the pose map (e.g., the corresponding camera images have the same or approximately the same timestamp as the 3D point cloud). One or more cameras on the vehicle (e.g., Figure 2 the camera 202a) capture camera images while the vehicle is moving.
[0109] At block 1106, the processor obtains image pixel labels for the camera images. For example, a pre-trained image segmentation network (e.g., Figure 8The image segmentation network at block 808) can segment each camera image and generate pixel labels or pixel colors for each pixel of each camera image.
[0110] At block 1108, the processor projects the 3D point cloud in the point cloud coordinate system into the image coordinate system through, for example, a pinhole camera model. The point cloud coordinate system is three-dimensional (X, Y, Z), while the image coordinate system is two-dimensional (X, Y). The Z coordinate is ignored.
[0111] At block 1110, the processor generates a 6D colored point cloud by combining the 3D point cloud and color information (e.g., RGB data) from the camera image. The point cloud is projected into the image coordinate system at block 1108, and thus color information can be added to or combined with the 3D point cloud to generate a 6D colored point cloud (X, Y, Z, R, G, B).
[0112] At block 1112, the processor transforms the 6D colored point cloud into a 5D (X, Y, R, G, B) map tile to form a colored map layer. The 6D colored point cloud can be transformed into a 5D map tile by ignoring the Z coordinate value. In some embodiments, the processor removes points with a height greater than a specific height value (e.g., 1.8 meters) from the 6D colored point cloud. The device 300 then transforms the truncated 6D colored point cloud (i.e., the 6D colored point cloud including the remaining points) into a 5D map tile by ignoring the Z coordinate value. The 5D map tile (i.e., the colored map layer) is provided to a map annotation tool (e.g., a map rasterizer) for map annotation.
[0113] The technology of the present invention can provide accurate alignment between colors and points. A colored map layer is provided to accelerate the map annotation process. By providing baseline multimodal fusion between the point cloud and the camera images and incorporating multiple surrounding environment camera images for sensor fusion, the technology of the present disclosure can benefit downstream semantic tasks such as lane extractor networks.
[0114] According to some non-limiting embodiments or examples, a method is provided, including: using at least one processor to receive a point cloud from a pose graph; using the at least one processor to receive an image from a driving log of a vehicle, the image corresponding to the point cloud from the pose graph; using the at least one processor to obtain an image pixel label for the image; using the at least one processor to project the point cloud in the point cloud coordinate system into the image coordinate system based on the image pixel label; using the at least one processor to generate a six-dimensional colored point cloud by combining the point cloud and color information from the image; and using the at least one processor to transform the six-dimensional colored point cloud into a five-dimensional map tile to form a colored map layer.
[0115] According to some non - limiting embodiments or examples, a system is provided, including: at least one processor; and a memory storing instructions thereon, which when executed by the at least one processor cause the at least one processor to perform operations, the operations including: receiving a point cloud from a pose graph; receiving an image from a vehicle's driving log, the image corresponding to the point cloud from the pose graph; obtaining an image pixel label for the image; projecting the point cloud in the point cloud coordinate system to the image coordinate system based on the image pixel label; generating a six - dimensional colored point cloud by combining the point cloud and color information from the image; and transforming the six - dimensional colored point cloud into a five - dimensional map tile to form a colored map layer.
[0116] According to some non - limiting embodiments or examples, a non - transitory computer - readable storage medium is provided, storing instructions thereon, which when executed by at least one processor cause the at least one processor to perform operations, the operations including: receiving a point cloud from a pose graph; receiving an image from a vehicle's driving log, the image corresponding to the point cloud from the pose graph; obtaining an image pixel label for the image; projecting the point cloud in the point cloud coordinate system to the image coordinate system based on the image pixel label; generating a six - dimensional colored point cloud by combining the point cloud and color information from the image; and transforming the six - dimensional colored point cloud into a five - dimensional map tile to form a colored map layer.
[0117] Clause 1: A method includes: using at least one processor to receive a point cloud from a pose graph; using the at least one processor to receive an image from a vehicle's driving log, the image corresponding to the point cloud from the pose graph; using the at least one processor to obtain an image pixel label for the image; using the at least one processor to project the point cloud in the point cloud coordinate system to the image coordinate system based on the image pixel label; using the at least one processor to generate a six - dimensional colored point cloud by combining the point cloud and color information from the image; and using the at least one processor to transform the six - dimensional colored point cloud into a five - dimensional map tile to form a colored map layer.
[0118] Clause 2: The method according to Clause 1 further includes: receiving calibration data from the vehicle's driving log, where the calibration data includes a transformation matrix, and the transformation matrix includes an intrinsic matrix of a camera, a rotation matrix, and a translation vector; and projecting the point cloud to the image coordinate system based on the calibration data.
[0119] Clause 3: The method according to Clause 1 or 2, wherein the image pixel label is obtained from a pre - trained image segmentation network.
[0120] Clause 4: The method according to any one of Clauses 1-3 further includes: interpolating the pose of the image according to the poses of two adjacent images, wherein the two adjacent images correspond to two point clouds, and wherein the two point clouds include a first point cloud having a timestamp immediately earlier than the timestamp of the image and a second point cloud having a timestamp immediately later than the timestamp of the image.
[0121] Clause 5: The method according to any one of Clauses 1-4 further includes: filtering out the overlapping images according to a specific driving distance indicated in the overlapping images.
[0122] Clause 6: The method according to any one of Clauses 1-5, wherein projecting the point cloud includes: projecting the point cloud through a pinhole camera model.
[0123] Clause 7: The method according to any one of Clauses 1-6 further includes: removing the point having different labels between the point and the corresponding image pixel from the point cloud.
[0124] Clause 8: The method according to any one of Clauses 1-7 further includes: removing the points representing dynamic objects from the point cloud, the dynamic objects including one or more of pedestrians, bicycles, and vehicles.
[0125] Clause 9: The method according to any one of Clauses 1-8, wherein projecting the point cloud further includes: when the point in the point cloud has at least two different image pixels, assigning the pixel color from the closest image pixel to the point color of the point.
[0126] Clause 10: The method according to any one of Clauses 1-9 further includes: enhancing the road marking color by adding intensity values to each point in the point cloud.
[0127] Clause 11: The method according to any one of Clauses 1-10 further includes: updating the pose graph using the six-dimensional colored point cloud.
[0128] Clause 12: The method according to any one of Clauses 1-11, wherein transforming the six-dimensional colored point cloud into the five-dimensional map tile includes: removing the points corresponding to the objects having a height greater than a specific height value from the six-dimensional colored point cloud; and transforming the six-dimensional colored point cloud into the five-dimensional map tile by ignoring the Z coordinate value.
[0129] Clause 13: A system includes: at least one processor; and a memory storing instructions that, when executed by the at least one processor, cause the at least one processor to perform operations including: receiving a point cloud from a pose graph; receiving an image from a vehicle's driving log, the image corresponding to the point cloud from the pose graph; obtaining an image pixel label for the image; projecting the point cloud in a point cloud coordinate system onto an image coordinate system based on the image pixel label; generating a six-dimensional colored point cloud by combining the point cloud and color information from the image; and transforming the six-dimensional colored point cloud into a five-dimensional map tile to form a colored map layer.
[0130] Clause 14: The system according to Clause 13, wherein the operations further include: receiving calibration data from the vehicle's driving log, wherein the calibration data includes a transformation matrix including an intrinsic matrix, a rotation matrix, and a translation vector of a camera; and projecting the point cloud onto the image coordinate system based on the calibration data.
[0131] Clause 15: The system according to Clause 13 or 14, wherein the operations further include: interpolating the pose of the image based on the poses of two adjacent images, wherein the two adjacent images correspond to two point clouds, and wherein the two point clouds include a first point cloud having a timestamp immediately earlier than the timestamp of the image and a second point cloud having a timestamp immediately later than the timestamp of the image.
[0132] Clause 16: The system according to any one of Clauses 13-15 further includes: removing from the point cloud a point having different labels between the point and the corresponding image pixel.
[0133] Clause 17: The system according to any one of Clauses 13-16 further includes: removing from the point cloud points representing dynamic objects, the dynamic objects including one or more of pedestrians, bicycles, and vehicles.
[0134] Clause 18: The system according to any one of Clauses 13-17, wherein projecting the point cloud further includes: when a point in the point cloud has at least two different image pixels, assigning the pixel color from the closest image pixel to the point color of the point.
[0135] Clause 19: The system according to any one of Clauses 13-18, wherein transforming the six-dimensional colored point cloud into the five-dimensional map tile includes: removing from the six-dimensional colored point cloud points corresponding to objects having a height greater than a specific height value; and transforming the six-dimensional colored point cloud into the five-dimensional map tile by ignoring the Z coordinate value.
[0136] Clause 20: A non-transitory computer-readable storage medium having instructions stored thereon that, when executed by at least one processor, cause the at least one processor to perform operations including: receiving a point cloud from a pose map; receiving an image from a vehicle's driving log, the image corresponding to the point cloud from the pose map; obtaining an image pixel label for the image; projecting the point cloud in a point cloud coordinate system onto an image coordinate system based on the image pixel label; generating a six-dimensional colored point cloud by combining the point cloud and color information from the image; and transforming the six-dimensional colored point cloud into a five-dimensional map tile to form a colored map layer.
[0137] In the foregoing description, aspects and embodiments of the present disclosure have been described with reference to numerous specific details, which may vary depending on the implementation. Accordingly, the specification and drawings are to be regarded as illustrative rather than restrictive in nature. The sole and exclusive indication of the scope of the invention, and what the applicant desires to be the scope of the invention, is the literal and equivalent scope of the claims as issued from this application in the specific form of the issued claims, including any subsequent amendments. Any definition of terms expressly set forth herein for inclusion in such claims shall be construed to have the meaning such terms have when used in the claims. Additionally, when the term "further comprising" is used in the foregoing specification or the appended claims, the text following that phrase may be additional steps or entities, or sub-steps / sub-entities of the previously recited steps or entities.
Claims
1. A method, comprising: Receiving, by means of at least one processor, a point cloud from a pose graph; Receiving, by means of the at least one processor, an image from a driving log of a vehicle, the image corresponding to the point cloud from the pose graph; Obtaining, by means of the at least one processor, an image pixel label for the image; Projecting, by means of the at least one processor, the point cloud in a point cloud coordinate system onto an image coordinate system based on the image pixel label; Generating, by means of the at least one processor, a six-dimensional colored point cloud by combining the point cloud and color information from the image; And Transforming, by means of the at least one processor, the six-dimensional colored point cloud into a five-dimensional map tile to form a colored map layer.
2. The method according to claim 1, further comprising: Receiving calibration data from the driving log of the vehicle, wherein the calibration data includes a transformation matrix, the transformation matrix including an intrinsic matrix of a camera, a rotation matrix, and a translation vector; and Projecting the point cloud onto the image coordinate system based on the calibration data.
3. The method according to claim 1 or 2, wherein The image pixel label is obtained from a pre-trained image segmentation network.
4. The method according to any one of claims 1-3, further comprising: Interpolating the pose of the image according to the poses of two adjacent images, wherein the two adjacent images correspond to two point clouds, and wherein the two point clouds include a first point cloud having a timestamp immediately earlier than the timestamp of the image and a second point cloud having a timestamp immediately later than the timestamp of the image.
5. The method according to any one of claims 1-4, further comprising: Filtering out the overlapping images according to a specific driving distance indicated in the overlapping images.
6. The method according to any one of claims 1-5, wherein Projecting the point cloud includes: Projecting the point cloud through a pinhole camera model.
7. The method according to any one of claims 1-6, further comprising: Removing from the point cloud a point having different labels between the point and the corresponding image pixel.
8. The method according to any one of claims 1-7, further comprising: Removing from the point cloud points representing dynamic objects, the dynamic objects including one or more of pedestrians, bicycles, and vehicles.
9. The method according to any one of claims 1-8, wherein Projecting the point cloud further includes: Assigning, in the case where a point in the point cloud has at least two different image pixels, the pixel color from the closest image pixel to the point color of the point.
10. The method according to any one of claims 1-9, further comprising: Enhancing the road marking color by adding intensity values to each point in the point cloud.
11. The method according to any one of claims 1-10, further comprising: Updating the pose graph using the six-dimensional colored point cloud.
12. The method according to any one of claims 1 to 11, wherein, Transforming the six-dimensional colored point cloud into the five-dimensional map tile includes: Removing from the six-dimensional colored point cloud points corresponding to objects having a height greater than a specific height value; and Transforming the six-dimensional colored point cloud into the five-dimensional map tile by ignoring the Z coordinate value.
13. A system, comprising: At least one processor; And A memory storing instructions that, when executed by the at least one processor, cause the at least one processor to perform operations, the operations including: Receiving a point cloud from a pose graph; Receiving an image from a vehicle's driving log, the image corresponding to the point cloud from the pose graph; Obtaining an image pixel label for the image; Projecting the point cloud in the point cloud coordinate system onto the image coordinate system based on the image pixel label; Generating a six-dimensional colored point cloud by combining the point cloud and color information from the image; and Transforming the six-dimensional colored point cloud into a five-dimensional map tile to form a colored map layer.
14. The system according to claim 13, wherein the operations further include: Receiving calibration data from the vehicle's driving log, wherein the calibration data includes a transformation matrix, the transformation matrix including an intrinsic matrix, a rotation matrix, and a translation vector of a camera; and Projecting the point cloud onto the image coordinate system based on the calibration data.
15. The system according to claim 13 or 14, wherein the operations further include: Interpolating the pose of the image based on the poses of two adjacent images, wherein the two adjacent images correspond to two point clouds, wherein the two point clouds include a first point cloud having a timestamp that is immediately earlier than the timestamp of the image and a second point cloud having a timestamp that is immediately later than the timestamp of the image.
16. The system according to any one of claims 13-15, further comprising: Removing a point from the point cloud that has a different label between the point and the corresponding image pixel.
17. The system according to any one of claims 13-16, further comprising: Removing points representing dynamic objects from the point cloud, the dynamic objects including one or more of pedestrians, bicycles, and vehicles.
18. The system according to any one of claims 13 - 17, wherein, Projecting the point cloud further includes: Assigning the pixel color from the closest image pixel to the point color of the point in the case where the point in the point cloud has at least two different image pixels.
19. The system according to any one of claims 13-18, wherein, Transforming the six-dimensional colored point cloud into the five-dimensional map tile includes: Removing points corresponding to objects having a height greater than a specific height value from the six-dimensional colored point cloud; and Transforming the six-dimensional colored point cloud into the five-dimensional map tile by ignoring the Z coordinate value.
20. A non-transitory computer-readable storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform operations, the operations including: Receiving a point cloud from a pose graph; Receiving an image from a vehicle's driving log, the image corresponding to the point cloud from the pose graph; Obtaining an image pixel label for the image; Projecting the point cloud in the point cloud coordinate system onto the image coordinate system based on the image pixel label; Generating a six-dimensional colored point cloud by combining the point cloud and color information from the image; And Transforming the six-dimensional colored point cloud into a five-dimensional map tile to form a colored map layer.