Camera-assisted LiDAR data verification
By combining the data of the camera and LiDAR sensor, the LiDAR confidence score is updated using the image semantic network, which solves the accuracy of LiDAR data verification, and improves the accuracy of object detection and the system performance of autonomous vehicles.
Patent Information
- Application Number
- CN202380084829.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-14
- Filing Date
- 2023-10-02
- Publication Date
- 2025-08-08
AI Technical Summary
In the prior art, there is insufficient accuracy in the verification of LiDAR data of autonomous vehicles, which affects the accuracy of object detection and the downstream performance of autonomous systems.
By combining the data of the camera and LiDAR sensor, using image semantic networks and LiDAR semantic networks for data verification, the confidence score of LiDAR is updated using the confidence score of the camera to improve the accuracy of LiDAR data.
It improves the data accuracy of LiDAR semantic network output, improves the accuracy of object detection, and thus improves the performance of the perception, planning and control system of autonomous vehicles.
Smart Images

Figure CN120457461A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to U.S. patent application No. 63,416 / 487, filed on October 14, 2022, entitled “Camera-Assisted LiDAR Data Verification for Self-Driving Vehicles,” which is incorporated herein by reference in its entirety. Background Art
[0003] Autonomous vehicles use sensor data to perform object detection. In object detection, sensor data is analyzed to determine the presence of instances of an object class. Autonomous vehicles navigate their environment based on the detected objects. For example, they generate routes to avoid collisions with detected objects. BRIEF DESCRIPTION OF THE DRAWINGS
[0004] Figure 1 is an example environment in which a vehicle including one or more components of an autonomous system may be implemented;
[0005] Figure 2 is a diagram of one or more systems of a vehicle including an autonomous system;
[0006] Figure 3 yes Figure 1 and Figure 2 a diagram of one or more devices and / or components of one or more systems;
[0007] Figure 4 is a diagram of some components of an autonomous system;
[0008] Figure 5 A diagram illustrating an implementation of camera-assisted LiDAR data validation;
[0009] Figure 6A Shows the impact of ambient lighting on a vehicle's LiDAR sensor and camera;
[0010] Figure 6B Shows the effect of objects of different colors on the vehicle's LiDAR sensor and camera;
[0011] Figure 6C Shows the impact of objects with different properties on the vehicle's LiDAR sensor and camera;
[0012] Figure 7 Shows the workflow for camera-assisted LiDAR data validation;
[0013] Figure 8Demonstrates the application of camera-assisted LiDAR data verification;
[0014] Figure 9 is a flow chart of a process that enables camera-assisted LiDAR data verification;
[0015] Figure 10 A bird's-eye view showing the light emitted by a LiDAR sensor and reflected at different detection angles;
[0016] Figure 11 Shows the workflow for camera-assisted LiDAR data validation;
[0017] Figure 12 A flow chart illustrating a process for camera-assisted LiDAR data validation; and
[0018] Figure 13 A flow chart illustrating a process for LiDAR semantic network confidence scoring based on angle information. DETAILED DESCRIPTION
[0019] In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it will be apparent that the embodiments described herein can be practiced without these specific details. In some instances, well-known configurations and devices are illustrated in block diagram form to avoid unnecessarily obscuring aspects of the present disclosure.
[0020] In the accompanying drawings, for ease of description, a specific arrangement or order of schematic elements (such as those representing systems, devices, modules, instruction blocks and / or data elements, etc.) is illustrated. However, those skilled in the art will understand that, unless expressly described, the specific order or arrangement of schematic elements in the accompanying drawings is not intended to require a specific processing order or sequence, or separation of processes. Furthermore, unless expressly described, the inclusion of a schematic element in a drawing is not intended to mean that such element is required in all embodiments, nor is it intended to mean that features represented by such element cannot be included in some embodiments or cannot be combined with other elements in some embodiments.
[0021] In addition, in the accompanying drawings, connecting elements (such as solid or dotted lines or arrows) are used to illustrate the connection, relationship or association between or among two or more other schematic elements, and there is no such connecting element and is not intended to mean that there can be no connection, relationship or association. In other words, some connections, relationships or associations between elements are not illustrated in the accompanying drawings, so as not to obscure the present disclosure. In addition, for ease of illustration, a single connecting element can be used to represent multiple connections, relationships or associations between elements. For example, if a connecting element represents the communication of a signal, data or instruction (for example, "software instruction"), it will be understood by those skilled in the art that this element can represent one or more signal paths (for example, bus) that may be needed to affect communication.
[0022] Although the terms "first," "second," and / or "third," etc. are used to describe various elements, these elements should not be limited by these terms. The terms "first," "second," and / or "third" are only used to distinguish one element from another. For example, a first contact may be referred to as a second contact, and similarly, a second contact may be referred to as a first contact without departing from the scope of the described embodiments. Both the first contact and the second contact are contacts, but they are not the same contact.
[0023] The terms used in the description of the various embodiments described herein are included only for the purpose of describing specific embodiments and are not intended to be limiting. As used in the description of the various embodiments described and the appended claims, the singular forms "a", "an", and "the" are also intended to include the plural forms and can be used interchangeably with "one or more than one" or "at least one" unless the context clearly indicates otherwise. It will also be understood that the term "and / or" as used herein refers to and includes any and all possible combinations of one or more than one of the associated listed items. It will also be understood that when the terms "comprises", "comprising", "having", and / or "having" are used in this specification, the presence of the stated features, integers, steps, operations, elements, and / or components is specified, but the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof is not excluded.
[0024] As used herein, the terms "communication" and "communicating" refer to at least one of receiving, receiving, transmitting, transferring, and / or providing information (or information represented by, for example, data, signals, messages, instructions, and / or commands). For a unit (e.g., a device, a system, a component of a device or system, and / or a combination thereof) to communicate with another unit, this means that the unit is able to directly or indirectly receive information from the other unit and / or send (e.g., transmit) information to the other unit. This can refer to a direct or indirect connection that is wired and / or wireless in nature. In addition, two units can communicate with each other even if the transmitted information can be modified, processed, relayed, and / or routed between the first unit and the second unit. For example, a first unit can communicate with a second unit even if the first unit passively receives information and does not actively transmit information to the second unit. As another example, a first unit can communicate with a second unit if at least one intermediary unit (e.g., a third unit located between the first unit and the second unit) processes information received from the first unit and transmits the processed information to the second unit. In some embodiments, a message may refer to a network packet (eg, a data packet, etc.) that includes data.
[0025] As used herein, the term "if" is optionally interpreted to mean "when," "at the time of," "in response to being determined to be," and / or "in response to being detected," etc., depending on the context. Similarly, the phrases "if it is determined" or "if [the stated condition or event] is detected" are optionally interpreted to mean "upon determining," "in response to being determined to be" or "upon detecting [the stated condition or event]," and / or "in response to detecting [the stated condition or event]," etc., depending on the context. Furthermore, as used herein, the terms "have," "have," or "possess," etc. are intended to be open-ended terms. Furthermore, unless expressly stated otherwise, the phrase "based on" is intended to mean "based at least in part on."
[0026] Reference will now be made in detail to the embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the various embodiments described. However, it will be apparent to one of ordinary skill in the art that the various embodiments described may be practiced without these specific details. In other instances, well-known methods, processes, components, circuits, and networks have not been described in detail so as not to unnecessarily obscure aspects of the embodiments.
[0027] General Overview
[0028] In some aspects and / or embodiments, the systems, methods, and computer program products described herein include and / or implement camera-assisted LiDAR data validation. A vehicle (such as an autonomous vehicle, etc.) has multiple sensors mounted at various locations on the vehicle. Data from these sensors can be used for object detection. In object detection, sensor data is analyzed to annotate portions of the sensor data with confidence scores, wherein the confidence scores indicate the presence of instances of a particular object class within the corresponding portions of the data captured by the sensor. In some embodiments, a first semantic segmentation network (e.g., a LiDAR semantic network) and a second semantic segmentation network (e.g., an image semantic network) are executed with the sensor data as input. A difference is determined between a first location output by the first semantic segmentation network (e.g., a 3D location from LiDAR) and a second location output by the second semantic segmentation network (e.g., a 2D location from a camera). When the difference between the first location and the second location meets a first predetermined threshold (e.g., a distance threshold), the first confidence score (e.g., LiDAR confidence data) output by the first semantic segmentation network is updated based on the second confidence score (e.g., camera confidence data) output by the second semantic segmentation network. In an example, an output of the first semantic segmentation network including an updated first confidence score and an output of the second semantic segmentation network are concatenated and used to detect objects.
[0029] In some embodiments, a first semantic segmentation network (e.g., an image semantic network) obtains camera data as input and outputs first object space information of an object, object attribute information of an object, object color information of an object, and a first detection confidence score associated with the object. Angle information (e.g., the angle of the detected surface of the object relative to the vehicle) is determined based on the first object space information. A second semantic segmentation network (e.g., a LiDAR semantic network) obtains second sensor data (e.g., LiDAR data, point cloud data), angle information, object attribute information, object color information, and a first detection confidence score as input and outputs a second object space location and a second detection confidence score. In an example, the output of the second semantic segmentation network is used for object detection.
[0030] By means of implementation of the systems, methods, and computer program products described herein, techniques for camera-assisted LiDAR data validation enable improved accuracy of data output by a LiDAR semantic network. Consequently, the resulting object detections are more accurate than object detections without LiDAR data validation. The present techniques use the existing output of an image semantic network to improve the accuracy of confidence scores output by a LiDAR semantic network in real time. The improved accuracy of the confidence scores enables autonomous vehicle (AV) stacks such as Figure 4The downstream performance of AV stacks (e.g., the AV stack) is improved. In particular, the performance of AV systems such as perception systems, planning systems, localization systems, and / or control systems is improved based on the use of accurate confidence scores output by the LiDAR semantic network.
[0031] Now refer to Figure 1 , illustrates an example environment 100 in which vehicles including autonomous systems and vehicles not including autonomous systems operate. As illustrated, environment 100 includes vehicles 102a-102n, objects 104a-104n, routes 106a-106n, area 108, vehicle-to-infrastructure (V2I) devices 110, network 112, remote autonomous vehicle (AV) systems 114, queue management system 116, and V2I system 118. Vehicles 102a-102n, vehicle-to-infrastructure (V2I) devices 110, network 112, autonomous vehicle (AV) systems 114, queue management system 116, and V2I system 118 are interconnected (e.g., establish connections for communication, etc.) via wired connections, wireless connections, or a combination of wired or wireless connections. In some embodiments, objects 104a-104n are interconnected with at least one of vehicles 102a-102n, vehicle-to-infrastructure (V2I) devices 110, network 112, autonomous vehicle (AV) system 114, fleet management system 116, and V2I system 118 via a wired connection, a wireless connection, or a combination of wired or wireless connections.
[0032] Vehicles 102a-102n (individually referred to as vehicles 102 and collectively referred to as vehicles 102) include at least one device configured to transport goods and / or people. In some embodiments, vehicles 102 are configured to communicate with V2I devices 110, remote AV systems 114, fleet management systems 116, and / or V2I systems 118 via network 112. In some embodiments, vehicles 102 include cars, buses, trucks, and / or trains. In some embodiments, vehicles 102 are similar to vehicles 200 described herein (see Figure 2 ) are the same or similar. In some embodiments, vehicles 200 in the set of vehicles 200 are associated with an autonomous queue manager. In some embodiments, as described herein, vehicles 102 travel along corresponding routes 106a-106n (individually referred to as routes 106 and collectively referred to as routes 106). In some embodiments, one or more vehicles 102 include an autonomous system (e.g., an autonomous system that is the same or similar to autonomous system 202).
[0033] Objects 104a-104n (individually referred to as object 104 and collectively referred to as objects 104) include, for example, at least one vehicle, at least one pedestrian, at least one cyclist, and / or at least one structure (e.g., a building, a sign, a fire hydrant, etc.). Each object 104 is stationary (e.g., located at a fixed location and over a period of time) or moving (e.g., having a velocity and associated with at least one trajectory). In some embodiments, objects 104 are associated with corresponding locations in area 108.
[0034] Routes 106a-106n (individually referred to as routes 106 and collectively referred to as routes 106) are each associated with (e.g., specifying) a series of actions (also referred to as trajectories) connecting states along which an AV can navigate. Each route 106 begins at an initial state (e.g., a state corresponding to a first spatiotemporal location and / or speed, etc.) and ends at a final target state (e.g., a state corresponding to a second spatiotemporal location different from the first spatiotemporal location) or a target zone (e.g., a subspace of acceptable states (e.g., terminal states)). In some embodiments, the first state includes a location where one or more individuals will board the AV, and the second state or zone includes one or more locations where the one or more individuals boarding the AV will disembark. In some embodiments, routes 106 include multiple acceptable state sequences (e.g., multiple spatiotemporal location sequences) that are associated with (e.g., define) multiple trajectories. In examples, routes 106 include only high-level actions or imprecise state locations, such as a series of connecting roads indicating a change of direction at a roadway intersection. Additionally or alternatively, the route 106 may include more precise actions or states, such as, for example, a specific target lane or precise locations within a lane zone and target speeds at those locations. In an example, the route 106 includes multiple precise state sequences along at least one high-level action with a limited look-ahead horizon to an intermediate goal, where the combination of consecutive iterations of the limited-horizon state sequences cumulatively corresponds to multiple trajectories that collectively form a high-level route terminating at a final target state or zone.
[0035] The area 108 includes a physical area (e.g., a geographic region) that the vehicle 102 can navigate. In an example, the area 108 includes at least one state (e.g., a country, a province, a separate state within a plurality of states included in a country, etc.), at least a portion of a state, at least one city, at least a portion of a city, etc. In some embodiments, the area 108 includes at least one named thoroughfare (referred to herein as a "road"), such as a highway, an interstate highway, a parkway, a city street, etc. Additionally or alternatively, in some examples, the area 108 includes at least one unnamed road, such as a driveway, a section of a parking lot, a section of an open space and / or undeveloped area, a dirt road, etc. In some embodiments, the road includes at least one lane (e.g., a portion of the road that the vehicle 102 can traverse). In an example, the road includes at least one lane associated with (e.g., identified based on) at least one lane marking line.
[0036] Vehicle-to-infrastructure (V2I) devices 110 (sometimes referred to as vehicle-to-infrastructure or vehicle-to-everything (V2X) devices) include at least one device configured to communicate with vehicle 102 and / or V2I system 118. In some embodiments, V2I devices 110 are configured to communicate with vehicle 102, remote AV system 114, fleet management system 116, and / or V2I system 118 via network 112. In some embodiments, V2I devices 110 include radio frequency identification (RFID) devices, signs, cameras (e.g., two-dimensional (2D) and / or three-dimensional (3D) cameras), lane markings, streetlights, parking meters, and the like. In some embodiments, V2I devices 110 are configured to communicate directly with vehicle 102. Additionally or alternatively, in some embodiments, the V2I device 110 is configured to communicate with the vehicle 102, the remote AV system 114, and / or the fleet management system 116 via the V2I system 118. In some embodiments, the V2I device 110 is configured to communicate with the V2I system 118 via the network 112.
[0037] The network 112 includes one or more wired and / or wireless networks. In an example, the network 112 includes a cellular network (e.g., a long-term evolution (LTE) network, a third-generation (3G) network, a fourth-generation (4G) network, a fifth-generation (5G) network, a code division multiple access (CDMA) network, etc.), a public land mobile network (PLMN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), a telephone network (e.g., a public switched telephone network (PSTN)), a private network, an ad hoc network, an intranet, the Internet, a fiber-optic-based network, a cloud computing network, etc., and / or a combination of some or all of these networks.
[0038] Remote AV system 114 includes at least one device configured to communicate with vehicle 102, V2I device 110, network 112, fleet management system 116, and / or V2I system 118 via network 112. In an example, remote AV system 114 includes a server, a server group, and / or other similar devices. In some embodiments, remote AV system 114 is co-located with fleet management system 116. In some embodiments, remote AV system 114 participates in the installation of some or all of the vehicle's components (including autonomous systems, autonomous vehicle computing, and / or software implemented by autonomous vehicle computing). In some embodiments, remote AV system 114 maintains (e.g., updates and / or replaces) these components and / or software during the vehicle's lifetime.
[0039] The queue management system 116 includes at least one device configured to communicate with the vehicles 102, the V2I devices 110, the remote AV system 114, and / or the V2I system 118. In an example, the queue management system 116 includes a server, a server group, and / or other similar devices. In some embodiments, the queue management system 116 is associated with a ride-sharing company (e.g., an organization that controls the operation of multiple vehicles (e.g., vehicles that include autonomous systems and / or vehicles that do not include autonomous systems).
[0040] In some embodiments, the V2I system 118 includes at least one device configured to communicate with the vehicle 102, the V2I device 110, the remote AV system 114, and / or the fleet management system 116 via the network 112. In some examples, the V2I system 118 is configured to communicate with the V2I device 110 via a connection other than the network 112. In some embodiments, the V2I system 118 includes a server, a server group, and / or other similar devices. In some embodiments, the V2I system 118 is associated with a municipality or a private entity (e.g., a private entity that maintains the V2I device 110).
[0041] supply Figure 1 The number and arrangement of elements illustrated are examples. Figure 1 There may be additional elements, fewer elements, different elements, and / or differently arranged elements than those illustrated. Additionally or alternatively, at least one element of the environment 100 may be described as being Figure 1 Additionally or alternatively, at least one set of elements of environment 100 may perform one or more functions described as being performed by at least one different set of elements of environment 100.
[0042] Now refer to Figure 2 , vehicle 200 (which can be Figure 1 102) includes or is associated with autonomous system 202, powertrain control system 204, steering control system 206, and braking system 208. In some embodiments, vehicle 200 is similar to vehicle 102 (see Figure 1 ) are the same or similar. In some embodiments, the autonomous system 202 is configured to give the vehicle 200 autonomous driving capabilities (e.g., implementing at least one driving automatic or maneuver-based function, feature and / or device, etc., which enables the vehicle 200 to operate partially or completely without human intervention, including but not limited to fully autonomous vehicles (e.g., vehicles that abandon reliance on human intervention, such as Level 5 ADS operating vehicles, etc.), highly autonomous vehicles (e.g., vehicles that abandon reliance on human intervention in certain situations, such as Level 4 ADS operating vehicles, etc.), and / or conditionally autonomous vehicles (e.g., vehicles that abandon reliance on human intervention in limited situations, such as Level 3 ADS operating vehicles, etc.). In one embodiment, the autonomous system 202 includes the operational or tactical functionality required to enable the vehicle 200 to operate in traffic on the road and continuously perform part or all of a dynamic driving task (DDT). In another embodiment, the autonomous system 202 includes an advanced driver assistance system (ADAS) that includes driver support features. The autonomous system 202 supports various levels of driving automation ranging from no driving automation (e.g., Level 0) to full driving automation (e.g., Level 5). For a detailed description of fully autonomous vehicles and highly autonomous vehicles, reference can be made to SAE International's standard J3016: Taxonomy and Definitions for Terms Related to On-Road Motor Vehicle Automated Driving Systems, the entire contents of which are incorporated by reference. In some embodiments, the vehicle 200 is associated with an autonomous queue manager and / or a ridesharing company.
[0043] Autonomous system 202 includes a sensor suite comprising one or more devices, such as a camera 202a, a LiDAR sensor 202b, a Radar sensor 202c, and a microphone 202d. In some embodiments, autonomous system 202 may include more, fewer, and / or different devices (e.g., ultrasonic sensors, inertial sensors, a GPS receiver (discussed below), and / or an odometer sensor for generating data associated with an indication of the distance traveled by vehicle 200). In some embodiments, autonomous system 202 uses one or more devices included in autonomous system 202 to generate data associated with environment 100, as described herein. The data generated by one or more devices of autonomous system 202 may be used by one or more systems described herein to observe the environment in which vehicle 200 is located (e.g., environment 100). In some embodiments, autonomous system 202 includes a communication device 202e, autonomous vehicle computing 202f, a drive-by-wire (DBW) system 202h, and a safety controller 202g.
[0044] The camera 202a includes a camera configured to communicate with the communication device 202e, the autonomous vehicle computer 202f, and / or the safety controller 202g via a bus (e.g., Figure 3 The camera 202a includes at least one device for communicating with the bus 302 (the same or similar bus as the bus 302). The camera 202a includes at least one camera (e.g., a digital camera using a light sensor such as a charge coupled device (CCD), a thermal camera, an infrared (IR) camera, and / or an event camera, etc.) to capture images including physical objects (e.g., cars, buses, curbs and / or people, etc.). In some embodiments, the camera 202a generates camera data as output. In some examples, the camera 202a generates camera data including image data associated with the image. In this example, the image data may specify at least one parameter corresponding to the image (e.g., image characteristics such as exposure, brightness, and / or image timestamp, etc.). In such an example, the image may be in a format (e.g., RAW, JPEG, and / or PNG, etc.). In some embodiments, the camera 202a includes a plurality of independent cameras configured (e.g., positioned) on the vehicle to capture images for the purpose of stereoscopic imaging (stereo vision). In some examples, the camera 202a includes a computer system that generates image data and transmits the image data to the autonomous vehicle computing 202f and / or a fleet management system (e.g., with Figure 1The autonomous vehicle computing system 202f may be configured to include multiple cameras (e.g., a fleet management system similar to or similar to the fleet management system 116 of the plurality of cameras). In such an example, the autonomous vehicle computing system 202f determines a depth to one or more objects in the field of view of at least two of the plurality of cameras based on image data from the at least two cameras. In some embodiments, the camera 202a is configured to capture images of objects within a distance relative to the camera 202a (e.g., up to 100 meters and / or up to 1 kilometer, etc.). Accordingly, the camera 202a includes features, such as a sensor and a lens, that are optimized for sensing objects at one or more distances relative to the camera 202a.
[0045] In embodiments, camera 202a includes at least one camera configured to capture one or more images associated with one or more traffic lights, street signs, and / or other physical objects that provide visual navigation information. In some embodiments, camera 202a generates traffic light data associated with the one or more images. In some examples, camera 202a generates TLD (traffic light detection) data associated with the one or more images in a format such as RAW, JPEG, and / or PNG. In some embodiments, camera 202a that generates TLD data differs from other systems incorporating cameras described herein in that camera 202a may include one or more cameras with a wide field of view (e.g., a wide-angle lens, a fisheye lens, and / or a lens with a viewing angle of approximately 120 degrees or greater) to generate images associated with as many physical objects as possible.
[0046] The light detection and ranging (LiDAR) sensor 202b includes a sensor configured to communicate with the communication device 202e, the autonomous vehicle computing 202f and / or the safety controller 202g via a bus (e.g., Figure 3The LiDAR sensor 202b includes at least one device that communicates with a bus (the same or similar bus as the bus 302) that is connected to the LiDAR sensor 202b. The LiDAR sensor 202b includes a system configured to emit light from a light emitter (e.g., a laser emitter). The light emitted by the LiDAR sensor 202b includes light outside the visible spectrum (e.g., infrared light, etc.). In some embodiments, during operation, the light emitted by the LiDAR sensor 202b encounters a physical object (e.g., a vehicle) and is reflected back to the LiDAR sensor 202b. In some embodiments, the light emitted by the LiDAR sensor 202b does not penetrate the physical object encountered by the light. The LiDAR sensor 202b also includes at least one light detector that detects the light emitted from the light emitter after it encounters the physical object. In some embodiments, at least one data processing system associated with the LiDAR sensor 202b generates an image (e.g., a point cloud and / or a combined point cloud, etc.) representing objects included in the field of view of the LiDAR sensor 202b. In some examples, at least one data processing system associated with LiDAR sensor 202b generates an image representing the boundaries of a physical object and / or the surface of the physical object (e.g., the topology of the surface), etc. In such examples, the image is used to determine the boundaries of the physical object in the field of view of LiDAR sensor 202b.
[0047] The radio detection and ranging (Radar) sensor 202c includes a sensor configured to communicate with the communication device 202e, the autonomous vehicle computing 202f and / or the safety controller 202g via a bus (e.g., Figure 3 The radar sensor 202c includes at least one device that communicates with a bus (same or similar to the bus 302) that is connected to the radar sensor 202c. The radar sensor 202c includes a system configured to transmit (pulsed or continuous) radio waves. The radio waves transmitted by the radar sensor 202c include radio waves within a predetermined frequency spectrum. In some embodiments, during operation, the radio waves transmitted by the radar sensor 202c encounter physical objects and are reflected back to the radar sensor 202c. In some embodiments, the radio waves transmitted by the radar sensor 202c are not reflected by some objects. In some embodiments, at least one data processing system associated with the radar sensor 202c generates a signal representing an object included in the field of view of the radar sensor 202c. For example, the at least one data processing system associated with the radar sensor 202c generates an image representing the boundaries of the physical object and / or the surface of the physical object (e.g., the topology of the surface). In some examples, the image is used to determine the boundaries of the physical object in the field of view of the radar sensor 202c.
[0048] The microphone 202d includes a microphone configured to communicate with the communication device 202e, the autonomous vehicle computing device 202f, and / or the safety controller 202g via a bus (e.g., Figure 3 At least one device that communicates with the vehicle 200 (e.g., a bus similar to or similar to bus 302). Microphone 202d includes one or more microphones (e.g., an array microphone and / or an external microphone, etc.) that capture audio signals and generate data associated with (e.g., representing) the audio signals. In some examples, microphone 202d includes a transducer device and / or the like. In some embodiments, one or more systems described herein can receive the data generated by microphone 202d and determine the location (e.g., distance, etc.) of an object relative to the vehicle 200 based on the audio signal associated with the data.
[0049] The communication device 202e includes at least one device configured to communicate with the camera 202a, the LiDAR sensor 202b, the Radar sensor 202c, the microphone 202d, the autonomous vehicle computing 202f, the safety controller 202g, and / or the DBW (drive-by-wire) system 202h. For example, the communication device 202e may include at least one device configured to communicate with the camera 202a, the LiDAR sensor 202b, the Radar sensor 202c, the microphone 202d, the autonomous vehicle computing 202f, the safety controller 202g, and / or the DBW (drive-by-wire) system 202h. Figure 3 In some embodiments, the communication device 202e includes a vehicle-to-vehicle (V2V) communication device (eg, a device for enabling wireless communication of data between vehicles).
[0050] Autonomous vehicle computing 202f includes at least one device configured to communicate with camera 202a, LiDAR sensor 202b, Radar sensor 202c, microphone 202d, communication device 202e, safety controller 202g, and / or DBW system 202h. In some examples, autonomous vehicle computing 202f includes devices such as client devices, mobile devices (e.g., cellular phones and / or tablet computers, etc.), and / or servers (e.g., computing devices including one or more central processing units and / or graphics processing units, etc.). In some embodiments, autonomous vehicle computing 202f is the same as or similar to autonomous vehicle computing 400 described herein. Additionally or alternatively, in some embodiments, autonomous vehicle computing 202f is configured to communicate with an autonomous vehicle system (e.g., with Figure 1 Remote AV system 114 of the same or similar autonomous vehicle system), a queue management system (e.g., Figure 1 The same or similar queue management system as the queue management system 116 of FIG), V2I devices (e.g., Figure 1 V2I device 110 that is the same as or similar to the V2I device 110) and / or a V2I system (e.g., Figure 1The V2I system 118 may communicate with the same or similar V2I system.
[0051] Safety controller 202g includes at least one device configured to communicate with camera 202a, LiDAR sensor 202b, Radar sensor 202c, microphone 202d, communication device 202e, autonomous vehicle computing 202f, and / or DBW system 202h. In some examples, safety controller 202g includes one or more controllers (electrical controllers and / or electromechanical controllers, etc.) configured to generate and / or transmit control signals to operate one or more devices of vehicle 200 (e.g., powertrain control system 204, steering control system 206, and / or braking system 208, etc.). In some embodiments, safety controller 202g is configured to generate control signals that take precedence over (e.g., override) control signals generated and / or transmitted by autonomous vehicle computing 202f.
[0052] The DBW system 202h includes at least one device configured to communicate with the communication device 202e and / or the autonomous vehicle computing device 202f. In some examples, the DBW system 202h includes one or more controllers (e.g., electrical controllers and / or electromechanical controllers, etc.) configured to generate and / or transmit control signals to operate one or more devices of the vehicle 200 (e.g., the powertrain control system 204, the steering control system 206, and / or the braking system 208, etc.). Additionally or alternatively, the one or more controllers of the DBW system 202h are configured to generate and / or transmit control signals to operate at least one different device of the vehicle 200 (e.g., turn signals, headlights, door locks, and / or windshield wipers, etc.).
[0053] The powertrain control system 204 includes at least one device configured to communicate with the DBW system 202h. In some examples, the powertrain control system 204 includes at least one controller and / or actuator. In some embodiments, the powertrain control system 204 receives control signals from the DBW system 202h and causes the vehicle 200 to perform longitudinal vehicle motion (such as starting forward movement, stopping forward movement, starting rearward movement, stopping rearward movement, accelerating in a certain direction, decelerating in a certain direction, etc.) or perform lateral vehicle motion (such as performing a left turn and / or performing a right turn, etc.). In examples, the powertrain control system 204 increases, maintains the same, or decreases the energy (e.g., fuel and / or electricity, etc.) provided to the vehicle's motor, thereby causing at least one wheel of the vehicle 200 to rotate or not rotate.
[0054] Steering control system 206 includes at least one device configured to rotate one or more wheels of vehicle 200. In some examples, steering control system 206 includes at least one controller and / or actuator, etc. In some embodiments, steering control system 206 rotates the two front wheels and / or the two rear wheels of vehicle 200 to the left or right to turn vehicle 200 left or right. In other words, steering control system 206 causes the movement required to regulate the y-axis component of the vehicle's motion.
[0055] Braking system 208 includes at least one device configured to actuate one or more brakes to slow down and / or hold vehicle 200 stationary. In some examples, braking system 208 includes at least one controller and / or actuator configured to cause one or more calipers associated with one or more wheels of vehicle 200 to close on the corresponding rotors of vehicle 200. Additionally or alternatively, in some examples, braking system 208 includes an automatic emergency braking (AEB) system and / or a regenerative braking system, among other things.
[0056] In some embodiments, vehicle 200 includes at least one platform sensor (not explicitly illustrated) for measuring or inferring a property of a state or condition of vehicle 200. In some examples, vehicle 200 includes platform sensors such as a global positioning system (GPS) receiver, an inertial measurement unit (IMU), wheel rate sensors, wheel brake pressure sensors, wheel torque sensors, engine torque sensors, and / or steering angle sensors. Although brake system 208 is illustrated as being located Figure 2 The braking system 208 is located on the proximal side of the vehicle 200 , but the braking system 208 can be located anywhere in the vehicle 200 .
[0057] Now refer to Figure 3, a schematic diagram illustrating device 300. As illustrated, device 300 includes a processor 304, a memory 306, a storage component 308, an input interface 310, an output interface 312, a communication interface 314, and a bus 302. In some embodiments, device 300 corresponds to: at least one device of vehicle 102 (e.g., at least one device of a system of vehicle 102); at least one device of camera 202a (e.g., at least one device of a system of camera 202a); at least one device of LiDAR sensor 202b (e.g., at least one device of a system of LiDAR sensor 202b); and / or one or more devices of network 112 (e.g., one or more devices of a system of network 112). In some embodiments, one or more devices of vehicle 102 (e.g., one or more devices of a system of vehicle 102), one or more devices of camera 202a (e.g., one or more devices of a system of camera 202a), one or more devices of LiDAR sensor 202b (e.g., one or more devices of a system of LiDAR sensor 202b), and / or one or more devices of network 112 (e.g., one or more devices of a system of network 112) include at least one device 300 and / or at least one component of device 300. Figure 3 As shown, apparatus 300 includes a bus 302 , a processor 304 , a memory 306 , a storage component 308 , an input interface 310 , an output interface 312 , and a communication interface 314 .
[0058] Bus 302 includes components that enable communication between components of device 300. In some cases, processor 304 includes a processor (e.g., a central processing unit (CPU), a graphics processing unit (GPU), and / or an accelerated processing unit (APU), etc.), a microphone, a digital signal processor (DSP), and / or any processing component that can be programmed to perform at least one function (e.g., a field programmable gate array (FPGA) and / or an application-specific integrated circuit (ASIC), etc.). Memory 306 includes random access memory (RAM), read-only memory (ROM), and / or another type of dynamic and / or static storage device (e.g., flash memory, magnetic memory, and / or optical memory, etc.) that stores data and / or instructions for use by processor 304.
[0059] The storage component 308 stores data and / or software related to the operation and use of the device 300. In some examples, the storage component 308 includes a hard disk (e.g., a magnetic disk, an optical disk, a magneto-optical disk, and / or a solid-state disk, etc.), a compact disk (CD), a digital versatile disk (DVD), a floppy disk, a cassette, a magnetic tape, a CD-ROM, a RAM, a PROM, an EPROM, a FLASH-EPROM, an NV-RAM, and / or another type of computer-readable medium, and a corresponding drive.
[0060] The input interface 310 includes components that permit the device 300 to receive information, such as via user input (e.g., a touch screen display, a keyboard, a keypad, a mouse, buttons, switches, a microphone, and / or a camera, etc.). Additionally or alternatively, in some embodiments, the input interface 310 includes a sensor for sensing information (e.g., a global positioning system (GPS) receiver, an accelerometer, a gyroscope, and / or an actuator, etc.). The output interface 312 includes components for providing output information from the device 300 (e.g., a display, a speaker, and / or one or more light emitting diodes (LEDs), etc.).
[0061] In some embodiments, the communication interface 314 includes a transceiver-like component (e.g., a transceiver and / or a separate receiver and transmitter, etc.) that allows the device 300 to communicate with other devices via a wired connection, a wireless connection, or a combination of a wired connection and a wireless connection. In some examples, the communication interface 314 allows the device 300 to receive information from another device and / or provide information to another device. In some examples, the communication interface 314 includes an Ethernet interface, an optical interface, a coaxial interface, an infrared interface, a radio frequency (RF) interface, a universal serial bus (USB) interface, interface and / or cellular network interface, etc.
[0062] In some embodiments, the device 300 performs one or more processes described herein. The device 300 performs these processes based on the processor 304 executing software instructions stored by a computer-readable medium such as a memory 306 and / or a storage component 308. A computer-readable medium (e.g., a non-transitory computer-readable medium) is defined herein as a non-transitory memory device. A non-transitory memory device includes a storage space located within a single physical storage device or a storage space distributed across multiple physical storage devices.
[0063] In some embodiments, software instructions are read into memory 306 and / or storage component 308 from another computer-readable medium or from another device via communication interface 314. When executed, the software instructions stored in memory 306 and / or storage component 308 cause processor 304 to perform one or more of the processes described herein. Additionally or alternatively, hardwired circuitry may be used in place of or in combination with the software instructions to perform one or more of the processes described herein. Therefore, unless expressly stated otherwise, the embodiments described herein are not limited to any specific combination of hardware circuitry and software.
[0064] Memory 306 and / or storage component 308 include a data store or at least one data structure (e.g., a database, etc.). Device 300 can receive information from, store information in, communicate information to, or search for information stored in the data store or at least one data structure in memory 306 or storage component 308. In some examples, the information includes network data, input data, output data, or any combination thereof.
[0065] In some embodiments, device 300 is configured to execute software instructions stored in memory 306 and / or a memory of another device (e.g., another device that is the same as or similar to device 300). As used herein, the term "module" refers to at least one instruction stored in memory 306 and / or a memory of another device that, when executed by processor 304 and / or a processor of another device (e.g., another device that is the same as or similar to device 300), causes device 300 (e.g., at least one component of device 300) to perform one or more processes described herein. In some embodiments, a module is implemented in software, firmware, and / or hardware.
[0066] supply Figure 3 The number and arrangement of components illustrated are examples. In some embodiments, Figure 3 The apparatus 300 may include additional components, fewer components, different components, or components arranged differently than those illustrated. Additionally or alternatively, a collection of components (e.g., one or more components) of the apparatus 300 may perform one or more functions described as being performed by another component or collection of components of the apparatus 300.
[0067] Now refer to Figure 4, illustrates an example block diagram of an autonomous vehicle computing system 400 (sometimes referred to as an "AV stack"). As illustrated, autonomous vehicle computing system 400 includes a perception system 402 (sometimes referred to as a perception module), a planning system 404 (sometimes referred to as a planning module), a positioning system 406 (sometimes referred to as a positioning module), a control system 408 (sometimes referred to as a control module), and a database 410. In some embodiments, perception system 402, planning system 404, positioning system 406, control system 408, and database 410 are included in and / or implemented within an autonomous navigation system of a vehicle (e.g., autonomous vehicle computing system 202f of vehicle 200). Additionally or alternatively, in some embodiments, perception system 402, planning system 404, positioning system 406, control system 408, and database 410 are included in one or more independent systems (e.g., one or more systems that are the same as or similar to autonomous vehicle computing system 400, etc.). In some examples, perception system 402, planning system 404, positioning system 406, control system 408, and database 410 are included in one or more independent systems located in the vehicle and / or at least one remote system as described herein. In some embodiments, any and / or all of the systems included in autonomous vehicle computing 400 are implemented in software (e.g., software instructions stored in a memory), computer hardware (e.g., via a microprocessor, microcontroller, application specific integrated circuit (ASIC) and / or field programmable gate array (FPGA)), or a combination of computer software and computer hardware. It will also be understood that in some embodiments, autonomous vehicle computing 400 is configured to communicate with a remote system (e.g., an autonomous vehicle system that is the same as or similar to remote AV system 114, a fleet management system 116 that is the same as or similar to fleet management system 116, and / or a V2I system that is the same as or similar to V2I system 118, etc.).
[0068] In some embodiments, perception system 402 receives data associated with at least one physical object in an environment (e.g., data used by perception system 402 to detect at least one physical object) and classifies the at least one physical object. In some examples, perception system 402 receives image data captured by at least one camera (e.g., camera 202a), the image being associated with (e.g., representing) one or more physical objects within the field of view of the at least one camera. In such examples, perception system 402 classifies at least one physical object based on one or more groups of physical objects (e.g., bicycles, vehicles, traffic signs, and / or pedestrians, etc.). In some embodiments, based on perception system 402 classifying the physical object, perception system 402 transmits data associated with the classification of the physical object to planning system 404.
[0069] In some embodiments, planning system 404 receives data associated with a destination and generates data associated with at least one route (e.g., route 106) along which a vehicle (e.g., vehicle 102) can travel toward the destination. In some embodiments, planning system 404 periodically or continuously receives data (e.g., the data associated with the classification of physical objects described above) from perception system 402, and planning system 404 updates at least one trajectory or generates at least one different trajectory based on the data generated by perception system 402. In other words, planning system 404 can perform tasks related to the tactical functions required to operate vehicle 102 in traffic on the road. Tactical efforts involve maneuvering the vehicle in traffic during the journey, including, but not limited to, deciding whether and when to overtake another vehicle, change lanes, or select an appropriate speed, acceleration, deceleration, etc. In some embodiments, planning system 404 receives data associated with the updated position of the vehicle (e.g., vehicle 102) from positioning system 406, and planning system 404 updates at least one trajectory or generates at least one different trajectory based on the data generated by positioning system 406.
[0070] In some embodiments, positioning system 406 receives data associated with (e.g., representing) a location of a vehicle (e.g., vehicle 102) in an area. In some examples, positioning system 406 receives LiDAR data associated with at least one point cloud generated by at least one LiDAR sensor (e.g., LiDAR sensor 202b). In some examples, positioning system 406 receives data associated with at least one point cloud from multiple LiDAR sensors, and positioning system 406 generates a combined point cloud based on the individual point clouds. In these examples, positioning system 406 compares the at least one point cloud or the combined point cloud with a two-dimensional (2D) and / or three-dimensional (3D) map of the area stored in database 410. Then, based on positioning system 406 comparing the at least one point cloud or the combined point cloud with the map, positioning system 406 determines the position of the vehicle in the area. In some embodiments, the map includes a combined point cloud of the area generated prior to navigation of the vehicle. In some embodiments, the map includes, but is not limited to, a high-precision map of roadway geometry, a map describing road network connectivity, a map describing roadway physical properties (such as traffic speed, traffic volume, number of vehicle and bicycle lanes, lane width, lane traffic direction, or type and location of lane markings, or a combination thereof), and a map describing the spatial location of road features (such as crosswalks, traffic signs, or various other types of traffic signals). In some embodiments, the map is generated in real time based on data received by the perception system.
[0071] In another example, positioning system 406 receives global navigation satellite system (GNSS) data generated by a global positioning system (GPS) receiver. In some examples, positioning system 406 receives GNSS data associated with the location of the vehicle in the area, and positioning system 406 determines the latitude and longitude of the vehicle in the area. In such an example, positioning system 406 determines the position of the vehicle in the area based on the latitude and longitude of the vehicle. In some embodiments, positioning system 406 generates data associated with the position of the vehicle. In some examples, based on positioning system 406 determining the position of the vehicle, positioning system 406 generates data associated with the position of the vehicle. In such an example, the data associated with the position of the vehicle include data associated with one or more semantic properties corresponding to the position of the vehicle.
[0072] In some embodiments, the control system 408 receives data associated with at least one trajectory from the planning system 404, and the control system 408 controls the operation of the vehicle. In some examples, the control system 408 receives data associated with at least one trajectory from the planning system 404, and the control system 408 controls the operation of the vehicle by generating and transmitting control signals to operate the powertrain control system (e.g., the DBW system 202h and / or the powertrain control system 204), the steering control system (e.g., the steering control system 206), and / or the braking system (e.g., the braking system 208). For example, the control system 408 is configured to perform operational functions such as lateral vehicle motion control or longitudinal vehicle motion control. Lateral vehicle motion control causes the necessary actions to regulate the y-axis component of the vehicle's motion. Longitudinal vehicle motion control causes the necessary actions to regulate the x-axis component of the vehicle's motion. In an example, if the trajectory includes a left turn, the control system 408 transmits a control signal to cause the steering control system 206 to adjust the steering angle of the vehicle 200, thereby causing the vehicle 200 to turn left. Additionally or alternatively, the control system 408 generates and transmits control signals to cause other devices of the vehicle 200 (eg, headlights, turn signals, door locks, and / or windshield wipers, etc.) to change states.
[0073] In some embodiments, perception system 402, planning system 404, positioning system 406, and / or control system 408 implement at least one machine learning model (e.g., at least one multilayer perceptron (MLP), at least one convolutional neural network (CNN), at least one recurrent neural network (RNN), at least one autoencoder, and / or at least one transformer, etc.). In some examples, perception system 402, planning system 404, positioning system 406, and / or control system 408, alone or in combination with one or more of the above systems, implement at least one machine learning model. In some examples, perception system 402, planning system 404, positioning system 406, and / or control system 408 implement at least one machine learning model as part of a pipeline (e.g., a pipeline for identifying one or more objects located in an environment, etc.).
[0074] Database 410 stores data transmitted to, received from, and / or updated by perception system 402, planning system 404, positioning system 406, and / or control system 408. In some examples, database 410 includes a storage component for storing data and / or software related to operations and using at least one system of autonomous vehicle computing 400 (e.g., Figure 3In some embodiments, database 410 stores data associated with a 2D and / or 3D map of at least one area. In some examples, database 410 stores data associated with a 2D and / or 3D map of a portion of a city, portions of multiple cities, multiple cities, a county, a state, and / or a country (e.g., a country), etc. In such an example, a vehicle (e.g., a vehicle that is the same as or similar to vehicle 102 and / or vehicle 200) can drive along one or more drivable areas (e.g., a single-lane road, a multi-lane road, a highway, a back road, and / or an off-road road, etc.) and cause at least one LiDAR sensor (e.g., a LiDAR sensor that is the same as or similar to LiDAR sensor 202b) to generate data associated with an image representing objects included in the field of view of the at least one LiDAR sensor.
[0075] In some embodiments, database 410 can be implemented across multiple devices. In some examples, database 410 includes a vehicle (e.g., a vehicle that is the same as or similar to vehicle 102 and / or vehicle 200), an autonomous vehicle system (e.g., an autonomous vehicle system that is the same as or similar to remote AV system 114), a fleet management system (e.g., a vehicle ... Figure 1 The same or similar queue management system as the queue management system 116 of FIG) and / or the V2I system (e.g., Figure 1 The V2I system 118 is the same or similar V2I system) and the like.
[0076] Now refer to Figure 5 , a diagram illustrating an implementation 500 of camera-assisted LiDAR data validation at a vehicle 502. In some embodiments, the implementation 500 includes a camera 504, a LiDAR sensor 506, and an AV computation 510. In some embodiments, the camera 504, the LiDAR sensor 506, and the AV computation 510 are each associated with Figure 2 The camera 202a, LiDAR sensor 202b, and AV computing 202f of the illustrated system 202 are the same or similar.
[0077] exist Figure 5In the example of , camera 504 generates camera data that forms an image of the environment. In some embodiments, the image is a two-dimensional (2D) representation of the environment. The image includes multiple pixels that specify the color and intensity at each pixel of the image. In the example, the color is represented by the component intensities such as red, green and blue or cyan, yellow and white. The LiDAR sensor 506 includes a transmitter and a receiver. The transmitter transmits light into the environment, which is reflected by objects in the environment. The receiver captures the light reflected by the objects in the environment, and the captured reflected light (e.g., LiDAR data) is used to generate a point cloud. In the example, the point cloud is a discrete set of three-dimensional (3D) data points in the environment. The camera 504 and LiDAR sensor 506 capture data within corresponding fields of view (FOV). The corresponding FOVs overlap, so that 2D image data and 3D point cloud data are captured for the same location in the environment.
[0078] Images created from camera data captured by camera 504 and point clouds created from LiDAR data captured by LiDAR sensor 506 are used to detect and classify features of the environment. Therefore, the outputs of camera 504 and LiDAR sensor 506 are provided to AV computation 510 for further processing, such as object detection and classification. In examples, camera data provides accurate measurements of edges, color, and lighting, which ultimately leads to accurate object classification in the resulting image. Compared to image data from camera 504, LiDAR data typically contains less semantic information and, instead, enables highly accurate 3D localization. In some examples, LiDAR data is sparse due to low reflectivity from objects in the environment. Low reflectivity corresponds to low confidence in the LiDAR data, which negatively impacts AV functionality based on LiDAR data, such as perception and localization. Confidence in LiDAR data is affected by several factors, such as lighting, object color, and object material type. The present technology uses camera data corresponding to the LiDAR data to determine a confidence level associated with the LiDAR data. In an example, camera data is used to analyze lighting, object color, and object material type, and confidence scores in corresponding LiDAR data are updated based on the analysis. For example, when camera data such as lighting, object color, and object material type indicates low reflectivity, the confidence score associated with the LiDAR data is updated based on the confidence score associated with the camera data.
[0079] Figure 6A The effect of ambient lighting on the vehicle's LiDAR sensor and camera is shown. Vehicle 602A includes camera 604A and LiDAR sensor 606A. In this example, camera 604A and Figure 5The camera 504 is the same as or similar to the camera 504, and the LiDAR sensor 606A is the same as Figure 5 The LiDAR sensor 506 is the same as or similar to the LiDAR sensor 506. Camera 604A and LiDAR sensor 606A capture data from the surrounding environment, which enables detection of object 610A. As shown, LiDAR sensor 606A emits infrared light 612A, which is reflected by the object, and reflected light 616A is captured by the receiver of LiDAR 606A. Similarly, camera 604A captures illumination 618A reflected by the object. Ambient illumination 622A in the environment is reflected from object 610A. Ambient light 622A includes, for example, sunlight, moonlight, and other light sources such as traffic lights, light from other buildings, and light from vehicles. Ambient light 622A affects the function of the LiDAR because the LiDAR receiver captures ambient light 622A along with reflected light 616A. The camera includes an infrared (IR) filter 620A. IR filter 620A can filter out ambient illumination 622A from other illumination 618A reflected by the object. As a result, the ambient illumination does not corrupt the camera data. However, ambient lighting may cause noise or other artifacts in LiDAR data.
[0080] Figure 6B The effect of objects of different colors on the vehicle's LiDAR sensor and camera is shown. In the example, camera 604B is Figure 5 The camera 504 is the same as or similar to the camera 504, and the LiDAR sensor 606B is the same as Figure 5 The same or similar LiDAR sensor 506 is used. Figure 6B In the example, object 610B is a light-colored object, such as white. Object 611B is a dark-colored object, such as black. Different colors are associated with different light reflectivity values. For example, on a scale from 0% to 100%, a reflectivity value of 0% corresponds to pure black, where almost no light is reflected by the pure black surface. A reflectivity value of 100% corresponds to pure white, where most or all light is reflected by the pure white surface. In the example, reflectivity values of less than 50% correspond to darker colors that absorb more light than they reflect. Reflectivity values of greater than 50% correspond to lighter colors that reflect more light than they absorb. In Figure 6B In the example shown in FIG, LiDAR 606B emits infrared light 612B and 613B, which are reflected by light-colored object 610B and dark-colored object 611B, respectively. Reflected light 616B and reflected light 617B are captured by the receiver of LiDAR 606B. Similarly, camera 604A captures illumination 618B reflected by light-colored object 610B and illumination 619B reflected by dark-colored object 611B.
[0081] Dark object 611B reflects a smaller portion of light when compared to light object 610B. Therefore, reflected light 617B is sparser than reflected light 616B. Dark object 611B has low reflectivity, and there is low confidence in the LiDAR data corresponding to object 611B. Additionally, the low reflectivity results in a lower detection rate associated with dark object 611B based on the LiDAR data. In some examples, when an actual, true dark object is present, the data associated with dark object 611B is incorrectly classified as noise or artifact. However, camera 604B is able to capture both light object 610B and dark object 611B. In the example, camera 604B captures dark object 611B with a higher confidence level when compared to the confidence level associated with the corresponding LiDAR data captured by LiDAR sensor 606B.
[0082] Figure 6C The effect of objects with different properties on the vehicle's LiDAR sensor and camera is shown. In the example, camera 604C is Figure 5 The camera 504 is the same as or similar to the camera 504, and the LiDAR sensor 606C is the same as Figure 5 The same or similar LiDAR sensor 506 is used. Figure 6C In the example shown in FIG6 , object 630 includes a metal component 632 and a tire component 634 . In the example, object 630 is a bicycle. Different materials are associated with different light reflectivity values. For example, metals (e.g., stainless steel, aluminum, zinc, brass, galvanized steel, etc.), plastics (e.g., polyurethane, polypropylene, polyvinyl chloride, acrylonitrile butadiene styrene (ABS), polyamide (PA), polystyrene (PS), polyethylene (PE), polyoxymethylene (POM), polycarbonate (PC), acrylic (PMMA), etc.), wood, glass, and rubber are material types with different reflectivity values. Therefore, the material type is a property of the object that affects the object's reflectivity and subsequent detection by the LiDAR sensor.
[0083] exist Figure 6C In the example shown, LiDAR sensor 606C emits infrared light 612C and 613C, which are reflected by metal component 632 and tire assembly 634, respectively. Reflected light 616C and reflected light 617C are captured by a receiver of LiDAR sensor 606C. Similarly, camera 604A captures illumination 618C reflected by metal component 632 and illumination 619C reflected by tire assembly 634. In the example shown, LiDAR sensor 606C receives more reflected light 616C from metal component 632 than reflected light 617C from tire assembly 634. Camera 604C captures the entire object 630, and the resulting image data is used to detect both metal component 632 and tire assembly 634 of object 630.
[0084] exist Figures 6A to 6C In an example, corresponding camera data and LiDAR data are used to update or generate a confidence score associated with the LiDAR data. In the example, lighting, object color, and object material type are analyzed from the camera data, and a confidence score in the corresponding LiDAR data is updated or generated based on the analysis. In this way, the camera data is used to verify the LiDAR data.
[0085] Figure 7 A workflow 700 for camera-assisted LiDAR data validation is shown. In some embodiments, one or more of the steps described with respect to workflow 700 is performed by Figure 1 The autonomous vehicle (AV) system 114, the fleet management system 116, Figure 2 200 vehicles, Figure 3 device 300, Figure 4 Autonomous delivery vehicle calculations 400 or Figure 5 The AV calculation 510 is performed (e.g., completely and / or partially, etc.). Figure 7 In the example of FIG. 7 , one or more cameras 704 and one or more LiDAR sensors 706 are shown. In some embodiments, the camera 704 and Figure 2 camera 202a, Figure 5 Camera 504 or Figures 6A to 6C In some embodiments, the LiDAR sensor 706 is the same as or similar to the cameras 604A-604C. Figure 2 LiDAR sensor 202b, Figure 5 LiDAR sensor 506 or Figures 6A to 6C The LiDAR606A-606C are the same or similar.
[0086] The workflow 700 includes two object detection networks, an image semantic segmentation network (ISN) 714 and a LiDAR semantic segmentation network (LSN) 716. Generally, an object detection neural network is configured to receive sensor data and process the sensor data to detect at least one object in the environment (e.g., Figure 1 Object 104, Figure 6A Object 610A, Figure 6B objects 610B and 611B, Figure 6C630). In an embodiment, the object detection neural network is a feed-forward convolutional neural network that, given sensor data (e.g., image data, LiDAR data, and / or Radar data, etc.), generates a set of bounding boxes for potential objects in a 3D space (e.g., an environment) and a confidence score for the presence of an object class instance (e.g., a car, a pedestrian, or a bicycle) within the bounding box. The higher the classification score, the more likely the corresponding object class instance is to be present in the box.
[0087] ISN 714 takes camera data from camera 704 as input and outputs a set of predicted 2D or 3D bounding boxes for objects in the environment (e.g., object space information) and a corresponding confidence score for the presence of an object class instance within the bounding box. ISN 714 also outputs color information and attribute information associated with the detected objects. For example, ISN 714 takes camera data as input, predicts the class of each pixel in the camera data, and outputs semantic segmentation data (e.g., a confidence score) for each pixel in the image. For example, each pixel is associated with a 2D spatial coordinate (e.g., x, y coordinate). ISN 714 is trained using an image dataset that includes images augmented with bounding boxes for classes and segmentation labels from the image dataset. In this example, the confidence score is a probability value indicating the probability that the pixel's class was correctly predicted. Similarly, LSN 716 takes LiDAR data as input and outputs a set of predicted 3D bounding boxes for potential objects in 3D space and a confidence score for the presence of an object class instance within the bounding box. In an example, the LSN receives a plurality of data points representing a 3D space. For example, each data point in the plurality of data points is a set of 3D space coordinates (e.g., x, y, z coordinates). The predicted 3D set of bounding boxes (e.g., object space information) also includes a confidence score for the presence of an object class instance within the bounding box.
[0088] exist Figure 7 In the example shown, output 724 of ISN 714 includes an object spatial location, object color information, object attribute information, and a detection confidence score. The object spatial location provides location information and dimensions associated with the object. The location information indicates the specific location or position of the object. The object color information refers to the specific color of the object. In the example, the object color information is specified based on color model values such as the RGB color model, the RYB color model, the CMY color model, the CMYK color model, or the cylindrical coordinate color model. The object attribute information indicates the characteristics of the object, such as the type of metal, plastic, rubber, wood, or other material. Output 726 of LSN 716 includes the object spatial location and the detection confidence score.
[0089] ISN 714 and LSN 716 each output a corresponding detection confidence score based on the data captured by the corresponding sensor. As described above, the detection confidence score can vary based on the data captured by the sensor. In some embodiments, the ISN detection confidence score is used to validate the LSN detection confidence score. For example, when analysis of the camera data determines that an object is associated with a low reflectivity value or that the ambient lighting has corrupted the reflectivity value captured by the LiDAR, the LSN detection confidence score is updated based on the ISN-based detection confidence score.
[0090] Figure 8 The application of camera-assisted LiDAR data verification is shown. In some embodiments, ISN 814 is connected to Figure 7 ISN 714 is the same as or similar to the ISN 714; and LSN 816 is the same as the Figure 7 The ISN 814 receives camera data as input and outputs seven channels of output (collectively, output 818) including spatial coordinates x 818A, spatial coordinates y 818B, ambient light 818C, color information 818D, attribute information 818E, detection angle 818F, and confidence score 818G. The LSN 816 receives LiDAR data as input and outputs five channels of output (collectively, output 820) including spatial coordinates x 820A, spatial coordinates y 820B, spatial coordinates z 820C, intensity 820D, and confidence score 820E.
[0091] The output 820 of the LSN 816 is projected onto a 2D plane via a 3D to 2D projection 822. The output 820 of the LSN is projected onto corresponding pixels arranged in the 2D plane. Figure 8 As shown, LSN 816 outputs 3D spatial information (x, y, and z), while ISN outputs 2D spatial information (x, y). Projection 822 transforms the LSN spatial information into a 2D format. Rendering manager 824 associates data from ISN 814 and projection data from LSN 816 with pixels on a 2D plane to generate rendered pixels. The rendered pixels have additional context that enables the generation of labels for objects with improved accuracy.
[0092] like Figure 8As shown, the rendering manager 824 outputs rendered pixels. A concatenator 826 takes the rendered pixels and the output 820 of the LSN 816 and concatenates corresponding points (e.g., points that include information associated with the same or approximately the same real-world location). In an example, the concatenator combines the rendered pixels with the LiDAR output. The rendered pixels output by the rendering manager include color and attribute information that is missing from the LiDAR output 820. The spatial location information associated with the rendered pixels is 2D spatial information (e.g., spatial coordinate x 818A and spatial coordinate y 818B). However, the LiDAR output 820 includes 3D spatial information (e.g., spatial coordinate x 820A, spatial coordinate y 820B, spatial coordinate z 820C). The rendered pixels and output 820 are concatenated and input to a neural network, where the input to the neural network includes both 2D and 3D spatial information. In an example, the concatenated information is input to a neural network, such as an object detection network. In some embodiments, the object detection network is an enhanced bird's-eye view network (BEVN) 828. The BEVN 828 obtains the 13-channel output from the cascade 826 and outputs an object detection result 830. In an example, the object detection result 830 includes spatial information (x, y, z, and size) associated with the object and a classification result.
[0093] In some embodiments, the BEVN 828 is trained to determine how to learn from the difference in confidence between the ISN 814 and the LSN 816. For example, the confidence 818G associated with the camera data and the confidence 820E associated with the LiDAR data are input to the BEVN 826, which learns from the confidence to output a detection result 830.
[0094] Figure 9 is a flow chart of a process that can implement camera-assisted LiDAR data verification. In some embodiments, one or more of the steps described with respect to process 900 are performed by Figure 1 The autonomous vehicle (AV) system 114, the fleet management system 116, Figure 2 200 vehicles, Figure 3 device 300, Figure 4 Autonomous delivery vehicle calculations 400 or Figure 5 In some embodiments, the ISN 914 is associated with the AV calculation 510 (e.g., completely and / or partially, etc.). Figure 8 ISN814 or Figure 7 ISN 714 is the same as or similar to ISN 714; and LSN 916 is the same as Figure 8 LSN 816 or Figure 7 The same or similar to LSN 716.
[0095] exist Figure 9 In the example of FIG, a series of decisions 902, 904 and 906 are used to analyze the camera data and make a decision based on the ISN confidence score θ ISN To determine the corresponding LSN confidence score θ associated with the LiDAR data LSN Is it updated (θ' LSN At decision block 902, determine the LSN location information L LSN 918B and ISN location information L ISN 920B. In this example, the LSN location information L LSN 918B is about Figure 8 The 3D spatial information described (eg, spatial coordinate x 820A, spatial coordinate y 820B, spatial coordinate z 820C). In addition, in the example, the ISN location information L ISN 920B is about Figure 8 The 2D spatial information described (eg, spatial coordinate x 818A and spatial coordinate y 818B). LSN location information L LSN 918B is projected onto a 2D plane (e.g., Figure 8 At decision block 902, the projected LSN location information L is compared. LSN 918B and ISN location information L ISN The difference between 920B.
[0096] If LSN location information L LSN 918B and ISN location information L ISN The difference between 920B satisfies the distance threshold Th location , then evaluate the second decision box 904. If the LSN location information L LSN 918B and ISN location information L ISN The difference between 920B does not meet the distance threshold Th location , then at block 908 the updated LSN confidence score θ' LSN Set to zero. For example, when the LSN location information L LSN 918B and ISN location information L ISN The difference between 920B is less than the distance threshold Th location When the difference meets the distance threshold Th location . Evaluate LSN location information L LSN 918B and ISN location information L ISN 920B to ensure that the camera data and LiDAR data are based on the same object or the same place in the environment. LSN 918B and ISN location information L ISNThe distance between 920B exceeds the distance threshold Th location , then the camera data and LiDAR data can represent different objects. Therefore, when the LSN location information L LSN 918B and ISN location information L ISN The distance between 920B exceeds the distance threshold Th location When θ is not based on the ISN confidence score ISN Update LSN confidence score θ LSN In an example, the threshold is the number of pixels in 2D camera coordinates, such as 5 pixels, 10 pixels, etc. For example, the threshold is determined during camera-LiDAR calibration.
[0097] At decision block 904, the color information C output by ISN 914 is evaluated. ISN 920C. If color information C ISN 920C and reference color C ref The difference satisfies the color threshold TH colordif , then in block 912, the confidence score θ is calculated based on the ISN confidence score θ. ISN Update LSN confidence score θ' LSN In this example, the LSN confidence score θ' LSN The function f(θ) is updated to be the ISN confidence score LSN ). If the color information C ISN 920C and reference color C ref The difference does not meet the color threshold TH colordif , then at block 910 as determined by LSN 916, the updated LSN confidence score θ′ LSN is set to the LSN confidence score θ LSN 918A. In the example, when the color information C ISN 920C and reference color C ref The difference between them is less than the color threshold TH colordif When the difference meets the color threshold TH colordif In an example, the reference color is a predetermined color detected by the LiDAR that is known to have a low reflectivity value. In some examples, the reference color is a dark color with a reflectivity value less than 50%. If the color information C determined by the ISN ISN 920C is close to the reference color (within the threshold of the reference color), it is a similar dark color, which results in sparse LiDAR detections and low confidence in such detections. Therefore, at box 912, for the color C relative to the reference color C ref At the color threshold TH colordif Update the LSN confidence score θ' LSN .
[0098] At decision block 906, the attribute information A output by the LSN is evaluated. ISN 920D. If attribute information A ISN 920D and reference attribute A ref The difference satisfies the attribute threshold TH attribute , then at box 912, according to the ISN confidence score θ ISN Update LSN confidence score θ' LSN If attribute information A ISN 920D and reference attribute A ref The difference does not meet the attribute threshold TH attribute , then at block 910 as determined by the LSN 916, the LSN confidence score θ′ LSN is set to the LSN confidence score θ LSN 918A. In the example, when attribute information A ISN 920D and reference attribute A ref The difference between them is less than the attribute threshold TH attribute When the difference satisfies the attribute threshold TH attribute In addition, in the example, the reference attribute A ref is one or more material types known to have low reflectivity values. In some examples, the reference attributes are plastic, rubber, and wood. In examples, attributes such as plastic have weak reflectivity, while attributes such as metal have strong reflectivity. If the attribute information A determined by the ISN ISN 920D is close to the reference attribute (within a threshold of the reference color), it is a similar attribute type that results in sparse LiDAR detections and low confidence in such detections. Therefore, at block 912, for each attribute A relative to the reference attribute A, ref 920D at attribute threshold TH attribute Update the LSN confidence score θ' LSN .
[0099] exist Figure 9 In the example, location information from both the LSN and ISN data is used to ensure that the same object is being compared between the LSN and ISN data. Once the same object is being compared, features of the ISN data are evaluated to determine if the corresponding LiDAR data is noisy or has low confidence. If the features indicate the presence of known negative influences on LiDAR detection, the LiDAR confidence score is updated based on the ISN confidence score.
[0100] In the example, black objects or substantially black objects are processed differently due to their low reflectivity. In the example, when LiDAR is used to detect black objects (such as black vehicles, pedestrians wearing black clothes, bicycles, etc.), holes exist in the point cloud data. For example, for a black vehicle, there are holes in the point cloud data of the black body, however, in some instances, the reflected windows are accurately represented in the point cloud. In the example, a pedestrian wearing black clothes produces a hole in the point cloud data where the dark clothing is located, however, in some instances, the pedestrian's visible skin is accurately represented in the point cloud. In the example, for a bicycle, there are holes in the point cloud data reflected by the bicycle tire, but the non-black bicycle frame is accurately represented in the point cloud data. In some embodiments, image-based neural network detection results and camera / vehicle calibration data are used to assist LiDAR detection. The LiDAR semantic network receives object color information, object angle information related to the AV, and vehicle model information (if applicable) as input and outputs a more robust spatial location and confidence score. Therefore, in some embodiments, data output by the ISN is used to evaluate LiDAR detections of objects with low or weak reflectivity values.
[0101] Figure 10 A bird's-eye view of a LiDAR sensor emitting light reflected at different detection angles is shown. Three examples 1020, 1030, and 1040 illustrate LiDAR sensors 1006A, 1006B, and 1006C (collectively, LiDAR 1006) of vehicles 1002A, 1002B, and 1002C (collectively, vehicles 1002) emitting light 1012A, 1012B, and 1012C (collectively, light 1012) toward respective objects 1010A, 1010B, and 1010C (collectively, objects 1010). Detection angles 1014A, 1014B, and 1014C (collectively, detection angles 1014) illustrate different reflected light 1016A, 1016B, and 1016C (collectively, reflected light 1016).
[0102] exist Figure 10 In the example shown in FIG, a LiDAR 1006 emits light 1012 from a light emitter (e.g., a laser emitter). The light emitted by a LiDAR system is typically not in the visible spectrum; for example, infrared light is often used. A portion of the emitted light 1012 encounters a physical object 1010 (e.g., a cyclist) and reflects back to the LiDAR 1006. The LiDAR 1006 also includes one or more receivers that capture the reflected light 1016.
[0103] In the example, the LiDAR sensor 1006 detects the boundaries of the object 1010 based on the characteristics of the reflected light 1016 captured by the LiDAR system 1002. As shown in example 1020, when the object 1010A is positioned substantially parallel to the light 1012A emitted from the LiDAR sensor 1006A with its narrowest field of view, the LiDAR sensor 1006A obtains a small amount of reflected light 1016A. In other words, because the object 1010A is oriented perpendicular to the LiDAR's line of sight (corresponding to light 1012A) with its narrowest field of view, the object's smallest surface is oriented to reflect light 1016A to the LiDAR sensor 1006A, thereby capturing the least amount of data at the LiDAR sensor 1006A.
[0104] In example 1030, an object is positioned such that the narrowest portion of the object in the XY plane is rotated so that the narrowest portion is not substantially perpendicular to light 1012B emitted from LiDAR sensor 1006B, where light 1012B corresponds to the LiDAR's line of sight. LiDAR sensor 1006B captures a greater amount of reflected light 1016B than reflected light 1016A captured by LiDAR sensor 1006A. Because object 1010A is oriented with its narrowest portion angled away from the LiDAR's line of sight, additional surface area of the object is oriented to reflect light 1016B to LiDAR sensor 1006B. This results in an additional amount of data being captured at LiDAR sensor 1006B compared to the data captured at LiDAR sensor 1006A.
[0105] In example 1040, the LiDAR sensor 1006C obtains a maximum amount of reflected light 1016C when the object is positioned substantially perpendicular to the light 1012C emitted from the LiDAR sensor 1006C with the widest field of view of the object. Since the object 1010C is oriented in the line of sight of the LiDAR sensor 1006C with its widest field of view, the largest surface area of the object is oriented to reflect light 1016C to the LiDAR sensor 1006C, thereby capturing a maximum amount of data at the LiDAR sensor 1006C when compared to other orientations of the object.
[0106] In an example, the strongest reflection from an object is observed when more of the object's surface area is positioned to intersect the LiDAR's line of sight. In an example, as the object's detection angle changes, the amount of reflection from the object also changes. In an example, the detection angle can be calculated based on the object's spatial location, vehicle / map information, and camera calibration information from the ISN. For example, known data points, object position, and calibration information are used to calculate the detection angle. In an example, the detection angle is determined using triangulation, multi-point positioning, or localization. When the object is a dark color (e.g., a color with a reflectivity value less than 50%), the detection angle information is used to determine a confidence score associated with the object. In an example, if the object is a dark color and is positioned at an angle such that a small surface area of the object is in the LiDAR's line of sight, the LSN confidence score is updated when the object is positioned to result in low reflected light and the object's color reflects light poorly. In an example, the LSN is trained to output the object's spatial location based on LiDAR data, ISN output data, and angle information as input.
[0107] Figure 11 A workflow 1100 for camera-assisted LiDAR data validation is shown. In some embodiments, one or more of the steps described with respect to workflow 1100 is performed by Figure 1 The autonomous vehicle (AV) system 114, the fleet management system 116, Figure 2 200 vehicles, Figure 3 device 300, Figure 4 Autonomous delivery vehicle calculations 400 or Figure 5 In some embodiments, the ISN 1114 is associated with the AV calculation 510 (e.g., completely and / or partially, etc.). Figure 9 ISN 914, Figure 8 ISN 814 or Figure 7 ISN 714 is the same as or similar to ISN 714; and LSN 1116 is the same as Figure 9 LSN 916, Figure 8 LSN 816, Figure 7 The same or similar to LSN 716.
[0108] The ISN 1114 takes as input camera data 1120 from the camera 1104. The output of the ISN 1114 includes an object space location 1122, object attribute information 1124, object color information 1126, and a detection confidence score 1128. The object space location 1122, vehicle / map information 1130, and camera calibration information 1132 are used to derive angle information 1134. In an example, the detection angle information 1134 is calculated by the angle of a triangle formed by the camera location (e.g., extracted from the vehicle / map information 1130 and the camera calibration information 1132) and the object space location 1122.
[0109] In some embodiments, object attribute information 1124 includes a vehicle model. In one example, object attribute information 1124 includes the object's material type. In examples where the object is a vehicle, the vehicle model is used. In one example, the vehicle model is a database storing 3D vehicle structure information for popular vehicles. In one example, the vehicle model is used for the vehicle. That is, when the ISN detects a vehicle in camera data, the vehicle model information in the database is used to match the vehicle in the camera image. When the object is a pedestrian, object attribute information is not used in the LSN.
[0110] exist Figure 11 In the example of FIG. 1 , LSN 1116 receives as input LiDAR data 1136 from LiDAR sensor 1106. LSN 1116 also receives as input object attribute information 1124, object color information 1126, detection confidence score 1128, and angle information 1134. LSN 1116 outputs object spatial location information 1140 and detection confidence score 1142. In an embodiment, LiDAR is trained to determine a confidence score based on the LiDAR data, the input object attribute information 1124, the input object color information 1126, the detection confidence score 1128, and the angle information 1134.
[0111] exist Figure 11In the example of , information generated from the camera is input to LSN 1116, and LSN 1116 uses the angle information to determine a confidence score associated with a detected object. In the example, when the angle causes a narrow portion of the object to reflect light emitted by the LiDAR, LSN 1116 outputs a reduced confidence in the LiDAR data including the object. When the angle causes a wider portion of the object to reflect light emitted by the LiDAR, a stronger reflection is observed, and LSN 1116 outputs a higher confidence score associated with the LiDAR data including the object. In some embodiments, the present technology uses machine learning to enable LSN 1116 to use the camera data and the angle information to determine its own improved confidence score. The ISN 1114 output plus the derived angle information are input to LSN 1116, resulting in a better, more robust confidence score for LSN 1116.
[0112] Now refer to Figure 12 , which illustrates a flow chart of a process 1200 for camera-assisted LiDAR data validation. In some embodiments, one or more of the steps described with respect to process 1200 is performed by Figure 2 1200 (e.g., completely and / or partially, etc.). Additionally or alternatively, in some embodiments, one or more steps described with respect to process 1200 are performed by another device or group of devices (such as a processor) that is separate from or includes autonomous system 202. Figure 11 ISN 1114 or LSN 1116, Figure 9 ISN 914 or LSN 916, Figure 8 ISN 814 or LSN 816, or Figure 7 ISN714 or LSN 716, etc.) (e.g., completely and / or partially, etc.).
[0113] At block 1202, a first semantic segmentation network and a second semantic segmentation network generate corresponding spatial locations and corresponding confidence scores for detected objects. In an example, the first semantic segmentation network is an LSN and the second semantic segmentation network is an ISN. In an example, the first semantic segmentation network receives first sensor data (e.g., LiDAR data) as input, and the second semantic segmentation network receives second sensor data (e.g., camera data) as input. The ISN outputs object spatial locations, object color information, object attribute information, and a detection confidence score. The LSN outputs object spatial information and a detection confidence score.
[0114] At block 1204, a difference is determined between a first spatial location (e.g., a 3D location from LiDAR) output by the first semantic segmentation network and a second spatial location (e.g., a 2D location from a camera) output by the second semantic segmentation network. In an example, the first spatial location and the second spatial location are compared to a predetermined threshold.
[0115] At box 1206, the data generated by the second semantic network is evaluated to determine whether to update the first confidence score output by the first semantic network. In some embodiments, the data generated by the second semantic network is evaluated when the difference between the spatial locations meets a distance threshold. When the data output by the second semantic segmentation network indicates a known low reflectivity at the second location, the first confidence score is modified by updating the first confidence score based on the second confidence score. In some embodiments, the updating includes setting the first confidence score equal to the second confidence score. The data output by the second semantic segmentation network includes color data or attribute data. At box 1208, the color data or the attribute data is selected for evaluation. In some embodiments, both the color data and the attribute data are selected for evaluation.
[0116] At block 1210, when the color data at the second spatial location is associated with low reflectivity, the first confidence score is updated based on the second confidence score. In an example, the color data is associated with low reflectivity when the color data is within a predetermined threshold of a reference color. At block 1212, when the attribute type at the second spatial location is associated with low reflectivity, the first confidence score is updated based on the second confidence score. In an example, the attribute type is associated with low reflectivity when the attribute type is within a predetermined threshold of a reference attribute.
[0117] In some embodiments, the object type is determined based on the data output by the first semantic segmentation network and the second semantic segmentation network. The corresponding sensor data and confidence score are obtained by the AV stack, and the AV stack performs various functions based on the sensor data and confidence score. For example, the perception system 402, the planning system 404, the positioning system 406, the control system 408, or the database 410 (e.g., Figure 4 ) obtains corresponding sensor data and confidence scores, and navigates through an environment (e.g., environment 100) based on the corresponding sensor data and confidence scores.
[0118] Now refer to Figure 13 , illustrating a flow chart of a process 1300 for LiDAR semantic network confidence scoring based on angle information. In some embodiments, one or more of the steps described with respect to process 1300 is performed by Figure 21300 (e.g., completely and / or partially, etc.). Additionally or alternatively, in some embodiments, one or more steps described with respect to process 1300 are performed by another device or group of devices (such as a processor) that is separate from or includes autonomous system 202. Figure 11 ISN 1114 or LSN 1116, Figure 9 ISN 914 or LSN 916, Figure 8 ISN 814 or LSN 816, or Figure 7 's ISN 714 or LSN 716, etc.) (e.g., completely and / or partially, etc.).
[0119] At block 1302, output data from a trained image semantic network (ISN) is obtained. In an example, the ISN data includes object spatial locations, object attribute information, object color information, and a detection confidence score 1128. At block 1304, angle information is determined based on the object spatial information, vehicle information, map information, and calibration information from the ISN.
[0120] At block 1306, the LiDAR data, ISN object attribute information, ISN object color information, ISN detection confidence score, and angle information are input to a trained LiDAR semantic network. The trained LiDAR semantic network is trained based on the input data to output LSN object spatial information and LSN confidence score.
[0121] In an example, the present technology enables accurate confidence scores associated with LiDAR data. Based on the accurate confidence scores, updating or generating LiDAR data enables the generation of accurate images representing the boundaries of physical objects and / or the surfaces of physical objects (e.g., the topology of the surface). The updated LiDAR information enables better positioning, wherein a positioning system (e.g., Figure 4 The positioning system 406) determines the AV's position in the area based on comparing the updated LiDAR information with the map.
[0122] According to some non-limiting embodiments or examples, a system is provided, comprising: at least one processor; and at least one non-transitory storage medium storing instructions. The instructions, when executed by the at least one processor, cause the at least one processor to execute a first semantic segmentation network and a second semantic segmentation network, wherein, when executed, the first semantic segmentation network receives first sensor data as input, and the second semantic segmentation network receives second sensor data as input. The instructions, when executed by the at least one processor, cause the at least one processor to determine a difference between a first location output by the first semantic segmentation network and a second location output by the second semantic segmentation network. The instructions, when executed by the at least one processor, cause the at least one processor to update a first confidence score output by the first semantic segmentation network if the difference between the first location and the second location satisfies a first predetermined threshold and the data output by the second semantic segmentation network indicates known low reflectivity at the second location, wherein the update modifies the first confidence score based on the second confidence score. The instructions, when executed by the at least one processor, cause the at least one processor to determine an object type based on the data output by the first semantic segmentation network and the second semantic segmentation network.
[0123] According to some non-limiting embodiments or examples, a method is provided. The method includes executing, using at least one processor, a first semantic segmentation network and a second semantic segmentation network, wherein, when executed, the first semantic segmentation network obtains first sensor data as input and the second semantic segmentation network obtains second sensor data as input. The method includes determining, using the at least one processor, a difference between a first location output by the first semantic segmentation network and a second location output by the second semantic segmentation network. The method includes updating, using the at least one processor, a first confidence score output by the first semantic segmentation network when the difference between the first location and the second location satisfies a first predetermined threshold and the data output by the second semantic segmentation network indicates a known low reflectivity at the second location, wherein the update modifies the first confidence score based on the second confidence score. The method includes determining, using the at least one processor, an object type based on the data output by the first semantic segmentation network and the second semantic segmentation network.
[0124] According to some non-limiting embodiments or examples, at least one non-transitory storage medium is provided that stores instructions. The instructions, when executed by at least one processor, cause the at least one processor to execute a first semantic segmentation network and a second semantic segmentation network, wherein, when executed, the first semantic segmentation network obtains first sensor data as input and the second semantic segmentation network obtains second sensor data as input. The instructions, when executed by at least one processor, cause the at least one processor to determine the difference between a first location output by the first semantic segmentation network and a second location output by the second semantic segmentation network. The instructions, when executed by at least one processor, cause the at least one processor to update a first confidence score output by the first semantic segmentation network when the difference between the first location and the second location meets a first predetermined threshold and the data output by the second semantic segmentation network indicates a known low reflectivity at the second location, wherein the update modifies the first confidence score based on the second confidence score. The instructions, when executed by at least one processor, cause the at least one processor to determine an object type based on the data output by the first semantic segmentation network and the second semantic segmentation network.
[0125] According to some non-limiting embodiments or examples, a system is provided, comprising: at least one processor; and at least one non-transitory storage medium storing instructions. The instructions, when executed by the at least one processor, cause the at least one processor to execute a first semantic segmentation network, wherein, upon execution, the first semantic segmentation network receives camera data as input and outputs data associated with a first spatial location of an object, object attribute information associated with the object, object color information associated with the object, and a first detection confidence score associated with the object. The instructions, when executed by the at least one processor, cause the at least one processor to determine angle information based on the first spatial location of the object, map information, and camera calibration information. The instructions, when executed by the at least one processor, cause the at least one processor to execute a second semantic segmentation network, wherein, upon execution, the second semantic segmentation network receives second sensor data, the angle information, the object attribute information, the object color information, and the first detection confidence score as input, and outputs a second detection confidence score and data associated with a second spatial location. The instructions, when executed by the at least one processor, cause the at least one processor to cause control of a vehicle based on the first spatial location, the first detection confidence score, the second spatial location, and the second detection confidence score.
[0126] According to some non-limiting embodiments or examples, a method is provided. The method includes executing, using at least one processor, a first semantic segmentation network, wherein, when executed, the first semantic segmentation network receives camera data as input and outputs data associated with a first spatial location of an object, object attribute information associated with the object, object color information associated with the object, and a first detection confidence score associated with the object. The method includes determining, using the at least one processor, angle information based on the first spatial location of the object, map information, and camera calibration information. The method includes executing, using the at least one processor, a second semantic segmentation network, wherein, when executed, the second semantic segmentation network receives second sensor data, the angle information, the object attribute information, the object color information, and the first detection confidence score as input and outputs a second detection confidence score and data associated with a second spatial location. The method includes controlling, using the at least one processor, a vehicle based on the first spatial location, the first detection confidence score, the second spatial location, and the second detection confidence score.
[0127] According to some non-limiting embodiments or examples, at least one non-transitory storage medium is provided that stores instructions. The instructions, when executed by at least one processor, cause the at least one processor to execute a first semantic segmentation network, wherein, upon execution, the first semantic segmentation network receives camera data as input and outputs data associated with a first spatial location of an object, object attribute information associated with the object, object color information associated with the object, and a first detection confidence score associated with the object. The instructions, when executed by the at least one processor, cause the at least one processor to determine angle information based on the first spatial location of the object, map information, and camera calibration information. The instructions, when executed by the at least one processor, cause the at least one processor to execute a second semantic segmentation network, wherein, upon execution, the second semantic segmentation network receives second sensor data, the angle information, the object attribute information, the object color information, and the first detection confidence score as input and outputs a second detection confidence score and data associated with a second spatial location. The instructions, when executed by the at least one processor, cause the at least one processor to control a vehicle based on the first spatial location, the first detection confidence score, the second spatial location, and the second detection confidence score.
[0128] Further non-limiting aspects or embodiments are set forth in the following numbered clauses:
[0129] Item 1: A system comprising: at least one processor; and at least one non-transitory storage medium storing instructions that, when executed by the at least one processor, cause the at least one processor to: execute a first semantic segmentation network and a second semantic segmentation network, wherein, upon execution, the first semantic segmentation network obtains first sensor data as input and the second semantic segmentation network obtains second sensor data as input; determine a difference between a first location output by the first semantic segmentation network and a second location output by the second semantic segmentation network; update a first confidence score output by the first semantic segmentation network when the difference between the first location and the second location satisfies a first predetermined threshold and the data output by the second semantic segmentation network indicates a known low reflectivity at the second location, wherein the update modifies the first confidence score based on the second confidence score; and determine an object type based on the data output by the first semantic segmentation network and the second semantic segmentation network.
[0130] Clause 2: The system of clause 1, wherein updating the first confidence score comprises setting the first confidence score equal to a second confidence score output by the second semantic segmentation network.
[0131] Clause 3: A system according to clause 1 or 2, wherein the data is color data, further comprising: evaluating the color data by determining the difference between the color data output by the second semantic segmentation network and reference color data, and updating the first confidence score output by the first semantic segmentation network when the difference between the color data and the reference color data satisfies a second predetermined threshold.
[0132] Clause 4: A system according to any one of clauses 1 to 3, wherein the data is attribute data, and further comprising: evaluating the attribute data by determining the difference between the attribute information output by the second semantic segmentation network and reference attribute information, and updating the first confidence score output by the first semantic segmentation network when the difference between the attribute information and the reference attribute information is less than a third predetermined threshold.
[0133] Clause 5: A system according to any one of clauses 1 to 4, wherein the first semantic segmentation network is a LiDAR semantic network, and the LiDAR semantic network is used to generate a first set of bounding boxes associated with objects in the environment, the first set of bounding boxes including the first location and the first confidence score indicating the presence of an object class instance within the first set of bounding boxes.
[0134] Clause 6: A system according to any one of clauses 1 to 5, wherein the second semantic segmentation network is an image semantic network, and the image semantic network is used to generate a second set of bounding boxes associated with objects in the environment, the second set of bounding boxes including the second location and the second confidence score indicating the presence of object class instances within the second set of bounding boxes.
[0135] Clause 7: A system according to any one of clauses 1 to 6, wherein the first location output by the first semantic segmentation network is a three-dimensional location; and wherein determining the difference between the first location and the second location comprises: projecting the first location onto a two-dimensional surface and comparing the projected first location with the second location.
[0136] Clause 8: The system of any of clauses 1 to 7, wherein a bird's eye view network receives as input the outputs of the first and second semantic segmentation networks and outputs an indication of a detected object.
[0137] Clause 9: The system of any one of clauses 1 to 8, wherein the second semantic segmentation network outputs ambient light information for detecting objects.
[0138] Clause 10: The system of clause 3, wherein the reference color is selected to correspond to a color associated with the known low reflectivity.
[0139] Clause 11: The system of clause 4, wherein the reference property is selected to correspond to a property type associated with the known low reflectivity.
[0140] Item 12: A method comprising: executing, using at least one processor, a first semantic segmentation network and a second semantic segmentation network, wherein, upon execution, the first semantic segmentation network obtains first sensor data as input and the second semantic segmentation network obtains second sensor data as input; determining, using the at least one processor, a difference between a first location output by the first semantic segmentation network and a second location output by the second semantic segmentation network; updating, using the at least one processor, a first confidence score output by the first semantic segmentation network when the difference between the first location and the second location satisfies a first predetermined threshold and the data output by the second semantic segmentation network indicates a known low reflectivity at the second location, wherein the update modifies the first confidence score based on the second confidence score; and determining, using the at least one processor, an object type based on the data output by the first semantic segmentation network and the second semantic segmentation network.
[0141] Clause 13: The method of clause 12, wherein updating the first confidence score comprises setting the first confidence score equal to a second confidence score output by the second semantic segmentation network.
[0142] Clause 14: A method according to clause 12 or 13, wherein the data is color data, and the method further comprises: evaluating the color data by determining the difference between the color data output by the second semantic segmentation network and reference color data, and updating the first confidence score output by the first semantic segmentation network when the difference between the color data and the reference color data satisfies a second predetermined threshold.
[0143] Clause 15: A method according to any one of clauses 12 to 14, wherein the data is attribute data, and the method further comprises: evaluating the attribute data by determining the difference between the attribute information output by the second semantic segmentation network and reference attribute information, and updating the first confidence score output by the first semantic segmentation network when the difference between the attribute information and the reference attribute information is less than a third predetermined threshold.
[0144] Clause 16: A method according to any one of clauses 12 to 15, wherein the first semantic segmentation network is a LiDAR semantic network, and the LiDAR semantic network is used to generate a first set of bounding boxes associated with objects in the environment, the first set of bounding boxes including the first location and the first confidence score indicating the presence of an object class instance within the first set of bounding boxes.
[0145] Clause 17: A method according to any one of clauses 12 to 16, wherein the second semantic segmentation network is an image semantic network, and the image semantic network is used to generate a second set of bounding boxes associated with objects in the environment, the second set of bounding boxes including the second location and the second confidence score indicating the presence of object class instances within the second set of bounding boxes.
[0146] Clause 18: A method according to any one of clauses 12 to 17, wherein the first location output by the first semantic segmentation network is a three-dimensional location; and wherein determining the difference between the first location and the second location comprises: projecting the first location onto a two-dimensional surface and comparing the projected first location with the second location.
[0147] Clause 19: A method according to any of clauses 12 to 18, wherein a bird's eye view network receives as input the outputs of the first semantic segmentation network and the second semantic segmentation network and outputs an indication of a detected object.
[0148] Clause 20: The method of any one of clauses 12 to 19, wherein the second semantic segmentation network outputs ambient light information for detecting objects.
[0149] Clause 21: The method of clause 14, wherein the reference color is selected to correspond to a color associated with the known low reflectivity.
[0150] Clause 22: The method of clause 15, wherein the reference property is selected to correspond to a property type associated with the known low reflectivity.
[0151] Item 23: At least one non-transitory storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to: execute a first semantic segmentation network and a second semantic segmentation network, wherein, upon execution, the first semantic segmentation network obtains first sensor data as input and the second semantic segmentation network obtains second sensor data as input; determine a difference between a first location output by the first semantic segmentation network and a second location output by the second semantic segmentation network; update a first confidence score output by the first semantic segmentation network if the difference between the first location and the second location satisfies a first predetermined threshold and the data output by the second semantic segmentation network indicates a known low reflectivity at the second location, wherein the update modifies the first confidence score based on the second confidence score; and determine an object type based on the data output by the first semantic segmentation network and the second semantic segmentation network.
[0152] Clause 24: At least one non-transitory storage medium according to clause 23, wherein updating the first confidence score comprises setting the first confidence score equal to a second confidence score output by the second semantic segmentation network.
[0153] Clause 25: At least one non-transitory storage medium according to clause 23 or 24, wherein the data is color data, further comprising: evaluating the color data by determining the difference between the color data output by the second semantic segmentation network and reference color data, and updating the first confidence score output by the first semantic segmentation network when the difference between the color data and the reference color data satisfies a second predetermined threshold.
[0154] Clause 26: At least one non-transitory storage medium according to any one of clauses 23 to 25, wherein the data is attribute data, further comprising: evaluating the attribute data by determining the difference between the attribute information output by the second semantic segmentation network and reference attribute information, and updating the first confidence score output by the first semantic segmentation network when the difference between the attribute information and the reference attribute information is less than a third predetermined threshold.
[0155] Clause 27: At least one non-transitory storage medium according to any one of clauses 23 to 26, wherein the first semantic segmentation network is a LiDAR semantic network, and the LiDAR semantic network is used to generate a first set of bounding boxes associated with objects in the environment, the first set of bounding boxes including the first location and the first confidence score indicating the presence of an object class instance within the first set of bounding boxes.
[0156] Clause 28: At least one non-transitory storage medium according to any one of clauses 23 to 27, wherein the second semantic segmentation network is an image semantic network, and the image semantic network is used to generate a second set of bounding boxes associated with objects in the environment, the second set of bounding boxes including the second location and the second confidence score indicating the presence of an object class instance within the second set of bounding boxes.
[0157] Clause 29: At least one non-transitory storage medium according to any one of clauses 23 to 28, wherein the first location output by the first semantic segmentation network is a three-dimensional location; and wherein determining the difference between the first location and the second location comprises: projecting the first location onto a two-dimensional surface and comparing the projected first location with the second location.
[0158] Clause 30: At least one non-transitory storage medium according to any one of clauses 23 to 29, wherein a bird's eye view network receives as input the output of the first semantic segmentation network and the second semantic segmentation network and outputs an indication of a detected object.
[0159] Clause 31: At least one non-transitory storage medium according to any one of clauses 23 to 30, wherein the second semantic segmentation network outputs ambient light information for detecting an object.
[0160] Clause 32: The at least one non-transitory storage medium of clause 25, wherein the reference color is selected to correspond to a color associated with the known low reflectivity.
[0161] Clause 33: The at least one non-transitory storage medium of clause 26, wherein the reference property is selected to correspond to a property type associated with the known low reflectivity.
[0162] Item 34: A system comprising: at least one processor; and at least one non-transitory storage medium storing instructions that, when executed by the at least one processor, cause the at least one processor to: execute a first semantic segmentation network, wherein, upon execution, the first semantic segmentation network obtains camera data as input and outputs data associated with a first spatial location of an object, object attribute information associated with the object, object color information associated with the object, and a first detection confidence score associated with the object; determine angle information based on the first spatial location of the object, map information, and camera calibration information; execute a second semantic segmentation network, wherein, upon execution, the second semantic segmentation network obtains second sensor data, the angle information, the object attribute information, the object color information, and the first detection confidence score as input and outputs a second detection confidence score and data associated with a second spatial location; and cause a vehicle to be controlled based on the first spatial location, the first detection confidence score, the second spatial location, and the second detection confidence score.
[0163] Clause 35: The system of clause 34, wherein the angle information is determined based on the map information and the camera calibration information.
[0164] Clause 36: The system of clause 34 or 35, wherein the object attribute information is based on a model associated with an object classification of the object.
[0165] Clause 37: The system of any of clauses 34 to 36, wherein the object color information is red, green, and blue values of the object as captured in the camera data.
[0166] Clause 38: The system of any one of clauses 34 to 37, wherein the object property information of the object or the object color information of the object indicates a low reflectivity associated with the object.
[0167] Clause 39: A system according to clause 38, wherein a first detection confidence score associated with the object is more accurate than a second detection confidence score based on low reflectivity associated with the object, and wherein the second semantic segmentation network updates the second detection confidence score based on the first detection confidence score.
[0168] Clause 40: The system of any one of clauses 34 to 39, wherein the object is a black vehicle and the attribute information is based on a vehicle model associated with a black vehicle.
[0169] Clause 41: A system according to any one of clauses 34 to 40, wherein the object is a pedestrian wearing dark clothing, and the attribute information is based on a pedestrian model.
[0170] Item 42: A method comprising: executing, using at least one processor, a first semantic segmentation network, wherein, when executed, the first semantic segmentation network obtains camera data as input and outputs data associated with a first spatial location of an object, object attribute information associated with the object, object color information associated with the object, and a first detection confidence score associated with the object; determining, using the at least one processor, angle information based on the first spatial location of the object, map information, and camera calibration information; executing, using the at least one processor, a second semantic segmentation network, wherein, when executed, the second semantic segmentation network obtains second sensor data, the angle information, the object attribute information, the object color information, and the first detection confidence score as input and outputs a second detection confidence score and data associated with the second spatial location; and controlling, using the at least one processor, a vehicle based on the first spatial location, the first detection confidence score, the second spatial location, and the second detection confidence score.
[0171] Clause 43: The method of clause 42, wherein the angle information is determined based on the map information and the camera calibration information.
[0172] Clause 44: The method of clause 42 or 43, wherein the object attribute information is based on a model associated with an object classification of the object.
[0173] Clause 45: The method of any one of clauses 42 to 44, wherein the object color information is red, green, and blue values of the object as captured in the camera data.
[0174] Clause 46: The method of any one of clauses 42 to 45, wherein the object property information of the object or the object color information of the object indicates a low reflectivity associated with the object.
[0175] Clause 47: A method according to clause 46, wherein a first detection confidence score associated with the object is more accurate than a second detection confidence score based on a low reflectivity associated with the object, and wherein the second semantic segmentation network updates the second detection confidence score based on the first detection confidence score.
[0176] Clause 48: The method of any one of clauses 42 to 47, wherein the object is a black vehicle and the attribute information is based on a vehicle model associated with a black vehicle.
[0177] Clause 49: A method according to any one of clauses 42 to 48, wherein the object is a pedestrian wearing dark clothing, and the attribute information is based on a pedestrian model.
[0178] Item 50: At least one non-transitory storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to: execute a first semantic segmentation network, wherein, upon execution, the first semantic segmentation network obtains camera data as input and outputs data associated with a first spatial location of an object, object attribute information associated with the object, object color information associated with the object, and a first detection confidence score associated with the object; determine angle information based on the first spatial location of the object, map information, and camera calibration information; execute a second semantic segmentation network, wherein, upon execution, the second semantic segmentation network obtains second sensor data, the angle information, the object attribute information, the object color information, and the first detection confidence score as input and outputs a second detection confidence score and data associated with a second spatial location; and cause a vehicle to be controlled based on the first spatial location, the first detection confidence score, the second spatial location, and the second detection confidence score.
[0179] Clause 51: The at least one non-transitory storage medium of clause 50, wherein the angle information is determined based on the map information and the camera calibration information.
[0180] Clause 52: At least one non-transitory storage medium according to clause 50 or 51, wherein the object attribute information is based on a model associated with an object classification of the object.
[0181] Clause 53: At least one non-transitory storage medium according to any one of clauses 50 to 52, wherein the object color information is red, green, and blue values of the object as captured in the camera data.
[0182] Clause 54: At least one non-transitory storage medium according to any one of clauses 50 to 53, wherein the object property information of the object or the object color information of the object indicates a low reflectivity associated with the object.
[0183] Item 55: At least one non-transitory storage medium according to Item 54, wherein a first detection confidence score associated with the object is more accurate than a second detection confidence score based on a low reflectivity associated with the object, and wherein the second semantic segmentation network updates the second detection confidence score based on the first detection confidence score.
[0184] Clause 56: At least one non-transitory storage medium according to any one of clauses 50 to 55, wherein the object is a black vehicle and the attribute information is based on a vehicle model associated with a black vehicle.
[0185] Clause 57: At least one non-transitory storage medium according to any one of clauses 50 to 56, wherein the object is a pedestrian wearing dark clothing, and the attribute information is based on a pedestrian model.
[0186] In the foregoing description, aspects and embodiments of the present disclosure have been described with reference to many specific details, which may vary from implementation to implementation. Therefore, the description and drawings should be regarded as illustrative, not restrictive. The sole and exclusive indication of the scope of the invention, and what the applicants intend to be the scope of the invention, is the literal and equivalent scope of the claims authorized from this application in the specific form of the claims of the grant announcement, including any subsequent amendments. Any definitions of terms expressly set forth herein for inclusion in such claims should be taken in the sense of such terms as used in the claims. In addition, when the term "also comprising" is used in the foregoing description or the appended claims, the phrase may be followed by additional steps or entities, or sub-steps / sub-entities of the steps or entities previously described.
Claims
1. A system comprising: at least one processor; as well as at least one non-transitory storage medium storing instructions that, when executed by the at least one processor, cause the at least one processor to: executing a first semantic segmentation network and a second semantic segmentation network, wherein, upon execution, the first semantic segmentation network receives as input the first sensor data and the second semantic segmentation network receives as input the second sensor data; determining a difference between a first location output by the first semantic segmentation network and a second location output by the second semantic segmentation network; updating a first confidence score output by the first semantic segmentation network if a difference between the first location and the second location satisfies a first predetermined threshold and the data output by the second semantic segmentation network indicates a known low reflectivity at the second location, wherein the updating modifies the first confidence score based on the second confidence score; and An object type is determined based on data output by the first semantic segmentation network and the second semantic segmentation network.
2. The system according to claim 1, wherein: Updating the first confidence score includes setting the first confidence score equal to a second confidence score output by the second semantic segmentation network.
3. The system according to claim 1 or 2, wherein: The data is color data and also includes: evaluating the color data by determining a difference between the color data output by the second semantic segmentation network and reference color data, and When the difference between the color data and the reference color data satisfies a second predetermined threshold, the first confidence score output by the first semantic segmentation network is updated.
4. The system according to any one of claims 1 to 3, wherein: The data is attribute data and also includes: evaluating the attribute data by determining a difference between the attribute information output by the second semantic segmentation network and reference attribute information, and When the difference between the attribute information and the reference attribute information is less than a third predetermined threshold, the first confidence score output by the first semantic segmentation network is updated.
5. The system according to any one of claims 1 to 4, wherein: The first semantic segmentation network is a LiDAR semantic network configured to generate a first set of bounding boxes associated with objects in an environment, the first set of bounding boxes including the first location and the first confidence score indicating the presence of an object class instance within the first set of bounding boxes.
6. The system according to any one of claims 1 to 5, wherein: The second semantic segmentation network is an image semantic network configured to generate a second set of bounding boxes associated with objects in the environment, the second set of bounding boxes including the second locations and the second confidence scores indicating the presence of object class instances within the second set of bounding boxes.
7. The system according to any one of claims 1 to 6, wherein: The first location output by the first semantic segmentation network is a three-dimensional location; as well as Wherein determining the difference between the first location and the second location comprises projecting the first location onto a two-dimensional surface and comparing the projected first location with the second location.
8. The system according to any one of claims 1 to 7, wherein: A bird's eye view network receives as input the outputs of the first semantic segmentation network and the second semantic segmentation network and outputs an indication of detected objects.
9. The system according to any one of claims 1 to 8, wherein: The second semantic segmentation network outputs ambient light information for detecting objects.
10. The system according to claim 3, wherein: The reference color is selected to correspond to a color associated with the known low reflectivity.
11. The system according to claim 4, wherein: The reference property is selected to correspond to a property type associated with the known low reflectivity.
12. A method comprising: executing, with at least one processor, a first semantic segmentation network and a second semantic segmentation network, wherein, upon execution, the first semantic segmentation network receives first sensor data as input and the second semantic segmentation network receives second sensor data as input; determining, with the at least one processor, a difference between a first location output by the first semantic segmentation network and a second location output by the second semantic segmentation network; updating, with the at least one processor, a first confidence score output by the first semantic segmentation network if a difference between the first location and the second location satisfies a first predetermined threshold and the data output by the second semantic segmentation network indicates a known low reflectivity at the second location, wherein the updating modifies the first confidence score based on the second confidence score; and Determine, using the at least one processor, an object type based on data output by the first semantic segmentation network and the second semantic segmentation network.
13. The method according to claim 12, wherein: Updating the first confidence score includes setting the first confidence score equal to a second confidence score output by the second semantic segmentation network.
14. The method according to claim 12 or 13, wherein: The data is color data, and the method further includes: evaluating the color data by determining a difference between the color data output by the second semantic segmentation network and reference color data, and When the difference between the color data and the reference color data satisfies a second predetermined threshold, the first confidence score output by the first semantic segmentation network is updated.
15. The method according to any one of claims 12 to 14, wherein The data is attribute data, and the method further includes: evaluating the attribute data by determining a difference between the attribute information output by the second semantic segmentation network and reference attribute information, and When the difference between the attribute information and the reference attribute information is less than a third predetermined threshold, the first confidence score output by the first semantic segmentation network is updated.
16. The method according to any one of claims 12 to 15, wherein The first semantic segmentation network is a LiDAR semantic network configured to generate a first set of bounding boxes associated with objects in an environment, the first set of bounding boxes including the first location and the first confidence score indicating the presence of an object class instance within the first set of bounding boxes.
17. The method according to any one of claims 12 to 16, wherein The second semantic segmentation network is an image semantic network configured to generate a second set of bounding boxes associated with objects in the environment, the second set of bounding boxes including the second locations and the second confidence scores indicating the presence of object class instances within the second set of bounding boxes.
18. The method according to any one of claims 12 to 17, wherein The first location output by the first semantic segmentation network is a three-dimensional location; as well as Wherein determining the difference between the first location and the second location comprises projecting the first location onto a two-dimensional surface and comparing the projected first location with the second location.
19. The method according to any one of claims 12 to 18, wherein A bird's eye view network receives as input the outputs of the first semantic segmentation network and the second semantic segmentation network and outputs an indication of detected objects.
20. The method according to any one of claims 12 to 19, wherein The second semantic segmentation network outputs ambient light information for detecting objects.
21. The method according to claim 14, wherein The reference color is selected to correspond to a color associated with the known low reflectivity.
22. The method according to claim 15, wherein The reference property is selected to correspond to a property type associated with the known low reflectivity.
23. At least one non-transitory storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to: Execute the first semantic segmentation network and the second semantic segmentation network, wherein, When executed, the first semantic segmentation network obtains first sensor data as input, and the second semantic segmentation network obtains second sensor data as input; determining a difference between a first location output by the first semantic segmentation network and a second location output by the second semantic segmentation network; updating a first confidence score output by the first semantic segmentation network if a difference between the first location and the second location satisfies a first predetermined threshold and the data output by the second semantic segmentation network indicates a known low reflectivity at the second location, wherein the updating modifies the first confidence score based on the second confidence score; and An object type is determined based on data output by the first semantic segmentation network and the second semantic segmentation network.
24. The at least one non-transitory storage medium of claim 23, wherein: Updating the first confidence score includes setting the first confidence score equal to a second confidence score output by the second semantic segmentation network.
25. At least one non-transitory storage medium according to claim 23 or 24, wherein: The data is color data and also includes: evaluating the color data by determining a difference between the color data output by the second semantic segmentation network and reference color data, and When the difference between the color data and the reference color data satisfies a second predetermined threshold, the first confidence score output by the first semantic segmentation network is updated.
26. At least one non-transitory storage medium according to any one of claims 23 to 25, wherein: The data is attribute data and also includes: evaluating the attribute data by determining a difference between the attribute information output by the second semantic segmentation network and reference attribute information, and When the difference between the attribute information and the reference attribute information is less than a third predetermined threshold, the first confidence score output by the first semantic segmentation network is updated.
27. At least one non-transitory storage medium according to any one of claims 23 to 26, wherein: The first semantic segmentation network is a LiDAR semantic network configured to generate a first set of bounding boxes associated with objects in an environment, the first set of bounding boxes including the first location and the first confidence score indicating the presence of an object class instance within the first set of bounding boxes.
28. At least one non-transitory storage medium according to any one of claims 23 to 27, wherein: The second semantic segmentation network is an image semantic network configured to generate a second set of bounding boxes associated with objects in the environment, the second set of bounding boxes including the second locations and the second confidence scores indicating the presence of object class instances within the second set of bounding boxes.
29. At least one non-transitory storage medium according to any one of claims 23 to 28, wherein: The first location output by the first semantic segmentation network is a three-dimensional location; as well as Wherein determining the difference between the first location and the second location comprises projecting the first location onto a two-dimensional surface and comparing the projected first location with the second location.
30. At least one non-transitory storage medium according to any one of claims 23 to 29, wherein: A bird's eye view network receives as input the outputs of the first semantic segmentation network and the second semantic segmentation network and outputs an indication of detected objects.
31. At least one non-transitory storage medium according to any one of claims 23 to 30, wherein: The second semantic segmentation network outputs ambient light information for detecting objects.
32. The at least one non-transitory storage medium of claim 25, wherein: The reference color is selected to correspond to a color associated with the known low reflectivity.
33. The at least one non-transitory storage medium of claim 26, wherein: The reference property is selected to correspond to a property type associated with the known low reflectivity.
34. A system comprising: at least one processor; as well as at least one non-transitory storage medium storing instructions that, when executed by the at least one processor, cause the at least one processor to: executing a first semantic segmentation network, wherein, upon execution, the first semantic segmentation network obtains as input the camera data and outputs data associated with a first spatial location of an object, object attribute information associated with the object, object color information associated with the object, and a first detection confidence score associated with the object; determining angle information based on the first spatial location of the object, map information, and camera calibration information; executing a second semantic segmentation network, wherein, upon execution, the second semantic segmentation network receives as input the second sensor data, the angle information, the object attribute information, the object color information, and the first detection confidence score, and outputs a second detection confidence score and data associated with a second spatial location; and A vehicle is caused to be controlled based on the first spatial location, the first detection confidence score, the second spatial location, and the second detection confidence score.
35. The system of claim 34, wherein: The angle information is determined based on the map information and the camera calibration information.
36. The system of claim 34 or 35, wherein: The object attribute information is based on a model associated with an object classification of the object.
37. A system according to any one of claims 34 to 36, wherein: The object color information is the red value, green value, and blue value of the object as captured in the camera data.
38. A system according to any one of claims 34 to 37, wherein The object attribute information of the object or the object color information of the object indicates low reflectivity associated with the object.
39. The system of claim 38, wherein: A first detection confidence score associated with the object is more accurate than a second detection confidence score based on low reflectivity associated with the object, and wherein the second semantic segmentation network updates the second detection confidence score based on the first detection confidence score.
40. The system of any one of claims 34 to 39, wherein: The object is a black vehicle, and the attribute information is based on a vehicle model associated with the black vehicle.
41. A system according to any one of claims 34 to 40, wherein The object is a pedestrian wearing dark clothes, and the attribute information is based on a pedestrian model.
42. A method comprising: executing, with at least one processor, a first semantic segmentation network, wherein, upon execution, the first semantic segmentation network takes camera data as input and outputs data associated with a first spatial location of an object, object attribute information associated with the object, object color information associated with the object, and a first detection confidence score associated with the object; determining, with the at least one processor, angle information based on the first spatial location of the object, map information, and camera calibration information; executing, with the at least one processor, a second semantic segmentation network, wherein, upon execution, the second semantic segmentation network receives as input the second sensor data, the angle information, the object attribute information, the object color information, and the first detection confidence score, and outputs a second detection confidence score and data associated with a second spatial location; and The at least one processor causes control of a vehicle based on the first spatial location, the first detection confidence score, the second spatial location, and the second detection confidence score.
43. The method according to claim 42, wherein The angle information is determined based on the map information and the camera calibration information.
44. The method according to claim 42 or 43, wherein The object attribute information is based on a model associated with an object classification of the object.
45. The method according to any one of claims 42 to 44, wherein The object color information is the red value, green value, and blue value of the object as captured in the camera data.
46. The method according to any one of claims 42 to 45, wherein The object attribute information of the object or the object color information of the object indicates low reflectivity associated with the object.
47. The method of claim 46, wherein A first detection confidence score associated with the object is more accurate than a second detection confidence score based on low reflectivity associated with the object, and wherein the second semantic segmentation network updates the second detection confidence score based on the first detection confidence score.
48. The method according to any one of claims 42 to 47, wherein The object is a black vehicle, and the attribute information is based on a vehicle model associated with the black vehicle.
49. The method according to any one of claims 42 to 48, wherein The object is a pedestrian wearing dark clothes, and the attribute information is based on a pedestrian model.
50. At least one non-transitory storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to: Execute the first semantic segmentation network, where When executed, the first semantic segmentation network obtains camera data as input and outputs data associated with a first spatial location of an object, object attribute information associated with the object, object color information associated with the object, and a first detection confidence score associated with the object; determining angle information based on the first spatial location of the object, map information, and camera calibration information; executing a second semantic segmentation network, wherein, upon execution, the second semantic segmentation network receives as input the second sensor data, the angle information, the object attribute information, the object color information, and the first detection confidence score, and outputs a second detection confidence score and data associated with a second spatial location; as well as A vehicle is caused to be controlled based on the first spatial location, the first detection confidence score, the second spatial location, and the second detection confidence score.
51. The at least one non-transitory storage medium of claim 50, wherein: The angle information is determined based on the map information and the camera calibration information.
52. At least one non-transitory storage medium according to claim 50 or 51, wherein: The object attribute information is based on a model associated with an object classification of the object.
53. At least one non-transitory storage medium according to any one of claims 50 to 52, wherein: The object color information is the red value, green value, and blue value of the object as captured in the camera data.
54. At least one non-transitory storage medium according to any one of claims 50 to 53, wherein: The object attribute information of the object or the object color information of the object indicates low reflectivity associated with the object.
55. The at least one non-transitory storage medium of claim 54, wherein: A first detection confidence score associated with the object is more accurate than a second detection confidence score based on low reflectivity associated with the object, and wherein the second semantic segmentation network updates the second detection confidence score based on the first detection confidence score.
56. At least one non-transitory storage medium according to any one of claims 50 to 55, wherein: The object is a black vehicle, and the attribute information is based on a vehicle model associated with the black vehicle.
57. At least one non-transitory storage medium according to any one of claims 50 to 56, wherein: The object is a pedestrian wearing dark clothes, and the attribute information is based on a pedestrian model.