Determining object orientation from map parameters and group parameters

By combining map parameters, group parameters and sensor data, the autonomous vehicle can more accurately determine the orientation of the object under sparse LiDAR conditions, solving the problem of difficult object detection in the prior art and improving detection and tracking performance.

CN120035850APending Publication Date: 2025-05-23MOTIONAL AD LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380070247.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-08-02
Filing Date
2023-07-31
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

Autonomous vehicles have difficulty accurately determining the location, dimension, and orientation of objects in the environment when dealing with sparse LiDAR point clouds, especially when objects are far away or partially obstructed.

Method used

By obtaining map parameters and group parameters using at least one processor, and combining sensor data, the orientation data of the object is determined. Map parameters provide the predetermined location information of the object in the environment, the group parameters describe the predetermined connection between the objects, and the sensor data supplement and verification of the orientation data.

Benefits of technology

Improved object detection and tracking performance of autonomous vehicles under sparse LiDAR conditions, and improved the accuracy and reliability of object orientation determination.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120035850A_ABST
    Figure CN120035850A_ABST
Patent Text Reader

Abstract

A method for object orientation determination is provided that may include obtaining map parameters and group parameters and determining orientation data using the map parameters and group parameters. Some described methods also include obtaining sensor data and using the sensor data to determine orientation data. Systems and computer program products are also provided.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to U.S. Provisional Application No. 63 / 394,316, filed on August 2, 2022, entitled “OBJECT ORIENTATION DETERMINATION FROM MAP AND GROUP PARAMETERS,” which is incorporated herein by reference in its entirety. BRIEF DESCRIPTION OF THE DRAWINGS

[0002] Figure 1 is an example environment in which a vehicle including one or more components of an autonomous system may be implemented;

[0003] Figure 2 is a diagram of one or more example systems of a vehicle including an autonomous system;

[0004] Figure 3 yes Figure 1 and Figure 2 diagrams of components of one or more example devices and / or one or more example systems;

[0005] Figure 4A is a diagram of certain components of an example autonomous system;

[0006] Figure 4B is a diagram of an example implementation of a neural network;

[0007] Figure 4C and Figure 4D is a diagram of an example operation of a convolutional neural network (CNN);

[0008] Figure 5 is a diagram of an example implementation of an object orientation determination process.

[0009] Figure 6 is a diagram of an example implementation of an object orientation determination process.

[0010] Fig. 7A and Figure 7B is a diagram of an example implementation of an object orientation determination process.

[0011] Figure 8 is a flow chart of an example process for object orientation determination.

[0012] Fig. 9 is a block diagram of example processing within an AV computing system and the flow of data between AV computing and related systems and components.

[0013] Fig. 10A It shows that it can be Figure 5 or Fig. 9 A diagram showing examples of map parameters and group parameters that an AV computing system uses to determine an orientation and / or position of one or more objects is shown.

[0014] Fig. 10B is a diagram showing the perception of the environment around the AV based on sensor data and object arrangement before generating orientation data. DETAILED DESCRIPTION

[0015] In the following description, for the purpose of explanation, many specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it will be apparent that the embodiments described in the present disclosure can be implemented without these specific details. In some instances, well-known configurations and devices are illustrated in block diagram form to avoid unnecessarily obscuring aspects of the present disclosure.

[0016] In the accompanying drawings, for ease of description, the specific arrangement or order of schematic elements (such as those representing systems, devices, modules, instruction blocks and / or data elements, etc.) is illustrated. However, those skilled in the art will understand that, unless explicitly described, the specific order or arrangement of schematic elements in the accompanying drawings is not intended to mean that a specific processing order or sequence, or separation of processing is required. In addition, unless explicitly described, the inclusion of schematic elements in the accompanying drawings is not intended to mean that such elements are required in all embodiments, nor is it intended to mean that the features represented by such elements cannot be included in some embodiments or cannot be combined with other elements in some embodiments.

[0017] In addition, in the accompanying drawings, connecting elements (such as solid or dotted lines or arrows, etc.) are used to illustrate the connection, relationship or association between or among two or more other schematic elements, and the absence of any such connecting elements is not intended to mean that there can be no connection, relationship or association. In other words, some connections, relationships or associations between elements are not illustrated in the accompanying drawings so as not to obscure the present disclosure. In addition, for ease of illustration, a single connecting element can be used to represent multiple connections, relationships or associations between elements. For example, if the connecting element represents the communication of a signal, data or instruction (e.g., "software instruction"), it will be understood by those skilled in the art that such an element can represent one or more than one signal path (e.g., bus) that may be needed to affect the communication.

[0018] Although the terms "first", "second", and / or "third", etc. are used to describe various elements, these elements should not be limited by these terms. The terms "first", "second", and / or "third" are only used to distinguish one element from another element. For example, a first contact may be referred to as a second contact, and similarly, a second contact may be referred to as a first contact without departing from the scope of the described embodiments. Both the first contact and the second contact are contacts, but they are not the same contact.

[0019] The terms used in the specification of the various embodiments described herein are included only for the purpose of describing specific embodiments and are not intended to be limiting. As used in the specification of the various embodiments described and the appended claims, the singular forms "a", "an" and "the" are also intended to include plural forms and can be used interchangeably with "one or more than one" or "at least one", unless the context clearly indicates otherwise. It will also be understood that the term "and / or" as used herein refers to and includes any and all possible combinations of one or more than one of the associated listed items. It will also be understood that when the terms "include", "comprise", "have" and / or "have" are used in this specification, the stated features, integers, steps, operations, elements and / or components are specifically stated, but the presence or addition of one or more than one other features, integers, steps, operations, elements, components and / or groups thereof are not excluded.

[0020] As used herein, the terms "communication" and "communicating" refer to at least one of receiving, receiving, transmitting, transmitting and / or providing information (or information represented by, for example, data, signals, messages, instructions and / or commands, etc.). For a unit (e.g., a device, a system, a component of a device or system, and / or a combination thereof) to communicate with another unit, this means that the unit is able to directly or indirectly receive information from the other unit and / or send (e.g., transmit) information to the other unit. This can refer to a direct or indirect connection that is wired and / or wireless in nature. In addition, even if the transmitted information can be modified, processed, relayed and / or routed between the first unit and the second unit, the two units can communicate with each other. For example, even if the first unit passively receives information and does not actively transmit information to the second unit, the first unit can communicate with the second unit. As another example, if at least one intermediary unit (e.g., a third unit located between the first unit and the second unit) processes the information received from the first unit and transmits the processed information to the second unit, the first unit can communicate with the second unit. In some embodiments, a message may refer to a network packet (eg, a data packet, etc.) that includes data.

[0021] As used herein, the term "if" is optionally interpreted to mean "when," "at," "in response to being determined to be," and / or "in response to being detected," etc., depending on the context. Similarly, the phrases "if it is determined" or "if [the stated condition or event] is detected" are optionally interpreted to mean "when determining," "in response to being determined to be" or "when [the stated condition or event] is detected," and / or "in response to being detected," etc., depending on the context. In addition, as used herein, the terms "have," "have," or "possess," etc. are intended to be open-ended terms. Furthermore, unless expressly stated otherwise, the phrase "based on" is intended to mean "based at least in part on."

[0022] “At least one” and “one or more than one” include a function being performed by one element, a function being performed by more than one element (eg, in a distributed fashion), several functions being performed by one element, several functions being performed by several elements, or any combination of the above.

[0023] Some embodiments of the present disclosure are described herein in conjunction with threshold values. As described herein, satisfying (such as meeting, etc.) a threshold value may refer to: a value greater than a threshold value, a value more than a threshold value, a value higher than a threshold value, a value greater than or equal to a threshold value, a value less than a threshold value, a value less than a threshold value, a value lower than a threshold value, a value less than or equal to a threshold value, and a value / or equal to a threshold value, etc.

[0024] Reference will now be made in detail to embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the various embodiments described. However, it will be apparent to one of ordinary skill in the art that the various embodiments described may be implemented without these specific details. In other cases, well-known methods, processes, components, circuits, and networks have not yet been described in detail in order not to unnecessarily obscure aspects of the embodiments.

[0025] General Overview

[0026] In an autonomous vehicle, a LiDAR-based object detector may be used to predict bounding boxes for objects, agents, and actors (such as vehicles) in the environment that the autonomous vehicle is operating in. Autonomous vehicles can perform this task relatively easily with dense LiDAR point clouds, but have difficulty when sparse point clouds are received, for example because objects are far away or partially occluded.

[0027] In particular, it may be difficult for an autonomous vehicle to determine the precise location and orientation of other vehicles in the environment, and identifying what is in front of or behind a vehicle is particularly challenging. For dynamic vehicles, autonomous vehicles can infer what is in front of and behind the vehicle from object motion, but still may not have the required accuracy. For static vehicles, autonomous vehicles cannot use object motion, which makes determination even more difficult.

[0028] In some aspects and / or embodiments, the systems, methods, and computer program products described herein include and / or implement object orientation determination. The method includes obtaining, using at least one processor, a map parameter indicating a predetermined position of a first object in an environment in which the autonomous vehicle is configured to operate. The method includes obtaining, using at least one processor, a group parameter indicating a predetermined connection between objects in a group, wherein the environment includes the object. The method includes determining, using at least one processor, orientation data indicating an orientation of at least one object among the first object and the objects in the group based on the map parameter and the group parameter. The method includes providing, using at least one processor, object detection data associated with at least one object to a device based on the orientation data, wherein the object detection data indicates detection of one or more spatial features of the at least one object.

[0029] Advantages of techniques for object orientation determination, by implementation of the systems, methods, and computer program products described herein, include improved object detection and / or object tracking for autonomous vehicles. For example, the disclosed techniques can improve sparse LiDAR cloud detection when determining the position, dimension, and orientation of an object. Advantageously, the techniques can be incorporated into one or more points in an object labeling pipeline that includes object detection, object tracking, and post-processing steps. The disclosed techniques can be advantageously used to improve the performance (e.g., output quality) of neural network models when estimating object orientation, dimension, and location.

[0030] Reference now Figure 1, illustrates an example environment 100 in which vehicles including autonomous systems and vehicles not including autonomous systems operate. As illustrated, the environment 100 includes vehicles 102a-102n, objects 104a-104n, routes 106a-106n, areas 108, vehicle-to-infrastructure (V2I) devices 110, a network 112, a remote autonomous vehicle (AV) system 114, a fleet management system 116, and a V2I system 118. The vehicles 102a-102n, the vehicle-to-infrastructure (V2I) devices 110, the network 112, the autonomous vehicle (AV) system 114, the fleet management system 116, and the V2I system 118 are interconnected (e.g., establish connections for communication, etc.) via wired connections, wireless connections, or a combination of wired or wireless connections. In some embodiments, objects 104a-104n are interconnected with at least one of vehicles 102a-102n, vehicle-to-infrastructure (V2I) devices 110, networks 112, autonomous vehicle (AV) systems 114, fleet management systems 116, and V2I systems 118 via wired connections, wireless connections, or a combination of wired or wireless connections.

[0031] Vehicles 102a-102n (individually referred to as vehicles 102 and collectively referred to as vehicles 102) include at least one device configured to transport goods and / or people. In some embodiments, vehicles 102 are configured to communicate with V2I devices 110, remote AV systems 114, fleet management systems 116, and / or V2I systems 118 via network 112. In some embodiments, vehicles 102 include cars, buses, trucks, and / or trains, etc. In some embodiments, vehicles 102 are similar to vehicles 200 described herein (see Figure 2 ). In some embodiments, vehicles 200 in the set of vehicles 200 are associated with an autonomous queue manager. In some embodiments, vehicles 102 travel along respective routes 106a-106n (individually referred to as routes 106 and collectively referred to as routes 106) as described herein. In some embodiments, one or more vehicles 102 include an autonomous system (e.g., an autonomous system that is the same as or similar to autonomous system 202).

[0032] Objects 104a-104n (individually referred to as objects 104 and collectively referred to as objects 104) include, for example, at least one vehicle, at least one pedestrian, at least one cyclist, and / or at least one structure (e.g., a building, a sign, a fire hydrant, etc.), etc. Each object 104 is stationary (e.g., located at a fixed location and over a period of time) or moves (e.g., has a speed and is associated with at least one trajectory). In some embodiments, objects 104 are associated with corresponding locations in area 108.

[0033] Routes 106a-106n (individually referred to as routes 106 and collectively referred to as routes 106) are each associated with (e.g., specify a series of actions (also referred to as trajectories) connecting states along which the AV can navigate. Each route 106 begins at an initial state (e.g., a state corresponding to a first spatiotemporal location and / or speed, etc.) and ends at a final target state (e.g., a state corresponding to a second spatiotemporal location different from the first spatiotemporal location) or a target zone (e.g., a subspace of acceptable states (e.g., terminal states)). In some embodiments, the first state includes a location where one or more individuals will board the AV, and the second state or zone includes one or more locations where one or more individuals boarding the AV will disembark. In some embodiments, routes 106 include multiple acceptable state sequences (e.g., multiple spatiotemporal location sequences) that are associated with (e.g., define multiple trajectories). In an example, routes 106 include only high-level actions or imprecise state locations, such as a series of connecting roads indicating a change of direction at a roadway intersection, etc. Additionally or alternatively, the route 106 may include more precise actions or states, such as, for example, specific target lanes or precise locations within lane regions and target speeds at those locations, etc. In an example, the route 106 includes a plurality of precise state sequences along at least one high-level action with a limited look-ahead horizon to an intermediate target, wherein a combination of consecutive iterations of the limited horizon state sequences cumulatively correspond to a plurality of trajectories that collectively form a high-level route terminating at a final target state or region.

[0034] The area 108 includes a physical area (e.g., a geographic region) in which the vehicle 102 can navigate. In an example, the area 108 includes at least one state (e.g., a country, a province, a separate state of a plurality of states included in a country, etc.), at least a portion of a state, at least one city, at least a portion of a city, etc. In some embodiments, the area 108 includes at least one named thoroughfare (referred to herein as a "road"), such as a highway, an interstate highway, a parkway, a city street, etc. Additionally or alternatively, in some examples, the area 108 includes at least one unnamed road, such as a driveway, a section of a parking lot, a section of an open space and / or undeveloped area, a dirt road, etc. In some embodiments, the road includes at least one lane (e.g., a portion of the road that the vehicle 102 can traverse). In an example, the road includes at least one lane associated with (e.g., identified based on) at least one lane marking line.

[0035] The vehicle-to-infrastructure (V2I) device 110 (sometimes referred to as a vehicle-to-infrastructure or vehicle-to-everything (V2X) device) includes at least one device configured to communicate with the vehicle 102 and / or the V2I system 118. In some embodiments, the V2I device 110 is configured to communicate with the vehicle 102, the remote AV system 114, the queue management system 116, and / or the V2I system 118 via the network 112. In some embodiments, the V2I device 110 includes a radio frequency identification (RFID) device, a sign, a camera (e.g., a two-dimensional (2D) and / or three-dimensional (3D) camera), lane markings, street lights, parking meters, etc. In some embodiments, the V2I device 110 is configured to communicate directly with the vehicle 102. Additionally or alternatively, in some embodiments, the V2I device 110 is configured to communicate with the vehicle 102, the remote AV system 114, and / or the fleet management system 116 via the V2I system 118. In some embodiments, the V2I device 110 is configured to communicate with the V2I system 118 via the network 112.

[0036] The network 112 includes one or more wired and / or wireless networks. In an example, the network 112 includes a cellular network (e.g., a long-term evolution (LTE) network, a third generation (3G) network, a fourth generation (4G) network, a fifth generation (5G) network, a code division multiple access (CDMA) network, etc.), a public land mobile network (PLMN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), a telephone network (e.g., a public switched telephone network (PSTN)), a private network, an ad hoc network, an intranet, the Internet, a fiber-based network, a cloud computing network, etc., and / or a combination of some or all of these networks, etc.

[0037] The remote AV system 114 includes at least one device configured to communicate with the vehicle 102, the V2I device 110, the network 112, the fleet management system 116, and / or the V2I system 118 via the network 112. In an example, the remote AV system 114 includes a server, a server group, and / or other similar devices. In some embodiments, the remote AV system 114 is co-located with the fleet management system 116. In some embodiments, the remote AV system 114 participates in the installation of some or all of the components of the vehicle (including autonomous systems, autonomous vehicle computing, and / or software implemented by autonomous vehicle computing, etc.). In some embodiments, the remote AV system 114 maintains (e.g., updates and / or replaces) these components and / or software during the life of the vehicle.

[0038] The queue management system 116 includes at least one device configured to communicate with the vehicles 102, the V2I devices 110, the remote AV system 114, and / or the V2I system 118. In an example, the queue management system 116 includes a server, a server group, and / or other similar devices. In some embodiments, the queue management system 116 is associated with a ridesharing company (e.g., an organization for controlling the operation of multiple vehicles (e.g., vehicles including autonomous systems and / or vehicles not including autonomous systems)).

[0039] In some embodiments, the V2I system 118 includes at least one device configured to communicate with the vehicle 102, the V2I device 110, the remote AV system 114, and / or the fleet management system 116 via the network 112. In some examples, the V2I system 118 is configured to communicate with the V2I device 110 via a connection other than the network 112. In some embodiments, the V2I system 118 includes a server, a server group, and / or other similar devices. In some embodiments, the V2I system 118 is associated with a municipality or a private agency (e.g., a private agency for maintaining the V2I device 110, etc.).

[0040] In some embodiments, the apparatus 300 is configured to perform the following Figure 8 Software instructions for one or more steps of the disclosed method are shown.

[0041] supply Figure 1 The number and arrangement of elements illustrated are examples. Figure 1 There may be additional elements, fewer elements, different elements, and / or differently arranged elements than those illustrated. Additionally or alternatively, at least one element of environment 100 may be described as being Figure 1 Additionally or alternatively, at least one set of elements of environment 100 may perform one or more functions described as being performed by at least one different set of elements of environment 100.

[0042] Reference now Figure 2 , vehicle 200 (which is Figure 1 102) includes, or is associated with, autonomous system 202, powertrain control system 204, steering control system 206, and braking system 208. In some embodiments, vehicle 200 is similar to vehicle 102 (see Figure 1) is the same or similar. In some embodiments, the autonomous system 202 is configured to give the vehicle 200 autonomous driving capabilities (e.g., implementing at least one of the following driving automatic or maneuver-based functions, features and / or devices, etc., which enable the vehicle 200 to operate partially or completely without human intervention, including but not limited to fully autonomous vehicles (e.g., vehicles that abandon reliance on human intervention, such as Level 5 ADS-operated vehicles, etc.), highly autonomous vehicles (e.g., vehicles that abandon reliance on human intervention in certain situations, such as Level 4 ADS-operated vehicles, etc.), and / or conditionally autonomous vehicles (e.g., vehicles that abandon reliance on human intervention in limited situations, such as Level 3 ADS-operated vehicles, etc.), etc.). In one embodiment, the autonomous system 202 includes the operational or tactical functions required to enable the vehicle 200 to operate in road traffic and continuously perform part or all of a dynamic driving task (DDT). In another embodiment, the autonomous system 202 includes an advanced driver assistance system (ADAS) including driver support features. The autonomous system 202 supports various levels of driving automation ranging from no driving automation (e.g., level 0) to full driving automation (e.g., level 5). For a detailed description of fully autonomous vehicles and highly autonomous vehicles, reference may be made to SAE International's standard J3016: Taxonomy and Definitions for Terms Related to On-Road Motor Vehicle Automated Driving Systems, the entire contents of which are incorporated by reference. In some embodiments, the vehicle 200 is associated with an autonomous queue manager and / or a ridesharing company.

[0043] Autonomous system 202 includes a sensor suite that includes one or more devices such as camera 202a, LiDAR sensor 202b, Radar sensor 202c, and microphone 202d. In some embodiments, autonomous system 202 may include more or fewer devices and / or different devices (e.g., ultrasonic sensors, inertial sensors, GPS receivers (discussed below), and / or odometer sensors for generating data associated with an indication of the distance that vehicle 200 has traveled, etc.). In some embodiments, autonomous system 202 uses one or more devices included in autonomous system 202 to generate data associated with environment 100 described herein. Data generated by one or more devices of autonomous system 202 can be used by one or more systems described herein to observe the environment (e.g., environment 100) in which vehicle 200 is located. In some embodiments, autonomous system 202 includes communication device 202e, autonomous vehicle computing 202f, drive-by-wire (DBW) system 202h, and safety controller 202g.

[0044] The camera 202a includes a communication device 202e, an autonomous vehicle computer 202f, and / or a safety controller 202g configured to communicate with the communication device 202e via a bus (e.g., Figure 3 The camera 202a includes at least one device for communicating with the autonomous vehicle computing 202f (e.g., an image processing unit 202a, a bus ... Figure 1In some embodiments, the autonomous vehicle computing 202f determines a depth to one or more objects in a field of view of at least two of the plurality of cameras based on image data from the at least two cameras. In some embodiments, the camera 202a is configured to capture images of objects within a distance relative to the camera 202a (e.g., up to 100 meters and / or up to 1 kilometer, etc.). Thus, the camera 202a includes features such as sensors and lenses that are optimized for sensing objects at one or more distances relative to the camera 202a.

[0045] In an embodiment, the camera 202a includes at least one camera configured to capture one or more images associated with one or more traffic lights, street signs, and / or other physical objects that provide visual navigation information. In some embodiments, the camera 202a generates traffic light data associated with the one or more images. In some examples, the camera 202a generates TLD (traffic light detection) data associated with one or more images including a format (e.g., RAW, JPEG, and / or PNG, etc.). In some embodiments, the camera 202a that generates TLD data differs from other systems incorporating cameras described herein in that the camera 202a may include one or more cameras with a wide field of view (e.g., a wide-angle lens, a fisheye lens, and / or a lens with a viewing angle of approximately 120 degrees or greater, etc.) to generate images related to as many physical objects as possible.

[0046] The LiDAR sensor 202b includes a communication device 202e, an autonomous vehicle computing device 202f, and / or a safety controller 202g via a bus (e.g., Figure 3The LiDAR sensor 202b includes at least one device that communicates with a bus (the same or similar bus as the bus 302 of the embodiment of the present invention). The LiDAR sensor 202b includes a system configured to emit light from a light emitter (e.g., a laser emitter). The light emitted by the LiDAR sensor 202b includes light outside the visible spectrum (e.g., infrared light, etc.). In some embodiments, during operation, the light emitted by the LiDAR sensor 202b encounters a physical object (e.g., a vehicle) and is reflected back to the LiDAR sensor 202b. In some embodiments, the light emitted by the LiDAR sensor 202b does not penetrate the physical object encountered by the light. The LiDAR sensor 202b also includes at least one light detector that detects the light emitted from the light emitter after encountering the physical object. In some embodiments, at least one data processing system associated with the LiDAR sensor 202b generates an image (e.g., a point cloud and / or a combined point cloud, etc.) representing objects included in the field of view of the LiDAR sensor 202b. In some examples, at least one data processing system associated with the LiDAR sensor 202b generates an image representing the boundaries of the physical object and / or the surface of the physical object (e.g., the topology of the surface), etc. In such examples, the image is used to determine the boundaries of the physical object in the field of view of the LiDAR sensor 202b.

[0047] The radio detection and ranging (Radar) sensor 202c includes a sensor configured to communicate with the communication device 202e, the autonomous vehicle computing 202f, and / or the safety controller 202g via a bus (e.g., Figure 3 At least one device for communicating with a bus (same or similar bus as bus 302 of the embodiment of the present invention). Radar sensor 202c includes a system configured to transmit (pulsed or continuous) radio waves. The radio waves transmitted by Radar sensor 202c include radio waves within a predetermined spectrum. In some embodiments, during operation, the radio waves transmitted by Radar sensor 202c encounter physical objects and are reflected back to Radar sensor 202c. In some embodiments, the radio waves transmitted by Radar sensor 202c are not reflected by some objects. In some embodiments, at least one data processing system associated with Radar sensor 202c generates a signal representing an object included in the field of view of Radar sensor 202c. For example, at least one data processing system associated with Radar sensor 202c generates an image representing the boundary of a physical object and / or the surface of a physical object (e.g., the topology of the surface), etc. In some examples, the image is used to determine the boundary of a physical object in the field of view of Radar sensor 202c.

[0048] The microphone 202d includes a microphone configured to communicate with the communication device 202e, the autonomous vehicle computing device 202f, and / or the safety controller 202g via a bus (e.g., Figure 3 At least one device for communicating with the vehicle 200 (the same or similar bus as bus 302 of FIG. 1 ). Microphone 202d includes one or more microphones (e.g., an array microphone and / or an external microphone, etc.) that capture an audio signal and generate data associated with (e.g., representing) the audio signal. In some examples, microphone 202d includes a transducer device and / or the like. In some embodiments, one or more systems described herein can receive the data generated by microphone 202d and determine the position (e.g., distance, etc.) of an object relative to vehicle 200 based on an audio signal associated with the data.

[0049] The communication device 202e includes at least one device configured to communicate with the camera 202a, the LiDAR sensor 202b, the Radar sensor 202c, the microphone 202d, the autonomous vehicle computing 202f, the safety controller 202g, and / or the drive-by-wire (DBW) system 202h. For example, the communication device 202e may include at least one device configured to communicate with the camera 202a, the LiDAR sensor 202b, the Radar sensor 202c, the microphone 202d, the autonomous vehicle computing 202f, the safety controller 202g, and / or the drive-by-wire (DBW) system 202h. Figure 3 The communication device 202e may be a device that is the same as or similar to the communication interface 314 of the vehicle. In some embodiments, the communication device 202e includes a vehicle-to-vehicle (V2V) communication device (eg, a device for enabling wireless communication of data between vehicles).

[0050] Autonomous vehicle computing 202f includes at least one device configured to communicate with camera 202a, LiDAR sensor 202b, Radar sensor 202c, microphone 202d, communication device 202e, safety controller 202g, and / or DBW system 202h. In some examples, autonomous vehicle computing 202f includes devices such as client devices, mobile devices (e.g., cellular phones and / or tablet computers, etc.), and / or servers (e.g., computing devices including one or more central processing units and / or graphics processing units, etc.). In some embodiments, autonomous vehicle computing 202f is the same or similar to autonomous vehicle computing 400 described herein. Additionally or alternatively, in some embodiments, autonomous vehicle computing 202f is configured to communicate with an autonomous vehicle system (e.g., with Figure 1 remote AV system 114 of the same or similar autonomous vehicle system), a fleet management system (e.g., Figure 1 of the same or similar queue management system as the queue management system 116), V2I devices (e.g., Figure 1 V2I device 110 that is the same as or similar to V2I device 110) and / or V2I system (e.g., Figure 1The V2I system 118 may communicate with the same or similar V2I system.

[0051] Safety controller 202g includes at least one device configured to communicate with camera 202a, LiDAR sensor 202b, Radar sensor 202c, microphone 202d, communication device 202e, autonomous vehicle computing 202f, and / or DBW system 202h. In some examples, safety controller 202g includes one or more controllers (electrical controllers and / or electromechanical controllers, etc.) configured to generate and / or transmit control signals to operate one or more devices of vehicle 200 (e.g., powertrain control system 204, steering control system 206, and / or braking system 208, etc.). In some embodiments, safety controller 202g is configured to generate control signals that take precedence over (e.g., override) control signals generated and / or transmitted by autonomous vehicle computing 202f.

[0052] The DBW system 202h includes at least one device configured to communicate with the communication device 202e and / or the autonomous vehicle computing 202f. In some examples, the DBW system 202h includes one or more controllers (e.g., electrical controllers and / or electromechanical controllers, etc.) configured to generate and / or transmit control signals to operate one or more devices of the vehicle 200 (e.g., powertrain control system 204, steering control system 206, and / or braking system 208, etc.). Additionally or alternatively, one or more controllers of the DBW system 202h are configured to generate and / or transmit control signals to operate at least one different device of the vehicle 200 (e.g., turn signals, headlights, door locks, and / or windshield wipers, etc.).

[0053] The powertrain control system 204 includes at least one device configured to communicate with the DBW system 202h. In some examples, the powertrain control system 204 includes at least one controller and / or actuator, etc. In some embodiments, the powertrain control system 204 receives control signals from the DBW system 202h, and the powertrain control system 204 causes the vehicle 200 to perform longitudinal vehicle motion such as starting forward movement, stopping forward movement, starting backward movement, stopping backward movement, accelerating in a certain direction, decelerating in a certain direction, etc., or lateral vehicle motion such as making a left turn and / or making a right turn. In an example, the powertrain control system 204 increases, maintains the same, or decreases the energy (e.g., fuel and / or electricity, etc.) provided to the motor of the vehicle, thereby rotating or not rotating at least one wheel of the vehicle 200. In other words, the steering control system 206 causes the activities required to adjust the y-axis component of the vehicle motion.

[0054] The steering control system 206 includes at least one device configured to rotate one or more wheels of the vehicle 200. In some examples, the steering control system 206 includes at least one controller and / or actuator, etc. In some embodiments, the steering control system 206 rotates the two front wheels and / or the two rear wheels of the vehicle 200 to the left or right to turn the vehicle 200 left or right.

[0055] Braking system 208 includes at least one device configured to actuate one or more brakes to decelerate and / or hold vehicle 200 stationary. In some examples, braking system 208 includes at least one controller and / or actuator configured to cause one or more calipers associated with one or more wheels of vehicle 200 to close on respective rotors of vehicle 200. Additionally or alternatively, in some examples, braking system 208 includes an automatic emergency braking (AEB) system and / or a regenerative braking system, etc.

[0056] In some embodiments, vehicle 200 includes at least one platform sensor (not explicitly illustrated) for measuring or inferring a property of a state or condition of vehicle 200. In some examples, vehicle 200 includes platform sensors such as a global positioning system (GPS) receiver, an inertial measurement unit (IMU), wheel rate sensors, wheel brake pressure sensors, wheel torque sensors, engine torque sensors, and / or steering angle sensors. Although in Figure 2 Braking system 208 is shown located on the proximal side of vehicle 200 , but braking system 208 may be located anywhere in vehicle 200 .

[0057] Reference now Figure 3, a schematic diagram of an exemplary device 300. As illustrated, the device 300 includes a processor 304, a memory 306, a storage component 308, an input interface 310, an output interface 312, a communication interface 314, and a bus 302. In some embodiments, the device 300 corresponds to: at least one device of the vehicle 102 (e.g., at least one device of a system of the vehicle 102); at least one device of the remote AV system 114, the fleet management system 116, the V2I system 118; and / or one or more devices of the network 112 (e.g., one or more devices of a system of the network 112). In some embodiments, one or more devices of vehicle 102 (e.g., one or more devices of a system of vehicle 102, such as remote AV system 114, fleet management system 116, at least one device of V2I system 118, etc.), and / or one or more devices of network 112 (e.g., one or more devices of a system of network 112) include at least one device 300 and / or at least one component of device 300. Figure 3 As shown, apparatus 300 includes a bus 302 , a processor 304 , a memory 306 , a storage component 308 , an input interface 310 , an output interface 312 , and a communication interface 314 .

[0058] The bus 302 includes components that permit communication between components of the device 300. In some cases, the processor 304 includes a processor (e.g., a central processing unit (CPU), a graphics processing unit (GPU), and / or an accelerated processing unit (APU), etc.), a microphone, a digital signal processor (DSP), and / or any processing component that can be programmed to perform at least one function (e.g., a field programmable gate array (FPGA) and / or an application specific integrated circuit (ASIC), etc.). The memory 306 includes a random access memory (RAM), a read-only memory (ROM), and / or another type of dynamic and / or static storage device (e.g., flash memory, magnetic memory, and / or optical memory, etc.) that stores data and / or instructions for use by the processor 304.

[0059] Storage component 308 stores data and / or software related to the operation and use of device 300. In some examples, storage component 308 includes a hard disk (e.g., a magnetic disk, an optical disk, a magneto-optical disk, and / or a solid-state disk, etc.), a compact disk (CD), a digital versatile disk (DVD), a floppy disk, a cassette, a tape, a CD-ROM, a RAM, a PROM, an EPROM, a FLASH-EPROM, an NV-RAM, and / or another type of computer-readable medium, and a corresponding drive.

[0060] The input interface 310 includes components that permit the device 300 to receive information, such as via user input (e.g., a touch screen display, a keyboard, a keypad, a mouse, a button, a switch, a microphone, and / or a camera, etc.). Additionally or alternatively, in some embodiments, the input interface 310 includes a sensor for sensing information (e.g., a global positioning system (GPS) receiver, an accelerometer, a gyroscope, and / or an actuator, etc.). The output interface 312 includes components for providing output information from the device 300 (e.g., a display, a speaker, and / or one or more light emitting diodes (LEDs), etc.).

[0061] In some embodiments, communication interface 314 includes a transceiver-like component (e.g., a transceiver and / or a separate receiver and transmitter, etc.) that permits device 300 to communicate with other devices via a wired connection, a wireless connection, or a combination of a wired connection and a wireless connection. In some examples, communication interface 314 permits device 300 to receive information from another device and / or provide information to another device. In some examples, communication interface 314 includes an Ethernet interface, an optical interface, a coaxial interface, an infrared interface, a radio frequency (RF) interface, a universal serial bus (USB) interface, a Wi-Fi interface, a wireless ... interface and / or cellular network interface, etc.

[0062] In some embodiments, the device 300 performs one or more processes described herein. The device 300 performs these processes based on the processor 304 executing software instructions stored by a computer-readable medium such as a memory 306 and / or a storage component 308. Computer-readable media (e.g., non-transitory computer-readable media) are defined herein as non-transitory memory devices. Non-transitory memory devices include storage space located within a single physical storage device or storage space distributed across multiple physical storage devices.

[0063] In some embodiments, the software instructions are read into the memory 306 and / or storage component 308 from another computer-readable medium or from another device via the communication interface 314. The software instructions stored in the memory 306 and / or storage component 308, when executed, cause the processor 304 to perform one or more processes described herein. Additionally or alternatively, hardwired circuitry is used in place of or in combination with the software instructions to perform one or more processes described herein. Therefore, unless expressly stated otherwise, the embodiments described herein are not limited to any specific combination of hardware circuitry and software.

[0064] The memory 306 and / or the storage component 308 include a data storage unit or at least one data structure (e.g., a database, etc.). The device 300 can receive information from the data storage unit or at least one data structure in the memory 306 or the storage component 308, store information in the data storage unit or at least one data structure, communicate information to the data storage unit or at least one data structure, or search for information stored in the data storage unit or at least one data structure. In some examples, the information includes network data, input data, output data, or any combination thereof.

[0065] In some embodiments, the device 300 is configured to execute software instructions stored in the memory 306 and / or the memory of another device (e.g., another device that is the same as or similar to the device 300). As used herein, the term "module" refers to at least one instruction stored in the memory 306 and / or the memory of another device, which, when executed by the processor 304 and / or the processor of another device (e.g., another device that is the same as or similar to the device 300), causes the device 300 (e.g., at least one component of the device 300) to perform one or more processes described herein. In some embodiments, the module is implemented in software, firmware, and / or hardware, etc.

[0066] supply Figure 3 The number and arrangement of components illustrated are examples. Figure 3 The device 300 may include additional components, fewer components, different components, or differently arranged components than those illustrated. Additionally or alternatively, a set of components (e.g., one or more components) of the device 300 may perform one or more functions described as being performed by another component or set of components of the device 300.

[0067] Reference now Figure 4A, illustrates an example block diagram of an autonomous vehicle computing 400 (sometimes referred to as an "AV stack"). As illustrated, the autonomous vehicle computing 400 includes a perception system 402 (sometimes referred to as a perception module), a planning system 404 (sometimes referred to as a planning module), a positioning system 406 (sometimes referred to as a positioning module), a control system 408 (sometimes referred to as a control module), and a database 410. In some embodiments, the perception system 402, the planning system 404, the positioning system 406, the control system 408, and the database 410 are included in and / or implemented in an automatic navigation system of a vehicle (e.g., the autonomous vehicle computing 202f of the vehicle 200). Additionally or alternatively, in some embodiments, the perception system 402, the planning system 404, the positioning system 406, the control system 408, and the database 410 are included in one or more independent systems (e.g., one or more systems that are the same or similar to the autonomous vehicle computing 400, etc.). In some examples, the perception system 402, planning system 404, positioning system 406, control system 408, and database 410 are included in one or more independent systems located in the vehicle and / or at least one remote system as described herein. In some embodiments, any and / or all of the systems included in the autonomous vehicle computing 400 are implemented in software (e.g., software instructions stored in a memory), computer hardware (e.g., by a microprocessor, microcontroller, application specific integrated circuit (ASIC) and / or field programmable gate array (FPGA), etc.), or a combination of computer software and computer hardware. It will also be understood that in some embodiments, the autonomous vehicle computing 400 is configured to communicate with a remote system (e.g., an autonomous vehicle system that is the same or similar to the remote AV system 114, a fleet management system 116 that is the same or similar to the fleet management system 116, and / or a V2I system that is the same or similar to the V2I system 118, etc.).

[0068] In some embodiments, the perception system 402 receives data associated with at least one physical object in the environment (e.g., data used by the perception system 402 to detect at least one physical object) and classifies the at least one physical object. In some examples, the perception system 402 receives image data captured by at least one camera (e.g., camera 202a), the image being associated with (e.g., representing) one or more physical objects within the field of view of the at least one camera. In such examples, the perception system 402 classifies the at least one physical object based on one or more groups of physical objects (e.g., bicycles, vehicles, traffic signs, and / or pedestrians, etc.). In some embodiments, based on the classification of the physical object by the perception system 402, the perception system 402 transmits data associated with the classification of the physical object to the planning system 404.

[0069] In some embodiments, the planning system 404 receives data associated with a destination and generates data associated with at least one route (e.g., route 106) along which a vehicle (e.g., vehicle 102) can travel toward the destination. In some embodiments, the planning system 404 periodically or continuously receives data (e.g., data associated with the classification of physical objects described above) from the perception system 402, and the planning system 404 updates at least one trajectory or generates at least one different trajectory based on the data generated by the perception system 402. In other words, the planning system 404 can perform tasks related to the tactical functions required to operate the vehicle 101 in traffic on the road. Tactical efforts involve maneuvering the vehicle in traffic during the journey, which includes but is not limited to deciding whether and when to overtake another vehicle, change lanes, or select an appropriate rate, acceleration, deceleration, etc. In some embodiments, the planning system 404 receives data associated with an updated position of the vehicle (e.g., vehicle 102) from the positioning system 406, and the planning system 404 updates at least one trajectory or generates at least one different trajectory based on the data generated by the positioning system 406.

[0070] In some embodiments, the positioning system 406 receives data associated with (e.g., representing) a location of a vehicle (e.g., vehicle 102) in an area. In some examples, the positioning system 406 receives LiDAR data associated with at least one point cloud generated by at least one LiDAR sensor (e.g., LiDAR sensor 202b). In some examples, the positioning system 406 receives data associated with at least one point cloud from multiple LiDAR sensors, and the positioning system 406 generates a combined point cloud based on each point cloud. In these examples, the positioning system 406 compares the at least one point cloud or the combined point cloud with a two-dimensional (2D) and / or three-dimensional (3D) map of the area stored in the database 410. Then, based on the positioning system 406 comparing the at least one point cloud or the combined point cloud with the map, the positioning system 406 determines the position of the vehicle in the area. In some embodiments, the map includes a combined point cloud of the area generated before the navigation of the vehicle. In some embodiments, the map includes, but is not limited to, a high-precision map of roadway geometry, a map describing the road network connectivity, a map describing the physical properties of the roadway (such as traffic speed, traffic volume, the number of vehicle and bicycle traffic lanes, lane width, lane traffic direction or the type and location of lane markings, or a combination thereof, etc.), and a map describing the spatial location of road features (such as crosswalks, traffic signs, or various types of other driving lights, etc.). In some embodiments, the map is generated in real time based on data received by the perception system.

[0071] In another example, positioning system 406 receives global navigation satellite system (GNSS) data generated by a global positioning system (GPS) receiver. In some examples, positioning system 406 receives GNSS data associated with a location of a vehicle in an area, and positioning system 406 determines the latitude and longitude of the vehicle in the area. In such an example, positioning system 406 determines the position of the vehicle in the area based on the latitude and longitude of the vehicle. In some embodiments, positioning system 406 generates data associated with the position of the vehicle. In some examples, based on the location of the vehicle determined by positioning system 406, positioning system 406 generates data associated with the position of the vehicle. In such an example, the data associated with the position of the vehicle include data associated with one or more semantic properties corresponding to the position of the vehicle.

[0072] In some embodiments, the control system 408 receives data associated with at least one trajectory from the planning system 404, and the control system 408 controls the operation of the vehicle. In some examples, the control system 408 receives data associated with at least one trajectory from the planning system 404, and the control system 408 controls the operation of the vehicle by generating and transmitting control signals to operate the powertrain control system (e.g., DBW system 202h and / or powertrain control system 204, etc.), the steering control system (e.g., steering control system 206) and / or the braking system (e.g., braking system 208). For example, the control system 408 is configured to perform operational functions such as lateral vehicle motion control or longitudinal vehicle motion control. Lateral vehicle motion control causes the activity required to adjust the y-axis component of the vehicle motion. Longitudinal vehicle motion control causes the activity required to adjust the x-axis component of the vehicle motion. In the example, in the case where the trajectory includes a left turn, the control system 408 transmits a control signal to cause the steering control system 206 to adjust the steering angle of the vehicle 200, thereby turning the vehicle 200 left. Additionally or alternatively, the control system 408 generates and transmits control signals to cause other devices of the vehicle 200 (eg, headlights, turn signals, door locks, and / or windshield wipers, etc.) to change states.

[0073] In some embodiments, perception system 402, planning system 404, positioning system 406, and / or control system 408 implement at least one machine learning model (e.g., at least one multi-layer perceptron (MLP), at least one convolutional neural network (CNN), at least one recurrent neural network (RNN), at least one autoencoder, and / or at least one transformer, etc.). In some examples, perception system 402, planning system 404, positioning system 406, and / or control system 408 implement at least one machine learning model alone or in combination with one or more of the above systems. In some examples, perception system 402, planning system 404, positioning system 406, and / or control system 408 implement at least one machine learning model as part of a pipeline (e.g., a pipeline for identifying one or more objects located in an environment, etc.). The following is about FIG. 4B to FIG. 4D Includes examples of implementations of machine learning models.

[0074] Database 410 stores data transmitted to, received from, and / or updated by perception system 402, planning system 404, positioning system 406, and / or control system 408. In some examples, database 410 includes a storage component (e.g., a storage component) for storing data and / or software related to operations and using at least one system of autonomous vehicle computing 400. Figure 3In some embodiments, database 410 stores data associated with a 2D and / or 3D map of at least one area. In some examples, database 410 stores data associated with a 2D and / or 3D map of a portion of a city, portions of multiple cities, multiple cities, counties, states, and / or countries (State) (e.g., a country), etc. In such an example, a vehicle (e.g., a vehicle that is the same or similar to vehicle 102 and / or vehicle 200) can be driven along one or more drivable areas (e.g., single-lane roads, multi-lane roads, highways, remote roads, and / or off-road roads, etc.) and cause at least one LiDAR sensor (e.g., a LiDAR sensor that is the same or similar to LiDAR sensor 202b) to generate data associated with an image representing an object included in the field of view of the at least one LiDAR sensor.

[0075] In some embodiments, database 410 can be implemented across multiple devices. In some examples, database 410 includes a vehicle (e.g., a vehicle that is the same or similar to vehicle 102 and / or vehicle 200), an autonomous vehicle system (e.g., an autonomous vehicle system that is the same or similar to remote AV system 114), a fleet management system (e.g., a vehicle that is the same or similar to remote AV system 114), and a fleet management system (e.g., a vehicle that is the same or similar to remote AV system 114). Figure 1 The same or similar queue management system as the queue management system 116 of FIG. 1 and / or the V2I system (e.g., Figure 1 The V2I system 118 is the same as or similar to the V2I system 118).

[0076] Reference now Figure 4B , a diagram illustrating an implementation of a machine learning model. More specifically, a diagram illustrating an implementation of a convolutional neural network (CNN) 420. For purposes of illustration, the following description of CNN 420 will be with respect to implementing CNN 420 by perception system 402. However, it will be understood that in some examples, CNN 420 (e.g., one or more components of CNN 420) is implemented by other systems (such as planning system 404, positioning system 406, and / or control system 408, etc.) other than or in addition to perception system 402. Although CNN 420 includes certain features as described herein, these features are provided for purposes of illustration and are not intended to limit the present disclosure.

[0077] CNN 420 includes a plurality of convolutional layers including a first convolutional layer 422, a second convolutional layer 424, and a convolutional layer 426. In some embodiments, CNN 420 includes a subsampling layer 428 (sometimes referred to as a pooling layer). In some embodiments, subsampling layer 428 and / or other subsampling layers have a dimension that is smaller than the dimension of the upstream system (i.e., the number of nodes). With the subsampling layer 428 having a dimension that is smaller than the dimension of the upstream layer, CNN 420 merges the amount of data associated with the initial input and / or output of the upstream layer, thereby reducing the amount of computation required for CNN 420 to perform downstream convolution operations. Additionally or alternatively, with the subsampling layer 428 being associated with (e.g., configured to perform) at least one subsampling function (as described below with respect to Figure 4C and Figure 4D As described above, CNN 420 incorporates the amount of data associated with the initial input.

[0078] The perception system 402 performs the convolution operation based on the perception system 402 providing respective inputs and / or outputs associated with each of the first convolution layer 422, the second convolution layer 424, and the convolution layer 426 to generate respective outputs. In some examples, the perception system 402 implements the CNN 420 based on the perception system 402 providing data as input to the first convolution layer 422, the second convolution layer 424, and the convolution layer 426. In such examples, the perception system 402 provides data as input to the first convolution layer 422, the second convolution layer 424, and the convolution layer 426 based on the perception system 402 receiving data from one or more different systems (e.g., one or more systems of a vehicle that is the same or similar to the vehicle 102, a remote AV system that is the same or similar to the remote AV system 114, a queue management system that is the same or similar to the queue management system 116, and / or a V2I system that is the same or similar to the V2I system 118, etc.). The following is about Figure 4C Includes a detailed description of the convolution operation.

[0079] In some embodiments, the perception system 402 provides data associated with the input (referred to as the initial input) to the first convolutional layer 422, and the perception system 402 generates data associated with the output using the first convolutional layer 422. In some embodiments, the perception system 402 provides the output generated by the convolutional layer as input to a different convolutional layer. For example, the perception system 402 provides the output of the first convolutional layer 422 as input to the subsampling layer 428, the second convolutional layer 424, and / or the convolutional layer 426. In such an example, the first convolutional layer 422 is referred to as an upstream layer, and the subsampling layer 428, the second convolutional layer 424, and / or the convolutional layer 426 are referred to as downstream layers. Similarly, in some embodiments, the perception system 402 provides the output of the subsampling layer 428 to the second convolutional layer 424 and / or the convolutional layer 426, and in this example, the subsampling layer 428 will be referred to as the upstream layer, and the second convolutional layer 424 and / or the convolutional layer 426 will be referred to as the downstream layer.

[0080] In some embodiments, before the perception system 402 provides the input to the CNN 420, the perception system 402 processes the data associated with the input provided to the CNN 420. For example, the perception system 402 processes the data associated with the input provided to the CNN 420 based on the perception system 402 normalizing the sensor data (e.g., image data, LiDAR data, and / or Radar data, etc.).

[0081] In some embodiments, CNN 420 generates an output based on perception system 402 performing convolution operations associated with each convolution layer. In some examples, CNN 420 generates an output based on perception system 402 performing convolution operations associated with each convolution layer and the initial input. In some embodiments, perception system 402 generates an output and provides the output to fully connected layer 430. In some examples, perception system 402 provides the output of convolution layer 426 to fully connected layer 430, wherein fully connected layer 430 includes data associated with multiple feature values ​​referred to as F1, F2, ..., FN. In this example, the output of convolution layer 426 includes data associated with multiple output feature values ​​representing predictions.

[0082] In some embodiments, perception system 402 identifies a prediction from the plurality of predictions based on perception system 402 identifying a feature value associated with a highest likelihood of being a correct prediction from the plurality of predictions. For example, where fully connected layer 430 includes feature values ​​F1, F2, ..., FN and F1 is the largest feature value, perception system 402 identifies the prediction associated with F1 as the correct prediction from the plurality of predictions. In some embodiments, perception system 402 trains CNN 420 to generate the predictions. In some examples, perception system 402 trains CNN 420 to generate the predictions based on perception system 402 providing training data associated with the predictions to CNN 420.

[0083] Reference now Figure 4C and Figure 4D , a diagram illustrating an example operation of CNN 440 utilizing perception system 402. In some embodiments, CNN 440 (e.g., one or more components of CNN 440) is coupled to CNN 420 (e.g., one or more components of CNN 420) (see Figure 4B ) are the same or similar.

[0084] At step 450, the perception system 402 provides data associated with the image as input to the CNN 440 (step 450). For example, as illustrated, the perception system 402 provides data associated with the image to the CNN 440, where the image is a grayscale image represented as values ​​stored in a two-dimensional (2D) array. In some embodiments, the data associated with the image may include data associated with a color image represented as values ​​stored in a three-dimensional (3D) array. Additionally or alternatively, the data associated with the image may include data associated with an infrared image and / or a Radar image, etc.

[0085] At step 455, CNN 440 performs a first convolution function. For example, CNN 440 performs a first convolution function based on CNN 440 providing a value representing an image as an input to one or more neurons (not explicitly illustrated) included in first convolution layer 442. In this example, the value representing the image may correspond to a value of a region (sometimes referred to as a receptive field) representing the image. In some embodiments, each neuron is associated with a filter (not explicitly illustrated). The filter (sometimes referred to as a kernel) may be represented as an array of values ​​corresponding in size to the value provided as input to the neuron. In one example, the filter may be configured to identify edges (e.g., horizontal lines, vertical lines, and / or straight lines, etc.). In successive convolution layers, the filters associated with the neurons may be configured to continuously identify more complex patterns (e.g., arcs and / or objects, etc.).

[0086] In some embodiments, CNN 440 performs a first convolution function based on CNN 440 multiplying the values ​​of each neuron provided as input to one or more neurons included in the first convolution layer 442 by the values ​​of the filters corresponding to each neuron in the same or more neurons. For example, CNN 440 may multiply the values ​​of each neuron provided as input to one or more neurons included in the first convolution layer 442 by the values ​​of the filters corresponding to each neuron in the one or more neurons to generate a single value or an array of values ​​as output. In some embodiments, the collective output of the neurons of the first convolution layer 442 is referred to as a convolution output. In some embodiments, when each neuron has the same filter, the convolution output is referred to as a feature map.

[0087] In some embodiments, CNN 440 provides the output of each neuron of the first convolutional layer 442 to the neurons of the downstream layer. For clarity, the upstream layer may be a layer that transmits data to a different layer (referred to as the downstream layer). For example, CNN 440 may provide the output of each neuron of the first convolutional layer 442 to the corresponding neurons of the subsampling layer. In the example, CNN 440 provides the output of each neuron of the first convolutional layer 442 to the corresponding neurons of the first subsampling layer 444. In some embodiments, CNN 440 adds a bias value to the aggregate set of all values ​​provided to each neuron of the downstream layer. For example, CNN 440 adds a bias value to the aggregate set of all values ​​provided to each neuron of the first subsampling layer 444. In such an example, CNN 440 determines the final value to be provided to each neuron of the first subsampling layer 444 based on the aggregate set of all values ​​provided to each neuron and the activation function associated with each neuron of the first subsampling layer 444.

[0088] At step 460, CNN 440 performs a first subsampling function. For example, based on CNN 440 providing the values ​​output by first convolutional layer 442 to the corresponding neurons of first subsampling layer 444, CNN 440 may perform the first subsampling function. In some embodiments, CNN 440 performs the first subsampling function based on an aggregation function. In an example, CNN 440 performs the first subsampling function based on CNN 440 determining the maximum input (referred to as a maximum pooling function) among the values ​​provided to a given neuron. In another example, CNN 440 performs the first subsampling function based on CNN 440 determining the average input (referred to as an average pooling function) among the values ​​provided to a given neuron. In some embodiments, based on CNN 440 providing values ​​to the respective neurons of first subsampling layer 444, CNN 440 generates an output, which is sometimes referred to as a subsampled convolution output.

[0089] At step 465, CNN 440 performs a second convolution function. In some embodiments, CNN 440 performs the second convolution function in a manner similar to how CNN 440 performs the first convolution function described above. In some embodiments, CNN 440 performs the second convolution function based on CNN 440 providing the value output by first subsampling layer 444 as input to one or more neurons (not explicitly illustrated) included in second convolution layer 446. In some embodiments, as described above, each neuron of second convolution layer 446 is associated with a filter. As described above, the filter (one or more) associated with second convolution layer 446 can be configured to recognize more complex patterns than the filter associated with first convolution layer 442.

[0090] In some embodiments, the CNN 440 performs a second convolution function based on the CNN 440 multiplying the value of each neuron provided as input to the one or more neurons included in the second convolution layer 446 by the value of the filter corresponding to each neuron of the one or more neurons. For example, the CNN 440 may multiply the value of each neuron provided as input to the one or more neurons included in the second convolution layer 446 by the value of the filter corresponding to each neuron of the one or more neurons to generate a single value or a value array as an output.

[0091] In some embodiments, the CNN 440 provides the output of each neuron of the second convolutional layer 446 to the neurons of the downstream layer. For example, the CNN 440 may provide the output of each neuron of the first convolutional layer 442 to the corresponding neurons of the subsampling layer. In an example, the CNN 440 provides the output of each neuron of the first convolutional layer 442 to the corresponding neurons of the second subsampling layer 448. In some embodiments, the CNN 440 adds a bias value to the aggregate set of all values ​​provided to each neuron of the downstream layer. For example, the CNN 440 adds a bias value to the aggregate set of all values ​​provided to each neuron of the second subsampling layer 448. In such an example, the CNN 440 determines the final value provided to each neuron of the second subsampling layer 448 based on the aggregate set of all values ​​provided to each neuron and the activation function associated with each neuron of the second subsampling layer 448.

[0092] At step 470, CNN 440 performs a second subsampling function. For example, based on CNN 440 providing the values ​​output by second convolutional layer 446 to corresponding neurons of second subsampling layer 448, CNN 440 may perform a second subsampling function. In some embodiments, based on CNN 440 using an aggregation function, CNN 440 performs a second subsampling function. In an example, as described above, based on CNN 440 determining the maximum input or average input among the values ​​provided to a given neuron, CNN 440 performs a first subsampling function. In some embodiments, based on CNN 440 providing values ​​to respective neurons of second subsampling layer 448, CNN 440 generates an output.

[0093] At step 475, CNN 440 provides the output of each neuron of second subsampling layer 448 to fully connected layer 449. For example, CNN 440 provides the output of each neuron of second subsampling layer 448 to fully connected layer 449 so that fully connected layer 449 generates output 480. In some embodiments, fully connected layer 449 is configured to generate output 480 associated with a prediction (sometimes referred to as a classification). The prediction may include an indication that the objects included in the image provided as input to CNN 440 include objects and / or sets of objects, etc. In some embodiments, perception system 402 performs one or more operations and / or provides data associated with the prediction to various systems described herein.

[0094] The present disclosure relates to systems, methods, and computer program products for providing orientation determination of objects (e.g., static objects, dynamic objects, agents, vehicles) by autonomous vehicles. In particular, the present disclosure can be used for offline purposes (such as for training machine learning models and / or neural networks, etc.), and / or for online purposes (such as for use by autonomous vehicles for real-time object detection and navigation, etc.). The disclosed systems, methods, and computer program products can be integrated at many different points in the labeling pipeline of an autonomous vehicle.

[0095] Reference now Figure 5 , a diagram showing a system 500 for object orientation determination. In some embodiments, the system 500 is coupled to a vehicle (e.g., Figure 1 The vehicle 102 or Figure 2 200) is connected and / or incorporated into the vehicle. In one or more embodiments or examples, for an AV (e.g., such as Figure 2 The autonomous system 202 shown, Figure 3 300, etc.), AV system, AV computing 540 (such as Figure 2 AV Computing 202F and / or Figure 4AAV computing 400, etc.), remote AV systems (such as Figure 1 remote AV system 114, etc.), queue management systems (such as Figure 1 queue management system 116, etc.) and V2I systems (such as Figure 1 The system 500 may be used to operate an autonomous vehicle. The system 500 may not be used to operate an autonomous vehicle.

[0096] In one or more embodiments or examples, system 500 includes an apparatus (such as Figure 3 device 300, etc.), a positioning system (such as Figure 4A positioning system 406, etc.), planning systems (such as Figure 4A planning system 404, etc.), perception systems (such as Figure 4A sensing system 402, etc.) and control systems (such as Figure 4A one or more of the control systems 408, etc.

[0097] A system 500 is disclosed herein. The system 500 includes at least one processor. The system 500 includes at least one memory storing instructions that, when executed by the at least one processor, cause the at least one processor to perform operations including obtaining a map parameter 502 indicating a predetermined position of a first object in an environment in which the autonomous vehicle is configured to operate. The operation includes obtaining a group parameter 503 indicating a predetermined connection between objects in a group, wherein the environment includes the object. The operation includes determining orientation data 510 indicating an orientation of at least one object among the first object and the objects in the group based on the map parameter 502 and the group parameter 503. The operation includes causing object detection data 512 associated with at least one object to be provided to a device based on the orientation data 510, wherein the object detection data indicates detection of one or more spatial features of the at least one object.

[0098] In one or more examples or embodiments, system 500 includes an object detection system 504 , a tracker 506 , and optionally a non-maximum suppression scheme (NMS) 505 , and optionally a post-processing system 507 .

[0099] In one or more examples or embodiments, the system 500 is configured to use (one or more) map parameters 502 and (one or more) group parameters 503, for example, via an object detection system 504, to determine objects and / or object interactions in an environment. For example, the system 500 uses (one or more) map parameters 502 and (one or more) group parameters 503 as additional clues to correct object detection in terms of the location, orientation, and size of the object. The group parameters and map parameters can be obtained by heuristics and / or data analysis. Because the system 500 can supplement the point cloud used to determine the orientation of the object, it can be advantageous to use (one or more) map parameters 502 and (one or more) group parameters 503 in the case where the sensor data 501 has a sparse point cloud. The system 500 can be used offline, for example, for training a machine learning model to improve recognition and labeling in an environment, and the system 500 can be used online for the operation of an autonomous vehicle.

[0100] Advantageously, in some instances, the disclosed system 500 is particularly useful for monocular vision methods that have high uncertainty in the depth dimension. The system 500 can be used (e.g., online) on an autonomous vehicle, for example, to improve detection performance, e.g., via an object detection system 504, and / or for tracking purposes, e.g., via a tracker 506. The system 500 can be used offline for perception and / or automatic labeling, which can significantly improve detection performance, e.g., for automatic labeling and / or semi-supervised learning, for example, to improve the quality of labeling.

[0101] In one or more examples or embodiments, the system 500 utilizes the map parameters 502 and / or the group parameters 503 to perform the following operations: Figure 5 Different operations of AV computation 540 are shown. For example, system 500 utilizes map parameters 502 and / or group parameters 503 at one or more of object detection system 504, non-maximum suppression scheme (NMS) 505, tracker 506, and post-processing system 507.

[0102] In one or more examples or embodiments, the system 500 is configured to obtain and use map parameters 502 (eg, map priors). The system 500 obtains, for example, from a memory or storage component (such as Figure 3 306 and storage component 308, etc.) and / or from a database (such as Figure 4A The map parameters 502 are obtained from a database 410 of the AV, etc. The map parameters 502 may be stored locally on the AV and / or may be stored on a distributed network such as a cloud network.

[0103] In one or more examples or embodiments, the map parameters 502 indicate a predetermined position of a first object in the environment. In other words, in some examples, the map parameters 502 provide positioning information related to the first object in the environment. The map parameters 502 may indicate assumptions about possible positions and orientations of objects in a particular area, such as the size and orientation of objects in a parking lot. The map parameters may be determined to represent a rectangular vehicle, but may be adjusted to other polygons, such as more complex polygons.

[0104] In one or more examples or embodiments, the environment includes one or more objects, which include a first object and optionally a second object. In one or more examples or embodiments, the second object is different from the first object. In some examples, the second object is considered to be an object in a group. In this article, the terms second object and object in a group can be used interchangeably. In one or more examples or embodiments, the first object is an object with a specific positioning in the environment. For example, the first object can be an environmental feature that can generally guide the position of other objects relative to the first object. Example first objects include agents, vehicles, pedestrians, parking lots, roadsides, stop signs and stop lights. In some examples, parking lots include many parking spaces for parking vehicles. If the vehicle is properly parked, the position of the vehicle in the parking lot will be that the front end or rear end faces outward. Similarly, a vehicle parked on the side of the road can have a specific positioning as follows: the vehicle is generally aligned with the traffic flow (e.g., parallel). Map parameters 502 can indicate such a predetermined position (e.g., a predetermined orientation). In some examples, the map parameters 502 include predetermined positions for a number of different first objects. In some examples, the predetermined positions as indicated by the map parameters 502 include one or more of size, orientation, alignment, horizontal alignment, vertical alignment, etc.

[0105] The environment is the environment in which the autonomous vehicle is configured to operate. Thus, the environment may be larger than the specific area in which the autonomous vehicle is operating and may extend beyond any sensor capabilities of the autonomous vehicle. In one or more examples or embodiments, the environment is a specific area, and the map parameters 502 indicate how objects may be positioned in the specific area. The system 500, for example, filters out areas in which the autonomous vehicle is not configured to operate, such as off-road areas, etc.

[0106] In one or more examples or embodiments, the system 500 is configured to obtain and use group parameters 503 (e.g., group priors). The system 500 retrieves, for example, a group parameter 503 from a memory or storage component such as Figure 3 306 and storage component 308, etc.) and / or from a database (such as Figure 4AThe group parameters 503 are obtained by using a database 410 of the AV, etc. The group parameters 503 may be stored locally on the AV and / or may be stored on a distributed network such as a cloud network. The group parameters may be obtained implicitly through a data-driven approach, or may be obtained from explicit group assignments. For example, explicit group assignment means marking objects as being grouped together through a set of heuristics or some geometric algorithms. For example, by using a typical clustering algorithm deployed with respect to the positions of pedestrians, pedestrians are clustered together and marked as belonging to the same group. Then, in some examples, on top of other attributes of these pedestrians, attributes indicating that these pedestrians are part of the same group / cluster are also fed to the neural network to help refine their detection. As another example, a data-driven approach means that the neural network receives inputs that can help it discover group objects by itself, rather than receiving inputs in which objects are explicitly marked as being part of the same group. For example, the system 500 is configured to calculate pairwise distances between all pedestrians or cars. When these pairwise distances are used as additional inputs to the network, it can implicitly use these pairwise distances to discover object groups by itself without requiring us to explicitly create groups.

[0107] In one or more than one example or embodiment, group parameter 503 indicates a predetermined connection (e.g., relationship) between objects in a group (e.g., a second object, a group of objects, a group of agents, a group of pedestrians, a group of vehicles). For example, group parameter indicates a hypothesis related to the interaction (e.g., relationship) of a group of objects, such as cars parked back to back along a road. The second object (e.g., an object in a group) can be a parked vehicle (e.g., a vehicle parked in a parking lot, a roadside, a lane), and the group can be a group of parked vehicles. In some examples, the orientation of vehicles (such as, a group, etc.) parked along the same road or parked in the same parking lot is associated by group parameter 503 (e.g., the vehicle faces the same direction, the vehicle faces one of the two directions, etc.). Similarly, vehicles that are stopped (e.g., due to traffic lights, traffic jams) can have a correlation indicated by group parameter 503. For example, in a traffic jam, all vehicles stay in their respective lanes and have the same orientation. System 500 can use group parameters 503 to cluster second objects by using implicit or explicit relationships between objects (e.g., second objects) in the group. The predetermined relationship can be a relationship based on position (e.g., interaction). For example, a group is formed based on a predetermined relationship between second objects in the group. In one embodiment, the first object can be the same object as the second object. In another embodiment, the first object can be an object different from the second object. In some examples, group parameters 503 include predetermined relationships for many different second objects and / or many different groups. Group parameters can be applied to static or dynamic vehicles and pedestrians (e.g., a group of pedestrians located at the same intersection has a direction aligned with the direction of the intersection).

[0108] In one or more examples or embodiments, the system 500 obtains different group parameters, some of which may be relevant only to specific map parameters. For example, the group parameter 503 indicates the trailing object (e.g., a trailing car) that has the minimum distance to the front object. The group parameter 503 may also indicate a specific alignment between classes, such as for pedestrians carrying vehicles.

[0109] Based on the map parameters 502 and the group parameters 503, the system 500 determines the orientation data 510. The orientation data 510 indicates the orientation (e.g., position) of at least one of the first object and the second object. For example, the orientation data 510 indicates that at least one object parked on the side of the road has an orientation aligned with the direction in which the vehicle is traveling along the road. Advantageously, the system 500 can determine the orientation data 510 to supplement the sparse point cloud, and can improve the detection and determination of objects in the environment by the autonomous vehicle.

[0110] The system 500 enables the device to provide object detection data 512 associated with at least one object based on the orientation data 510. The object detection data 512 includes, for example, one or more of a position, an orientation, and a size of the at least one object (e.g., a spatial feature). For example, the object detection data 512 is used to supplement the sparse data cloud for object detection, such as via the object detection system 504, and / or for tracking, such as via the tracker 506. In one or more examples or embodiments, the system 500 is configured to determine the object detection data 512 based on the orientation data 510.

[0111] In one or more examples or embodiments, the system 500 is configured for online operation, for example, using sensor data. In some examples, the system 500 uses the sensor data 501 to verify and / or check the determined heading data 510. In one or more examples or embodiments, the operation also includes obtaining sensor data 501 associated with the environment. In one or more examples or embodiments, the operation also includes determining the heading data 510 based on the map parameters 502, the group parameters 503, and the sensor data 501.

[0112] In one or more examples or embodiments, the system 500 can be used, for example, via a sensory system such as Figure 4A The system 500 may use the sensor data 501 to determine the orientation data 510. The sensor data 501 may be one or more of Radar sensor data, image sensor data (e.g., camera sensor data), and LiDAR sensor data. The specific type of the sensor data 501 is not limited. The sensor data 501 may indicate the environment surrounding the autonomous vehicle. For example, the sensor data 501 may indicate an object and / or multiple objects (e.g., a first object, a second object) in the environment in which the vehicle is operating.

[0113] The sensor may be one or more sensors such as onboard sensors. The sensor may be associated with a vehicle. The vehicle may include one or more sensors that may be configured to monitor the environment in which the vehicle is operating, for example, via the sensor through sensor data 501. For example, monitoring provides sensor data 501 indicating what is happening in the environment around the vehicle, for example, to determine orientation data 510. The sensor may be one or more of a Radar sensor, a camera sensor, an infrared sensor, an image sensor, and a LiDAR sensor. The sensor may include Figure 2 One or more of the sensors shown, such as camera 202a, LiDAR sensor 202b, and Radar sensor 202c, etc.

[0114] In one or more examples or embodiments, the system 500 uses the sensor data 501, the map parameters 502, and the group parameters 503 to determine the heading data 510. Since the sensor data 501 may include sparse data such as a sparse point cloud from a LiDAR, the map parameters 502 and the group parameters 503 may be complementary. In addition, the system 500 may utilize the sensor data 501 (such as real-time sensor data) to determine and / or verify the heading data 510. In addition, the sensor data 501 may be used to improve the accuracy of the determination of the heading data 510.

[0115] In one or more examples or embodiments, the system 500 obtains the sensor data 501 for other purposes such as obtaining group parameters 503. In one or more examples or embodiments, obtaining the group parameters 503 includes determining the distance between the second objects (e.g., objects in the group) based on the sensor data 501.

[0116] In one or more examples or embodiments, obtaining group parameters 503 includes clustering the second object based on the distance to form a group. System 500, for example, determines the paired distance between the second object. System 500, for example, determines the paired distance matrix of all pedestrians in the environment. In one or more examples or embodiments, system 500 uses clustering techniques to extract the group so that system 500 does not need to track each second object in the group separately. If the distance between the second objects meets the clustering threshold or is below the clustering threshold, system 500 is configured to cluster the second object to form a group. If the distance between the second objects does not meet the clustering threshold (for example, above the clustering threshold), system 500 is configured not to cluster the second object to form a group. For example, if the autonomous vehicle is located near a crosswalk, there may be many pedestrians crossing the crosswalk. Instead of determining and / or tracking each of the pedestrians separately as a separate object (which may require high computing power and lead to inconsistencies between the results associated with each pedestrian), system 500 is configured to cluster the pedestrians together as a single group.

[0117] In one or more examples or embodiments, the operation further includes discarding the group from the object detection data 512 based on the group parameters 503 and the map parameters 502. In other words, the system 500 can be configured to filter out irrelevant groups and / or objects, such as filtering out those groups and / or objects that are not in relevant map areas, etc. (e.g., objects in parking lots, roadside parking, and before parking lines are relevant map areas). The relevant map areas can be obtained by the system 500. For example, the system 500 determines whether the group meets the detection criteria. The detection criteria can indicate irrelevant areas based on one or more of the sensor data 501, the map parameters 502, and the group parameters 503, for example. For example, in response to determining that the group does not meet the detection criteria, the system 500 is configured to discard the group from the object detection data 512. For example, in response to determining that the group meets the detection criteria, the system 500 is configured not to discard the group from the object detection data 512. This can be beneficial to reduce the processing of groups that are not relevant to the operation of the autonomous vehicle.

[0118] In one or more examples or embodiments, determining the orientation data 510 based on the map parameters 502 and the group parameters 503 includes extracting one or more line patterns associated with the first object and / or the second object based on the sensor data 501 and the group parameters 503. In one or more examples or embodiments, the system 500 applies a Hough transform to identify one or more line patterns of objects in the environment. The Hough transform can be viewed as a feature extraction technique that can be used to find lines in general and can be used on vehicle locations to gather vehicles aligned on the same line together. In some examples, the line pattern allows the system 500 to further discard irrelevant objects in the environment. In one or more examples or embodiments, determining the orientation data 510 based on the map parameters 502 and the group parameters 503 includes discarding one or more lines associated with the first object and / or the second object (e.g., objects in the group) based on the one or more line patterns. In some examples, the system 500 discards lines that are irregularly spaced or incompatible with each other between vehicles. The line pattern can be one-dimensional (such as, roadside parking, etc.), two-dimensional or three-dimensional. The system 500 can be configured to detect a regular grid of lines (e.g., a parking lot). For example, the system 500 determines whether the line pattern meets the line standard. In response to determining that the group does not meet the line standard, the system 500 is configured not to consider the case where objects that are not on the same line are part of the same group (such as discarding one or more lines, etc.). In response to determining that the group meets the line standard, the system 500 is configured not to discard the one or more lines, and the corresponding objects can be marked as belonging to the same group. This can be beneficial to create a consistent group of objects that share similar features.

[0119] In one or more examples or embodiments, determining the orientation data 510 based on the map parameters 502 and the group parameters 503 includes aligning at least one of the first object and the second object based on one or more line patterns. In one or more examples or embodiments, determining the orientation data 510 based on the map parameters 502 and the group parameters 503 includes determining the orientation data 510 based on the alignment. For example, the system 500 can determine the alignment between the first object and the second object. This can be useful for improving the determination of the orientation parameter because some objects may have a particular alignment in an environment.

[0120] For example, the system 500 aligns objects that form a line (e.g., roadside parking) or a regular two-dimensional grid (e.g., a parking lot). The alignment can also be used to align curved roads with curves. For example, the system 500 defines the heading parameter as relative to the road direction.

[0121] In one or more examples or embodiments, the system 500 determines orientation parameters based on alignment. This can allow object orientation to be refined (particularly from pre-existing map parameters). For example, vehicles in a parking lot (as indicated by the first object and / or the second object) typically share the same orientation, with a potential variation of 180 degrees. However, a vehicle that is not in the correct position may affect other vehicles in the parking lot. Alignment can be used to correct orientation data 510 for objects within the map parameters 502 or group parameters 503 that may not be suitable. Similar situations may occur for vehicles parked along a road, where the vehicle should follow the road direction and orientation, but this may not actually be the case.

[0122] In one or more examples or embodiments, the system 500 may be used at various points in a pipeline, particularly a labeling pipeline. Figure 6 600 is a diagram of an example implementation of a process for object orientation determination along a labeling pipeline 600. The labeling pipeline may include an object detector 602 (e.g., Figure 5 similar to the object detection system 504 of FIG. 5 ), non-maximum suppression scheme 604 (NMS, such as Figure 5 ), tracker 606 (e.g., similar to NMS 505 of Figure 5 ), tracker refinement 608, and post-processing 610 (e.g., similar to tracker 506 of Figure 5 The post-processing system 507 is similar).

[0123] In one or more examples or embodiments, the operation also includes improving the object detection data 512 using map layer information based on the map parameters 502 and / or the group parameters 503 using at least one processor. In one or more examples or embodiments, the system 500 is improved in terms of the object detector 602. The map layer can be viewed as a semantic layer obtained by the system 500, such as a bird's-eye view of a particular environment. The map layer indicates, for example, the ground truth of the environment. In one or more examples or embodiments, the operation also includes improving the object detection data 512 using map layer information based on the map parameters 502 and the group parameters 503 using at least one processor. In one or more examples or embodiments, the operation also includes improving the object detection data 512 using map layer information based on the map parameters 502 or the group parameters 503 using at least one processor. The improvement can be a partial or complete fusion of different data.

[0124] Advantageously, the system 500 can be implemented at multiple levels. For example, during early fusion, the LiDAR point cloud (e.g., via sensor data 501) is improved using information about the map layer in which a particular point in the point code is located. During a mid-term fusion process, the system 500 can obtain one or more map layers (e.g., raster maps and / or vector maps), such as input to a machine learning model, etc. For example, a map layer can be used as an additional channel for pseudo-images in an encoder for detecting objects in a point cloud (e.g., PointPillar as described in the following publication: AHLang et al., "PointPillars: Fast Encoders for Object Detection from Point Clouds", arXiv:1812.05784v2, May 2019, incorporated herein by reference).

[0125] The system 500 can be configured to use any neural network that takes an organized point cloud as input. For example, on top of an organized LiDAR point cloud (which is a pseudo image, such as the input of PointPillars), the neural network can also take a pseudo image of the same dimensions as the LiDAR pseudo image, where the pseudo image depicts a scene map. The system 500 can be configured to use a map layer as a map area filter. For example, the system 500 uses a mask of the drivable area of ​​the map to remove any detections outside the drivable area, which can result in increased processing speed and can be done on boxes or points.

[0126] In one or more examples or embodiments, the operation further includes performing a non-maximum suppression scheme (NMS) 505, 604 on the object detection data 512 based on the group parameters 503 using at least one processor. For example, the system 500 uses NMS 505, 604 to remove duplicate detections made by the network. The group prior can be used to inform the NMS settings to, for example, avoid removing detections of vehicles parked close to each other. For example, the non-maximum suppression (NMS) algorithm will use a less restrictive intersection-over-union (IoU threshold) and will thereby retain more objects because boxes corresponding to vehicles close to each other may share a relatively high IoU score. The system can use the map prior and the group prior to offset the threshold associated with the detection score before NMS 505, 604 for regularly spaced or aligned boxes that conform (or do not conform) to the group, so that boxes with low confidence detections but consistency with respect to the group are still considered in the NMS step. In other words, before NMS, in locations where there are groups, the system 500 is configured to keep more boxes. In some examples, each box is associated with a confidence score, and boxes with low confidence scores are discarded, and thus when there are groups, the system 500 is configured to keep boxes with lower confidence scores. The list of boxes that are kept can then be refined via NMS.

[0127] In one or more examples or embodiments, the operation further includes tracking one or more objects in the environment based on the orientation data 510 and the sensor data 501 using at least one processor. The system 500 is configured to track in the tracker 606 and / or during tracker refinement 608, for example. For example, the system 500 uses a probability distribution over location and / or orientation in the tracker 606 to change a Kalman filter update step (e.g., to offset the vehicle toward the center of the lane). As another example, the system 500 refines the pose, size, and / or orientation of the entire track over time in the tracker refinement 608. The system 500 can implicitly incorporate them using group parameters and map parameters (e.g., in early fusion and / or mid-term fusion).

[0128] In one or more examples or embodiments, the operations further include determining, using at least one processor, control data for controlling the autonomous vehicle based on the sensor data 501 and the orientation data 510. The control data is used, for example, to control the operation of the vehicle. The system 500, for example, provides control data to a control system of the autonomous vehicle, such as Figure 4A The system 500 may transmit the control data to, for example, a control system of an autonomous vehicle and / or an external system. The system 500 may transmit the control data to Figure 6 Vehicle tracking 612 is shown.

[0129] In one or more examples or embodiments, the operation further comprises determining, using at least one processor, to label at least one object in the environment based on the orientation data 510 and the sensor data 501. In other words, the system 500 is configured, for example, to apply a label to an object in the environment. Advantageously, the system 500 can improve the internal labeling of the object offline or online, which in turn can improve the operation of the autonomous vehicle. Labeling can further be used for object detection and labeling, which can improve the quality of labeling of the system 500. For example, the system 500 labels at least one object via the object detection system 504.

[0130] In one or more examples or embodiments, the system 500 is configured to apply one or more post-processing 610, for example, prior to any vehicle tracking 612. In one or more examples or embodiments, the operations further include updating, using at least one processor, a machine learning model based on the heading data 510. The machine learning model may be a neural network, for example Figure 4B CNN 420 or Figure 4C and Figure 4D CNN 440, etc. The system 500 trains a machine learning model based on the heading data 510, for example. Such training can be performed offline. Alternatively or in combination, training can be performed online using sensor data 501. In one or more examples or embodiments, updating the machine learning model allows improvements to the tracker and / or tracker refinement network, and / or improvements in any post-processing functions. In one or more examples or embodiments, updating the machine learning model includes inputting one or more of the heading data 510, one or more object parameters, group parameters 503, and map parameters 502 into the machine learning model. In one or more examples or embodiments, updating the machine learning model includes outputting updated heading data by the machine learning model. In one or more examples or embodiments, updating the machine learning model includes recursively applying the updated heading data to replace the heading data 510. The updated heading data can be improved by the machine learning model (compared to the original heading data 510). Machine learning models such as fully connected neural networks can be used for implicit post-processing. The one or more object parameters include, for example, a box size, location, and / or score associated with the first object and / or the second object, and can be used as input into the machine learning model. The system 500, for example, determines one or more object parameters and / or obtains one or more object parameters.

[0131] The machine learning model may output updated object parameters. Recursively applied by the system 500 may include continuously refreshing and updating data in the machine learning model (e.g., iterating). The system 500 is applied recursively, for example, until convergence or until a set number of iterations. In one or more examples or embodiments, the system 500 uses a PointNet-like network for a variable number of objects in a group. A PointNet-like network can be viewed as a neural network for detecting LiDAR objects (e.g., a network that processes points or point neighborhoods individually and extracts global features from them). For example, PointNet (e.g., PointNet described in CRQi et al., "PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation" (April 10, 2017, ArXiv:1612.00593v2, incorporated herein by reference)) is a neural network architecture specifically designed to process unordered point clouds.

[0132] In one or more examples or embodiments, the operation also includes obtaining one or more estimated object parameters using at least one processor. In one or more examples or embodiments, the operation also includes comparing one or more estimated object parameters and orientation data 510 using at least one processor. In one or more examples or embodiments, the operation also includes determining, using at least one processor, a differential parameter indicating an orientation difference between the one or more estimated object parameters and the orientation data 510 based on the comparison. In one or more examples or embodiments, the operation also includes updating the orientation data 510 based on the differential parameter using at least one processor. This processing can be referred to as post-processing 610, such as rule-based post-processing or explicit post-processing. For example, the system 500 compares one or more estimated object parameters and the orientation data 510 to see if the "expected" object parameters match what the orientation data 510 actually indicates, for example, based on the map parameters 502. If the difference between the one or more estimated object parameters (e.g., orientation data 510) and the map parameters 502 is within a threshold, the system 500 updates the one or more estimated object parameters and / or orientation data 510 based on the map parameters 502. If the difference between the one or more estimated object parameters and the map parameters 502 is not within a threshold (e.g., a difference parameter) according to the map parameters 502, the system 500 does not update the one or more estimated object parameters and / or orientation data 510 based on the map parameters 502. As an example, if the orientation of a detected parked car located at the side of the road is close to the orientation of the road, the system 500 is configured to determine that the orientation data 510 of the parked car is substantially equal to the orientation of the angle.

[0133] In one or more examples or embodiments, the system 500 retrieves data from a database (e.g. Figure 4A The system 500 may obtain one or more estimated object parameters from a database 410 of the system 500. The one or more estimated object parameters may indicate the orientation of one or more objects (such as the first object or the second object, etc.). In some examples, the system 500 compares the one or more estimated object parameters with the orientation data 510. This is useful for determining the accuracy of the system 500. Based on the comparison, the system 500 may determine a differential parameter, which may be a numerical representation of the accuracy of the system 500 for the orientation data 510.

[0134] In one or more examples or embodiments, the operation also includes estimating a probability distribution of the orientation data 510 based on the map parameters 502 and / or the group parameters 503 using at least one processor. In one or more examples or embodiments, the operation also includes determining the object detection data 512 based on the probability distribution of the orientation data 510 using at least one processor. For example, the system 500 uses Bayesian interference in the post-processing 610. The system 500 modifies the object detector 602, for example, to output a probability distribution on the estimated box parameters (e.g., position, orientation, size). Such calculations can be simplified by modeling the position, orientation, and size separately. For example, the system 500 estimates the probability distribution of these parameters (e.g., the orientation of the vehicle at a specific position on the lane) for each type of map parameter 502. The system 500 can fuse the map parameters 502 and the estimated box parameters using Bayes' theorem to calculate the posterior probability. In one or more examples or embodiments, the system 500 uses the probability to determine the object detection data 512.

[0135] In one or more examples or embodiments, the operation also includes determining at least one ground truth object in the environment based on the sensor data 501 using at least one processor. In one or more examples or embodiments, the operation also includes determining object orientation data indicating the ground truth orientation of the at least one ground truth object based on the sensor data 501 using at least one processor. In one or more examples or embodiments, the operation also includes determining, using at least one processor, a confidence parameter indicating the orientation difference between the object orientation data and the orientation data 510 based on a comparison of the object orientation data and the orientation data 510. In one or more examples or embodiments, the operation also includes updating the orientation data 510 based on the confidence parameter using at least one processor. In other words, uncertainty modeling can be used, for example, by using confidence intervals (e.g., confidence parameters). For example, the system 500 uses the inconsistency of priors and observations to model the uncertainty of objects in the environment for active learning, for example, to improve future detectors. As an example, the system 500 determines that the vehicle is on the grass may be a false positive, and the system 500 determines that the gap in the detected vehicle line may indicate a false negative. In one or more examples or embodiments, the magnitude of the position, orientation, and / or confidence parameters increases with larger group sizes and / or longer time horizons. For example, if the system 500 determines that nine vehicles are perfectly aligned in a parking lot, the system 500 expects the vehicle in the tenth parking spot to have a similar orientation. The system 500 can update itself to correct one or more of object orientation assumptions, improve object detection, train a machine learning model, and / or improve data fusion (e.g., early and / or mid-term fusion).

[0136] In one or more examples or embodiments, a ground truth object is an accurate object in the environment, such as an object that actually exists in the environment. The system 500 can use sensor data 501 to make such a determination. The system, for example, determines object orientation data of a ground truth object that indicates the accurate orientation of the object in the environment. In addition, in some examples, the system is configured to compare the object orientation data with the orientation data 510 (e.g., such as comparing the "real" object orientation with the orientation that the system 500 has determined in the absence of sensor data 501). The comparison is represented by a confidence parameter, which the system 500 can then use to update the orientation data 510 if necessary.

[0137] In one or more examples or embodiments, the map parameter 502 indicates an area of ​​the environment. In one or more examples or embodiments, the group parameter 503 indicates a predetermined connection in the area. For example, the group parameter may be only relevant to a specific map area (e.g., area, location, boundary, interaction), such as a curved road, a loading area, and / or a pedestrian crossing, etc.

[0138] In one or more examples or embodiments, the first object is a first static object. In one or more examples or embodiments, the second object is a second static object. In one or more examples or embodiments, the first object is a first moving object. In one or more examples or embodiments, the second object is a second moving object. The first object may be a static object, and the second object may be a dynamic object. The first object may be a dynamic object, and the second object may be a static object.

[0139] Fig. 7A and Figure 7B is a graph of example group parameters that may be used for object orientation determination. Fig. 7A An example parking lot 700 is shown. As shown, vehicles 702 in parking lot 700 share the same orientation (+ / -180 degrees). This is an example of a group parameter where the system 500 determines the interaction between the second objects in the group, i.e., their relative positioning in parking lot 700. Figure 7B Road 750 is shown with vehicles 752 parked along the side of road 750 and vehicles 754 legally driving on road 750. Group parameters indicate, for example, that vehicles parked along road 750 should follow the road direction and orientation.

[0140] Reference now Figure 8 , a flow chart of a method or process 800 for object orientation determination, for example, for operating and / or controlling an AV, is shown. The method may be performed by a system (such as Figure 2 AV Computing 202F and Figure 4A AV Computing 400, Figure 1 102. Figure 2 200 vehicles, Figure 3 The device 300 and Figure 5 AV calculation 540, and Figure 6 , Fig. 7A and Figure 7B The disclosed system may include at least one processor, which may be configured to perform one or more operations of method 800. Method 800 may be performed by other devices or device groups (e.g., completely and / or partially, etc.) separate from the system disclosed herein or including the system disclosed herein.

[0141] The method 800 includes obtaining, at step 802, using at least one processor, a map parameter indicating a predetermined position of a first object in an environment in which the autonomous vehicle is configured to operate. In one or more embodiments or examples, the method 800 includes obtaining, at step 804, using at least one processor, a group parameter indicating a predetermined connection between objects in a group. In one or more embodiments or examples, the environment includes objects in the group. In one or more embodiments or examples, the method 800 includes determining, at step 806, using at least one processor, orientation data indicating an orientation of at least one object among the first object and the objects in the group based on the map parameter and the group parameter. In one or more embodiments or examples, the method 800 includes providing, at step 808, using at least one processor, object detection data associated with at least one object to the device based on the orientation data, wherein the object detection data indicates detection of one or more spatial features of the at least one object. In some cases, the device may include a control system of an AV, and providing the object detection data may cause the AV to be controlled based on the object detection data associated with at least one object (e.g., an object in the environment indicated by the orientation data). Thus, in some cases, method 800 may include causing the AV to be controlled based on object detection data associated with at least one object.

[0142] In one or more embodiments or examples, method 800 includes obtaining sensor data associated with the environment using at least one processor. In one or more embodiments or examples, method 800 includes determining, at step 806, using at least one processor, orientation data based on map parameters, group parameters, and sensor data.

[0143] In one or more embodiments or examples, obtaining group parameters at step 804 includes determining distances between objects in the group based on sensor data. In one or more embodiments or examples, obtaining group parameters at step 804 includes clustering the objects based on distance to form groups.

[0144] In one or more embodiments or examples, method 800 includes discarding groups from the object detection data based on the group parameters and the map parameters.

[0145] In one or more embodiments or examples, determining the orientation data based on the map parameters and the group parameters at step 806 includes extracting one or more line patterns associated with the first object and / or objects in the group based on the sensor data and the group parameters.

[0146] In one or more embodiments or examples, determining the orientation data based on the map parameters and the group parameters at step 806 includes discarding one or more lines associated with the first object and / or objects in the group based on one or more line patterns.

[0147] In one or more embodiments or examples, determining the orientation data based on the map parameters and the group parameters at step 806 includes aligning the first object and at least one of the objects in the group based on one or more line patterns. In one or more embodiments or examples, determining the orientation data based on the map parameters and the group parameters at step 806 includes determining the orientation data based on the alignment.

[0148] In one or more embodiments or examples, method 800 includes utilizing, using at least one processor, map layer information to improve object detection data based on map parameters and / or group parameters.

[0149] In one or more embodiments or examples, method 800 includes performing, using at least one processor, a non-maximum suppression scheme on object detection data based on a group parameter.

[0150] In one or more embodiments or examples, method 800 includes tracking, using at least one processor, one or more objects in an environment based on orientation data and sensor data.

[0151] In one or more embodiments or examples, method 800 includes determining, using at least one processor, control data for controlling an autonomous vehicle based on sensor data and orientation data.

[0152] In one or more embodiments or examples, method 800 includes determining, using at least one processor, to label at least one object in an environment based on orientation data and sensor data.

[0153] In one or more embodiments or examples, method 800 includes updating, using at least one processor, a machine learning model based on the heading data.

[0154] In one or more embodiments or examples, updating the machine learning model includes inputting one or more of the orientation data, one or more object parameters, group parameters, and map parameters into the machine learning model. In one or more embodiments or examples, updating the machine learning model includes outputting updated orientation data by the machine learning model. In one or more embodiments or examples, updating the machine learning model includes recursively applying the updated orientation data to replace the orientation data.

[0155] In one or more embodiments or examples, method 800 includes obtaining one or more estimated object parameters using at least one processor. In one or more embodiments or examples, method 800 includes comparing one or more estimated object parameters and orientation data using at least one processor. In one or more embodiments or examples, method 800 includes determining a differential parameter based on the comparison using at least one processor. In one or more embodiments or examples, the differential parameter indicates a direction difference between the one or more estimated object parameters and the orientation data. In one or more embodiments or examples, method 800 includes updating the orientation data based on the differential parameter using at least one processor.

[0156] In one or more embodiments or examples, method 800 includes using at least one processor to estimate a probability distribution of orientation data based on map parameters and / or group parameters. In one or more embodiments or examples, method 800 includes using at least one processor to determine object detection data based on the probability distribution of orientation data.

[0157] In one or more embodiments or examples, method 800 includes determining, using at least one processor, at least one ground truth object in the environment based on sensor data. In one or more embodiments or examples, method 800 includes determining, using at least one processor, object orientation data indicating a ground truth orientation of at least one ground truth object based on sensor data. In one or more embodiments or examples, method 800 includes determining, using at least one processor, a confidence parameter based on a comparison of the object orientation data and the orientation data. In one or more embodiments or examples, the confidence parameter indicates an orientation difference between the object orientation data and the orientation data. In one or more embodiments or examples, method 800 includes updating the orientation data based on the confidence parameter using at least one processor.

[0158] In one or more embodiments or examples, the map parameter indicates an area of ​​the environment. In one or more embodiments or examples, the group parameter indicates a predetermined contact in the area.

[0159] In one or more embodiments or examples, the first object is a first static object and the object is a second static object.

[0160] In one or more embodiments or examples, the first object is a first moving object, and the object is a second moving object.

[0161] In the previous description, aspects and embodiments of the present disclosure have been described with reference to many specific details, which may vary depending on the implementation. Therefore, the description and the accompanying drawings should be regarded as illustrative, not restrictive. The only and exclusive indication of the scope of the invention, and the applicant's expectation that the content of the scope of the invention is the literal and equivalent scope of the claims issued from this application in the specific form of the claims, including any subsequent amendments. Any definition of the terms used to be included in such claims that are clearly set forth herein should be based on the meaning of such terms as used in the claims. In addition, when the term "also includes" is used in the previous description or the attached claims, the following of the phrase may be an additional step or entity, or a sub-step / sub-entity of the previously described step or entity.

[0162] A non-transitory computer-readable medium is disclosed that includes instructions stored thereon that, when executed by at least one processor, cause the at least one processor to implement operations according to one or more of the methods disclosed herein.

[0163] Example Autonomous Vehicle (AV) Computing System

[0164] Fig. 9 502 to generate heading data 510 (indicative of the heading of at least one object) and object detection data 512 (associated with at least one object to be provided to the device based on the heading data 510) based at least in part on group parameters 503 and map parameters 502. AV computing 540 may include at least one processor and at least one memory storing instructions that, when executed by the at least one processor, cause the at least one processor to perform these processes. In some cases, the device provided with the object detection data 512 may be a control system 910 of an autonomous vehicle (AV), and the AV may be controlled based on the object detection data 512 enhanced and / or modified using the heading data 510.

[0165] In some implementations, AV computing 540 may perform processing using one or more of object detection system 504, NMS 505, tracker 506, and post-processing system 507. It will be appreciated that AV computing 540 may perform other processing, and generate and / or output other types of data and control commands, and that other implementation variations are possible.

[0166] In various implementations, the AV computing 540 may use (one or more) map parameters 502 and / or (one or more) group parameters 503 to perform these processes. For example, the AV computing 540 may use heuristics with (one or more) map parameters 502 and / or (one or more) group parameters 503 as additional clues or indicators that can be used for object detection in terms of the location, orientation, and size of the object. In some examples, one or both of the map parameters 502 and the group parameters 503 may be stored in a memory 916 (e.g., a non-transitory memory of the AV) or stored on a distributed network (such as a cloud network, etc.). In some examples, the memory 916 may be a memory of the AV computing 540. In some examples, the group parameters 503 may have been generated by using a set of heuristics or geometric algorithms (e.g., by a processor of the AV computing 540 or other system). In some examples, the set of heuristics or geometric algorithms may be used to extract and / or determine the group parameters, which may have been determined based on previously collected sensor data. In some other instances, the group parameter 503 may be generated by the AV computing 540, for example, using sensor data 501 received from one of the AV's sensors 904. In some cases, the map parameter 502 may have been extracted from and / or determined using previously collected sensor data. In some embodiments, the (one or more) map parameter 502 may indicate a predetermined position of an object (e.g., a first object or a primary object) or a group of objects in an environment 902 in which the autonomous vehicle is configured to operate. In some cases, the first object of the map parameter 502 may be a feature of the environment that may lead to the position and / or orientation of other objects associated with the map parameter 502 (e.g., relative to the first object). In some cases, the (one or more) group parameter 503 may indicate a predetermined relationship between two or more objects (e.g., a second object) that form a group. In some cases, an object (e.g., a primary object, or a first object such as a first object of a map parameter) may lead to the position and / or orientation of other objects in the group via its predetermined relationship.

[0167] In a non-limiting implementation, the processing performed by AV computing 540 may include an orientation data generation and evaluation process 908a, an object detection data generation process 908b, and a group parameter generation process 908c.

[0168] The heading data generation and evaluation process 908a may include generating heading data 510 based at least in part on map parameters 502 and / or group parameters 503, and / or evaluating the accuracy or reliability of heading data 510. In some cases, AV computing 540 may generate heading data 510 using map parameters 502, group parameters 503, and / or sensor data 501.

[0169] In some cases, AV computation 540 may generate orientation data 510 by determining at least an orientation of at least one object associated with a group associated with group parameter(s) 503. In some examples, AV computation 540 may select a group from group parameter(s) 502 based at least in part on map parameter 502. In some cases, AV computation 540 may generate orientation data 510 by determining at least an orientation of at least one object associated with a group of group parameter(s) 502 and / or at least an orientation of at least one object associated with map parameter(s) 502.

[0170] In some cases, AV computing 540 may also generate orientation data by extracting one or more line patterns associated with an object associated with map parameter 502 and / or an object associated with group parameter 503 based on sensor data 501 and group parameter 503. Subsequently, in some cases, AV computing 540 may use the one or more line patterns to align at least one of the object associated with map parameter 502 and the object associated with group parameter 503. In some examples, alignment may include determining alignment between the object associated with map parameter 502 and the object associated with group parameter 503.

[0171] In some cases, assessing the accuracy or reliability of orientation data 510 may include comparing orientation data 510 to a determined orientation of a ground truth object and generating a confidence parameter indicating a difference between the determined orientation and orientation data 510 .

[0172] In some cases, evaluating the accuracy or reliability of the heading data 510 may include comparing the heading data 510 with data from a database (e.g., Figure 4A The system 500 compares one or more estimated object parameters received from the database 410) and determines a differential parameter, which can be a numerical representation of the accuracy of the system 500 generating the data 510.

[0173] In some cases, object detection data generation process 908B may include receiving orientation data 510 and generating object detection data 512 based at least in part on orientation data 510. In some examples, object detection data 512 may be enhanced or modified based on orientation data 510. In some cases, AV computing 540 may generate orientation data 510 using sensor data 501 and orientation data 510. In some examples, orientation data 510 may be used to filter sensor data 501, organize sensor data 501 or label sensor data 501, determine the orientation and position of an object captured by sensor data 501, or supplement sensor data 501 (e.g., a point cloud generated by a LiDAR sensor). In some cases, object detection data generation process 908B may also include:

[0174] • Estimating a probability distribution of orientation data 510 . In some cases, AV computation 540 may determine object detection data 512 based at least in part on the estimated probability distribution of orientation data 510 .

[0175] Discarding objects or groups of objects from the object detection data 512 (eg, groups and / or objects in non-relevant map areas) based on the group parameters 503 and / or the map parameters 502 .

[0176] Label and identify objects in the environment.

[0177] Generate map layer information and use the map layer information to improve object detection data 512.

[0178] • Inform the NMS settings of the AV calculation 540 to avoid removing detections of objects in a group associated with the group parameter as duplicate detections of other objects in the group.

[0179] In some implementations, the AV computing 540 may perform the object detection data generation process 908 b based on one or both of the group parameters 503 and the map parameters 502 .

[0180] The group parameter generation process 908c may include generating group parameters 503 based at least in part on the sensor data 501. In some examples, the AV computing 540 generates the group parameters 503 using a clustering technique to select one or more objects to form a group.

[0181] In some cases, map parameters 502 may have been generated using sensor data obtained by sensor 904 (eg, sensor data 501 ) or other sensors, and may indicate a predetermined location of an object in environment 902 .

[0182] In some implementations, the AV computing 540 can be used online for the operation of an AV or other system that actively receives sensor data 501 and generates object detection data 512 indicating the location, orientation, and / or characteristics (e.g., size, shape, form factor, etc.) of one or more objects based on the sensor data 501. In these implementations, the AV computing 540 can transmit the object detection data 512 to the control system 910 of the AV (e.g., control system 404), where the control system 910 uses the object detection data 512 to navigate in the environment 902. In some cases, the AV computing 540 can cause the AV to be controlled based on the object detection data 512 associated with at least one object, where the orientation of the at least one object is indicated by the orientation data 510. In some cases, the sensor data 501 can be generated by one or more sensors 904 such as a camera (e.g., camera 202a), a LiDAR sensor (e.g., LiDAR sensor 202b), a Radar sensor (e.g., Radar sensor 202c), and other types of sensors. Sensor 904 may receive sound waves, electromagnetic radiation (e.g., light beams, optical signals, radio frequency (RF) waves or signals, microwaves or signals), or other types of waves or signals from environment 902 (e.g., the environment in which an AV or other system navigates or operates).

[0183] In some embodiments, AV computing 540 can be used online for operation of an AV or other system that actively receives sensor data 501 and generates object detection data 512 indicating the location, orientation, and / or characteristics (e.g., size, shape, form factor, etc.) of one or more objects based on the sensor data 501. In some examples, when used online, AV computing 540 can use orientation data 510 to update a machine learning model 914 used by the AV. In some cases, the machine learning model 914 can be recursively updated by applying the updated orientation data in place of the orientation data 510. The updated orientation data can be improved by the machine learning model 914 relative to the original orientation data 510.

[0184] In some embodiments, the AV computing 540 can be used offline, where the AV computing 540 receives and processes offline data. In some cases, the offline data 912 can include data generated by other systems or data pre-generated by the sensors 904 during one or more operating periods of the AV. In some cases, in the offline mode, the AV computing 540 can be used to generate the heading data 510 and the object detection data using the offline data 912, and provide the resulting object detection data 512 to other applications 918. In some cases, in the offline mode, the AV computing 540 can be used to train the machine learning model 914 of the AV or other machine learning model based on the offline data 912. In some cases, training the machine learning model can include updating the heading data 510 and recursively applying the updated heading data 510 instead of the heading data 510.

[0185] Fig. 10A It shows that it can be Figure 5 or Fig. 9 A diagram of example map parameters and example group parameters used by AV computing 540 to determine the orientation and / or position of one or more objects in an environment is shown. In some examples, group parameters 1002 may include assumptions about the connections and / or interactions (depicted as solid lines connecting objects) between objects 1002a-1002f that form the group of group parameters 1002. In some cases, a group of objects may be identified based on the assumed connections and / or interactions between objects 1002a-1002f. In some cases, a primary object 1004 (e.g., a static object such as a point of interest) may be linked to the group and may indicate a high probability of forming a group near the primary object 1004. In some cases, detection of the primary object 1004 may be used to verify group identification and / or guide determination of the position and / or orientation of one or more objects 1002a-1002f (e.g., relative to the primary object 1004) in the group. In some cases, objects 1006a-1006b may form map parameters 1006 that include assumptions about possible positions and orientations of objects 1006a-1006c relative to each other and relative to a predetermined location (e.g., the location of first object 1006a in the environment). In some such cases, detection of first object 1006a of map parameters 1006 may indicate an alignment, orientation, or position of objects 1006a and 1006c.

[0186] Fig. 10B 1 is a diagram showing an initial perception of the environment around the AV 1010 based on sensor data (e.g., based on sensor data 501) and before considering map parameters and group parameters. The initial perception of the environment may include objects whose orientation and / or position may be detected using Fig. 10A1006 and / or group parameters 1002 shown in the figure. In some cases, AV 1010 includes a sensor 1012 (e.g., LiDAR or a camera) configured to monitor the environment and detect objects in the environment. In some cases, AV 1010 may include AV computing similar to AV computing 540, which is configured to generate heading and object detection data based on the map parameters and group parameters stored in the memory of AV 1010. The map parameters and group parameters may include the above-mentioned Fig. 10A In some cases, the sensor 1012 may generate sensor data indicating the presence of vehicles 1014, 1016, and 1018 parked to the side of the road along which the AV 1010 is moving. Fig. 10B As shown, the orientation of vehicle 1016 initially perceived based on sensor data may be opposite to the orientation of vehicle 1014, and the orientation of vehicle 1018 may be significantly tilted relative to first vehicle 1014 and the roadside. In some examples, AV computing 540 of AV 1010 may identify vehicle 1014 as the first vehicle associated with map parameters 1006, and determine the orientation and / or position of vehicles 1016 and 1018 (e.g., relative to vehicle 1014 and / or the road direction or roadside) based on map parameters 1006 indicating that the orientation of second object 1006b and third object 1006c should be similar to the orientation of first object 1006a and that objects 1006a-1006c should be substantially parallel to each other. Thus, in the resulting orientation data, the orientation of the second vehicle 1016 (which is perceived as being opposite to the orientation of the first vehicle 1014) can be flipped based on the expected or assumed orientation of the second object 1006b relative to the first object 1006a. Similarly, in the resulting orientation data, the orientation of the third vehicle 1018 (which is initially perceived as not being substantially parallel to the first vehicle 1014) can be adjusted based on the expected or assumed orientation of the third object 1006c so that the orientation becomes parallel to the first object 1006 and, thereby, parallel to the first vehicle 1014.

[0187] Continue to refer Fig. 10BAs another example, the AV computing 540 of the AV 1010 may identify the pedestrians 1022a-1022d as separate dynamic objects that move independently in different (e.g., random) directions. In some cases, the AV computing 540 may use the relationship or relative dynamics between the group parameters 1002 and the detected pedestrians 1022a-1022d (based on the sensor data) to determine that the pedestrians form a group (e.g., corresponding to the group parameters 1002) and that there is a high probability that the pedestrians are moving in the same direction or oriented in a common direction. In some examples, the AV computing system may assume that the detected object 1020 has the characteristics (e.g., geometric and spatial characteristics) of the primary object 1004 to further confirm that the pedestrians 1022a-1022d form a group associated with the group parameters 1002. For example, the object 1020 may be a traffic light, and once the AV computing 540 identifies it as the primary object 1004, it may be assumed that the pedestrians 1022a-1022d form a group corresponding to the group parameters 1002. Thus, in the resulting orientation data, the orientation and movement direction of pedestrians 1022a-1022d (initially perceived as random and / or independent) may be adjusted based on the expected or assumed orientation of objects 1002a-1002f relative to each other (e.g., generally walking in a common direction). In some cases, the orientation and movement direction of pedestrians 1022a-1022d may be additionally adjusted relative to the primary object 1004, and thereby relative to the curb relative to which the primary object 1004 (corresponding to the detected object 1020) is statically positioned.

[0188] Example Embodiments

[0189] The example embodiments described herein have several features, no single one of which is essential or solely responsible for its desirable attributes.Various example systems and methods are provided below.

[0190] Also disclosed are methods, non-transitory computer readable media, and systems according to any of the following:

[0191] Example 1. A method comprising:

[0192] obtaining, using at least one processor, map parameters indicating a predetermined location of a first object in an environment in which the autonomous vehicle is configured to operate;

[0193] obtaining, using the at least one processor, a group parameter indicating a predetermined connection between objects in a group, wherein the environment includes the objects in the group;

[0194] determining, using the at least one processor, orientation data indicating an orientation of at least one of the objects in the group and the first object based on the map parameter and the group parameter; and

[0195] Using the at least one processor, object detection data associated with the at least one object is provided to an apparatus based on the orientation data, wherein the object detection data indicates detection of one or more spatial features of the at least one object.

[0196] Example 2. The method according to Example 1, further comprising:

[0197] obtaining, using the at least one processor, sensor data associated with the environment,

[0198] Wherein, determining the orientation data based on the map parameter and the group parameter includes: also determining the orientation data based on the sensor data.

[0199] Example 3. The method according to Example 2, wherein obtaining the group parameter comprises:

[0200] determining distances between objects in the group based on the sensor data; and

[0201] The objects are clustered based on the distance to form the group.

[0202] Example 4. The method according to any of the preceding examples, further comprising:

[0203] Groups are discarded from the object detection data based on the group parameters and the map parameters.

[0204] Example 5. A method according to any of the preceding examples, wherein determining the orientation data based on the map parameters and the group parameters comprises: extracting one or more line patterns associated with the first object and / or objects in the group based on the sensor data and the group parameters.

[0205] Example 6. The method according to Example 5, wherein determining the orientation data based on the map parameters and the group parameters comprises discarding one or more lines associated with the first object and / or objects in the group based on the one or more line patterns.

[0206] Example 7. The method of any one of Examples 5 and 6, wherein determining the orientation data based on the map parameter and the group parameter comprises:

[0207] aligning objects in the group and at least one of the first objects based on the one or more line patterns; and

[0208] The orientation data is determined based on the alignment.

[0209] Example 8. The method according to any of the preceding examples, further comprising:

[0210] Using the at least one processor, map layer information is used to improve the object detection data based on the map parameters and / or the group parameters.

[0211] Example 9. The method according to any of the preceding examples, further comprising:

[0212] Using the at least one processor, a non-maximum suppression scheme is performed on the object detection data based on the group parameters.

[0213] Example 10. The method according to any of the preceding examples, further comprising:

[0214] Using the at least one processor, one or more objects in the environment are tracked based on the orientation data and the sensor data.

[0215] Example 11. The method according to any one of Examples 2 to 10, further comprising:

[0216] Using the at least one processor, control data for controlling the autonomous vehicle is determined based on the sensor data and the orientation data.

[0217] Example 12. The method according to any one of Examples 2 to 11, further comprising:

[0218] Using the at least one processor, a determination is made to label the at least one object in the environment based on the orientation data and the sensor data.

[0219] Example 13. The method according to any of the preceding examples, further comprising:

[0220] Using the at least one processor, updating a machine learning model based on the orientation data.

[0221] Example 14. The method of Example 13, wherein updating the machine learning model comprises:

[0222] inputting one or more of the orientation data, one or more object parameters, the group parameters, and the map parameters into the machine learning model;

[0223] outputting updated orientation data from the machine learning model; and

[0224] Instead of the orientation data, the updated orientation data is recursively applied.

[0225] Example 15. The method according to any of the preceding examples, further comprising:

[0226] obtaining, using the at least one processor, one or more estimated object parameters;

[0227] using the at least one processor, comparing the one or more estimated object parameters to the orientation data;

[0228] determining, using the at least one processor, a difference parameter indicative of an orientation difference between the one or more estimated object parameters and the orientation data based on the comparing; and

[0229] Using the at least one processor, the orientation data is updated based on the differential parameter.

[0230] Example 16. The method according to any of the preceding examples, further comprising:

[0231] estimating, using the at least one processor, a probability distribution of the orientation data based on the map parameters and / or the group parameters; and

[0232] The object detection data is determined, using the at least one processor, based on a probability distribution of the orientation data.

[0233] Example 17. The method according to any of the preceding examples, further comprising:

[0234] determining, using the at least one processor, at least one ground truth object in the environment based on the sensor data;

[0235] determining, using the at least one processor, object orientation data indicative of a ground truth orientation of the at least one ground truth object based on the sensor data;

[0236] determining, using the at least one processor, a confidence parameter indicative of an orientation difference between the object orientation data and the orientation data based on a comparison of the object orientation data and the orientation data; and

[0237] Using the at least one processor, the orientation data is updated based on the confidence parameter.

[0238] Example 18. The method of any of the preceding examples, wherein the map parameter indicates an area of ​​the environment, and wherein the group parameter indicates the predetermined contact in the area.

[0239] Example 19. The method of any of the preceding examples, wherein the first object is a first static object and the object is a second static object.

[0240] Example 20. The method of any one of Examples 1 to 18, wherein the first object is a first moving object and the object is a second moving object.

[0241] Example 21. The method of any of the preceding examples further comprises: using the at least one processor to determine the object detection data based on the orientation data.

[0242] Example 22. A non-transitory computer-readable medium comprising instructions stored thereon, which when executed by at least one processor causes the at least one processor to perform operations comprising:

[0243] obtaining map parameters indicating a predetermined location of a first object in an environment in which the autonomous vehicle is configured to operate;

[0244] obtaining a group parameter indicating a predetermined connection between objects in a group, wherein the environment includes the objects in the group;

[0245] determining, based on the map parameter and the group parameter, orientation data indicating an orientation of at least one of the objects in the group and the first object; and

[0246] The method causes object detection data associated with the at least one object to be provided to a device based on the orientation data, wherein the object detection data indicates detection of one or more spatial features of the at least one object.

[0247] Example 23. The non-transitory computer readable medium of Example 22, wherein the operations further include:

[0248] obtaining, using the at least one processor, sensor data associated with the environment;

[0249] Wherein, determining the orientation data based on the map parameter and the group parameter includes: also determining the orientation data based on the sensor data.

[0250] Example 24. The non-transitory computer-readable medium of Example 23, wherein obtaining the group parameter comprises:

[0251] determining distances between objects in the group based on the sensor data; and

[0252] The objects are clustered based on the distance to form the group.

[0253] Example 25. The non-transitory computer-readable medium of any one of Examples 22 to 24, wherein the operations further comprise:

[0254] Groups are discarded from the object detection data based on the group parameters and the map parameters.

[0255] Example 26. A non-transitory computer-readable medium according to any one of Examples 22 to 25, wherein determining the orientation data based on the map parameters and the group parameters includes: extracting one or more line patterns associated with the first object and / or objects in the group based on the sensor data and the group parameters.

[0256] Example 27. According to the non-transitory computer-readable medium of Example 26, determining the orientation data based on the map parameters and the group parameters includes: discarding one or more lines associated with the first object and / or objects in the group based on the one or more line patterns.

[0257] Example 28. The non-transitory computer-readable medium of any one of Examples 26 to 27, wherein determining the orientation data based on the map parameter and the group parameter comprises:

[0258] aligning objects in the group and at least one of the first objects based on the one or more line patterns; and

[0259] The orientation data is determined based on the alignment.

[0260] Example 29. The non-transitory computer-readable medium of any one of Examples 22 to 28, wherein the operations further comprise:

[0261] Using the at least one processor, map layer information is used to improve the object detection data based on the map parameters and / or the group parameters.

[0262] Example 30. The non-transitory computer readable medium of any one of Examples 22 to 29, wherein the operations further comprise:

[0263] Using the at least one processor, a non-maximum suppression scheme is performed on the object detection data based on the group parameters.

[0264] Example 31. The non-transitory computer readable medium of any one of Examples 22 to 30, wherein the operations further comprise:

[0265] Using the at least one processor, one or more objects in the environment are tracked based on the orientation data and the sensor data.

[0266] Example 32. The non-transitory computer readable medium of any one of Examples 23 to 31, wherein the operations further comprise:

[0267] Using the at least one processor, control data for controlling the autonomous vehicle is determined based on the sensor data and the orientation data.

[0268] Example 33. The non-transitory computer readable medium of any one of Examples 23 to 32, wherein the operations further comprise:

[0269] Using the at least one processor, a determination is made to label the at least one object in the environment based on the orientation data and the sensor data.

[0270] Example 34. The non-transitory computer-readable medium of any one of Examples 22 to 33, wherein the operations further comprise:

[0271] Using the at least one processor, updating a machine learning model based on the orientation data.

[0272] Example 35. The non-transitory computer-readable medium of Example 34, wherein updating the machine learning model comprises:

[0273] inputting one or more of the orientation data, one or more object parameters, the group parameters, and the map parameters into the machine learning model;

[0274] outputting updated orientation data from the machine learning model; and

[0275] Instead of the orientation data, the updated orientation data is recursively applied.

[0276] Example 36. The non-transitory computer-readable medium of any one of Examples 22 to 35, wherein the operations further comprise:

[0277] obtaining, using the at least one processor, one or more estimated object parameters;

[0278] using the at least one processor, comparing the one or more estimated object parameters to the orientation data;

[0279] determining, using the at least one processor, a difference parameter indicative of an orientation difference between the one or more estimated object parameters and the orientation data based on the comparing; and

[0280] Using the at least one processor, the orientation data is updated based on the differential parameter.

[0281] Example 37. The non-transitory computer-readable medium of any one of Examples 22 to 36, wherein the operations further comprise:

[0282] estimating, using the at least one processor, a probability distribution of the orientation data based on the map parameters and / or the group parameters; and

[0283] The object detection data is determined, using the at least one processor, based on a probability distribution of the orientation data.

[0284] Example 38. The non-transitory computer-readable medium of any one of Examples 22 to 37, wherein the operations further comprise:

[0285] determining, using the at least one processor, at least one ground truth object in the environment based on the sensor data;

[0286] determining, using the at least one processor, object orientation data indicative of a ground truth orientation of the at least one ground truth object based on the sensor data;

[0287] determining, using the at least one processor, a confidence parameter indicative of an orientation difference between the object orientation data and the orientation data based on a comparison of the object orientation data and the orientation data; and

[0288] Using the at least one processor, the orientation data is updated based on the confidence parameter.

[0289] Example 39. The non-transitory computer-readable medium of any one of Examples 22 to 38, wherein the map parameter indicates an area of ​​the environment, and wherein the group parameter indicates the predetermined contact in the area.

[0290] Example 40. The non-transitory computer-readable medium of any one of Examples 22 to 39, wherein the first object is a first static object and the object is a second static object.

[0291] Example 41. The non-transitory computer-readable medium of any one of Examples 22 to 39, wherein the first object is a first moving object and the object is a second moving object.

[0292] Example 42. According to the non-transitory computer readable medium of any one of Examples 22 to 41, the operation further includes: determining the object detection data based on the orientation data.

[0293] Example 43. A system comprising at least one processor and at least one memory storing instructions on the at least one memory that, when executed by the at least one processor, cause the at least one processor to perform operations comprising:

[0294] obtaining map parameters indicating a predetermined location of a first object in an environment in which the autonomous vehicle is configured to operate;

[0295] obtaining a group parameter indicating a predetermined connection between objects in a group, wherein the environment includes the objects in the group;

[0296] determining, based on the map parameter and the group parameter, orientation data indicating an orientation of at least one of the objects in the group and the first object; and

[0297] The method causes object detection data associated with the at least one object to be provided to a device based on the orientation data, wherein the object detection data indicates detection of one or more spatial features of the at least one object.

[0298] Example 44. The system of Example 43, wherein the operations further comprise:

[0299] obtaining, using the at least one processor, sensor data associated with the environment;

[0300] Wherein, determining the orientation data based on the map parameter and the group parameter includes: also determining the orientation data based on the sensor data.

[0301] Example 45. The system of Example 44, wherein obtaining the group parameter comprises:

[0302] determining distances between objects in the group based on the sensor data; and

[0303] The objects are clustered based on the distance to form the group.

[0304] Example 46. The system of any one of Examples 43 to 45, wherein the operations further comprise:

[0305] Groups are discarded from the object detection data based on the group parameters and the map parameters.

[0306] Example 47. A system according to any one of Examples 43 to 46, wherein determining the orientation data based on the map parameters and the group parameters includes: extracting one or more line patterns associated with the first object and / or objects in the group based on the sensor data and the group parameters.

[0307] Example 48. A system according to Example 47, wherein determining the orientation data based on the map parameters and the group parameters includes: discarding one or more lines associated with the first object and / or objects in the group based on the one or more line patterns.

[0308] Example 49. The system of any one of Examples 47 to 48, wherein determining the orientation data based on the map parameter and the group parameter comprises:

[0309] aligning objects in the group and at least one of the first objects based on the one or more line patterns; and

[0310] The orientation data is determined based on the alignment.

[0311] Example 50. The system of any one of Examples 43 to 49, wherein the operations further comprise:

[0312] Using the at least one processor, map layer information is used to improve the object detection data based on the map parameters and / or the group parameters.

[0313] Example 51. The system of any one of Examples 43 to 50, wherein the operations further comprise:

[0314] Using the at least one processor, a non-maximum suppression scheme is performed on the object detection data based on the group parameters.

[0315] Example 52. The system of any one of Examples 43 to 51, wherein the operations further comprise:

[0316] Using the at least one processor, one or more objects in the environment are tracked based on the orientation data and the sensor data.

[0317] Example 53. The system of any one of Examples 44 to 52, wherein the operations further comprise:

[0318] Using the at least one processor, control data for controlling the autonomous vehicle is determined based on the sensor data and the orientation data.

[0319] Example 54. The system of any one of Examples 44 to 53, wherein the operations further comprise:

[0320] Using the at least one processor, a determination is made to label the at least one object in the environment based on the orientation data and the sensor data.

[0321] Example 55. The system of any one of Examples 43 to 54, wherein the operations further comprise:

[0322] Using the at least one processor, updating a machine learning model based on the orientation data.

[0323] Example 56. The system of Example 55, wherein updating the machine learning model comprises:

[0324] inputting one or more of the orientation data, one or more object parameters, the group parameters, and the map parameters into the machine learning model;

[0325] outputting updated orientation data from the machine learning model; and

[0326] Instead of the orientation data, the updated orientation data is recursively applied.

[0327] Example 57. The system of any one of Examples 43 to 56, wherein the operations further comprise:

[0328] obtaining, using the at least one processor, one or more estimated object parameters;

[0329] using the at least one processor, comparing the one or more estimated object parameters to the orientation data;

[0330] determining, using the at least one processor, a difference parameter indicative of an orientation difference between the one or more estimated object parameters and the orientation data based on the comparing; and

[0331] Using the at least one processor, the orientation data is updated based on the differential parameter.

[0332] Example 58. The system of any one of Examples 43 to 57, wherein the operations further comprise:

[0333] estimating, using the at least one processor, a probability distribution of the orientation data based on the map parameters and / or the group parameters; and

[0334] The object detection data is determined, using the at least one processor, based on a probability distribution of the orientation data.

[0335] Example 59. The system of any one of Examples 43 to 58, wherein the operations further comprise:

[0336] determining, using the at least one processor, at least one ground truth object in the environment based on the sensor data;

[0337] determining, using the at least one processor, object orientation data indicative of a ground truth orientation of the at least one ground truth object based on the sensor data;

[0338] determining, using the at least one processor, a confidence parameter indicative of an orientation difference between the object orientation data and the orientation data based on a comparison of the object orientation data and the orientation data; and

[0339] Using the at least one processor, the orientation data is updated based on the confidence parameter.

[0340] Example 60. A system according to any one of Examples 43 to 59, wherein the map parameter indicates an area of ​​the environment, and wherein the group parameter indicates the predetermined contact in the area.

[0341] Example 61. A system according to any one of Examples 43 to 60, wherein the first object is a first static object and the object is a second static object.

[0342] Example 62. A system according to any one of Examples 43 to 60, wherein the first object is a first moving object and the object is a second moving object.

[0343] Example 63. According to the system of any one of Examples 43 to 62, the operation also includes: determining the object detection data based on the orientation data.

Claims

1. A method, include: obtaining, using at least one processor, map parameters indicating a predetermined location of a first object in an environment in which the autonomous vehicle is configured to operate; obtaining, using the at least one processor, a group parameter indicating a predetermined connection between objects in a group, wherein the environment includes the objects in the group; determining, using the at least one processor, orientation data indicating an orientation of at least one of the objects in the group and the first object based on the map parameter and the group parameter; as well as Using the at least one processor, object detection data associated with the at least one object is provided to an apparatus based on the orientation data, wherein the object detection data indicates detection of one or more spatial features of the at least one object.

2. The method according to claim 1, further comprising: include: obtaining, using the at least one processor, sensor data associated with the environment, Wherein, determining the orientation data based on the map parameter and the group parameter includes: determining the orientation data based on the sensor data.

3. The method according to claim 2, in, Obtaining the group parameters includes: determining distances between objects in the group based on the sensor data; and The objects are clustered based on distances between the objects in the group to form the group.

4. The method according to any one of claims 2 and 3, in, Determining the orientation data based on the map parameter and the group parameter comprises: Based on the sensor data and the group parameters, one or more line patterns associated with the first object and / or objects in the group are extracted.

5. The method according to claim 4, in, Determining the orientation data based on the map parameter and the group parameter comprises: Based on the one or more line patterns, one or more lines associated with the first object and / or objects in the group are discarded.

6. The method according to any one of claims 4 and 5, in, Determining the orientation data based on the map parameter and the group parameter comprises: aligning objects in the group and at least one of the first objects based on the one or more line patterns; and The orientation data is determined based on the alignment.

7. The method according to any one of claims 2 to 6, further comprising: include: determining, using the at least one processor, at least one ground truth object in the environment based on the sensor data; determining, using the at least one processor, object orientation data indicative of a ground truth orientation of the at least one ground truth object based on the sensor data; determining, using the at least one processor, a confidence parameter indicative of an orientation difference between the object orientation data and the orientation data based on a comparison of the object orientation data and the orientation data; as well as Using the at least one processor, the orientation data is updated based on the confidence parameter.

8. The method according to any one of the preceding claims, further comprising: include: Groups are discarded from the object detection data based on the group parameters and the map parameters.

9. The method according to any one of the preceding claims, further comprising: include: Using the at least one processor, map layer information is used to improve the object detection data based on the map parameters and / or the group parameters.

10. The method according to any one of the preceding claims, further comprising: include: Using the at least one processor, a non-maximum suppression scheme is performed on the object detection data based on the group parameters.

11. The method according to any one of claims 2 to 10, further comprising: include: Using the at least one processor, one or more objects in the environment are tracked based on the orientation data and the sensor data.

12. The method according to any one of claims 2 to 11, further comprising: include: Using the at least one processor, control data for controlling the autonomous vehicle is determined based on the sensor data and the orientation data.

13. The method according to any one of claims 2 to 12, further comprising: include: Determining, using the at least one processor, to label the at least one object in the environment based on the orientation data and the sensor data.

14. The method according to any one of the preceding claims, further comprising: include: Using the at least one processor, updating a machine learning model based on the orientation data.

15. The method according to claim 14, in, Updating the machine learning model includes: inputting one or more of the orientation data, one or more object parameters, the group parameters, and the map parameters into the machine learning model; outputting updated orientation data from the machine learning model; and Instead of the orientation data, the updated orientation data is recursively applied.

16. The method according to any one of the preceding claims, further comprising: include: obtaining, using the at least one processor, one or more estimated object parameters; using the at least one processor, comparing the one or more estimated object parameters to the orientation data; determining, using the at least one processor, a difference parameter indicative of an orientation difference between the one or more estimated object parameters and the orientation data based on the comparing; as well as Using the at least one processor, the orientation data is updated based on the differential parameter.

17. The method according to any one of the preceding claims, further comprising: include: estimating, using the at least one processor, a probability distribution of the orientation data based on the map parameters and / or the group parameters; as well as The object detection data is determined, using the at least one processor, based on a probability distribution of the orientation data.

18. The method according to any one of the preceding claims, further comprising: include: The object detection data is determined based on the orientation data using the at least one processor.

19. A non-transitory computer readable medium comprising instructions stored on the medium, wherein when the instructions are executed by at least one processor, the at least one processor performs an operation, wherein the operation include: obtaining map parameters indicating a predetermined location of a first object in an environment in which the autonomous vehicle is configured to operate; obtaining a group parameter indicating a predetermined connection between objects in a group, wherein the environment includes the objects in the group; determining, based on the map parameter and the group parameter, orientation data indicating an orientation of at least one of the objects in the group and the first object; as well as The method causes object detection data associated with the at least one object to be provided to a device based on the orientation data, wherein the object detection data indicates detection of one or more spatial features of the at least one object.

20. A system comprising at least one processor and at least one memory, wherein instructions are stored on the at least one memory, and when the instructions are executed by the at least one processor, the at least one processor performs an operation, wherein the operation include: obtaining map parameters indicating a predetermined location of a first object in an environment in which the autonomous vehicle is configured to operate; obtaining a group parameter indicating a predetermined connection between second objects in the group, wherein the environment includes the second objects; determining, based on the map parameter and the group parameter, orientation data indicating an orientation of at least one of the first object and the second object; as well as The method causes object detection data associated with the at least one object to be provided to a device based on the orientation data, wherein the object detection data indicates detection of one or more spatial features of the at least one object.