Trace refinement network

Through the trace refinement network, the object bounding box is refined, the problem of object detection error in autonomous vehicles is solved, the detection accuracy and stability are improved, and the perception ability is enhanced.

CN120345009APending Publication Date: 2025-07-18MOTIONAL AD LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380085211.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-01
Filing Date
2023-10-10
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Object detection systems in autonomous vehicles are prone to errors, affecting driving performance.

Method used

The trace refinement network is adopted, and point cloud and trajectory features are extracted by generating the central box and trace window, the object bounding box is refined using a single model, and the central box is adjusted using residuals, which is suitable for multiple object classifications and tasks.

Benefits of technology

It improves the accuracy and stability of object detection, enhances the perception ability of autonomous vehicles, and reduces the impact of errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120345009A_ABST
    Figure CN120345009A_ABST
Patent Text Reader

Abstract

A method for refining a network for traces is provided. In an example, a center frame is obtained from a record of driving data, where the center frame is a center of a sequence of frames along a trace, and the trace is associated with a tracked object detected within the sequence of frames, each respective frame including a center, a size, and an orientation. Trace windows are generated around the respective center frames, wherein the trace windows correspond to the respective center frames along the trace. The trace window is cropped and normalized relative to the center frame to implement a single refinement model for multiple object classes. Point cloud features and trajectory features are extracted from the cut and normalized trace window. The point cloud features and trajectory features are input into a trace refinement network, where the trace refinement network uses features from the entire trace to output a refined center, refined size, and refined orientation for each respective center frame.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims the benefit of priority of U.S. Provisional Application 63 / 416,473, filed Oct. 14, 2022, and U.S. Application 18 / 073,104, filed Dec. 1, 2022, the entire contents of which are incorporated herein by reference. Background Art

[0003] Autonomous vehicles rely on perception of the surrounding environment to ensure safe and robust driving performance. Perception systems use object detection techniques to identify and locate objects in the environment. Errors can be introduced into object detection, which has a negative impact on driving performance. Brief Description of the Drawings

[0004] Figure 1 is an example environment of a vehicle that can implement one or more components of an autonomous system;

[0005] Figure 2 is a diagram of one or more systems of a vehicle including an autonomous system;

[0006] Figure 3 is Figure 1 and Figure 2 a diagram of one or more devices and / or components of one or more systems;

[0007] Figure 4A is a diagram of certain components of an autonomous system;

[0008] Figure 4B is a diagram of an implementation of a neural network;

[0009] Figure 4C and Figure 4D is a diagram illustrating an example operation of a CNN;

[0010] Figure 5 is a diagram of an implementation of a trace refinement network;

[0011] Figure 6 is an illustration of the architecture of a trace refinement network;

[0012] Figure 7A illustrates points associated with multiple boxes for encoding;

[0013] Figure 7B illustrates multiple boxes;

[0014] Figure 8 is an illustration of a vehicle reference coordinate system and a canvas;

[0015] Figure 9 is an illustration of classes and corresponding points for each class;

[0016] Figure 10 is a block diagram of a multi-task trace refinement network;

[0017] Figure 11 is an illustration of an auxiliary loss function for enhancing trace movement; and

[0018] Figure 12 is a flowchart enabling the processing of the trace refinement network. DETAILED DESCRIPTION

[0019] In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. It will be apparent, however, that the embodiments described herein may be practiced without these specific details. In some instances, well-known structures and devices are illustrated in block diagram form in order to avoid unnecessarily obscuring aspects of the present disclosure.

[0020] In the drawings, for ease of description, a specific arrangement or order of schematic elements (such as those representing systems, devices, modules, instruction blocks, and / or data elements, etc.) is illustrated. However, those skilled in the art will understand that, unless explicitly described, the specific order or arrangement of schematic elements in the drawings is not intended to imply a required processing order or sequence, or a separation of processes. Additionally, unless explicitly described, the inclusion of schematic elements in the drawings is not intended to imply that such elements are required in all embodiments, nor that the features represented by such elements cannot be included in some embodiments or combined with other elements in some embodiments.

[0021] Furthermore, in the drawings, connecting elements (such as solid lines, dashed lines, or arrows, etc.) are used to illustrate connections, relationships, or associations between two or more other schematic elements or within them. The absence of any such connecting element is not intended to imply that connections, relationships, or associations cannot exist. In other words, some connections, relationships, or associations between elements are not illustrated in the drawings so as not to obscure the present disclosure. Additionally, for ease of illustration, a single connecting element may be used to represent multiple connections, relationships, or associations between elements. For example, if a connecting element represents the communication of a signal, data, or instruction (e.g., "software instruction"), those skilled in the art should understand that such an element may represent one or more signal paths (e.g., a bus) that may be required to affect the communication.

[0022] Although terms such as "first", "second", and / or "third" are used to describe various elements, these elements should not be limited by these terms. The terms "first", "second", and / or "third" are only used to distinguish one element from another. For example, without departing from the scope of the described embodiments, a first contact may be referred to as a second contact, and similarly, a second contact may be referred to as a first contact. Both the first contact and the second contact are contacts, but they are not the same contact.

[0023] The terms used in the description of the various embodiments herein are included only for the purpose of describing specific embodiments and are not intended to be limiting. As used in the description of the various embodiments and the appended claims, the singular forms "a", "an", and "the" are also intended to include the plural forms and may be used interchangeably with "one or more than one" or "at least one", unless the context clearly indicates otherwise. It will also be understood that the term "and / or" as used herein refers to and includes any and all possible combinations of one or more of the associated listed items. It will also be understood that when the terms "comprises", "comprising", "includes", and / or "having" are used in this specification, it specifies the presence of the stated features, integers, steps, operations, elements, and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0024] As used herein, the terms "communicate" and "communicating" refer to at least one of receiving, receiving, transmitting, conveying, and / or providing information (or information represented by, for example, data, signals, messages, instructions, and / or commands, etc.). For a unit (e.g., a device, a system, a component of a device or system, and / or a combination thereof, etc.) that is to communicate with another unit, this means that the unit can directly or indirectly receive information from the other unit and / or send (e.g., transmit) information to the other unit. This may refer to a direct or indirect connection that is essentially wired and / or wireless. Additionally, two units can communicate with each other even if the information transmitted between the first unit and the second unit is modified, processed, relayed, and / or routed. For example, even if the first unit receives information passively and does not actively transmit information to the second unit, the first unit can communicate with the second unit. As another example, if at least one intermediate unit (e.g., a third unit located between the first unit and the second unit) processes the information received from the first unit and transmits the processed information to the second unit, the first unit can communicate with the second unit. In some embodiments, a message may refer to a network packet (e.g., a data packet, etc.) that includes data.

[0025] As used herein, depending on the context, the term "if" is optionally interpreted to mean "when", "at the time of", "in response to determining that", and / or "in response to detecting", etc. Similarly, depending on the context, the phrase "if it has been determined" or "if [the stated condition or event] is detected" is optionally interpreted to mean "when determining...", "in response to determining that", or "at the time of detecting [the stated condition or event]" and / or "in response to detecting [the stated condition or event]", etc. Further, as used herein, terms such as "have", "having", or "possessing" are intended to be open-ended terms. Additionally, unless otherwise explicitly stated, the phrase "based on" is intended to mean "at least partially based on".

[0026] Reference will now be made in detail to the embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the various described embodiments. However, it will be apparent to those of ordinary skill in the art that the various described embodiments may be practiced without these specific details. In other instances, well-known methods, procedures, components, circuits, and networks have not been described in detail so as not to unnecessarily obscure aspects of the embodiments.

[0027] General Overview

[0028] In some aspects and / or embodiments, the systems, methods, and computer program products described herein include and / or implement a trace refinement network. A vehicle (such as an autonomous vehicle, etc.) includes a system that implements driverless functions, including a perception system. The perception system includes a machine learning system trained using a large-scale dataset. An offline perception system is used to generate an automatically labeled training dataset. In an example, a center box is obtained by the offline perception system. A trace window is generated around the corresponding center box. Features are extracted from the trace window using an encoded box alignment and a normalized canvas. The features are input into the trace refinement network, which outputs a refined box (e.g., refined box center, refined box size, refined box orientation) and trace attributes (e.g., speed, motion state). The refined box is deployed for use in the perception system. In an example, the refined center box is used in online perception. In an example, the refined center box is used to implement image-LiDAR fusion. In an example, the refined box is used to automatically generate a high-quality labeled dataset.

[0029] Implementations of the systems, methods, and computer program products described herein implement techniques for a trace refinement network. Some advantages of these techniques include generating a single model (e.g., a trace refinement network) to refine bounding boxes corresponding to objects with different classifications. Additionally, the trace refinement network uses the residuals between the original center box captured by the trace window and the canvas-aligned box with size normalization to determine the refinement. These techniques do not regress the extracted features from scratch (e.g., do not utilize or rely on any previously computed bounding boxes). Instead, the bounding boxes are used to determine the residuals that are regressed to refine the center box. Additionally, a single model is applicable to multiple object classifications and multiple tasks.

[0030] Now refer to Figure 1 , exemplary environment 100 is illustrated, in which vehicles including autonomous systems and vehicles not including autonomous systems operate. As illustrated, environment 100 includes vehicles 102a - 102n, objects 104a - 104n, routes 106a - 106n, area 108, vehicle-to-infrastructure (V2I) devices 110, network 112, remote autonomous vehicle (AV) system 114, queue management system 116, and V2I system 118. Vehicles 102a - 102n, vehicle-to-infrastructure (V2I) devices 110, network 112, autonomous vehicle (AV) system 114, queue management system 116, and V2I system 118 are interconnected via a wired connection, a wireless connection, or a combination of wired and wireless connections (e.g., establishing a connection for communication, etc.). In some embodiments, objects 104a - 104n are interconnected with at least one of vehicles 102a - 102n, vehicle-to-infrastructure (V2I) devices 110, network 112, autonomous vehicle (AV) system 114, queue management system 116, and V2I system 118 via a wired connection, a wireless connection, or a combination of wired and wireless connections.

[0031] Vehicles 102a - 102n (individually referred to as vehicle 102 and collectively as vehicles 102) include at least one device configured to transport goods and / or people. In some embodiments, vehicle 102 is configured to communicate with V2I device 110, remote AV system 114, queue management system 116, and / or V2I system 118 via network 112. In some embodiments, vehicle 102 includes cars, buses, trucks, and / or trains, etc. In some embodiments, vehicle 102 is the same as vehicle 200 described herein (see Figure 2) are the same or similar. In some embodiments, the vehicle 200 in the set of vehicles 200 is associated with an autonomous platoon manager. In some embodiments, as described herein, the vehicle 102 travels along corresponding routes 106a - 106n (individually referred to as route 106 and collectively referred to as routes 106). In some embodiments, one or more than one vehicle 102 includes an autonomous system (e.g., an autonomous system that is the same as or similar to the autonomous system 202).

[0032] The objects 104a - 104n (individually referred to as object 104 and collectively referred to as objects 104) include, for example, at least one vehicle, at least one pedestrian, at least one cyclist, and / or at least one structure (e.g., a building, a sign, a fire hydrant, etc.). Each object 104 is stationary (e.g., located at a fixed location and over a period of time) or moving (e.g., having a speed and associated with at least one trajectory). In some embodiments, the object 104 is associated with a corresponding location in the region 108.

[0033] The routes 106a - 106n (individually referred to as route 106 and collectively referred to as routes 106) are each associated with (e.g., define) a series of actions (also referred to as a trajectory) that connect the states along which the AV can navigate. Each route 106 begins at an initial state (e.g., a state corresponding to a first spatio - temporal location and / or speed, etc.) and ends at a final target state (e.g., a state corresponding to a second spatio - temporal location different from the first spatio - temporal location) or a target zone (e.g., a subspace of acceptable states (e.g., a termination state)). In some embodiments, the first state includes a location where one or more individuals will board the AV, and the second state or zone includes one or more locations where one or more individuals boarding the AV will disembark. In some embodiments, the route 106 includes multiple sequences of acceptable states (e.g., multiple sequences of spatio - temporal locations), which are associated with (e.g., define) multiple trajectories. In an example, the route 106 includes only high - level actions or imprecise state locations, such as a series of connected roads indicating a direction change at a roadway intersection. Additionally or alternatively, the route 106 can include more precise actions or states, such as, for example, a specific target lane or precise location within a lane area and a target rate at those locations. In an example, the route 106 includes multiple sequences of precise states along at least one high - level action with a finite look - ahead horizon to reach an intermediate target, where the combination of consecutive iterations of the finite - horizon state sequences cumulatively corresponds to multiple trajectories that together form a high - level route terminating at the final target state or zone.

[0034] Region 108 includes a physical region (e.g., a geographical region) that the vehicle 102 can navigate. In an example, region 108 includes at least one state (e.g., a country, a province, an individual state among multiple states included in a country, etc.), at least a portion of a state, at least one city, at least a portion of a city, etc. In some embodiments, region 108 includes at least one named thoroughfare (referred to herein as a "road"), such as a highway, an interstate highway, a parkway, a city street, etc. Additionally or alternatively, in some examples, region 108 includes at least one unnamed road, such as a lane, a section of a parking lot, a section of a vacant and / or undeveloped area, a dirt road, etc. In some embodiments, a road includes at least one lane (e.g., a portion of the road that the vehicle 102 can traverse). In an example, a road includes at least one lane associated with (e.g., identified based on) at least one lane marking line.

[0035] A vehicle-to-infrastructure (V2I) device 110 (sometimes referred to as a vehicle-to-infrastructure or vehicle-to-everything (V2X) device) includes at least one device configured to communicate with the vehicle 102 and / or the V2I system 118. In some embodiments, the V2I device 110 is configured to communicate with the vehicle 102, the remote AV system 114, the platoon management system 116, and / or the V2I system 118 via the network 112. In some embodiments, the V2I device 110 includes a radio frequency identification (RFID) device, a sign, a camera (e.g., a two-dimensional (2D) and / or three-dimensional (3D) camera), a lane marking, a street light, a parking meter, etc. In some embodiments, the V2I device 110 is configured to communicate directly with the vehicle 102. Additionally or alternatively, in some embodiments, the V2I device 110 is configured to communicate with the vehicle 102, the remote AV system 114, and / or the platoon management system 116 via the V2I system 118. In some embodiments, the V2I device 110 is configured to communicate with the V2I system 118 via the network 112.

[0036] The network 112 includes one or more wired and / or wireless networks. In an example, the network 112 includes a cellular network (e.g., a Long Term Evolution (LTE) network, a third generation (3G) network, a fourth generation (4G) network, a fifth generation (5G) network, a Code Division Multiple Access (CDMA) network, etc.), a Public Land Mobile Network (PLMN), a Local Area Network (LAN), a Wide Area Network (WAN), a Metropolitan Area Network (MAN), a telephone network (e.g., a Public Switched Telephone Network (PSTN)), a private network, an ad hoc network, an intranet, the Internet, a fiber-based network, a cloud computing network, etc., and / or a combination of some or all of these networks.

[0037] The remote AV system 114 includes at least one device configured to communicate with the vehicle 102, the V2I device 110, the network 112, the queue management system 116, and / or the V2I system 118 via the network 112. In an example, the remote AV system 114 includes a server, a server group, and / or other similar devices. In some embodiments, the remote AV system 114 is co-located with the queue management system 116. In some embodiments, the remote AV system 114 participates in the installation of some or all of the components of the vehicle (including autonomous systems, autonomous vehicle computing, and / or software implemented by autonomous vehicle computing, etc.). In some embodiments, the remote AV system 114 maintains (e.g., updates and / or replaces) these components and / or software during the life of the vehicle.

[0038] The queue management system 116 includes at least one device configured to communicate with the vehicle 102, the V2I device 110, the remote AV system 114, and / or the V2I system 118. In an example, the queue management system 116 includes a server, a server group, and / or other similar devices. In some embodiments, the queue management system 116 is associated with a ride-sharing company (e.g., an organization for controlling the operation of multiple vehicles (e.g., vehicles including autonomous systems and / or vehicles not including autonomous systems), etc.).

[0039] In some embodiments, the V2I system 118 includes at least one device configured to communicate with the vehicle 102, the V2I device 110, the remote AV system 114, and / or the queue management system 116 via the network 112. In some examples, the V2I system 118 is configured to communicate with the V2I device 110 via a connection different from the network 112. In some embodiments, the V2I system 118 includes a server, a server group, and / or other similar devices. In some embodiments, the V2I system 118 is associated with a municipal authority or a private institution (e.g., a private institution for maintaining the V2I device 110, etc.).

[0040] Provide Figure 1 The number and arrangement of the illustrated elements are provided as examples. Compared with Figure 1 the illustrated elements, there may be additional elements, fewer elements, different elements, and / or elements with different arrangements. Additionally or alternatively, at least one element of the environment 100 may perform one or more functions described as being performed by Figure 1 at least one different element. Additionally or alternatively, at least one set of elements of the environment 100 may perform one or more functions described as being performed by at least one different set of elements of the environment 100.

[0041] Now refer to Figure 2 , vehicle 200 (which may be the same as or similar to Figure 1 vehicle 102) includes an autonomous system 202, a powertrain control system 204, a steering control system 206, and a braking system 208, or is associated with the autonomous system 202, the powertrain control system 204, the steering control system 206, and the braking system 208. In some embodiments, vehicle 200 is the same as or similar to vehicle 102 (see Figure 1 ). In some embodiments, the autonomous system 202 is configured to endow vehicle 200 with autonomous driving capabilities (e.g., implement at least one of the following driving automation or maneuver-based functions, features, and / or devices, etc., the at least one driving automation or maneuver-based function, feature, and / or device enabling vehicle 200 to operate partially or fully without human intervention, including but not limited to fully autonomous vehicles (e.g., vehicles that abandon reliance on human intervention, such as level 5 ADS-operated vehicles, etc.), highly autonomous vehicles (e.g., vehicles that abandon reliance on human intervention in certain situations, such as level 4 ADS-operated vehicles, etc.), and / or conditionally autonomous vehicles (e.g., vehicles that abandon reliance on human intervention in limited situations, such as level 3 ADS-operated vehicles, etc.), etc.). In one embodiment, the autonomous system 202 includes the operations or tactical functionality required to operate vehicle 200 in on-road traffic and continuously perform a part or all of the dynamic driving task (DDT). In another embodiment, the autonomous system 202 includes an advanced driver assistance system (ADAS) that includes driver support features. The autonomous system 202 supports various levels of driving automation ranging from no driving automation (e.g., level 0) to full driving automation (e.g., level 5). For a detailed description of fully autonomous vehicles and highly autonomous vehicles, reference can be made to SAE International Standard J3016: Taxonomy and Definitions for Terms Related to On-Road Motor Vehicle Automated Driving Systems, the entire content of which is incorporated by reference. In some embodiments, vehicle 200 is associated with an autonomous queue manager and / or a ridesharing company.

[0042] The autonomous system 202 includes a sensor suite that includes one or more devices such as a camera 202a, a LiDAR sensor 202b, a Radar sensor 202c, and a microphone 202d. In some embodiments, the autonomous system 202 may include more or fewer devices and / or different devices (e.g., ultrasonic sensors, inertial sensors, GPS receivers (discussed below), and / or odometer sensors for generating data associated with an indication of the distance traveled by the vehicle 200, etc.). In some embodiments, the autonomous system 202 uses one or more devices included in the autonomous system 202 to generate data associated with the environment 100 described herein. The data generated by one or more devices of the autonomous system 202 can be used by one or more systems described herein to observe the environment (e.g., environment 100) in which the vehicle 200 is located. In some embodiments, the autonomous system 202 includes a communication device 202e, an autonomous vehicle computing 202f, a drive-by-wire (DBW) system 202h, and a safety controller 202g.

[0043] The camera 202a includes at least one device configured to communicate with the communication device 202e, the autonomous vehicle computing 202f, and / or the safety controller 202g via a bus (e.g., a bus 302 that is the same as or similar to Figure 3 the bus). The camera 202a includes at least one camera (e.g., a digital camera using an optical sensor such as a charge-coupled device (CCD), a thermal camera, an infrared (IR) camera, and / or an event camera, etc.) for capturing images including physical objects (e.g., cars, buses, curbs, and / or people, etc.). In some embodiments, the camera 202a generates camera data as output. In some examples, the camera 202a generates camera data including image data associated with the image. In this example, the image data may specify at least one parameter corresponding to the image (e.g., image characteristics such as exposure, brightness, etc., and / or an image timestamp, etc.). In such an example, the image may be in a format (e.g., RAW, JPEG, and / or PNG, etc.). In some embodiments, the camera 202a includes a plurality of independent cameras configured (e.g., positioned) on the vehicle for capturing images for the purpose of stereovision (stereo vision). In some examples, the camera 202a includes generating image data and transmitting the image data to the autonomous vehicle computing 202f and / or a queue management system (e.g., the same as Figure 1a plurality of cameras of the same or similar queue management system as queue management system 116. In such an example, the autonomous vehicle computing 202f determines the depth to one or more objects in the fields of view of at least two of the plurality of cameras based on image data from at least two cameras. In some embodiments, the camera 202a is configured to capture images of objects within a distance relative to the camera 202a (e.g., up to 100 meters and / or up to 1 kilometer, etc.). Thus, the camera 202a includes features such as sensors and lenses optimized for sensing objects at one or more distances relative to the camera 202a.

[0044] In an embodiment, the camera 202a includes at least one camera configured to capture one or more images associated with one or more traffic lights, street signs, and / or other physical objects that provide visual navigation information. In some embodiments, the camera 202a generates traffic light data associated with one or more images. In some examples, the camera 202a generates TLD (Traffic Light Detection) data associated with one or more images including a format (e.g., RAW, JPEG, and / or PNG, etc.). In some embodiments, the camera 202a that generates TLD data is different from other systems incorporating cameras described herein in that the camera 202a may include one or more cameras having a wide field of view (e.g., a wide-angle lens, a fish-eye lens, and / or a lens having a viewing angle of about 120 degrees or greater, etc.) to generate images related to as many physical objects as possible.

[0045] The Light Detection and Ranging (LiDAR) sensor 202b includes being configured to communicate with the communication device 202e, the autonomous vehicle computing 202f, and / or the safety controller 202g via a bus (e.g., with Figure 3at least one device configured to communicate via a bus (e.g., a bus identical or similar to bus 302). The LiDAR sensor 202b includes a system configured to emit light from a light emitter (e.g., a laser emitter). The light emitted by the LiDAR sensor 202b includes light outside the visible spectrum (e.g., infrared light, etc.). In some embodiments, during operation, the light emitted by the LiDAR sensor 202b encounters a physical object (e.g., a vehicle) and is reflected back to the LiDAR sensor 202b. In some embodiments, the light emitted by the LiDAR sensor 202b does not penetrate the physical object it encounters. The LiDAR sensor 202b further includes at least one light detector that detects the light after the light emitted from the light emitter encounters a physical object. In some embodiments, at least one data processing system associated with the LiDAR sensor 202b generates an image (e.g., a point cloud and / or a combined point cloud, etc.) representing the objects included in the field of view of the LiDAR sensor 202b. In some examples, at least one data processing system associated with the LiDAR sensor 202b generates an image representing the boundary of a physical object and / or the surface of a physical object (e.g., the topology of the surface), etc. In such examples, the image is used to determine the boundary of the physical object in the field of view of the LiDAR sensor 202b.

[0046] A Radio Detection and Ranging (Radar) sensor 202c includes at least one device configured to communicate with a communication device 202e, an autonomous vehicle computer 202f, and / or a safety controller 202g via a bus (e.g., a bus identical or similar to Figure 3 bus 302). The Radar sensor 202c includes a system configured to emit (pulsed or continuous) radio waves. The radio waves emitted by the Radar sensor 202c include radio waves within a predetermined spectrum. In some embodiments, during operation, the radio waves emitted by the Radar sensor 202c encounter a physical object and are reflected back to the Radar sensor 202c. In some embodiments, the radio waves emitted by the Radar sensor 202c are not reflected by some objects. In some embodiments, at least one data processing system associated with the Radar sensor 202c generates a signal representing the objects included in the field of view of the Radar sensor 202c. For example, at least one data processing system associated with the Radar sensor 202c generates an image representing the boundary of a physical object and / or the surface of a physical object (e.g., the topology of the surface), etc. In some examples, the image is used to determine the boundary of the physical object in the field of view of the Radar sensor 202c.

[0047] The microphone 202d includes at least one device configured to communicate with the communication device 202e, the autonomous vehicle computing 202f, and / or the safety controller 202g via a bus (e.g., a bus the same as or similar to the bus 302 of Figure 3 ). The microphone 202d includes one or more microphones (e.g., an array microphone and / or an external microphone, etc.) that capture an audio signal and generate data associated with (e.g., representing) the audio signal. In some examples, the microphone 202d includes a transducer device and / or a similar device. In some embodiments, one or more of the systems described herein may receive the data generated by the microphone 202d and determine the position (e.g., distance, etc.) of an object relative to the vehicle 200 based on the audio signal associated with the data.

[0048] The communication device 202e includes at least one device configured to communicate with the camera 202a, the LiDAR sensor 202b, the Radar sensor 202c, the microphone 202d, the autonomous vehicle computing 202f, the safety controller 202g, and / or the DBW (drive-by-wire) system 202h. For example, the communication device 202e may include a device the same as or similar to the communication interface 314 of Figure 3 . In some embodiments, the communication device 202e includes a vehicle-to-vehicle (V2V) communication device (e.g., a device for enabling wireless communication of data between vehicles).

[0049] The autonomous vehicle computing 202f includes at least one device configured to communicate with the camera 202a, the LiDAR sensor 202b, the Radar sensor 202c, the microphone 202d, the communication device 202e, the safety controller 202g, and / or the DBW system 202h. In some examples, the autonomous vehicle computing 202f includes devices such as a client device, a mobile device (e.g., a cellular phone and / or a tablet, etc.), and / or a server (e.g., a computing device including one or more central processing units and / or graphics processing units, etc.). In some embodiments, the autonomous vehicle computing 202f is the same as or similar to the autonomous vehicle computing 400 described herein. Additionally or alternatively, in some embodiments, the autonomous vehicle computing 202f is configured to communicate with an autonomous vehicle system (e.g., an autonomous vehicle system the same as or similar to the remote AV system 114 of Figure 1 ), a queue management system (e.g., a queue management system the same as or similar to the queue management system 116 of Figure 1 ), a V2I device (e.g., a V2I device the same as or similar to the V2I device 110 of Figure 1 ), and / or a V2I system (e.g., a V2I system the same as or similar to the Figure 1communicate with a V2I system 118 that is the same as or similar to the V2I system).

[0050] The safety controller 202g includes at least one device configured to communicate with the camera 202a, the LiDAR sensor 202b, the Radar sensor 202c, the microphone 202d, the communication device 202e, the autonomous vehicle computing 202f, and / or the DBW system 202h. In some examples, the safety controller 202g includes one or more controllers (such as an electrical controller and / or an electromechanical controller, etc.) configured to generate and / or transmit control signals to operate one or more devices of the vehicle 200 (such as the powertrain control system 204, the steering control system 206, and / or the braking system 208, etc.). In some embodiments, the safety controller 202g is configured to generate control signals that are prior to (e.g., override) the control signals generated and / or transmitted by the autonomous vehicle computing 202f.

[0051] The DBW system 202h includes at least one device configured to communicate with the communication device 202e and / or the autonomous vehicle computing 202f. In some examples, the DBW system 202h includes one or more controllers (such as an electrical controller and / or an electromechanical controller, etc.) configured to generate and / or transmit control signals to operate one or more devices of the vehicle 200 (such as the powertrain control system 204, the steering control system 206, and / or the braking system 208, etc.). Additionally or alternatively, one or more controllers of the DBW system 202h are configured to generate and / or transmit control signals to operate at least one different device of the vehicle 200 (such as turn signals, headlights, door locks, and / or windshield wipers, etc.).

[0052] The powertrain control system 204 includes at least one device configured to communicate with the DBW system 202h. In some examples, the powertrain control system 204 includes at least one controller and / or actuator, etc. In some embodiments, the powertrain control system 204 receives control signals from the DBW system 202h, and the powertrain control system 204 causes the vehicle 200 to perform longitudinal vehicle movement (such as starting to move forward, stopping moving forward, starting to move backward, stopping moving backward, accelerating in a certain direction, decelerating in a certain direction, etc.), or causes the vehicle 200 to perform lateral vehicle movement (such as making a left turn and / or making a right turn, etc.). In an example, the powertrain control system 204 increases, maintains the same, or decreases the energy (such as fuel and / or electricity, etc.) provided to the motor of the vehicle, thereby causing at least one wheel of the vehicle 200 to rotate or not rotate.

[0053] The steering control system 206 includes at least one device configured to rotate one or more wheels of the vehicle 200. In some examples, the steering control system 206 includes at least one controller and / or actuator, etc. In some embodiments, the steering control system 206 rotates two front wheels and / or two rear wheels of the vehicle 200 left or right to turn the vehicle 200 left or right. In other words, the steering control system 206 causes the activities required to regulate the y-axis component of the vehicle's movement.

[0054] The braking system 208 includes at least one device configured to actuate one or more brakes to decelerate the vehicle 200 and / or keep it stationary. In some examples, the braking system 208 includes at least one controller and / or actuator configured to close one or more calipers associated with one or more wheels of the vehicle 200 on the corresponding rotors of the vehicle 200. Additionally or alternatively, in some examples, the braking system 208 includes an automatic emergency braking (AEB) system and / or a regenerative braking system, etc.

[0055] In some embodiments, the vehicle 200 includes at least one platform sensor (not explicitly illustrated) for measuring or inferring the nature of the state or condition of the vehicle 200. In some examples, the vehicle 200 includes platform sensors such as a global positioning system (GPS) receiver, an inertial measurement unit (IMU), a wheel speed sensor, a wheel brake pressure sensor, a wheel torque sensor, an engine torque sensor, and / or a steering angle sensor. Although the braking system 208 is illustrated as being located Figure 2 proximal to the vehicle 200 in, the braking system 208 can be located anywhere in the vehicle 200.

[0056] Now refer to Figure 3, A schematic diagram of an exemplary device 300. As illustrated, device 300 includes a processor 304, a memory 306, a storage component 308, an input interface 310, an output interface 312, a communication interface 314, and a bus 302. In some embodiments, device 300 corresponds to: at least one device of vehicle 102 (e.g., at least one device of a system of vehicle 102), at least one device of a remote AV system 114, at least one device of a queue management system 116, at least one device of a V2I system 118, and / or one or more devices of network 112 (e.g., one or more devices of a system of network 112). In some embodiments, one or more devices of vehicle 102 (e.g., one or more devices of a system of vehicle 102, at least one device of a remote AV system 114, at least one device of a queue management system 116, at least one device of a V2I system 118, and / or one or more devices of network 112 (e.g., one or more devices of a system of network 112) include at least one device 300 and / or at least one component of device 300. As Figure 3 shown, device 300 includes a bus 302, a processor 304, a memory 306, a storage component 308, an input interface 310, an output interface 312, and a communication interface 314.

[0057] Bus 302 includes components that permit communication between components of device 300. In some cases, processor 304 includes a processor (e.g., a central processing unit (CPU), a graphics processing unit (GPU), and / or an accelerated processing unit (APU), etc.), a microphone, a digital signal processor (DSP), and / or any processing component that can be programmed to perform at least one function (e.g., a field programmable gate array (FPGA) and / or an application specific integrated circuit (ASIC), etc.). Memory 306 includes random access memory (RAM), read only memory (ROM), and / or another type of dynamic and / or static storage device that stores data and / or instructions for use by processor 304 (e.g., flash memory, magnetic memory, and / or optical memory, etc.).

[0058] Storage component 308 stores data and / or software related to the operation and use of device 300. In some examples, storage component 308 includes a hard disk (e.g., a magnetic disk, an optical disk, a magneto-optical disk, and / or a solid state disk, etc.), a compact disk (CD), a digital versatile disk (DVD), a floppy disk, a cassette tape, a magnetic tape, a CD-ROM, a RAM, a PROM, an EPROM, a FLASH-EPROM, an NV-RAM, and / or another type of computer readable medium, and a corresponding drive.

[0059] The input interface 310 includes components of the enabling device 300 that receive information via a user input (e.g., a touchscreen display, a keyboard, a keypad, a mouse, a button, a switch, a microphone, and / or a camera, etc.). Additionally or alternatively, in some embodiments, the input interface 310 includes sensors for sensing information (e.g., a Global Positioning System (GPS) receiver, an accelerometer, a gyroscope, and / or an actuator, etc.). The output interface 312 includes components for providing output information from the device 300 (e.g., a display, a speaker, and / or one or more Light Emitting Diodes (LEDs), etc.).

[0060] In some embodiments, the communication interface 314 includes transceiver-like components (e.g., a transceiver and / or separate receiver and transmitter, etc.) that enable the device 300 to communicate with other devices via a wired connection, a wireless connection, or a combination of a wired connection and a wireless connection. In some examples, the communication interface 314 enables the device 300 to receive information from another device and / or provide information to another device. In some examples, the communication interface 314 includes an Ethernet interface, an optical interface, a coaxial interface, an infrared interface, a Radio Frequency (RF) interface, a Universal Serial Bus (USB) interface, a Wi-Fi interface, and / or a cellular network interface, etc.

[0061] In some embodiments, the device 300 performs one or more processes described herein. The device 300 performs these processes based on software instructions stored by a computer-readable medium such as the memory 306 and / or the storage component 308, executed by the processor 304. A computer-readable medium (e.g., a non-transitory computer-readable medium) is defined herein as a non-transitory memory device. The non-transitory memory device includes a storage space located within a single physical storage device or a storage space distributed across multiple physical storage devices.

[0062] In some embodiments, software instructions are read into the memory 306 and / or the storage component 308 from another computer-readable medium or from another device via the communication interface 314. When executed, the software instructions stored in the memory 306 and / or the storage component 308 cause the processor 304 to perform one or more processes described herein. Additionally or alternatively, hardwired circuitry is used instead of or in combination with software instructions to perform one or more processes described herein. Thus, unless otherwise explicitly stated, the embodiments described herein are not limited to any particular combination of hardware circuitry and software.

[0063] The memory 306 and / or the storage component 308 includes a data store or at least one data structure (e.g., a database, etc.). The apparatus 300 is capable of receiving information from the data store or at least one data structure in the memory 306 or the storage component 308, storing the information in the data store or at least one data structure, communicating the information to the data store or at least one data structure, or searching for the information stored in the data store or at least one data structure. In some examples, the information includes network data, input data, output data, or any combination thereof.

[0064] In some embodiments, the apparatus 300 is configured to execute software instructions stored in the memory 306 and / or the memory of another apparatus (e.g., another apparatus that is the same as or similar to the apparatus 300). As used herein, the term "module" refers to at least one instruction stored in the memory 306 and / or the memory of another apparatus, which, when executed by the processor 304 and / or the processor of another apparatus (e.g., another apparatus that is the same as or similar to the apparatus 300), causes the apparatus 300 (e.g., at least one component of the apparatus 300) to perform one or more processes described herein. In some embodiments, the module is implemented in software, firmware, and / or hardware, etc.

[0065] Provided Figure 3 The number and arrangement of the illustrated components are provided as examples. In some embodiments, compared to Figure 3 the illustrated components, the apparatus 300 may include additional components, fewer components, different components, or components arranged differently. Additionally or alternatively, a set of components of the apparatus 300 (e.g., one or more components) may perform one or more functions described as being performed by another component or another set of components of the apparatus 300.

[0066] Now refer to Figure 4A, an example block diagram of an autonomous vehicle computing 400 (sometimes referred to as an "AV stack") is illustrated. As illustrated, the autonomous vehicle computing 400 includes a perception system 402 (sometimes referred to as a perception module), a planning system 404 (sometimes referred to as a planning module), a localization system 406 (sometimes referred to as a localization module), a control system 408 (sometimes referred to as a control module), and a database 410. In some embodiments, the perception system 402, the planning system 404, the localization system 406, the control system 408, and the database 410 are included in and / or implemented in the vehicle's automatic navigation system (e.g., the autonomous vehicle computing 202f of the vehicle 200). Additionally or alternatively, in some embodiments, the perception system 402, the planning system 404, the localization system 406, the control system 408, and the database 410 are included in one or more independent systems (e.g., one or more systems identical or similar to the autonomous vehicle computing 400, etc.). In some examples, the perception system 402, the planning system 404, the localization system 406, the control system 408, and the database 410 are included in one or more independent systems located in the vehicle and / or at least one remote system as described herein. In some embodiments, any and / or all of the systems included in the autonomous vehicle computing 400 are implemented in software (e.g., software instructions stored in memory), computer hardware (e.g., via a microprocessor, a microcontroller, an application-specific integrated circuit (ASIC), and / or a field-programmable gate array (FPGA), etc.), or a combination of computer software and computer hardware. It will also be understood that, in some embodiments, the autonomous vehicle computing 400 is configured to communicate with remote systems (e.g., an autonomous vehicle system identical or similar to the remote AV system 114, a queue management system identical or similar to the queue management system 116, and / or a V2I system identical or similar to the V2I system 118, etc.).

[0067] In some embodiments, the perception system 402 receives data associated with at least one physical object in the environment (e.g., data used by the perception system 402 to detect at least one physical object), and classifies the at least one physical object. In some examples, the perception system 402 receives image data captured by at least one camera (e.g., camera 202a), the image being associated with one or more physical objects within the field of view of the at least one camera (e.g., representing the one or more physical objects). In such examples, the perception system 402 classifies the at least one physical object based on one or more groupings of physical objects (e.g., bicycles, vehicles, traffic signs, and / or pedestrians, etc.). In some embodiments, based on the classification of the physical objects by the perception system 402, the perception system 402 transmits data associated with the classification of the physical objects to the planning system 404.

[0068] In some embodiments, the planning system 404 receives data associated with a destination, and generates data associated with at least one route (e.g., route 106) along which a vehicle (e.g., vehicle 102) can travel towards the destination. In some embodiments, the planning system 404 periodically or continuously receives data from the perception system 402 (e.g., the data associated with the classification of the physical objects described above), and the planning system 404 updates at least one trajectory or generates at least one different trajectory based on the data generated by the perception system 402. In other words, the planning system 404 can perform tasks related to the tactical functions required to operate the vehicle 102 in on-road traffic. Tactical efforts involve maneuvering the vehicle in traffic during the journey, which includes but is not limited to deciding whether and when to overtake another vehicle, change lanes, or select an appropriate speed, acceleration, deceleration, etc. In some embodiments, the planning system 404 receives data associated with the updated position of the vehicle (e.g., vehicle 102) from the positioning system 406, and the planning system 404 updates at least one trajectory or generates at least one different trajectory based on the data generated by the positioning system 406.

[0069] In some embodiments, the positioning system 406 receives data associated with (e.g., representing) the location of a vehicle (e.g., vehicle 102) in an area. In some examples, the positioning system 406 receives LiDAR data associated with at least one point cloud generated by at least one LiDAR sensor (e.g., LiDAR sensor 202b). In certain examples, the positioning system 406 receives data associated with at least one point cloud from multiple LiDAR sensors, and the positioning system 406 generates a combined point cloud based on the respective point clouds. In these examples, the positioning system 406 compares the at least one point cloud or the combined point cloud with a two-dimensional (2D) and / or three-dimensional (3D) map of the area stored in the database 410. Then, based on the positioning system 406 comparing the at least one point cloud or the combined point cloud with the map, the positioning system 406 determines the position of the vehicle in the area. In some embodiments, the map includes a combined point cloud of the area generated prior to the navigation of the vehicle. In some embodiments, the map includes, but is not limited to, a high-precision map of the roadway geometry, a map describing the connection nature of the road network, a map describing the physical properties of the roadways (such as traffic speed, traffic flow, the number of vehicle and bicycle traffic lanes, lane width, lane traffic direction or the type and location of lane markings, or a combination thereof, etc.), and a map describing the spatial location of road features (such as crosswalks, traffic signs or various types of other driving signals, etc.). In some embodiments, the map is generated in real time based on the data received by the perception system.

[0070] In another example, the positioning system 406 receives global navigation satellite system (GNSS) data generated by a global positioning system (GPS) receiver. In some examples, the positioning system 406 receives GNSS data associated with the location of a vehicle in an area, and the positioning system 406 determines the latitude and longitude of the vehicle in the area. In such examples, the positioning system 406 determines the position of the vehicle in the area based on the latitude and longitude of the vehicle. In some embodiments, the positioning system 406 generates data associated with the position of the vehicle. In some examples, based on the positioning system 406 determining the position of the vehicle, the positioning system 406 generates data associated with the position of the vehicle. In such examples, the data associated with the position of the vehicle includes data associated with one or more semantic properties corresponding to the position of the vehicle.

[0071] In some embodiments, the control system 408 receives data associated with at least one trajectory from the planning system 404, and the control system 408 controls the operation of the vehicle. In some examples, the control system 408 receives data associated with at least one trajectory from the planning system 404, and the control system 408 controls the operation of the vehicle by generating and transmitting control signals to cause the powertrain control system (e.g., the DBW system 202h and / or the powertrain control system 204, etc.), the steering control system (e.g., the steering control system 206), and / or the braking system (e.g., the braking system 208) to operate. For example, the control system 408 is configured to perform operating functions such as lateral vehicle motion control or longitudinal vehicle motion control. Lateral vehicle motion control causes the activities required to regulate the y-axis component of the vehicle motion. Longitudinal vehicle motion control causes the activities required to regulate the x-axis component of the vehicle motion. In an example, in the case where the trajectory includes a left turn, the control system 408 transmits a control signal to cause the steering control system 206 to adjust the steering angle of the vehicle 200, thereby causing the vehicle 200 to turn left. Additionally or alternatively, the control system 408 generates and transmits control signals to cause other devices of the vehicle 200 (e.g., headlights, turn signals, door locks, and / or windshield wipers, etc.) to change states.

[0072] In some embodiments, the perception system 402, the planning system 404, the positioning system 406, and / or the control system 408 implement at least one machine learning model (e.g., at least one multi-layer perceptron (MLP), at least one convolutional neural network (CNN), at least one recurrent neural network (RNN), at least one autoencoder, and / or at least one transformer, etc.). In some examples, the perception system 402, the planning system 404, the positioning system 406, and / or the control system 408 implement at least one machine learning model alone or in combination with one or more of the above systems. In some examples, the perception system 402, the planning system 404, the positioning system 406, and / or the control system 408 implement at least one machine learning model as part of a pipeline (e.g., a pipeline for identifying one or more objects located in the environment, etc.). Examples of the implementation of the machine learning model are described below. Figures 4B to 4D Examples of the implementation of the machine learning model are described below.

[0073] The database 410 stores data transmitted to, received from, and / or updated by the perception system 402, the planning system 404, the positioning system 406, and / or the control system 408. In some examples, the database 410 includes a storage component for storing data and / or software related to operations and using the autonomous vehicle computing 400 of at least one system (e.g., related to Figure 3the same or similar storage components as the storage component 308). In some embodiments, the database 410 stores data associated with 2D and / or 3D maps of at least one area. In some examples, the database 410 stores data associated with 2D and / or 3D maps of a part of a city, multiple parts of multiple cities, multiple cities, counties, states, and / or countries (e.g., nations), etc. In such examples, a vehicle (e.g., a vehicle the same or similar to the vehicle 102 and / or the vehicle 200) can drive along one or more drivable areas (e.g., single-lane roads, multi-lane roads, highways, back roads, and / or off-road paths, etc.), and cause at least one LiDAR sensor (e.g., a LiDAR sensor the same or similar to the LiDAR sensor 202b) to generate data associated with an image representing the objects included in the field of view of the at least one LiDAR sensor.

[0074] In some embodiments, the database 410 can be implemented across multiple devices. In some examples, the database 410 is included in a vehicle (e.g., a vehicle the same or similar to the vehicle 102 and / or the vehicle 200), an autonomous vehicle system (e.g., an autonomous vehicle system the same or similar to the remote AV system 114), a queue management system (e.g., a queue management system the same or similar to Figure 1 the queue management system 116) and / or a V2I system (e.g., a V2I system the same or similar to Figure 1 the V2I system 118), etc.

[0075] Now refer to Figure 4B , a diagram illustrating the implementation of a machine learning model. More specifically, a diagram illustrating the implementation of a convolutional neural network (CNN) 420. For illustrative purposes, the following description of the CNN 420 will be with respect to implementing the CNN 420 by the perception system 402. However, it will be understood that in some examples, the CNN 420 (e.g., one or more components of the CNN 420) is implemented by other systems different from or in addition to the perception system 402, such as the planning system 404, the positioning system 406, and / or the control system 408, etc. Although the CNN 420 includes certain features as described herein, these features are provided for illustrative purposes and are not intended to limit the present disclosure.

[0076] The CNN 420 includes a plurality of convolutional layers including a first convolutional layer 422, a second convolutional layer 424, and a convolutional layer 426. In some embodiments, the CNN 420 includes a subsampling layer 428 (sometimes referred to as a pooling layer). In some embodiments, the subsampling layer 428 and / or other subsampling layers have dimensions smaller than the dimensions of the upstream system (i.e., the amount of nodes). By means of the subsampling layer 428 having dimensions smaller than the dimensions of the upstream layer, the CNN 420 combines the amount of data associated with the initial input and / or output of the upstream layer, thereby reducing the amount of computation required for the CNN 420 to perform downstream convolutional operations. Additionally or alternatively, by means of the subsampling layer 428 being associated with at least one subsampling function (e.g., being configured to perform at least one subsampling function) (as described below with respect to Figure 4C and Figure 4D ), the CNN 420 combines the amount of data associated with the initial input.

[0077] Based on the perception system 402 providing corresponding inputs and / or outputs associated with the first convolutional layer 422, the second convolutional layer 424, and the convolutional layer 426 respectively to generate corresponding outputs, the perception system 402 performs convolutional operations. In some examples, based on the perception system 402 providing data as inputs to the first convolutional layer 422, the second convolutional layer 424, and the convolutional layer 426, the perception system 402 implements the CNN 420. In such examples, based on the perception system 402 receiving data from one or more different systems (e.g., one or more systems of a vehicle the same or similar to the vehicle 102, a remote AV system the same or similar to the remote AV system 114, a queue management system the same or similar to the queue management system 116, and / or a V2I system the same or similar to the V2I system 118, etc.), the perception system 402 provides the data as inputs to the first convolutional layer 422, the second convolutional layer 424, and the convolutional layer 426. The following is a detailed description of Figure 4C including convolutional operations.

[0078] In some embodiments, the sensing system 402 provides data associated with an input (referred to as an initial input) to the first convolutional layer 422, and the sensing system 402 uses the first convolutional layer 422 to generate data associated with an output. In some embodiments, the sensing system 402 provides the output generated by the convolutional layer as an input to a different convolutional layer. For example, the sensing system 402 provides the output of the first convolutional layer 422 as an input to the subsampling layer 428, the second convolutional layer 424, and / or the convolutional layer 426. In such an example, the first convolutional layer 422 is referred to as an upstream layer, and the subsampling layer 428, the second convolutional layer 424, and / or the convolutional layer 426 are referred to as downstream layers. Similarly, in some embodiments, the sensing system 402 provides the output of the subsampling layer 428 to the second convolutional layer 424 and / or the convolutional layer 426, and in this example, the subsampling layer 428 will be referred to as an upstream layer, and the second convolutional layer 424 and / or the convolutional layer 426 will be referred to as downstream layers.

[0079] In some embodiments, before the sensing system 402 provides an input to the CNN 420, the sensing system 402 processes data associated with the input provided to the CNN 420. For example, based on the sensing system 402 normalizing sensor data (such as image data, LiDAR data, and / or Radar data, etc.), the sensing system 402 processes data associated with the input provided to the CNN 420.

[0080] In some embodiments, based on the sensing system 402 performing convolution operations associated with each convolutional layer, the CNN 420 generates an output. In some examples, based on the sensing system 402 performing convolution operations associated with each convolutional layer and the initial input, the CNN 420 generates an output. In some embodiments, the sensing system 402 generates an output and provides the output as the fully connected layer 430. In some examples, the sensing system 402 provides the output of the convolutional layer 426 as the fully connected layer 430, where the fully connected layer 430 includes data associated with a plurality of eigenvalues referred to as F1, F2,..., FN. In this example, the output of the convolutional layer 426 includes data associated with a plurality of output eigenvalues representing predictions.

[0081] In some embodiments, based on the perception system 402 identifying the eigenvalue associated with the highest likelihood of being the correct prediction among a plurality of predictions, the perception system 402 identifies a prediction from among the plurality of predictions. For example, in the case where the fully connected layer 430 includes eigenvalues F1, F2, ..., FN and F1 is the largest eigenvalue, the perception system 402 identifies the prediction associated with F1 as the correct prediction among the plurality of predictions. In some embodiments, the perception system 402 trains the CNN 420 to generate predictions. In some examples, based on the perception system 402 providing training data associated with a prediction to the CNN 420, the perception system 402 trains the CNN 420 to generate predictions.

[0082] Now refer Figure 4C and Figure 4D , a diagram illustrating an example operation of the CNN 440 that utilizes the perception system 402. In some embodiments, the CNN 440 (e.g., one or more components of the CNN 440) is the same as or similar to the CNN 420 (e.g., one or more components of the CNN 420) (see Figure 4B ).

[0083] In step 450, the perception system 402 provides data associated with an image as an input to the CNN 440 (step 450). For example, as illustrated, the perception system 402 provides data associated with an image to the CNN 440, where the image is a grayscale image represented as values stored in a two-dimensional (2D) array. In some embodiments, the data associated with the image may include data associated with a color image, which is represented as values stored in a three-dimensional (3D) array. Additionally or alternatively, the data associated with the image may include data associated with an infrared image and / or a Radar image, etc.

[0084] In step 455, the CNN 440 performs a first convolution function. For example, based on the CNN 440 providing the values representing the image as an input to one or more neurons (not explicitly illustrated) included in the first convolutional layer 442, the CNN 440 performs the first convolution function. In this example, the values representing the image may correspond to the values of a region (sometimes referred to as a receptive field) representing the image. In some embodiments, each neuron is associated with a filter (not explicitly illustrated). The filter (sometimes referred to as a kernel) may be represented as an array of values corresponding in size to the values provided as an input to the neuron. In one example, the filter may be configured to identify edges (e.g., horizontal lines, vertical lines, and / or straight lines, etc.). In successive convolutional layers, the filters associated with the neurons may be configured to successively identify more complex patterns (e.g., arcs and / or objects, etc.).

[0085] In some embodiments, based on the CNN 440, the values provided as input to each neuron among one or more neurons included in the first convolutional layer 442 are multiplied by the values of the filters corresponding to each neuron among the same one or more neurons, and the CNN 440 performs a first convolutional function. For example, the CNN 440 may multiply the values provided as input to each neuron among one or more neurons included in the first convolutional layer 442 by the values of the filters corresponding to each neuron among the same one or more neurons to generate a single value or an array of values as output. In some embodiments, the collective output of the neurons of the first convolutional layer 442 is referred to as the convolutional output. In some embodiments, when each neuron has the same filter, the convolutional output is referred to as the feature map.

[0086] In some embodiments, the CNN 440 provides the output of each neuron of the first convolutional layer 442 to the neurons of the downstream layer. For clarity, the upstream layer may be a layer that transmits data to a different layer (referred to as the downstream layer). For example, the CNN 440 may provide the output of each neuron of the first convolutional layer 442 to the corresponding neurons of the subsampling layer. In an example, the CNN 440 provides the output of each neuron of the first convolutional layer 442 to the corresponding neurons of the first subsampling layer 444. In some embodiments, the CNN 440 adds a bias value to the aggregated set of all values provided to each neuron of the downstream layer. For example, the CNN 440 adds a bias value to the aggregated set of all values provided to each neuron of the first subsampling layer 444. In such an example, the CNN 440 determines the final value to be provided to each neuron of the first subsampling layer 444 based on the aggregated set of all values provided to each neuron and the activation function associated with each neuron of the first subsampling layer 444.

[0087] In step 460, the CNN 440 performs a first subsampling function. For example, based on the CNN 440 providing the values output by the first convolutional layer 442 to the corresponding neurons of the first subsampling layer 444, the CNN 440 may perform a first subsampling function. In some embodiments, the CNN 440 performs the first subsampling function based on an aggregation function. In an example, based on the CNN 440 determining the maximum input among the values provided to a given neuron (referred to as the max pooling function), the CNN 440 performs the first subsampling function. In another example, based on the CNN 440 determining the average input among the values provided to a given neuron (referred to as the average pooling function), the CNN 440 performs the first subsampling function. In some embodiments, based on the CNN 440 providing values to each neuron of the first subsampling layer 444, the CNN 440 generates an output, which is sometimes referred to as the subsampled convolutional output.

[0088] At step 465, CNN 440 performs a second convolution function. In some embodiments, CNN 440 performs the second convolution function in a manner similar to how CNN 440 performs the first convolution function as described above. In some embodiments, based on the values output by the first subsampling layer 444 being provided as inputs to one or more neurons (not explicitly illustrated) included in the second convolutional layer 446, CNN 440 performs the second convolution function. In some embodiments, as described above, each neuron of the second convolutional layer 446 is associated with a filter. As described above, the (one or more) filters associated with the second convolutional layer 446 may be configured to identify more complex patterns compared to the filters associated with the first convolutional layer 442.

[0089] In some embodiments, based on CNN 440 multiplying the values provided as inputs to each of the one or more neurons included in the second convolutional layer 446 by the values of the filters corresponding to each of the one or more neurons, CNN 440 performs the second convolution function. For example, CNN 440 may multiply the values provided as inputs to each of the one or more neurons included in the second convolutional layer 446 by the values of the filters corresponding to each of the one or more neurons to generate a single value or an array of values as output.

[0090] In some embodiments, CNN 440 provides the output of each neuron of the second convolutional layer 446 to the neurons of a downstream layer. For example, CNN 440 may provide the output of each neuron of the first convolutional layer 442 to the corresponding neurons of the subsampling layer. In an example, CNN 440 provides the output of each neuron of the first convolutional layer 442 to the corresponding neurons of the second subsampling layer 448. In some embodiments, CNN 440 adds a bias value to the aggregate set of all values provided to each neuron of the downstream layer. For example, CNN 440 adds a bias value to the aggregate set of all values provided to each neuron of the second subsampling layer 448. In such an example, CNN 440 determines the final value provided to each neuron of the second subsampling layer 448 based on the aggregate set of all values provided to each neuron and the activation function associated with each neuron of the second subsampling layer 448.

[0091] In step 470, the CNN 440 performs a second subsampling function. For example, based on the values output by the second convolutional layer 446 being provided to the corresponding neurons of the second subsampling layer 448 by the CNN 440, the CNN 440 may perform the second subsampling function. In some embodiments, based on the CNN 440 using an aggregation function, the CNN 440 performs the second subsampling function. In an example, as described above, based on the CNN 440 determining the maximum input or average input among the values provided to a given neuron, the CNN 440 performs the first subsampling function. In some embodiments, based on the CNN 440 providing values to the respective neurons of the second subsampling layer 448, the CNN 440 generates an output.

[0092] In step 475, the CNN 440 provides the outputs of the respective neurons of the second subsampling layer 448 to the fully connected layer 449. For example, the CNN 440 provides the outputs of the respective neurons of the second subsampling layer 448 to the fully connected layer 449 such that the fully connected layer 449 generates an output. In some embodiments, the fully connected layer 449 is configured to generate an output associated with a prediction (sometimes referred to as classification). The prediction may include an indication that the objects included in the image provided as input to the CNN 440 include objects and / or a collection of objects, etc. In some embodiments, the perception system 402 performs one or more operations and / or provides data associated with the prediction to different systems described herein.

[0093] Now refer to Figure 5 , which illustrates a diagram of an implementation 500 of a trace refinement network. In some embodiments, the implementation 500 includes a system 510, an offline perception system 504, a trace refinement network 506, a multi-class / multi-task deployment 508, auto-labeled data 520, and an online perception system 522. In some embodiments, the implementation 500 is implemented in one or more devices of the vehicle 102 (e.g., one or more devices of the systems of the vehicle 102), at least one device of the remote AV system 114, at least one device of the queue management system 116, at least one device of the V2I system 118, one or more devices of the network 112 (e.g., one or more devices of the systems of the network 112), one device 300, and / or Figure 4A the AV computing 400 of

[0094] In Figure 5In the example, the offline perception system 504 receives data 514. The perception system includes an online perception system and an offline perception system. The perception system receives data associated with at least one object in the environment (e.g., data used by the perception system to detect at least one physical object) and classifies at least one physical object. In an embodiment, the perception system 504 detects an object over a period of time and enables tracking of the object as the AV navigates through the environment. When the object moves in the environment, a trace of the movement of the object is generated. In particular, at least one bounding box of the detected object is generated. At each point along the trace, the detected object is associated with one or more than one parameter (such as orientation, the center position of at least one bounding box, and the size of at least one bounding box, etc.). In the example, multiple bounding boxes surrounding the object are generated.

[0095] The perception system 504 labels the data 514 with a confidence score that indicates the presence of an instance of a particular object class in the region of the environment associated with the data. In the example, the data 514 is a driving log that includes sensor data captured over a period of time (e.g., data from Figure 2 the camera 202a, LiDAR sensor 202b, Radar sensor 202c, and microphone 202d). The perception system 504 is an offline perception system that operates under minimal constraints. For example, the offline perception system operates non-real-time, such that additional time is used for computations that improve the ability of the perception system to detect objects. In the example, there is a delay between data generation and data processing in the offline perception system. The online perception system (e.g., the online perception system 522) operates in real-time. In the example, the online perception system is executed on the AV and enables real-time perception as the AV navigates through the environment. The online perception system is implemented via the AV stack, whose main role is to detect objects in real-time within a few milliseconds. In the example, when implemented via the AV stack, the computational power and memory of the computer system used to execute the online perception system on the AV are limited by cost and energy consumption. In an embodiment, the offline perception system is cloud-based and is not restricted by the constraints of the online perception system. For example, the offline perception system is a dedicated perception system with increased computational power and memory compared to the online perception system.

[0096] The data 514 is used to estimate a set of bounding boxes for objects in the environment. Each bounding box is labeled with the likelihood that the bounding box contains an object of a specific class. Objects include, but are not limited to, pedestrians, vehicles, bicycles, buses, etc. In an embodiment, the perception system includes an object detection neural network, which is a feed-forward convolutional neural network that generates, given an input (e.g., sensor data), tracking bounding boxes 516 of potential objects in a two-dimensional (2D) or three-dimensional (3D) space, as well as confidence scores that an object class instance (e.g., car, pedestrian, or bicycle) exists within the bounding box. The higher the confidence score, the more likely the corresponding object class instance exists in the box.

[0097] The bounding box 516 is a tracking 3D box with an associated orientation, size, and classification. At least one 3D object detection network semantically detects objects by identifying one or more object class instances in the sensor data. For example, an image 3D object detection network receives image data as input and outputs a set of predicted 3D bounding boxes for potential objects in 3D space and corresponding confidence scores that an object class instance exists within the bounding box. The 3D bounding box includes information about the size (e.g., dimensions), orientation, and position of the 3D bounding box of the object. In an example, the image 3D object detection network can also predict the class of each pixel in the image and output semantic segmentation data (e.g., confidence scores) for each pixel in the image. An example confidence score is a probability value indicating the probability that the class of the pixel is correctly predicted.

[0098] Similarly, in an example, a LiDAR 3D object detection network receives LiDAR data as input and outputs a set of predicted 3D bounding boxes for potential objects in 3D space and confidence scores that an object class instance exists within the bounding box. The LiDAR data includes at least one point cloud. In an example, the 3D object detection network receives a plurality of data points representing 3D space. For example, each data point in the plurality of data points is a set of 3D space coordinates (e.g., x, y, z coordinates). Using one or more point clouds, the 3D object detection network estimates oriented 3D bounding boxes for cars, pedestrians, and cyclists based on the point cloud. Similar to the image 3D object detection network, the predicted 3D bounding boxes output by the 3D object detection network include information related to the size, orientation, and position of the 3D bounding box of the object. The set of predicted 3D bounding boxes also includes confidence scores that an object class instance exists within the bounding box.

[0099] During object detection, per-pixel confidence scores create many high scores that are close to each other, generating a large number of predicted 3D bounding boxes that ultimately create a large number of projected bounding boxes with associated confidence scores associated with the same detected object. Non-maximum suppression suppresses scores that are not locally maximum within a local range. In some embodiments, non-maximum suppression is used to determine the final predicted bounding box or center box. In one or more embodiments, object detection includes determining a center box for each set of bounding boxes, where the center box is a region where the per-pixel confidence score of the sensor data is above a predetermined threshold. The center box has an associated center position (e.g., centroid), center box size (e.g., length, width, and height), and center box orientation.

[0100] The trace refinement network 506 receives as input a tracking bounding box 516 associated with a detected object and the trace of each detected object. The center box of the detected object is located along the trace of the object over time. In an example, the trace refinement network obtains the center box and object trace within a large time window and incorporates the center box at a future timestamp into the processing of the data at the current center box. Since the center box is a sequence of center boxes over a series of timestamps, the trace refinement network combines data from the entire trace window to refine each center box.

[0101] In an embodiment, the trace refinement network 506 outputs a refined center box and trace attributes 518. For example, the trace refinement network 506 outputs a refined center box with a refined center position, refined center box size, and refined center box orientation. The trace attributes include, but are not limited to, the speed and motion state of the object (e.g., whether the trace is static or dynamic). The refined center box and trace attributes are deployed for use in a multi-class / multi-task deployment 508. In an example, the refined center box is used to identify regions of the environment to interpret the associated per-pixel confidence scores as corresponding to specific object class instances across multiple classes. In an example, the refined center box and trace attributes are used to train an online perception system. The trace refinement network executes fast enough for real-time deployment and can also be used in an online perception system by refining the boxes in the current frame using only past boxes. For example, the output of an offline perception system generates automatically labeled speed data. The speed attribute can be refined according to the present technology and used to train and improve the speed predictions made by the online perception system. In some embodiments, the refined boxes are used to update the tracking model for improved future tracking or trajectory prediction. Additionally, the planning system uses the refined center box for future planning. In some embodiments, the refined center box is used in image-LiDAR fusion to ensure well-aligned fusion features.

[0102] Additionally, in some embodiments, deployment 508 includes automatically labeling data 514 with refined center boxes and trace attributes. The automatically labeled data 520 is used to train an online perception system 522 that executes on the AV 502. In an example, a driving log including recorded sensor data is collected from a database. The data is encoded using an adaptive canvas, and a trace refinement network outputs the refined center boxes and trace attributes for a first stage of automatically labeling the data 514. The automatically labeled data 520 is used to train the online perception system 522 in a second stage.

[0103] In an embodiment, offline perception implements a continuous learning framework where driving data is continuously mined for difficult scenarios by comparing detections (e.g., detected objects) from the online perception system and the offline perception system. In some cases, the input to the online perception system and the offline perception system is the same, and the outputs of the systems are inconsistent. For example, the online perception system may not be able to detect a pedestrian hidden behind a tree at one or more timestamps as the vehicle navigates through the environment. However, the offline perception system uses look-ahead and look-behind across one or more timestamps as the vehicle navigates through the environment. The offline perception system determines that a pedestrian observed in the past and in the future also exists at timestamps in the sensor data where the pedestrian is partially occluded by the tree and has poor sensor data. Thus, the offline perception system implements object persistence, where the AV knows the presence of an object even when it is difficult to observe the object well. In an example, when the offline perception system and the online perception system disagree on object detection, the output of the offline perception system is used as training data to improve the online perception system.

[0104] Figure 6 is an illustration of an architecture that includes a trace refinement network. The architecture 600 can be implemented using a perception system (e.g., fully, partially, etc.) that is the same or similar to the perception system 402 described with reference to Figure 4C Additionally, the architecture 600 can be implemented using a perception system (e.g., fully, partially, etc.) that is the same or similar to the perception systems 504 and 522 described with reference to Figure 5 In some embodiments, the architecture is implemented by other devices or systems or other groups of devices and / or systems (e.g., fully and / or partially, etc.) that are separate from or include the perception system. For example, the architecture can be implemented by the remote AV system 114 and / or the AV computing 400 (e.g., one or more systems of the AV computing 400) (e.g., fully and / or partially, etc.) described with reference to Figure 1 In some embodiments, the architecture is implemented by the device 300 (e.g., fully and / or partially, etc.) described with reference to Figure 3

[0105] ​In an example, the trace refinement network 630 refines the center box by leveraging the long-term temporal context captured by the trace of the object. A set of predicted bounding boxes is obtained. For example, the set of predicted bounding boxes is generated by an object detection neural network. The center boxes from the respective sets of bounding boxes at a series of timestamps form a sequence of center boxes along the trace. The trace is associated with the object detected within the center box, and each corresponding center box includes a center position, a center box size, and a center box orientation. A trace window is generated around the corresponding center box, where the trace window corresponds to: the corresponding center box at the timestamp along the trace, and a set of context boxes from adjacent timestamps of the same trace around the center box. For example, for a trace window with a radius of 50, it includes the center box at timestamp t, 50 boxes from the same object trace just before timestamp t, and 50 boxes from the same trace just after timestamp t. In the example, for a window radius of 50 and a total of 101 boxes in the trace window, N = 50. Features are extracted from the trace window, and the trace window is cropped and normalized to create a canvas corresponding to the respective center box. The canvas is encoded to obtain features. The features are input into the trace refinement network, where the trace refinement network uses the features from the entire trace window to output a refined center, a refined size, and a refined orientation for each corresponding center box. In some embodiments, the trace window 602 is used to create an adaptive canvas corresponding to each respective trace window. Regarding Figures 8 to 9 The canvas is further described. Data from the entire sequence of boxes within the trace window is used to refine the center box. The trace window is encoded using the adaptive canvas, and the adaptive canvas is input into the point feature encoder 612 and the trajectory feature encoder 622.

[0106] In Figure 6 the example, data is illustrated using blocks with dashed lines, while the processing applied to the data is illustrated using blocks with solid lines. In Figure 6 the example, data associated with the trace window 602 is used to generate a canvas for input into the point feature encoder 612 and the trajectory feature encoder 622. In the example, the point feature encoder 612 takes as input the adaptive canvas including the raw point features, which include x, y, z, intensity, and time increment. For example, the raw point features are data representing the 3D space around the vehicle captured by LiDAR. The point feature encoder 612 outputs a size of width (W), height (H), and the number of channels for each pixel of the canvas (W*H*C piThe column feature of (). In an embodiment, the point feature encoder 612 may include point columns, consistent with at least some implementations, as described in PointPillars: Fast Encoders for Object Detection from Point Clouds, arXiv:1812.05784v2 [cs.LG] on May 7, 2019. Figure 7A Illustrates points 702 associated with multiple boxes for encoding. In the example, points 702 are input into the trace refinement network. Points 702 represent the union of all points within the trace window, while Figure 7B represents the corresponding bounding box trajectory, which is also an input to the trace refinement network. In Figure 7A the example, aggregated points within the tracking box of a trace window of length 5 (or radius 2) are shown. A center box point appears at reference numeral 703, which corresponds to Figure 7B the box 706 shown in

[0107] In the example, the point feature encoder 612 applies point column encoding to the points of all boxes from the trace window (spatially extended to include more context along the trace of the object). For example, each data point among the multiple data points is a set of 3D spatial coordinates. The 3D space is divided into multiple columns, where each column among the multiple columns extends vertically (in the z direction) from a corresponding portion of the 2D ground plane of the 3D space. Each data point among the multiple data points is assigned to a column among the multiple columns. A pseudo-image is generated based on the multiple columns. For each column among the multiple columns, the pseudo-image includes a corresponding feature representation of the data points assigned to that column (e.g., W*H*C pi column features).

[0108] The point feature network 614 obtains W*H*C piColumn features are used as input. A CNN network (i.e., VGG or ResNet architecture) is used to extract a single feature vector representing the entire point cloud of each corresponding trace window around the center box (e.g., 10 boxes before the center box in time and 10 boxes after the center box in time). For example, as described above, the trace window is a time series of tracking boxes around the center box. In the example, the center box at box t with a trace window radius of 10 is identified (e.g., the center box). The trace window will include 10 tracking boxes from box t+1 to t+10, and 10 boxes from box t-10 to t-1, for a total of 21 boxes for tracking the same object. The points in these 21 boxes will be used for point feature encoding. The 21 boxes will be used for trajectory encoding. The box is represented by the center position and the size including length and width. For example, the features are encoded using the x coordinate of the box center (e.g., x), the y coordinate of the box center (e.g., y), the box width (e.g., w), the box length (e.g., w), the orientation of the box (e.g., yaw), and the timestamp associated with each corresponding box (e.g., the time difference between other boxes and the center box). In the example, other features of the box are specified, such as height and orientation, etc. The point feature network 614 outputs point features 616. The point features 616 are a 1×1*C point cloud structure representing the object trace p features.

[0109] In parallel with the point feature encoder 612, the trajectory feature encoder 622 takes as input an adaptive canvas including the original center box features, the original center box features including center_x, center_y, width, length, yaw angle, and delta_time for N t boxes in the trace window. In the example, N t = 101, where the predicted bounding box set includes 101 bounding boxes. The trajectory feature encoder 622 outputs trajectory features for N t boxes, each box with C ti channels, having a dimension of N t *C ti . The trajectory feature network 624 obtains N t *C t trajectory features extracted from N ti context boxes as input. In an embodiment, the trajectory feature encoder is a one-dimensional (1D) CNN or MLP network that extracts C t features representing the trajectory and motion of the object trace. The trajectory feature network 624 outputs trajectory features 626.

[0110] The point feature 616 and the trajectory feature 626 are input into the trace refinement network 630. In an example, the trace refinement network 630 is an MLP network that regresses multiple task heads from the concatenated point feature and trajectory feature. In an example, the point feature 616 is represented by a first feature vector, and the trajectory feature is represented by a second feature vector. The first feature vector and the second feature vector are concatenated and input into the trace refinement network 630. The trace refinement network outputs a refined and improved center box 604, and the output refined and improved center box 604 has a refined center position, a refined box size, and a refined box orientation. In an embodiment, trace attributes such as speed and motion state are output by the trace refinement network 630.

[0111] Figure 8 is an illustration of the vehicle reference coordinate system 800 and the canvas 802. In an example, the coordinates of the vehicle match the coordinates of the canvas used for encoding as described Figure 6 above. The center position 808 within the center box 804 represents the center of the center box 804. The box 806 on the canvas is the ground truth box position. The canvas is used to encode all the points within the trace window of the center box. The legend 820 shows that in Figure 8 the normalized ground truth box 806 is illustrated using a dashed line, and the center box (e.g., the normalized input box) 804 is illustrated using a solid line.

[0112] As shown in the legend 820, the regression target variables include the residual position, which indicates the distance by which the center box 804 is shifted to align with the ground truth box 806, shown as an arrow 810 on the canvas. In particular, the arrow 810 represents the following regression target variable, which indicates the change in the position (e.g., the residual position) of the center position 808 as the change along the x - dimension (e.g., Δx) divided by the width of the center box 804 (e.g., w) and the change along the y - dimension (e.g., Δy) divided by the length of the center box 804 (e.g., l). As shown in the legend 820, the regression target variables also include the residual orientation, which indicates the angle (e.g., yaw) by which the center box 804 is rotated to align with the ground truth box 806, shown as an angle 814 on the canvas. The regression target variables also include the residual size, which indicates the change in the size of the center box 804 to obtain the size of the ground truth box 806. The residual size (e.g., in terms of width and length) is calculated as the change in the width (e.g., Δw) between the boxes 804 and 806 divided by the width of the center box 804, and the change in the length (e.g., Δl) between the boxes 804 and 806 divided by the length of the center box 804.

[0113] In an embodiment, for the corresponding center boxes obtained from the object detection neural network, from Figure 6The trace window 602 crops out the canvas 802. The canvas is an encoded canvas whose center position is the same as the center of the corresponding central box of the trace window. The points for encoding within the trace window are cropped based on the canvas. The canvas 802 is a box-aligned canvas having a coordinate system corresponding to the coordinate system of the central box. The size of the canvas is based on the size of the corresponding central box.

[0114] The central box 804 is illustrated. A trace window (not illustrated) is generated around the central box 804, and the trace window is cropped to obtain the canvas 802. As described with respect to Figure 9 the encoded canvas is scaled according to the size of the input central box. For ease of illustration, the canvas 802 is depicted as having a square shape. However, the size of the canvas is determined based on the central box and can be of any shape and size. In the example, the input central box is used to scale and normalize the canvas. In the example, the input box is normalized to a unit square. The absolute size of the canvas in meters is related to the size of the central box. In the example, the canvas is chosen to be three times larger than the central box to capture the context around the central box. In the example, the canvas is chosen to be other sizes. For different central boxes, the absolute canvas size (in meters) will be different. By using different bin sizes, the characteristic size (i.e., W*H) of each canvas is always the same, where there are W bins along the width direction and H bins along the length direction. In the example, the context refers to the area outside the central box. The context enables the area outside the object (e.g., the road surface) to be considered, which supports a better understanding of the position of the box. Additionally, including more context enables the trace refinement network to see the entire object even if the central box does not capture the complete appearance of the object.

[0115] The canvas 802 is rotated to align with the corresponding central box 804, scaled according to the size of the corresponding central box 804, and normalized using the size of the corresponding central box. In this way, the canvas is adaptively scaled and normalized based on the input central box. In an embodiment, the input central box is used as an initial estimate to regress the central box size, orientation, travel direction, and position. The residual between the input central box and the normalized ground truth box is regressed to determine the refined central box. As described herein, the residual is the offset from the input box representing the distance between the input central box and the normalized ground truth box 806. In the example, the training database storing the traces includes pairs of central boxes and ground truth boxes. During the database generation phase, global association is performed between the ground truth traces and the traces generated from the offline perception system based on their proximity to each other. After training, the trace refinement network is deployed. During deployment, the ground truth box is not used. Instead, the central box (e.g., the box to be refined) is obtained together with the residual from the output of the trace refinement network to generate the refined central box. In the example, the refined central box is close to the ground truth.

[0116] The normalized ground truth box 806 is illustrated in the vehicle reference coordinate system 800. In the example, the regression target variable is the immediate output of the trace refinement network. The trace refinement network is trained to output the regression target variable. At deployment, the regression target variable will be denormalized by the input box to obtain the refined center box in the vehicle coordinate system. The normalized ground truth box 806 represents the regression target, where the input center box 804 is rotated and aligned with the normalized ground truth box 806. The center 808 of the center box 804 is shifted by the distance indicated by the arrow 810 to align with the center 812 of the normalized ground truth box 806. Additionally, the center box 804 is rotated by the angle indicated by the reference numeral 814 to align with the normalized ground truth box 806. After alignment and size normalization, the canvas is provided as an input for encoding, such as by a point feature encoder and a trajectory feature encoder. In the example, the encoding encodes the center box into a 2D shape for processing by a neural network, such as a 2D CNN, etc.

[0117] This technique enables the regression of the residuals of the size, orientation, and position of the center box, respectively. The input center box is used as a reasonably good estimate, and it is easier and computationally efficient to regress the residuals. In the example, scaling enables this technique to support multi-class trace refinement. This is due to size normalization, as the same canvas can be applied to different classes with different shape statistics. The size of the canvas will be enlarged / reduced to match the size of the center box. In some embodiments, the absolute size of the canvas in meters is related to the size of the center box. For example, the canvas is three times larger than the center box to capture sufficient context. For different center boxes, the absolute canvas size (in meters) will be different. By using different bin sizes, the feature size (i.e., W*H) of each canvas is always the same, so there are W bins along the width direction and H bins along the length direction. In an embodiment, the trace refinement network can be used across multiple object classifications of any size.

[0118] Figure 9 Examples of classes with different shape statistics and corresponding points for each class are shown. In the example, the point cloud data points are cropped to the points that appear within the respective canvas.

[0119] In Figure 9In the example, the bus class 902 and the pedestrian class 904 are shown after various processes for obtaining a canvas. At line 910, a single canvas size is used for both the bus class 902 and the pedestrian class 904. Traditionally, the single canvas size is a fixed size around a center box. At line 910, the points of the bus class are shown within the input center box in the cropped canvas 912. At line 910, the cropped canvas 914 of the same size is applied to the pedestrian class 904. As shown, the cropped canvas 912 at line 910 is too small to capture the entire center box representing the bus class 902. Additionally, the cropped canvas 914 at line 910 represents the points corresponding to the pedestrian at a coarse resolution in the center of the canvas, where the data points appear in a small portion of the canvas while the rest of the canvas is wasted.

[0120] At line 920, a single canvas size is used for both the bus class 902 and the pedestrian class 904. However, the canvas is rotated to align with the corresponding center box. At line 920, the points of the bus class 902 are shown within the input center box in the cropped canvas 922 that is rotated to align with the center box. At line 910, the cropped canvas 924 of the same size is applied to the pedestrian class 904, and the canvas is rotated to align with the center box. As shown, the cropped canvas 922 used at line 920 captures the entire center box representing the bus class 902. However, the cropped canvas 924 at line 920 still captures the points corresponding to the pedestrian class 904 at a coarse resolution in the center of the canvas, where the data points appear in a small portion of the canvas while the rest of the canvas is wasted.

[0121] As Figure 9 shown in the example, a single canvas setting (e.g., size) adversely affects the processing of objects with different shape statistics (such as size, etc.). When the canvas is too small, the canvas cannot capture the data points associated with the entire object and the appropriate parts of the surrounding environment. When the canvas is too large, the points are represented at a coarse resolution in a small area of the canvas, where the canvas characteristics will be dominated by the wasted background data positions on the rest of the canvas.

[0122] At line 930, an adaptive canvas is illustrated. The adaptive canvas is normalized according to the size of the input center box. The adaptive canvas is a box-aligned canvas with size normalization. The adaptive canvas 932 at line 930 is applied to the bus class 902. The canvas 932 is rotated to align with the center box and scaled according to the size of the center box. At line 930, the adaptive canvas 934 is applied to the pedestrian class 904, and the canvas 934 is rotated to align with the center box and scaled according to the size of the center box. As shown, the adaptive canvas 932 used at line 930 captures the entire center box representing the bus class 902. The adaptive canvas 934 at line 930 also captures the points corresponding to the pedestrian class 904 at a fine resolution at the center of the canvas, where data points appear throughout the canvas.

[0123] At line 930, although the physical size of the pedestrian is much smaller than that of the bus, since the encoded canvas normalized by the corresponding center box size is used to process each object, the encoded canvas can adapt to the corresponding shape and size of the detected object. If the object itself is large, a large encoded canvas is generated. If the object is small, a small encoded canvas is generated to ensure capturing the details of the small object. The use of the adaptive canvas reduces computational waste by reducing the area in the canvas without data. In addition, the adaptive canvas enables using a single trace refinement network instead of multiple class-specific trace refinement networks to refine multiple classes with different shapes and sizes. This can greatly save training and deployment efforts.

[0124] Figure 10 is a block diagram of a multi-task trace refinement network. A trace refinement network 1000 can be implemented using a trace refinement network (e.g., completely, partially, etc.) that is the same as or similar to the trace refinement network 630 described in reference Figure 6 In the example of Figure 10 , the trace refinement network 1000 performs multiple tasks in parallel. The cascaded point features and trajectory features 1010 are input into the shared layer 1020. In some examples, the cascaded point and trajectory features 1010 are the point features 616 and trajectory features 626 as described in reference Figure 6 .

[0125] This technology enables multiple tasks to be performed using a single model. The trace refinement network includes a shared layer 1020 and task heads, and the task heads include a center head 1022, a size head 1023, an orientation head 1024, a velocity head 1025, and a motion head 1026. The task heads receive data from the shared layer 1020. The center head 1022 outputs a refined box center 1032 (e.g., x, y, z). The size head 1023 outputs a refined box size 1033 (e.g., length, width, height). The orientation head 1024 outputs a refined box orientation 1034 (e.g., yaw angle). The velocity head 1025 outputs a refined box velocity 1035 (v x , v y ). The motion head 1026 outputs a trace window motion state 1036 (e.g., static or moving).

[0126] To perform multiple tasks, multiple different heads are placed in the last few layers of the trace refinement network. Thus, the trace refinement network outputs, for example, a refined center box center, a refined center box size, and a refined center box orientation. Additionally, trace attributes are estimated and output by the trace refinement network. Trace attributes include, but are not limited to, velocity and motion state (e.g., whether the trace is static or dynamic). In an example, the velocity head 1025 uses an object trace corresponding to the center box as input to predict velocity, and the object trace is used to encode features as described above. Similarly, the motion head 1026 uses an object trace corresponding to the center box as input to determine whether the object is static or moving.

[0127] Figure 11 is an illustration of an auxiliary loss function for enhancing the prediction smoothness of the box along the trace. In Figure 11 example, a trace refinement network 1110 is illustrated. The trace refinement network 1110 can be implemented using a trace refinement network (e.g., completely, partially, etc.) that is the same as or similar to the trace refinement network 630 described with reference to Figure 6 . In an example, the trace refinement network is trained using a ground truth marked center box. In an embodiment, a smoothness loss is introduced during the training of the trace refinement network to implement smooth trace refinement network center box prediction along the entire trace. This prevents the trace refinement network from predicting sharp box changes between nearby trace windows, such as large center position jumps, head-tail flips, and center box size changes, etc.

[0128] An auxiliary loss 1120 is used to enhance the smoothness of box predictions between adjacent boxes along a trace. As shown, a series of tracking boxes 1104A - 1104F create an input trace 1102. This series of tracking boxes is along a timeline 1106 according to their corresponding times, where time increases from left to right. Similarly, a series of refined boxes 1114A - 1114F create a refined trace 1112. This series of refined boxes 1114A - 1114F is along a timeline 1116 according to their corresponding times, where time increases from left to right. This series of refined boxes is generated using a series of tracking boxes from the entire input trace 1102.

[0129] To implement smooth prediction of frame sensor data for the same trace, for example, a penalty is imposed on jumps in the center box size. For example, box 1104C is smaller than box 1104D along the input trace 1102. For each detected object, the predicted input box size, yaw angle, and speed will be consistent along the trace. The auxiliary loss 1120 is implemented such that the trace refinement network learns how to predict a smooth output of boxes along the trace.

[0130] In an example, the auxiliary loss 1120 for smoothing will implement the center box size, center position, center box orientation, and speed. The auxiliary loss 1120 is based on the box size. If the trace refinement network predicts different object sizes for trace windows from the same trajectory, a penalty is applied during training. In the example, Figure 11 it is shown that box 1104C is smaller than box 1104D. The context of the center box along the trace is used to refine the input box size (such as the small box 1104C, etc.). However, the boxes along the trace are also used to refine the box size of the larger box 1104D. If the box size changes for sequentially occurring boxes, a penalty is applied to the final loss function during the training of the trace refinement network.

[0131] In an example, the auxiliary loss 1120 is based on the jitter of the box center. As mentioned above, motion is predicted for a trace window. For a static trace, a penalty is imposed on the training loss in response to the jitter of the center position. For a dynamic trace, some jumps in the center position are allowed and reflect the center position of the smoothly moving dynamic trace. In the example, a penalty is imposed on the training loss for large jumps in the center position on a dynamic trace. This ensures that the center position corresponding to the center box does not jump along the trace.

[0132] In the example, the auxiliary loss 1120 is based on the box orientation. For trace windows from the same static trace, differences in the orientation prediction across traces are penalized by applying a higher loss when the orientations are different. In particular, if the model predicts a flipped head-to-tail orientation for trace windows along the same trace, a higher penalty is imposed. For a static trace, if the boxes from the same trace have predictions of different orientations, a higher loss is imposed. If the model predicts a flipped head-to-tail orientation, the penalty will be even higher, where a flip refers to a 180-degree difference in the head or tail position between the center boxes. Using the entire trace, center boxes with flipped orientations are corrected by applying an auxiliary loss function during training. In the example, the auxiliary training loss 1120 is based on speed. Large jumps in the speed prediction of trace windows from the same trace are penalized with a higher loss. This adds some auxiliary losses to enhance trace smoothness.

[0133] In some embodiments, trace-level constraint enforcement is applied to the refined output of the trace refinement network. Trace-level constraint enforcement enhances the box quality of the entire trace. In the first trace-level constraint enforcement, a shared box size is obtained by applying median filtering to the predictions of the entire trace. In the example, median filtering eliminates variations in the box size predictions for the same trace. For example, using the shared box size means that the individual boxes are the same size along the trace. The shared box size results in zero difference in the box size along the entire trace.

[0134] In the second trace-level constraint enforcement, the trace-level static / dynamic state is obtained using the motion states of the trace refinement network output for all trace windows. In the third trace-level constraint enforcement, additional constraints are imposed on static traces. For example, for a static trace, zero velocity is used for all boxes. In the example, for a static trace, the shared box center is determined as the cluster center of the center predictions of the trace refinement network for all trace windows. In addition to clustering, the mean and median, etc., can also be used to determine the shared box center. In the example, for a static trace, the best orientation of the trace is shared by applying median filtering to the orientation predictions of all trace windows. This removes small jitters associated with static objects.

[0135] In some embodiments, the present technology achieves a more comprehensive model of an object by aggregating observations across an entire trace, which then benefits the localization accuracy of individual boxes along the trace. The trace refinement network utilizes trace-level information (including both point clouds and trajectories from past and future outputs of an offline perception system) to refine the center box. The use of trace-level information enables the trace refinement network to further significantly improve box quality (box center position, size, orientation, and velocity). The trace refinement network is a lightweight multi-class multi-task model that is combined with an object detection and perception system to enhance box quality. The trace refinement network can be scaled and efficiently deployed with a large amount of driving data. The present technology is very fast and faster than real-time. This results in scalability and efficiency in deploying a large amount of driving data. In an example, faster than real-time means that the model can process data faster than the data is generated. For example, LiDAR point clouds are generated at 20Hz (e.g., 20 frames per second), and the present technology processes approximately 1.5 * 20 frames in less than a second.

[0136] Now refer to Figure 12 , which illustrates a flowchart of a process 1200 for a trace refinement network. In some embodiments, one or more steps described with respect to process 1200 are performed by an autonomous system (e.g., Figure 2 autonomous system 202) (e.g., fully and / or partially, etc.). In some embodiments, process 1200 is implemented in one or more devices of vehicle 102 (e.g., one or more devices of the system of vehicle 102), at least one device of remote AV system 114, at least one device of queue management system 116, at least one device of V2I system 118, one or more devices of network 112 (e.g., one or more devices of the system of network 112), a device 300, and / or at least one component of device 300 and / or Figure 4A AV computing 400.

[0137] At block 1202, a center box is obtained. In some examples, the center box is the output of an object detection neural network. In an example, the object detection network outputs a set of predicted 3D bounding boxes of potential objects in 3D space and corresponding confidence scores that an object class instance exists within the bounding box. In an example, the center box is the center bounding box from a sequence of boxes along a trace. A record of driving data is provided as input to generate a sequence of boxes along the trace. In an embodiment, the trace is associated with the detected objects within the sequence of boxes, where each respective center box includes a center, size, and orientation. In an example, each box in the sequence of boxes is selected as the center box and is processed iteratively as described herein.

[0138] At block 1204, a trace window is generated around the corresponding center box, where the trace window corresponds to the corresponding center box along the trace. At block 1206, the trace window is cropped and normalized to create a canvas corresponding to the corresponding center box (e.g., a box-aligned canvas with normalized size). In an example, the canvas is encoded to obtain features. At block 1208, features are extracted from the cropped and normalized window (e.g., the canvas). In an example, the features include point cloud features and trajectory features regarding Figure 6 described.

[0139] At block 1210, the features are input into a trace refinement network, where the trace refinement network uses the features from the entire trace window around the center box to output the refined center, refined size, and refined orientation of each corresponding center box. In an example, the trace refinement network regresses the residuals between the input box obtained from an offline perception system and the actual object box. Then, the residuals are used to obtain the refined box based on the input box.

[0140] At block 1212, the refined boxes are deployed. In an example, the refined centers, refined sizes, and refined orientations of each corresponding center box are deployed.

[0141] According to some non-limiting embodiments or examples, a system is provided that includes at least one processor and at least one non-transitory storage medium. The at least one non-transitory storage medium stores instructions that, when executed by the at least one processor, cause the at least one processor to perform operations. The operations include: obtaining center boxes from a record of driving data, where the center boxes form a sequence of boxes along a trace, and the trace is associated with a tracked object detected within the sequence of boxes, and each corresponding center box includes a center, a size, and an orientation. The operations include generating a trace window around the corresponding center box, where the trace window corresponds to the corresponding center box along the trace. The operations include cropping and normalizing the trace window relative to the center box to enable a single refinement model for multiple object classes. The operations include extracting point cloud features and trajectory features from the cropped and normalized trace window. Additionally, the operations include inputting the point cloud features and trajectory features into a trace refinement network, where the trace refinement network uses the features from the trace to output the refined center, refined size, and refined orientation of each corresponding center box. Further, the operations include deploying the refined centers, refined sizes, and refined orientations of each corresponding center box.

[0142] According to some non - limiting embodiments or examples, a method is provided. The method includes: obtaining, by at least one processor, a center box from a record of driving data, where the center box forms a sequence of boxes along a trace, and the trace is associated with a tracked object detected within the sequence of boxes, and each corresponding center box includes a center, a size, and an orientation. The method includes generating, by at least one processor, a trace window around the corresponding center box, where the trace window corresponds to the corresponding center box along the trace. Additionally, the method includes cropping and normalizing the trace window with respect to the center box to achieve a single refined model for multiple object classes. The method includes extracting, by at least one processor, point cloud features and trajectory features from the cropped and normalized trace window. Additionally, the method includes inputting the point cloud features and the trajectory features into a trace refinement network, where the trace refinement network uses features from the trace to output a refined center, a refined size, and a refined orientation of each corresponding center box. Further, the method includes deploying the refined center, the refined size, and the refined orientation of each corresponding center box.

[0143] According to some non - limiting embodiments or examples, at least one non - transitory storage medium storing instructions is provided, the instructions when executed by at least one processor cause the at least one processor to perform operations. The operations include obtaining a center box from a record of driving data, where the center box forms a sequence of boxes along a trace, and the trace is associated with a tracked object detected within the sequence of boxes, and each corresponding center box includes a center, a size, and an orientation. The operations include generating a trace window around the corresponding center box, where the trace window corresponds to the corresponding center box along the trace. The operations include cropping and normalizing the trace window with respect to the center box to achieve a single refined model for multiple object classes. The operations include extracting point cloud features and trajectory features from the cropped and normalized trace window. Additionally, the operations include inputting the point cloud features and the trajectory features into a trace refinement network, where the trace refinement network uses features from the trace to output a refined center, a refined size, and a refined orientation of each corresponding center box. Further, the operations include deploying the refined center, the refined size, and the refined orientation of each corresponding center box.

[0144] Other non - limiting aspects or embodiments are set forth in the numbered clauses below:

[0145] Clause 1: A system, comprising: at least one processor, and at least one non-transitory storage medium storing instructions that, when executed by the at least one processor, cause the at least one processor to: obtain a center box from a record of driving data, wherein the center box forms a sequence of boxes along a trace, and the trace is associated with a tracked object detected within the sequence of boxes, each respective center box including a center, a size, and an orientation; generate a trace window around the respective center box, wherein the trace window corresponds to the respective center box along the trace; crop and normalize the trace window with respect to the center box to achieve a single refined model for multiple object classes; extract point cloud features and trajectory features from the cropped and normalized trace window; input the point cloud features and the trajectory features into a trace refinement network, wherein the trace refinement network uses features from the trace to output a refined center, a refined size, and a refined orientation of each respective center box; and deploy the refined center, the refined size, and the refined orientation of each respective center box.

[0146] Clause 2: The system according to Clause 1, wherein the trace refinement network regresses a residual between the respective center box obtained from an offline perception system and a ground truth box.

[0147] Clause 3: The system according to Clause 1 or 2, wherein the normalization scales a canvas based on the respective center box.

[0148] Clause 4: The system according to any one of Clauses 1 to 3, wherein the trace refinement network refines center boxes corresponding to multiple classifications associated with object detection.

[0149] Clause 5: The system according to any one of Clauses 1 to 4, wherein the trace refinement network includes a shared layer and multiple task heads, and the trace refinement network enables multiple tasks corresponding to the multiple task heads to be performed.

[0150] Clause 6: The system according to any one of Clauses 1 to 5, wherein the trace refinement network also outputs trace attributes.

[0151] Clause 7: The system according to any one of Clauses 1 to 6, wherein the trace refinement network is trained using an auxiliary loss function that smooths the output of the trained trace refinement network.

[0152] Clause 8: The system according to any one of Clauses 1 to 7, wherein deploying the driving data includes automatically generating a database of automatically labeled training data.

[0153] Clause 9: The system according to any one of Clauses 1 to 8, wherein deploying the driving data includes inputting a center box into an online tracker in online perception.

[0154] Clause 10: The system according to any one of Clauses 1 to 9, wherein deploying the driving data includes generating a refined box to achieve image-LiDAR fusion.

[0155] Clause 11: The system according to any one of Clauses 1 to 10, including imposing at least one constraint on the refined center, refined size, and refined orientation of each respective center box.

[0156] Clause 12: A method, comprising: obtaining, by using at least one processor, a center box from a record of driving data, wherein the center box forms a sequence of boxes along a trace, and the trace is associated with a tracked object detected within the sequence of boxes, each respective center box including a center, a size, and an orientation; generating, by using the at least one processor, a trace window around the respective center box, wherein the trace window corresponds to the respective center box along the trace; cropping and normalizing, by using the at least one processor, the trace window with respect to the center box to achieve a single refined model for multiple object classes; extracting, by using the at least one processor, point cloud features and trajectory features from the cropped and normalized trace window; inputting, by using the at least one processor, the point cloud features and the trajectory features into a trace refinement network, wherein the trace refinement network uses features from the trace to output a refined center, a refined size, and a refined orientation of each respective center box; and deploying, by using the at least one processor, the refined center, the refined size, and the refined orientation of each respective center box.

[0157] Clause 13: The method according to Clause 12, wherein the trace refinement network regresses a residual between the respective center box obtained from an offline perception system and a ground truth box.

[0158] Clause 14: The method according to Clause 12 or 13, wherein the normalization scales a canvas based on the respective center box.

[0159] Clause 15: The method according to any one of Clauses 12 to 14, wherein the trace refinement network refines center boxes corresponding to multiple classifications associated with object detection.

[0160] Clause 16: The method according to any one of Clauses 12 to 15, wherein the trace refinement network includes a shared layer having multiple task heads, and the trace refinement network enables performing multiple tasks corresponding to the multiple task heads.

[0161] Clause 17: At least one non-transitory storage medium storing instructions which, when executed by at least one processor, cause the at least one processor to: obtain a center box from a record of driving data, wherein the center box forms a sequence of boxes along a trace, and the trace is associated with a tracked object detected within the sequence of boxes, and each respective center box includes a center, a size, and an orientation; generate a trace window around the respective center box, wherein the trace window corresponds to the respective center box along the trace; crop and normalize the trace window relative to the center box to achieve a single refined model for multiple object classes; extract point cloud features and trajectory features from the cropped and normalized trace window; input the point cloud features and the trajectory features into a trace refinement network, wherein the trace refinement network uses features from the trace to output a refined center, a refined size, and a refined orientation of each respective center box; and deploy the refined center, the refined size, and the refined orientation of each respective center box.

[0162] Clause 18: The at least one non-transitory storage medium according to Clause 17, wherein the trace refinement network causes a regression of the residual between the respective center box obtained from an offline perception system and a ground truth box.

[0163] Clause 19: The at least one non-transitory storage medium according to Clause 17 or 18, wherein the normalization scales a canvas based on the respective center box.

[0164] Clause 20: The at least one non-transitory storage medium according to any one of Clauses 17 to 19, wherein the trace refinement network refines center boxes corresponding to multiple classifications associated with object detection.

[0165] In the foregoing description, aspects and embodiments of the present disclosure have been described with reference to numerous specific details, which may vary according to implementation. Accordingly, the specification and drawings are to be regarded as illustrative rather than in a limiting sense. The sole and exclusive indication of the scope of the invention, and what the applicant desires to be the scope of the invention, is the literal and equivalent scope of the claims as issued from this application in the specific form of the issued claims, including any subsequent amendments. Any definition explicitly set forth herein for terms to be included in such claims shall be construed in the sense such terms are used in the claims. Additionally, when the term "further comprises" is used in the foregoing specification or the appended claims, the text following such phrase may be additional steps or entities, or sub-steps / sub-entities of the previously recited steps or entities.

Claims

1. A system, comprising: at least one processor, and at least one non-transitory storage medium storing instructions that, when executed by the at least one processor, cause the at least one processor to: obtain a center box from a record of driving data, wherein the center box forms a sequence of boxes along a trace, and the trace is associated with a tracked object detected within the sequence of boxes, and each respective center box includes a center, a size, and an orientation; generate a trace window around the respective center box, wherein the trace window corresponds to the respective center box along the trace; crop and normalize the trace window with respect to the center box to implement a single refinement model for multiple object classes; extract point cloud features and trajectory features from the cropped and normalized trace window; input the point cloud features and the trajectory features into a trace refinement network, wherein the trace refinement network uses features from the trace to output a refined center, a refined size, and a refined orientation of each respective center box; and deploy the refined center, the refined size, and the refined orientation of each respective center box.

2. The system according to claim 1, wherein The trace refinement network regresses a residual between the respective center box obtained from an offline perception system and a ground truth box.

3. The system according to claim 1 or 2, wherein, Normalization scales the canvas based on the respective center box.

4. The system according to any one of claims 1 to 3, wherein, The trace refinement network refines center boxes corresponding to multiple classifications associated with object detection.

5. The system according to any one of claims 1 to 4, wherein, The trace refinement network includes a shared layer and multiple task heads, and the trace refinement network enables multiple tasks corresponding to the multiple task heads to be performed.

6. The system according to any one of claims 1 to 5, wherein The trace refinement network also outputs trace attributes.

7. The system according to any one of claims 1 to 6, wherein, The trace refinement network is trained using an auxiliary loss function that smooths the output of the trained trace refinement network.

8. The system according to any one of claims 1 to 7, wherein Deploying the driving data includes automatically generating a database of automatically labeled training data.

9. The system according to any one of claims 1 to 8, wherein, Deploying the driving data includes inputting the center box into an online tracker in online perception.

10. The system according to any one of claims 1 to 9, wherein, Deploying the driving data includes generating refined boxes to achieve image-LiDAR fusion.

11. The system according to any one of claims 1 to 10, comprising imposing at least one constraint on the refined center, the refined size, and the refined orientation of each respective center box.

12. A method, comprising: obtaining, by using at least one processor, a center box from a record of driving data, wherein the center box forms a sequence of boxes along a trace, and the trace is associated with a tracked object detected within the sequence of boxes, and each respective center box includes a center, a size, and an orientation; generating, by using the at least one processor, a trace window around the respective center box, wherein the trace window corresponds to the respective center box along the trace; cropping and normalizing, by using the at least one processor, the trace window with respect to the center box to implement a single refinement model for multiple object classes; extracting, by using the at least one processor, point cloud features and trajectory features from the cropped and normalized trace window; Input the point cloud features and the trajectory features into a trace refinement network using the at least one processor, wherein the trace refinement network uses features from the trace to output a refined center, a refined size, and a refined orientation of each corresponding center box; and Deploy the refined center, the refined size, and the refined orientation of each corresponding center box using the at least one processor.

13. The method according to claim 12, wherein, The trace refinement network regresses the residual between the corresponding center box obtained from an offline perception system and the ground truth box.

14. The method according to claim 12 or 13, wherein Normalize to scale the canvas based on the corresponding center box.

15. The method according to any one of claims 12 to 14, wherein The trace refinement network refines center boxes corresponding to multiple classifications associated with object detection.

16. The method according to any one of claims 12 to 15, wherein, The trace refinement network includes a shared layer and multiple task heads, and the trace refinement network enables multiple tasks corresponding to the multiple task heads to be performed.

17. At least one non-transitory storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to: Obtain a center box from the recording of driving data, where The center boxes form a sequence of boxes along a trace, and the trace is associated with a tracked object detected within the sequence of boxes, each corresponding center box including a center, a size, and an orientation; Generate a trace window around the corresponding center box, wherein the trace window corresponds to the corresponding center box along the trace; Crop and normalize the trace window with respect to the center box to implement a single refinement model for multiple object classes; Extract point cloud features and trajectory features from the cropped and normalized trace window; Input the point cloud features and the trajectory features into a trace refinement network, wherein the trace refinement network uses features from the trace to output a refined center, a refined size, and a refined orientation of each corresponding center box; and Deploy the refined center, the refined size, and the refined orientation of each corresponding center box.

18. The at least one non-transitory storage medium according to claim 17, wherein, The trace refinement network regresses the residual between the corresponding center box obtained from an offline perception system and the ground truth box.

19. The at least one non-transitory storage medium according to claim 17 or 18, wherein, Normalize to scale the canvas based on the corresponding center box.

20. The at least one non-transitory storage medium according to any one of claims 17 to 19, wherein, The trace refinement network refines center boxes corresponding to multiple classifications associated with object detection.