Enhancing later time feature maps using earlier time feature maps

By using earlier time feature maps to enhance late time feature maps in autonomous vehicles, the challenge of object characteristics and bounding box determination in real-time driving environments is solved, and more accurate object recognition and navigation is achieved.

CN120153404APending Publication Date: 2025-06-13MOTIONAL AD LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380076339.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-09-02
Filing Date
2023-08-31
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

In real-time driving environments, it is challenging to accurately determine object characteristics and generate accurate 3D enclosure boxes, especially since neural networks cannot obtain sufficient semantic and local information.

Method used

The ability of autonomous vehicles to determine object characteristics, bounding boxes, and object trajectories is improved by generating multiple feature maps using multiple images in the image stream from the same image sensor and enhancing the later time feature maps using earlier time feature maps.

Benefits of technology

By enhancing the feature map, the autonomous vehicle can more accurately identify the object, determine the depth, speed, center and other characteristics of the object, and generate more accurate bounding boxes, thereby improving the safety and comfort of navigation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120153404A_ABST
    Figure CN120153404A_ABST
Patent Text Reader

Abstract

The system may be used to determine object characteristics and / or generate bounding boxes for objects in a vehicle scene by enhancing a later time feature map using an earlier time feature map. The system may generate a feature map from the received. Using the earlier time feature map, the system may enhance semantic data of the generated feature map to form an enhanced feature map. The system may use the enhanced feature map to generate one or more object characteristics of the object in the scene.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims priority to U.S. Patent Application No. 17 / 929,405 (Attorney Docket No. MOTN.094A), filed on September 2, 2022, titled "ENRICHING LATER - IN - TIME FEATUREMAPS USING EARLIER - IN - TIME FEATURE MAPS", which is hereby incorporated by reference in its entirety for all purposes. Background Art

[0003] An autonomous vehicle can use images obtained from one or more image sensors to determine object characteristics and / or generate bounding boxes for objects in a vehicle scene. Brief Description of the Drawings

[0004] Figure 1 is an example environment of a vehicle that can implement one or more components of an autonomous system.

[0005] Figure 2 is a diagram of one or more systems of a vehicle including an autonomous system.

[0006] Figure 3 is Figure 1 and Figure 2 is a diagram of one or more devices and / or components of one or more systems.

[0007] Figure 4A is a diagram of certain components of an autonomous system.

[0008] Figure 4B is a diagram of an implementation of a neural network.

[0009] Figure 4C and Figure 4D are diagrams illustrating example operations of a CNN.

[0010] Figure 5 is a block diagram illustrating an example perception environment in which a perception system receives and processes images to determine one or more characteristics of an object in a vehicle scene.

[0011] Figure 6A and Figure 6B are data flow diagrams illustrating examples of a perception environment in which a perception system determines one or more characteristics of an object in a vehicle scene.

[0012] Figure 7It is a flowchart illustrating an example of a routine implemented by at least one processor to determine object characteristics of an object in a vehicle scenario. Detailed Description

[0013] In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. It will be apparent, however, that the embodiments described herein may be practiced without these specific details. In some instances, well-known structures and devices are illustrated in block diagram form to avoid unnecessarily obscuring aspects of the present disclosure.

[0014] In the drawings, for ease of description, the specific arrangements or orderings of schematic elements (such as those representing systems, devices, modules, instruction blocks, and / or data elements, etc.) are illustrated. However, those skilled in the art will understand that, unless explicitly described, the specific orderings or arrangements of schematic elements in the drawings are not intended to imply a required processing order or sequence, or a separation of processing. Additionally, unless explicitly described, the inclusion of schematic elements in the drawings is not intended to imply that such elements are required in all embodiments, nor that the features represented by such elements cannot be included in some embodiments or cannot be combined with other elements in some embodiments.

[0015] Furthermore, in the drawings, connecting elements (such as solid lines, dashed lines, or arrows, etc.) are used to illustrate connections, relationships, or associations between or among two or more other schematic elements. The absence of any such connecting element is not intended to imply that no connection, relationship, or association can exist. In other words, some connections, relationships, or associations between elements are not illustrated in the drawings so as not to obscure the present disclosure. Additionally, for ease of illustration, a single connecting element may be used to represent multiple connections, relationships, or associations between elements. For example, if a connecting element represents the communication of a signal, data, or instruction (e.g., "software instruction"), those skilled in the art should understand that such an element may represent one or more than one signal path (e.g., a bus) that may be required to affect the communication.

[0016] Although terms such as "first," "second," and / or "third," etc. are used to describe various elements, these elements should not be limited by these terms. The terms "first," "second," and / or "third" are only used to distinguish one element from another. For example, without departing from the scope of the described embodiments, a first contact may be referred to as a second contact, and similarly, a second contact may be referred to as a first contact. Both the first contact and the second contact are contacts, but they are not the same contact.

[0017] The terms used in the description of the various embodiments described herein are included only for the purpose of describing a particular embodiment and are not intended to be limiting. As used in the description of the various embodiments and the appended claims, the singular forms "a", "an", and "the" are also intended to include the plural forms and may be used interchangeably with "one or more than one" or "at least one", unless the context clearly dictates otherwise. It will also be understood that the term "and / or" as used herein refers to and includes any and all possible combinations of one or more of the associated listed items. In addition, the term "or" is used in its inclusive sense (and not in its exclusive sense) such that when used, for example, to connect a list of elements, the term "or" means one, some, or all of the elements in the list. Unless specifically stated otherwise, disjunctive language such as the phrase "at least one of X, Y, and Z" is generally understood, in the context in which it is used, to mean that the items, terms, etc. can be X, Y, or Z, or any combination thereof (e.g., X, Y, or Z). Thus, such disjunctive language is generally not intended to and should not imply that certain embodiments require the presence of at least one of each of X, at least one of Y, and at least one of Z. It will also be understood that when the terms "comprises", "comprising", "includes", and / or "including" are used in this specification, they specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0018] As used herein, the terms "communicate" and "communicating" refer to at least one of receiving, receipt, transmission, conveyance, and / or provision of information (or information represented by, for example, data, signals, messages, instructions, and / or commands, etc.). For a unit (e.g., a device, a system, a component of a device or system, and / or a combination thereof, etc.) that is to communicate with another unit, this means that the unit can receive information from and / or send (e.g., transmit) information to the other unit, either directly or indirectly. This can refer to a direct or indirect connection that is inherently wired and / or wireless. Additionally, two units can communicate with each other even if the information transmitted between the first unit and the second unit is modified, processed, relayed, and / or routed. For example, even if the first unit receives information passively and does not actively transmit information to the second unit, the first unit can communicate with the second unit. As another example, if at least one intermediary unit (e.g., a third unit located between the first unit and the second unit) processes the information received from the first unit and transmits the processed information to the second unit, the first unit can communicate with the second unit. In some embodiments, a message can refer to a network packet (e.g., a data packet, etc.) that includes data.

[0019] As used herein, conditional language such as "can", "could", "might", "may", "e.g.", and the like, unless specifically stated otherwise or otherwise understood within the context in which it is used, is generally intended to convey that certain embodiments include certain features, elements, or steps, while other embodiments do not include certain features, elements, or steps. Thus, such conditional language is generally not intended to imply that the features, elements, or steps are in any way required for one or more embodiments, or that one or more embodiments necessarily include logic for deciding whether to include or perform these features, elements, or steps in any particular embodiment in the presence or absence of other inputs or cues. As used herein, depending on the context, the term "if" is optionally interpreted to mean "when", "upon", "in response to determining that", and / or "in response to detecting", etc. Similarly, depending on the context, the phrases "if it has been determined" or "if [stated condition or event] has been detected" are optionally interpreted to mean "upon determining", "in response to determining that", or "upon detecting [stated condition or event]" and / or "in response to detecting [stated condition or event]", etc. Further, as used herein, the terms "has", "have", or "having", etc., are intended to be open-ended terms. Additionally, unless expressly stated otherwise, the phrase "based on" is intended to mean "at least partially based on".

[0020] Reference will now be made in detail to the embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the various described embodiments. However, it will be apparent to one of ordinary skill in the art that the various described embodiments may be practiced without these specific details. In other instances, well-known methods, procedures, components, circuits, and networks have not been described in detail so as not to unnecessarily obscure aspects of the embodiments.

[0021] Overview

[0022] To effectively navigate through various scenarios, an autonomous vehicle uses computer vision to identify objects in the scene and then navigates the scene based on the identified objects. As part of the navigation process, the autonomous vehicle can determine the characteristics of the objects in the image (e.g., depth, speed, direction, size, etc.) and / or draw (3D) bounding boxes around the objects to understand the spatial relationship between the objects and the autonomous vehicle and navigate the path through the scene.

[0023] Accurately determining object characteristics in a real-time driving environment and drawing accurate 3D bounding boxes around objects can be challenging, in some cases because neural networks are unable to obtain sufficient semantic and local information related to the objects.

[0024] To address these issues, an autonomous vehicle can use multiple images in an image stream from the same image sensor (e.g., from consecutive images) to generate multiple feature maps, and use a feature map generated from an earlier image (also referred to herein as an earlier feature map or an earlier-time feature map) to enrich a feature map of a later image (also referred to herein as a later feature map or a later-time feature map).

[0025] By using the earlier-time feature map to enrich the later-time feature map, the autonomous vehicle can generate a feature map with additional (or richer) data. This can improve the autonomous vehicle's ability to determine the characteristics, bounding boxes, and / or object trajectories of objects within the autonomous vehicle's scene. For example, the enriched later-time feature map can improve the autonomous vehicle's ability to determine the depth, speed, centerness, etc. of an object and to determine an accurate bounding box. In turn, the improved characteristics and / or bounding boxes can enhance the autonomous vehicle's ability to navigate a particular scene safely (e.g., without collisions) or more comfortably (e.g., avoiding large accelerations / decelerations).

[0026] General Overview

[0027] By virtue of the implementations of the systems, methods, and computer program products described herein, an autonomous vehicle can more accurately identify objects within an image, more accurately identify the locations of the identified objects within the image, more accurately predict the trajectories of the identified objects within the image, determine additional features of the identified objects, and infer additional information related to the scene of the image.

[0028] Now refer to Figure 1, an exemplary environment 100 is illustrated, in which vehicles including autonomous systems and vehicles not including autonomous systems operate. As illustrated, environment 100 includes vehicles 102a - 102n, objects 104a - 104n, routes 106a - 106n, area 108, vehicle - to - infrastructure (V2I) devices 110, network 112, remote autonomous vehicle (AV) system 114, queue management system 116, and V2I system 118. Vehicles 102a - 102n, vehicle - to - infrastructure (V2I) devices 110, network 112, autonomous vehicle (AV) system 114, queue management system 116, and V2I system 118 are interconnected via a wired connection, a wireless connection, or a combination of wired and wireless connections (e.g., establishing a connection for communication, etc.). In some embodiments, objects 104a - 104n are interconnected with at least one of vehicles 102a - 102n, vehicle - to - infrastructure (V2I) devices 110, network 112, autonomous vehicle (AV) system 114, queue management system 116, and V2I system 118 via a wired connection, a wireless connection, or a combination of wired and wireless connections.

[0029] Vehicles 102a - 102n (individually referred to as vehicle 102 and collectively referred to as vehicles 102) include at least one device configured to transport goods and / or people. In some embodiments, vehicle 102 is configured to communicate with V2I device 110, remote AV system 114, queue management system 116, and / or V2I system 118 via network 112. In some embodiments, vehicles 102 include cars, buses, trucks, and / or trains, etc. In some embodiments, vehicle 102 is the same as or similar to vehicle 200 described herein (see Figure 2 ). In some embodiments, vehicles 200 in the set of vehicles 200 are associated with an autonomous queue manager. In some embodiments, as described herein, vehicle 102 travels along corresponding routes 106a - 106n (individually referred to as route 106 and collectively referred to as routes 106). In some embodiments, one or more than one vehicle 102 includes an autonomous system (e.g., an autonomous system the same as or similar to autonomous system 202).

[0030] Objects 104a - 104n (individually referred to as object 104 and collectively as objects 104) include, for example, at least one vehicle, at least one pedestrian, at least one cyclist, and / or at least one structure (e.g., building, sign, fire hydrant, etc.). Each object 104 (e.g., located at a fixed location and over a period of time) is stationary or (e.g., having a speed and associated with at least one trajectory) moving. In some embodiments, object 104 is associated with a corresponding location in region 108.

[0031] Routes 106a - 106n (individually referred to as route 106 and collectively as routes 106) are each associated with (e.g., define) a sequence of actions (also referred to as a trajectory) that a connected AV can navigate along. Each route 106 begins at an initial state (e.g., a state corresponding to a first spatio - temporal location and / or speed, etc.) and ends at a final goal state (e.g., a state corresponding to a second spatio - temporal location different from the first spatio - temporal location) or a target zone (e.g., a subspace of acceptable states (e.g., termination states)). In some embodiments, the first state includes a location where one or more individuals will board the AV, and the second state or zone includes one or more locations where one or more individuals boarding the AV will disembark. In some embodiments, route 106 includes multiple acceptable sequences of states (e.g., multiple sequences of spatio - temporal locations) that are associated with (e.g., define) multiple trajectories. In an example, route 106 includes only high - level actions or imprecise state locations, such as a series of connected roads indicating a direction change at a roadway intersection. Additionally or alternatively, route 106 can include more precise actions or states, such as a specific target lane or precise location within a lane region and a target rate at those locations. In an example, route 106 includes multiple precise state sequences along at least one high - level action with a finite look - ahead horizon to reach an intermediate goal, where the combination of successive iterations of the finite - horizon state sequences cumulatively corresponds to multiple trajectories that together form a high - level route terminating at the final goal state or zone.

[0032] Region 108 includes a physical region (e.g., a geographical region) in which vehicle 102 can navigate. In an example, region 108 includes at least one state (e.g., a country, a province, an individual state among a plurality of states included in a country, etc.), at least a portion of a state, at least one city, at least a portion of a city, etc. In some embodiments, region 108 includes at least one named arterial road (referred to herein as a "road"), such as a highway, an interstate highway, a parkway, an urban street, etc. Additionally or alternatively, in some examples, region 108 includes at least one unnamed road, such as a lane, a section of a parking lot, a section of a vacant and / or undeveloped area, a dirt road, etc. In some embodiments, a road includes at least one lane (e.g., a portion of the road through which vehicle 102 can pass). In an example, a road includes at least one lane associated with (e.g., identified based on) at least one lane marking line.

[0033] A vehicle-to-infrastructure (V2I) device 110 (sometimes referred to as a vehicle-to-everything (V2X) device) includes at least one device configured to communicate with vehicle 102 and / or V2I system 118. In some embodiments, V2I device 110 is configured to communicate with vehicle 102, remote AV system 114, queue management system 116, and / or V2I system 118 via network 112. In some embodiments, V2I device 110 includes a radio frequency identification (RFID) device, a sign, a camera (e.g., a two-dimensional (2D) and / or three-dimensional (3D) camera), a lane marking, a street light, a parking meter, etc. In some embodiments, V2I device 110 is configured to communicate directly with vehicle 102. Additionally or alternatively, in some embodiments, V2I device 110 is configured to communicate with vehicle 102, remote AV system 114, and / or queue management system 116 via V2I system 118. In some embodiments, V2I device 110 is configured to communicate with V2I system 118 via network 112.

[0034] Network 112 includes one or more wired and / or wireless networks. In an example, network 112 includes a cellular network (e.g., a Long-Term Evolution (LTE) network, a third-generation (3G) network, a fourth-generation (4G) network, a fifth-generation (5G) network, a Code Division Multiple Access (CDMA) network, etc.), a Public Land Mobile Network (PLMN), a Local Area Network (LAN), a Wide Area Network (WAN), a Metropolitan Area Network (MAN), a telephone network (e.g., a Public Switched Telephone Network (PSTN)), a private network, an ad hoc network, an intranet, the Internet, a fiber-optic-based network, a cloud computing network, etc., and / or a combination of some or all of these networks.

[0035] The remote AV system 114 includes at least one device configured to communicate with the vehicle 102, the V2I device 110, the network 112, the remote AV system 114, the queue management system 116, and / or the V2I system 118 via the network 112. In an example, the remote AV system 114 includes a server, a server group, and / or other similar devices. In some embodiments, the remote AV system 114 is co-located with the queue management system 116. In some embodiments, the remote AV system 114 participates in the installation of some or all of the components of the vehicle (including autonomous systems, autonomous vehicle computing, and / or software implemented by autonomous vehicle computing, etc.). In some embodiments, the remote AV system 114 maintains (e.g., updates and / or replaces) these components and / or software during the life of the vehicle.

[0036] The queue management system 116 includes at least one device configured to communicate with the vehicle 102, the V2I device 110, the remote AV system 114, and / or the V2I system 118. In an example, the queue management system 116 includes a server, a server group, and / or other similar devices. In some embodiments, the queue management system 116 is associated with a ridesharing company (e.g., an organization for controlling the operation of multiple vehicles (e.g., vehicles including autonomous systems and / or vehicles not including autonomous systems), etc.).

[0037] In some embodiments, the V2I system 118 includes at least one device configured to communicate with the vehicle 102, the V2I device 110, the remote AV system 114, and / or the queue management system 116 via the network 112. In some examples, the V2I system 118 is configured to communicate with the V2I device 110 via a connection different from the network 112. In some embodiments, the V2I system 118 includes a server, a server group, and / or other similar devices. In some embodiments, the V2I system 118 is associated with a municipal authority or a private institution (e.g., a private institution for maintaining the V2I device 110, etc.).

[0038] Provide Figure 1 The number and arrangement of the illustrated elements are provided as examples. Compared with Figure 1 the illustrated elements, there may be additional elements, fewer elements, different elements, and / or elements with a different arrangement. Additionally or alternatively, at least one element of the environment 100 may perform one or more functions described as being performed by Figure 1 at least one different element. Additionally or alternatively, at least one set of elements of the environment 100 may perform one or more functions described as being performed by at least one different set of elements of the environment 100.

[0039] Now refer to Figure 2 , the vehicle 200 includes an autonomous system 202, a powertrain control system 204, a steering control system 206, and a braking system 208. In some embodiments, the vehicle 200 is the same as or similar to the vehicle 102 (see Figure 1 ). In some embodiments, the vehicle 200 has autonomous capabilities (e.g., implementing at least one of the following functions, features, and / or devices, etc., the at least one function, feature, and / or device enabling the vehicle 200 to operate partially or fully without human intervention, including but not limited to a fully autonomous vehicle (e.g., a vehicle that abandons reliance on human intervention) and / or a highly autonomous vehicle (e.g., a vehicle that abandons reliance on human intervention in certain situations), etc.). For a detailed description of fully autonomous vehicles and highly autonomous vehicles, reference may be made to SAE International Standard J3016: Taxonomy and Definitions for Terms Related to On-Road Motor Vehicle Automated Driving Systems, the entire content of which is incorporated herein by reference. In some embodiments, the vehicle 200 is associated with an autonomous queue manager and / or a ridesharing company.

[0040] The autonomous system 202 includes a sensor suite that includes one or more devices such as a camera 202a, a LiDAR sensor 202b, a Radar sensor 202c, and a microphone 202d. In some embodiments, the autonomous system 202 may include more or fewer devices and / or different devices (e.g., ultrasonic sensors, inertial sensors, GPS receivers (discussed below), and / or odometer sensors for generating data associated with an indication of the distance traveled by the vehicle 200, etc.). In some embodiments, the autonomous system 202 uses one or more devices included in the autonomous system 202 to generate data associated with the environment 100 described herein. The data generated by one or more devices of the autonomous system 202 may be used by one or more systems described herein to observe the environment (e.g., environment 100) in which the vehicle 200 is located. In some embodiments, the autonomous system 202 includes a communication device 202e, an autonomous vehicle computing 202f, and a drive-by-wire (DBW) system 202h.

[0041] The camera 202a includes being configured to communicate with the communication device 202e, the autonomous vehicle computing 202f, and / or the safety controller 202g via a bus (e.g., withFigure 3 at least one device that communicates over a bus 302 that is the same as or similar to the bus). The camera 202a includes at least one camera (e.g., a digital camera using an optical sensor such as a charge-coupled device (CCD), a thermal camera, an infrared (IR) camera, and / or an event camera, etc.) configured to capture an image including a physical object (e.g., a car, a bus, a curb, and / or a person, etc.). In some embodiments, the camera 202a generates camera data as output. In some examples, the camera 202a generates camera data including image data associated with the image. In this example, the image data may specify at least one parameter corresponding to the image (e.g., image characteristics such as exposure, brightness, etc., and / or an image timestamp, etc.). In such an example, the image may be in a format (e.g., RAW, JPEG, and / or PNG, etc.). In some embodiments, the camera 202a includes a plurality of independent cameras configured on (e.g., positioned on) a vehicle to capture images for the purpose of stereovision (stereo vision). In some examples, the camera 202a includes a plurality of cameras that generate image data and transmit the image data to the autonomous vehicle computing 202f and / or a queue management system (e.g., a queue management system that is the same as or similar to the Figure 1 queue management system 116). In such an example, the autonomous vehicle computing 202f determines the depth to one or more objects in the fields of view of at least two of the plurality of cameras based on the image data from at least two cameras. In some embodiments, the camera 202a is configured to capture images of objects within a distance (e.g., up to 100 meters and / or up to 1 kilometer, etc.) relative to the camera 202a. Thus, the camera 202a includes features such as sensors and lenses optimized to sense objects at one or more distances relative to the camera 202a.

[0042] In an embodiment, the camera 202a includes at least one camera configured to capture one or more images associated with one or more traffic lights, street signs, and / or other physical objects that provide visual navigation information. In some embodiments, the camera 202a generates traffic light data associated with one or more images. In some examples, the camera 202a generates TLD data associated with one or more images including a format (e.g., RAW, JPEG, and / or PNG, etc.). In some embodiments, the camera 202a that generates TLD data is different from other systems incorporating cameras described herein in that: the camera 202a may include one or more cameras having a wide field of view (e.g., a wide-angle lens, a fish-eye lens, and / or a lens having a viewing angle of about 120 degrees or greater, etc.) to generate images related to as many physical objects as possible.

[0043] A Light Detection and Ranging (LiDAR) sensor 202b includes at least one device configured to communicate with a communication device 202e, an autonomous vehicle computing 202f, and / or a safety controller 202g via a bus (e.g., a bus identical or similar to bus 302 of Figure 3 . The LiDAR sensor 202b includes a system configured to emit light from a light emitter (e.g., a laser emitter). The light emitted by the LiDAR sensor 202b includes light outside the visible spectrum (e.g., infrared light, etc.). In some embodiments, during operation, the light emitted by the LiDAR sensor 202b encounters a physical object (e.g., a vehicle) and is reflected back to the LiDAR sensor 202b. In some embodiments, the light emitted by the LiDAR sensor 202b does not penetrate the physical object it encounters. The LiDAR sensor 202b further includes at least one light detector that detects the light after the light emitted from the light emitter encounters a physical object. In some embodiments, at least one data processing system associated with the LiDAR sensor 202b generates an image (e.g., a point cloud and / or a combined point cloud, etc.) representing the objects included in the field of view of the LiDAR sensor 202b. In some examples, at least one data processing system associated with the LiDAR sensor 202b generates an image representing the boundary of a physical object and / or the surface of a physical object (e.g., the topology of the surface, etc.). In such examples, the image is used to determine the boundary of the physical object in the field of view of the LiDAR sensor 202b.

[0044] A Radio Detection and Ranging (Radar) sensor 202c includes at least one device configured to communicate with a communication device 202e, an autonomous vehicle computing 202f, and / or a safety controller 202g via a bus (e.g., a bus identical or similar to Figure 3at least one device that communicates via a bus 302 that is the same as or similar to the bus). The Radar sensor 202c includes a system configured to emit (pulsed or continuous) radio waves. The radio waves emitted by the Radar sensor 202c include radio waves within a predetermined spectrum. In some embodiments, during operation, the radio waves emitted by the Radar sensor 202c encounter physical objects and are reflected back to the Radar sensor 202c. In some embodiments, the radio waves emitted by the Radar sensor 202c are not reflected by some objects. In some embodiments, at least one data processing system associated with the Radar sensor 202c generates a signal representing the objects included in the field of view of the Radar sensor 202c. For example, at least one data processing system associated with the Radar sensor 202c generates an image representing the boundaries of the physical objects and / or the surface of the physical objects (e.g., the topology of the surface), etc. In some examples, the image is used to determine the boundaries of the physical objects in the field of view of the Radar sensor 202c.

[0045] The microphone 202d includes at least one device configured to communicate with the communication device 202e, the autonomous vehicle computing 202f, and / or the safety controller 202g via a bus (e.g., a bus that is the same as or similar to Figure 3 the bus 302). The microphone 202d includes one or more microphones (e.g., an array microphone and / or an external microphone, etc.) that capture an audio signal and generate data associated with (e.g., representing) the audio signal. In some examples, the microphone 202d includes a transducer device and / or a similar device. In some embodiments, one or more of the systems described herein may receive the data generated by the microphone 202d and determine the position (e.g., distance, etc.) of the object relative to the vehicle 200 based on the audio signal associated with the data.

[0046] The communication device 202e includes at least one device configured to communicate with the camera 202a, the LiDAR sensor 202b, the Radar sensor 202c, the microphone 202d, the autonomous vehicle computing 202f, the safety controller 202g, and / or the drive-by-wire (DBW) system 202h. For example, the communication device 202e may include a device that is the same as or similar to Figure 3 the communication interface 314. In some embodiments, the communication device 202e includes a vehicle-to-vehicle (V2V) communication device (e.g., a device for enabling wireless communication of data between vehicles).

[0047] The autonomous vehicle computing 202f includes at least one device configured to communicate with a camera 202a, a LiDAR sensor 202b, a Radar sensor 202c, a microphone 202d, a communication device 202e, a safety controller 202g, and / or a DBW system 202h. In some examples, the autonomous vehicle computing 202f includes devices such as client devices, mobile devices (e.g., cellular phones and / or tablets, etc.), and / or servers (e.g., computing devices including one or more central processing units and / or graphics processing units, etc.). In some embodiments, the autonomous vehicle computing 202f is the same as or similar to the autonomous vehicle computing 400 described herein. Additionally or alternatively, in some embodiments, the autonomous vehicle computing 202f is configured to communicate with an autonomous vehicle system (e.g., an autonomous vehicle system the same as or similar to the remote AV system 114 of Figure 1 , a queue management system (e.g., a queue management system the same as or similar to the queue management system 116 of Figure 1 ), a V2I device (e.g., a V2I device the same as or similar to the V2I device 110 of Figure 1 ), and / or a V2I system (e.g., a V2I system the same as or similar to the V2I system 118 of Figure 1 ).

[0048] The safety controller 202g includes at least one device configured to communicate with a camera 202a, a LiDAR sensor 202b, a Radar sensor 202c, a microphone 202d, a communication device 202e, the autonomous vehicle computing 202f, and / or a DBW system 202h. In some examples, the safety controller 202g includes one or more controllers (such as electrical controllers and / or electromechanical controllers, etc.) configured to generate and / or transmit control signals to operate one or more devices of the vehicle 200 (e.g., a powertrain control system 204, a steering control system 206, and / or a braking system 208, etc.). In some embodiments, the safety controller 202g is configured to generate control signals that are prior to (e.g., override) the control signals generated and / or transmitted by the autonomous vehicle computing 202f.

[0049] The DBW system 202h includes at least one device configured to communicate with the communication device 202e and / or the autonomous vehicle computing 202f. In some examples, the DBW system 202h includes one or more controllers (such as an electrical controller and / or an electromechanical controller, etc.) configured to generate and / or transmit control signals to operate one or more devices of the vehicle 200 (such as the powertrain control system 204, the steering control system 206, and / or the braking system 208, etc.). Additionally or alternatively, one or more controllers of the DBW system 202h are configured to generate and / or transmit control signals to operate at least one different device of the vehicle 200 (such as turn signals, headlights, door locks, and / or windshield wipers, etc.).

[0050] The powertrain control system 204 includes at least one device configured to communicate with the DBW system 202h. In some examples, the powertrain control system 204 includes at least one controller and / or actuator, etc. In some embodiments, the powertrain control system 204 receives control signals from the DBW system 202h, and the powertrain control system 204 causes the vehicle 200 to start moving forward, stop moving forward, start moving backward, stop moving backward, accelerate in a certain direction, decelerate in a certain direction, turn left, and / or turn right, etc. In an example, the powertrain control system 204 causes the energy (such as fuel and / or electricity, etc.) provided to the motor of the vehicle to increase, remain the same, or decrease, thereby causing at least one wheel of the vehicle 200 to rotate or not rotate.

[0051] The steering control system 206 includes at least one device configured to rotate one or more wheels of the vehicle 200. In some examples, the steering control system 206 includes at least one controller and / or actuator, etc. In some embodiments, the steering control system 206 causes two front wheels and / or two rear wheels of the vehicle 200 to rotate left or right, so that the vehicle 200 turns left or right.

[0052] The braking system 208 includes at least one device configured to actuate one or more brakes to decelerate the vehicle 200 and / or keep it stationary. In some examples, the braking system 208 includes at least one controller and / or actuator configured to close one or more calipers associated with one or more wheels of the vehicle 200 on the corresponding rotors of the vehicle 200. Additionally or alternatively, in some examples, the braking system 208 includes an automatic emergency braking (AEB) system and / or a regenerative braking system, etc.

[0053] In some embodiments, vehicle 200 includes at least one platform sensor (not explicitly illustrated) for measuring or inferring properties of the state or condition of vehicle 200. In some examples, vehicle 200 includes platform sensors such as a global positioning system (GPS) receiver, an inertial measurement unit (IMU), a wheel speed sensor, a wheel brake pressure sensor, a wheel torque sensor, an engine torque sensor, and / or a steering angle sensor.

[0054] Now refer to Figure 3 , a schematic diagram of illustrative device 300. As illustrated, device 300 includes processor 304, memory 306, storage component 308, input interface 310, output interface 312, communication interface 314, and bus 302. In some embodiments, device 300 corresponds to: at least one device of vehicle 102 (e.g., at least one device of a system of vehicle 102); and / or one or more than one device of network 112 (e.g., one or more than one device of a system of network 112). In some embodiments, one or more than one device of vehicle 102 (e.g., at least one device of a system of vehicle 102), and / or one or more than one device of network 112 (e.g., one or more than one device of a system of network 112) includes at least one device 300 and / or at least one component of device 300. As Figure 3 shown, device 300 includes bus 302, processor 304, memory 306, storage component 308, input interface 310, output interface 312, and communication interface 314.

[0055] Bus 302 includes components that permit communication between components of device 300. In some cases, processor 304 includes a processor (e.g., a central processing unit (CPU), a graphics processing unit (GPU), and / or an accelerated processing unit (APU), etc.), a microphone, a digital signal processor (DSP), and / or any processing component that can be programmed to perform at least one function (e.g., a field programmable gate array (FPGA) and / or an application specific integrated circuit (ASIC), etc.). Memory 306 includes random access memory (RAM), read only memory (ROM), and / or another type of dynamic and / or static storage device that stores data and / or instructions for use by processor 304 (e.g., flash memory, magnetic memory, and / or optical memory, etc.).

[0056] The storage component 308 stores data and / or software related to the operation and use of the device 300. In some examples, the storage component 308 includes a hard disk (e.g., a magnetic disk, an optical disk, a magneto-optical disk, and / or a solid state disk, etc.), a compact disk (CD), a digital versatile disk (DVD), a floppy disk, a cassette tape, a magnetic tape, a CD-ROM, a RAM, a PROM, an EPROM, a FLASH-EPROM, an NV-RAM, and / or another type of computer-readable medium, and corresponding drives.

[0057] The input interface 310 includes components that permit the device 300 to receive information such as via a user input (e.g., a touch screen display, a keyboard, a keypad, a mouse, a button, a switch, a microphone, and / or a camera, etc.). Additionally or alternatively, in some embodiments, the input interface 310 includes sensors for sensing information (e.g., a global positioning system (GPS) receiver, an accelerometer, a gyroscope, and / or an actuator, etc.). The output interface 312 includes components for providing output information from the device 300 (e.g., a display, a speaker, and / or one or more light emitting diodes (LEDs), etc.).

[0058] In some embodiments, the communication interface 314 includes transceiver-like components that permit the device 300 to communicate with other devices via a wired connection, a wireless connection, or a combination of a wired connection and a wireless connection (e.g., a transceiver and / or separate receivers and transmitters, etc.). In some examples, the communication interface 314 permits the device 300 to receive information from another device and / or provide information to another device. In some examples, the communication interface 314 includes an Ethernet interface, an optical interface, a coaxial interface, an infrared interface, a radio frequency (RF) interface, a universal serial bus (USB) interface, an interface, and / or a cellular network interface, etc.

[0059] In some embodiments, the device 300 performs one or more processes described herein. The device 300 performs these processes based on software instructions executed by a processor 304 that are stored on a computer-readable medium such as the memory 306 and / or the storage component 308. A computer-readable medium (e.g., a non-transitory computer-readable medium) is defined herein as a non-transitory memory device. A non-transitory memory device includes a storage space located within a single physical storage device or a storage space distributed across multiple physical storage devices.

[0060] In some embodiments, software instructions are read into memory 306 and / or storage component 308 from another computer-readable medium or from another device via communication interface 314. When executed, the software instructions stored in memory 306 and / or storage component 308 cause processor 304 to perform one or more processes described herein. Additionally or alternatively, hardwired circuitry is used in place of or in combination with software instructions to perform one or more processes described herein. Accordingly, unless otherwise expressly stated, the embodiments described herein are not limited to any specific combination of hardware circuitry and software.

[0061] Memory 306 and / or storage component 308 includes a data store or at least one data structure (e.g., a database, etc.). Device 300 is capable of receiving information from the data store or at least one data structure in memory 306 or storage component 308, storing information in the data store or at least one data structure, communicating information to the data store or at least one data structure, or searching for information stored in the data store or at least one data structure. In some examples, the information includes network data, input data, output data, or any combination thereof.

[0062] In some embodiments, device 300 is configured to execute software instructions stored in memory 306 and / or the memory of another device (e.g., another device that is the same as or similar to device 300). As used herein, the term “module” refers to at least one instruction stored in memory 306 and / or the memory of another device, which when executed by processor 304 and / or the processor of another device (e.g., another device that is the same as or similar to device 300), causes device 300 (e.g., at least one component of device 300) to perform one or more processes described herein. In some embodiments, the module is implemented in software, firmware, and / or hardware, etc.

[0063] Provided Figure 3 The number and arrangement of the illustrated components are provided as an example. In some embodiments, compared to Figure 3 the illustrated components, device 300 may include additional components, fewer components, different components, or components arranged differently. Additionally or alternatively, a set of components of device 300 (e.g., one or more components) may perform one or more functions described as being performed by another component or another set of components of device 300.

[0064] Now refer to Figure 4A, an example block diagram of an autonomous vehicle computing 400 (sometimes referred to as an "AV stack") is illustrated. As illustrated, the autonomous vehicle computing 400 includes a perception system 402 (sometimes referred to as a perception module), a planning system 404 (sometimes referred to as a planning module), a positioning system 406 (sometimes referred to as a positioning module), a control system 408 (sometimes referred to as a control module), and a database 410. In some embodiments, the perception system 402, the planning system 404, the positioning system 406, the control system 408, and the database 410 are included in and / or implemented in an automatic navigation system of the vehicle (e.g., the autonomous vehicle computing 202f of the vehicle 200). Additionally or alternatively, in some embodiments, the perception system 402, the planning system 404, the positioning system 406, the control system 408, and the database 410 are included in one or more independent systems (e.g., one or more systems the same or similar to the autonomous vehicle computing 400, etc.). In some examples, the perception system 402, the planning system 404, the positioning system 406, the control system 408, and the database 410 are included in one or more independent systems located in the vehicle and / or at least one remote system as described herein. In some embodiments, any and / or all of the systems included in the autonomous vehicle computing 400 are implemented in software (e.g., software instructions stored in a memory), computer hardware (e.g., via a microprocessor, a microcontroller, an application specific integrated circuit (ASIC), and / or a field programmable gate array (FPGA), etc.), or a combination of computer software and computer hardware. It will also be understood that, in some embodiments, the autonomous vehicle computing 400 is configured to communicate with remote systems (e.g., an autonomous vehicle system the same or similar to the remote AV system 114, a queue management system the same or similar to the queue management system 116, and / or a V2I system the same or similar to the V2I system 118, etc.).

[0065] In some embodiments, the perception system 402 receives data associated with at least one physical object in the environment (e.g., data used by the perception system 402 to detect the at least one physical object), and classifies the at least one physical object. In some examples, the perception system 402 receives image data captured by at least one camera (e.g., camera 202a), the image being associated with one or more physical objects within the field of view of the at least one camera (e.g., representing the one or more physical objects). In such examples, the perception system 402 classifies the at least one physical object based on one or more groupings of physical objects (e.g., bicycles, vehicles, traffic signs, and / or pedestrians, etc.). In some embodiments, based on the classification of the physical objects by the perception system 402, the perception system 402 transmits data associated with the classification of the physical objects to the planning system 404.

[0066] In some embodiments, the planning system 404 receives data associated with a destination, and generates data associated with at least one route (e.g., route 106) along which a vehicle (e.g., vehicle 102) can travel toward the destination. In some embodiments, the planning system 404 periodically or continuously receives data from the perception system 402 (e.g., the data associated with the classification of the physical objects described above), and the planning system 404 updates at least one trajectory or generates at least one different trajectory based on the data generated by the perception system 402. In some embodiments, the planning system 404 receives data associated with an updated position of a vehicle (e.g., vehicle 102) from the positioning system 406, and the planning system 404 updates at least one trajectory or generates at least one different trajectory based on the data generated by the positioning system 406.

[0067] In some embodiments, the positioning system 406 receives data associated with (e.g., representing) the location of a vehicle (e.g., vehicle 102) in an area. In some examples, the positioning system 406 receives LiDAR data associated with at least one point cloud generated by at least one LiDAR sensor (e.g., LiDAR sensor 202b). In certain examples, the positioning system 406 receives data associated with at least one point cloud from multiple LiDAR sensors, and the positioning system 406 generates a combined point cloud based on the respective point clouds. In these examples, the positioning system 406 compares the at least one point cloud or combined point cloud with a two-dimensional (2D) and / or three-dimensional (3D) map of the area stored in the database 410. Then, based on the positioning system 406 comparing the at least one point cloud or combined point cloud with the map, the positioning system 406 determines the position of the vehicle in the area. In some embodiments, the map includes a combined point cloud of the area generated prior to the navigation of the vehicle. In some embodiments, the map includes, but is not limited to, a high-precision map of the roadway geometry, a map describing the connectivity of the road network, a map describing the physical properties of the roadways (such as traffic speed, traffic flow, the number of vehicle and bicycle traffic lanes, lane width, lane traffic direction or the type and location of lane markings, or a combination thereof, etc.), and a map describing the spatial location of road features (such as crosswalks, traffic signs, or various other driving signals, etc.). In some embodiments, the map is generated in real time based on the data received by the perception system.

[0068] In another example, the positioning system 406 receives global navigation satellite system (GNSS) data generated by a global positioning system (GPS) receiver. In some examples, the positioning system 406 receives GNSS data associated with the location of a vehicle in an area, and the positioning system 406 determines the latitude and longitude of the vehicle in the area. In such examples, the positioning system 406 determines the position of the vehicle in the area based on the latitude and longitude of the vehicle. In some embodiments, the positioning system 406 generates data associated with the position of the vehicle. In some examples, based on the positioning system 406 determining the position of the vehicle, the positioning system 406 generates data associated with the position of the vehicle. In such examples, the data associated with the position of the vehicle includes data associated with one or more semantic properties corresponding to the position of the vehicle.

[0069] In some embodiments, the control system 408 receives data associated with at least one trajectory from the planning system 404, and the control system 408 controls the operation of the vehicle. In some examples, the control system 408 receives data associated with at least one trajectory from the planning system 404, and the control system 408 controls the operation of the vehicle by generating and transmitting control signals to cause the powertrain control system (e.g., the DBW system 202h and / or the powertrain control system 204, etc.), the steering control system (e.g., the steering control system 206), and / or the braking system (e.g., the braking system 208) to operate. In an example, when the trajectory includes a left turn, the control system 408 transmits a control signal to cause the steering control system 206 to adjust the steering angle of the vehicle 200, thereby causing the vehicle 200 to turn left. Additionally or alternatively, the control system 408 generates and transmits control signals to cause other devices of the vehicle 200 (e.g., headlights, turn signals, door locks, and / or windshield wipers, etc.) to change states.

[0070] In some embodiments, the perception system 402, the planning system 404, the positioning system 406, and / or the control system 408 implement at least one machine learning model (e.g., at least one multi-layer perceptron (MLP), at least one convolutional neural network (CNN), at least one recurrent neural network (RNN), at least one autoencoder, and / or at least one transformer, etc.). In some examples, the perception system 402, the planning system 404, the positioning system 406, and / or the control system 408 implement at least one machine learning model individually or in combination with one or more of the above systems. In some examples, the perception system 402, the planning system 404, the positioning system 406, and / or the control system 408 implement at least one machine learning model as part of a pipeline (e.g., a pipeline for identifying one or more objects located in the environment, etc.). The following are examples of Figures 4B to 4D the implementation of the machine learning model.

[0071] The database 410 stores data transmitted to, received from, and / or updated by the perception system 402, the planning system 404, the positioning system 406, and / or the control system 408. In some examples, the database 410 includes a storage component for storing data and / or software related to operations and using the autonomous vehicle computing 400 of at least one system (e.g., related to Figure 3the same or similar storage components as the storage component 308). In some embodiments, the database 410 stores data associated with 2D and / or 3D maps of at least one region. In some examples, the database 410 stores data associated with 2D and / or 3D maps of a part of a city, multiple parts of multiple cities, multiple cities, counties, states, and / or countries (e.g., nations), etc. In such examples, a vehicle (e.g., a vehicle the same or similar to the vehicle 102 and / or the vehicle 200) can drive along one or more drivable areas (e.g., single-lane roads, multi-lane roads, highways, back roads, and / or off-road paths, etc.), and cause at least one LiDAR sensor (e.g., a LiDAR sensor the same or similar to the LiDAR sensor 202b) to generate data associated with an image representing the objects included in the field of view of the at least one LiDAR sensor.

[0072] In some embodiments, the database 410 can be implemented across multiple devices. In some examples, the database 410 is included in a vehicle (e.g., a vehicle the same or similar to the vehicle 102 and / or the vehicle 200), an autonomous vehicle system (e.g., an autonomous vehicle system the same or similar to the remote AV system 114), a queue management system (e.g., a queue management system the same or similar to Figure 1 the queue management system 116), and / or a V2I system (e.g., a V2I system the same or similar to Figure 1 the V2I system 118), etc.

[0073] Now refer to Figure 4B , a diagram illustrating the implementation of a machine learning model. More specifically, a diagram illustrating the implementation of a convolutional neural network (CNN) 420. For illustrative purposes, the following description of the CNN 420 will be with respect to implementing the CNN 420 by the perception system 402. However, it will be understood that in some examples, the CNN 420 (e.g., one or more components of the CNN 420) is implemented by other systems different from or in addition to the perception system 402, such as the planning system 404, the positioning system 406, and / or the control system 408, etc. Although the CNN 420 includes certain features as described herein, these features are provided for illustrative purposes and are not intended to limit the present disclosure.

[0074] The CNN 420 includes a plurality of convolutional layers including a first convolutional layer 422, a second convolutional layer 424, and a convolutional layer 426. In some embodiments, the CNN 420 includes a subsampling layer 428 (sometimes referred to as a pooling layer). In some embodiments, the subsampling layer 428 and / or other subsampling layers have dimensions smaller than the dimensions of the upstream system (i.e., the amount of nodes). By means of the subsampling layer 428 having dimensions smaller than the dimensions of the upstream layer, the CNN 420 combines the amount of data associated with the initial input and / or output of the upstream layer, thereby reducing the amount of computation required for the CNN 420 to perform downstream convolutional operations. Additionally or alternatively, by means of the subsampling layer 428 being associated with at least one subsampling function (e.g., being configured to perform at least one subsampling function) (as described below with respect to Figure 4C and Figure 4D ), the CNN 420 combines the amount of data associated with the initial input.

[0075] Based on the perception system 402 providing corresponding inputs and / or outputs respectively associated with the first convolutional layer 422, the second convolutional layer 424, and the convolutional layer 426 to generate corresponding outputs, the perception system 402 performs convolutional operations. In some examples, based on the perception system 402 providing data as inputs to the first convolutional layer 422, the second convolutional layer 424, and the convolutional layer 426, the perception system 402 implements the CNN 420. In such examples, based on the perception system 402 receiving data from one or more different systems (e.g., one or more systems of a vehicle the same or similar to the vehicle 102, a remote AV system the same or similar to the remote AV system 114, a queue management system the same or similar to the queue management system 116, and / or a V2I system the same or similar to the V2I system 118, etc.), the perception system 402 provides the data as inputs to the first convolutional layer 422, the second convolutional layer 424, and the convolutional layer 426. The following is a detailed description of Figure 4C including convolutional operations.

[0076] In some embodiments, the perception system 402 provides data associated with an input (referred to as an initial input) to the first convolutional layer 422, and the perception system 402 uses the first convolutional layer 422 to generate data associated with an output. In some embodiments, the perception system 402 provides the output generated by the convolutional layer as an input to a different convolutional layer. For example, the perception system 402 provides the output of the first convolutional layer 422 as an input to the subsampling layer 428, the second convolutional layer 424, and / or the convolutional layer 426. In such an example, the first convolutional layer 422 is referred to as an upstream layer, and the subsampling layer 428, the second convolutional layer 424, and / or the convolutional layer 426 are referred to as downstream layers. Similarly, in some embodiments, the perception system 402 provides the output of the subsampling layer 428 to the second convolutional layer 424 and / or the convolutional layer 426, and in this example, the subsampling layer 428 will be referred to as an upstream layer, and the second convolutional layer 424 and / or the convolutional layer 426 will be referred to as downstream layers.

[0077] In some embodiments, before the perception system 402 provides an input to the CNN 420, the perception system 402 processes data associated with the input provided to the CNN 420. For example, based on the perception system 402 normalizing sensor data (such as image data, LiDAR data, and / or Radar data, etc.), the perception system 402 processes data associated with the input provided to the CNN 420.

[0078] In some embodiments, based on the perception system 402 performing convolution operations associated with each convolutional layer, the CNN 420 generates an output. In some examples, based on the perception system 402 performing convolution operations associated with each convolutional layer and the initial input, the CNN 420 generates an output. In some embodiments, the perception system 402 generates an output and provides the output to the fully connected layer 430. In some examples, the perception system 402 provides the output of the convolutional layer 426 to the fully connected layer 430, where the fully connected layer 430 includes data associated with a plurality of eigenvalues referred to as F1, F2,..., FN. In this example, the output of the convolutional layer 426 includes data associated with a plurality of output eigenvalues representing predictions.

[0079] In some embodiments, based on the perception system 402 identifying an eigenvalue associated with the highest likelihood of being the correct prediction among a plurality of predictions, the perception system 402 identifies a prediction from among the plurality of predictions. For example, in the case where the fully connected layer 430 includes eigenvalues F1, F2, ..., FN and F1 is the largest eigenvalue, the perception system 402 identifies the prediction associated with F1 as the correct prediction among the plurality of predictions. In some embodiments, the perception system 402 trains the CNN 420 to generate predictions. In some examples, based on the perception system 402 providing training data associated with the predictions to the CNN 420, the perception system 402 trains the CNN 420 to generate predictions.

[0080] Now refer to Figure 4C and Figure 4D , a diagram illustrating an example operation of the CNN 440 that utilizes the perception system 402. In some embodiments, the CNN 440 (e.g., one or more components of the CNN 440) is the same as or similar to the CNN 420 (e.g., one or more components of the CNN 420) (see Figure 4B ).

[0081] In step 450, the perception system 402 provides data associated with an image as an input to the CNN 440 (step 450). For example, as illustrated, the perception system 402 provides data associated with an image to the CNN 440, where the image is a grayscale image represented as values stored in a two-dimensional (2D) array. In some embodiments, the data associated with the image may include data associated with a color image, which is represented as values stored in a three-dimensional (3D) array. Additionally or alternatively, the data associated with the image may include data associated with an infrared image and / or a Radar image, etc.

[0082] In step 455, the CNN 440 performs a first convolution function. For example, based on the CNN 440 providing the values representing the image as an input to one or more neurons (not explicitly illustrated) included in the first convolutional layer 442, the CNN 440 performs the first convolution function. In this example, the values representing the image may correspond to the values of a region (sometimes referred to as a receptive field) representing the image. In some embodiments, each neuron is associated with a filter (not explicitly illustrated). The filter (sometimes referred to as a kernel) may be represented as an array of values corresponding in size to the values provided as an input to the neuron. In one example, the filter may be configured to identify edges (e.g., horizontal lines, vertical lines, and / or straight lines, etc.). In successive convolutional layers, the filters associated with the neurons may be configured to successively identify more complex patterns (e.g., arcs and / or objects, etc.).

[0083] In some embodiments, based on the CNN 440, the values provided as input to each neuron among one or more neurons included in the first convolutional layer 442 are multiplied by the values of the filters corresponding to each neuron among the same one or more neurons, and the CNN 440 performs a first convolutional function. For example, the CNN 440 may multiply the values provided as input to each neuron among one or more neurons included in the first convolutional layer 442 by the values of the filters corresponding to each neuron among the same one or more neurons to generate a single value or an array of values as output. In some embodiments, the collective output of the neurons of the first convolutional layer 442 is referred to as the convolutional output. In some embodiments, when each neuron has the same filter, the convolutional output is referred to as a feature map.

[0084] In some embodiments, the CNN 440 provides the output of each neuron of the first convolutional layer 442 to the neurons of a downstream layer. For clarity, an upstream layer may be a layer that transmits data to a different layer (referred to as a downstream layer). For example, the CNN 440 may provide the output of each neuron of the first convolutional layer 442 to the corresponding neurons of a subsampling layer. In an example, the CNN 440 provides the output of each neuron of the first convolutional layer 442 to the corresponding neurons of the first subsampling layer 444. In some embodiments, the CNN 440 adds a bias value to the aggregated set of all values provided to each neuron of the downstream layer. For example, the CNN 440 adds a bias value to the aggregated set of all values provided to each neuron of the first subsampling layer 444. In such an example, the CNN 440 determines the final value to be provided to each neuron of the first subsampling layer 444 based on the aggregated set of all values provided to each neuron and the activation function associated with each neuron of the first subsampling layer 444.

[0085] In step 460, the CNN 440 performs a first subsampling function. For example, based on the CNN 440 providing the values output by the first convolutional layer 442 to the corresponding neurons of the first subsampling layer 444, the CNN 440 may perform a first subsampling function. In some embodiments, the CNN 440 performs the first subsampling function based on an aggregation function. In an example, based on the CNN 440 determining the maximum input among the values provided to a given neuron (referred to as the max pooling function), the CNN 440 performs the first subsampling function. In another example, based on the CNN 440 determining the average input among the values provided to a given neuron (referred to as the average pooling function), the CNN 440 performs the first subsampling function. In some embodiments, based on the CNN 440 providing values to each neuron of the first subsampling layer 444, the CNN 440 generates an output, which is sometimes referred to as the subsampled convolutional output.

[0086] At step 465, CNN 440 performs a second convolution function. In some embodiments, CNN 440 performs the second convolution function in a manner similar to how CNN 440 performs the first convolution function as described above. In some embodiments, based on the values output by the first subsampling layer 444 being provided as inputs to one or more neurons (not explicitly illustrated) included in the second convolutional layer 446, CNN 440 performs the second convolution function. In some embodiments, as described above, each neuron of the second convolutional layer 446 is associated with a filter. As described above, the (one or more) filters associated with the second convolutional layer 446 may be configured to identify more complex patterns compared to the filters associated with the first convolutional layer 442.

[0087] In some embodiments, based on CNN 440 multiplying the values provided as inputs to each of the one or more neurons included in the second convolutional layer 446 by the values of the filters corresponding to each of the one or more neurons, CNN 440 performs the second convolution function. For example, CNN 440 may multiply the values provided as inputs to each of the one or more neurons included in the second convolutional layer 446 by the values of the filters corresponding to each of the one or more neurons to generate a single value or an array of values as output.

[0088] In some embodiments, CNN 440 provides the output of each neuron of the second convolutional layer 446 to the neurons of the downstream layer. For example, CNN 440 may provide the output of each neuron of the first convolutional layer 442 to the corresponding neurons of the subsampling layer. In an example, CNN 440 provides the output of each neuron of the first convolutional layer 442 to the corresponding neurons of the second subsampling layer 448. In some embodiments, CNN 440 adds a bias value to the aggregate set of all values provided to each neuron of the downstream layer. For example, CNN 440 adds a bias value to the aggregate set of all values provided to each neuron of the second subsampling layer 448. In such an example, CNN 440 determines the final value provided to each neuron of the second subsampling layer 448 based on the aggregate set of all values provided to each neuron and the activation function associated with each neuron of the second subsampling layer 448.

[0089] At step 470, the CNN 440 performs a second subsampling function. For example, based on the values output by the second convolutional layer 446 being provided to the respective neurons of the second subsampling layer 448 by the CNN 440, the CNN 440 may perform the second subsampling function. In some embodiments, based on the CNN 440 using an aggregation function, the CNN 440 performs the second subsampling function. In an example, as described above, based on the CNN 440 determining the maximum input or average input among the values provided to a given neuron, the CNN 440 performs the first subsampling function. In some embodiments, based on the CNN 440 providing values to the respective neurons of the second subsampling layer 448, the CNN 440 generates an output.

[0090] At step 475, the CNN 440 provides the output of each neuron of the second subsampling layer 448 to the fully connected layer 449. For example, the CNN 440 provides the output of each neuron of the second subsampling layer 448 to the fully connected layer 449 such that the fully connected layer 449 generates an output. In some embodiments, the fully connected layer 449 is configured to generate an output associated with a prediction (sometimes referred to as classification). The prediction may include an indication that the objects included in the image provided as input to the CNN 440 include an object and / or a collection of objects, etc. In some embodiments, the perception system 402 performs one or more operations and / or provides data associated with the prediction to different systems described herein.

[0091] Generate a bounding box for navigation

[0092] As described herein, to improve the functionality of an autonomous vehicle and its ability to determine the characteristics of objects in a scene, generate bounding boxes, and / or navigate an environment in real time, the autonomous vehicle may be configured to use an earlier time feature map to enhance a later time feature map. By using the earlier time feature map to enhance the later time feature map, the semantic data of the earlier time feature map can provide context for the later time feature map, which enables the autonomous vehicle to more accurately detect objects and determine the characteristics of these objects (e.g., depth, speed, classification, size, offset, rotation, orientation, etc.) and / or generate bounding boxes for these objects.

[0093] Figure 5FIG. 0 is a block diagram illustrating an example perception environment 500 in which a perception system 402 receives and processes an image 502 to provide one or more object characteristics or (3D) bounding boxes 512 for objects in a vehicle scene (corresponding to the image 502). In the illustrated example, the perception system 402 includes an image feature extractor 504, a task-based feature enhancement stage 506, and a detection stage 510 (having at least one task-based feature enhancement stage 508). However, it will be understood that the perception system 402 may include fewer or more components. In some cases, the perception system 402 may omit the task-based feature enhancement stage 506. For example, in some cases, the detection stage 510 may receive the output of the image feature extractor 504 as an input. In some such cases, as part of the detection stage 510, one or more task-based feature enhancement stages 508 may enhance one or more feature maps. In certain cases, the perception system 402 may omit the task-based feature enhancement stage 508.

[0094] The image 502 (also referred to herein as a collection of images 502, a stream of images, or an image stream) may include image data from a particular sensor in a sensor suite. The type of image may correspond to the image sensor used to generate the image 502. For example, the image 502 may be a camera image generated from one or more cameras (such as camera 202a, etc.), or may be a LiDAR image generated from one or more LiDAR sensors (such as LiDAR sensor 202b, etc.). Other image types may be used, such as a Radar image generated from one or more Radar sensors (e.g., generated from Radar sensor 202c).

[0095] In some cases, the collection of images may correspond to a stream of time-varying images from the same image sensor. Thus, the first image in the collection of images may be generated (or captured) by the image sensor at time t 0 and the second image in the collection of images may be generated (or captured) at time t 1 and so on. When the perception system 402 uses the image 502 to determine object characteristics and / or generate the bounding box 512 and navigate the vehicle, it will be understood that the perception system 402 may process the image 502 in real time or near real time to generate the object characteristics and / or the bounding box 512.

[0096] In addition, since there can be multiple image sensors, each image sensor can generate its own set of images (or image stream). Thus, images from different image streams can be generated at approximately the same time. Therefore, images taken at the same time from different image streams can represent the vehicle scene at that time.

[0097] The image feature extractor 504 can be implemented using one or more neural networks or layers of a neural network to extract features from the image 502. In some cases, the image feature extractor 504 can be implemented using a backbone with a Feature Pyramid Network (FPN), Residual Network (Resnet), or Swin Transformer, CSWin Transformer, etc.

[0098] The image feature extractor 504 can use the image 502 to generate one or more feature maps. In some cases, the image feature extractor 504 generates at least one feature map for each image in the image 502. For example, if the image feature extractor 504 receives six consecutive images from an image stream, the image feature extractor 504 can generate six feature maps respectively.

[0099] The feature maps can have the same or different shapes as the images used to generate them, and / or can have the same or different shapes from each other. For example, if each image in the image 502 has a shape of [900, 1600, 3], the corresponding feature maps can have a shape of [45, 80, 256], although it will be understood that the feature maps can have different shapes, or even different shapes from each other.

[0100] Each of the generated feature maps can include an array of grid cells with a specific channel depth. The grid cells can include semantic data (or features) extracted from the pixels in the image (or images) from which the feature map is generated. The features of the grid cells can be organized as a vector or some other tensor shape. For example, the features (or semantic data) of the grid cells can indicate the shape, light, texture, reflectivity, edges, object classification, location, etc. of something detected by the image feature extractor 504.

[0101] In some cases, the image feature extractor 504 may generate multiple feature maps for each image. For example, the image feature extractor 504 may include an FPN that generates multiple feature maps from a specific image. In some cases, some of the feature maps (e.g., among the multiple feature maps generated from a specific image) may be generated from each other. For example, the first feature map may be downsampled (or convolved) to generate the second feature map, and the second feature map may be downsampled (or convolved) to generate the third feature map, and so on. In some such cases, the second feature map may have a smaller height and width than the first feature map, and the third feature map may have a smaller height and width than the second feature map, and so on.

[0102] It will be understood that the description herein regarding the processing of one feature map generated from an image may also be performed on some or all of the feature maps generated from the image. For example, if the image feature extractor 504 outputs five feature maps generated from a specific image (e.g., corresponding to different feature levels), then as described herein, the task-based feature enhancement stage 506 and / or the detection stage 510 may use the corresponding feature maps generated from a previous image to enhance some or all of these five feature maps. Thus, it will be understood that the perception system 402 may perform multiple enhancement iterations on multiple feature maps generated from one image 502.

[0103] Task-based feature enhancement phase

[0104] The task-based feature enhancement stage 506 may enhance the feature maps. In some cases, the task-based feature enhancement stage 506 may use earlier feature maps to enhance later feature maps. For example, the task-based feature enhancement stage 506 may use one or more earlier-time feature maps to enhance a later-time feature map. The earlier feature maps may include the feature maps corresponding to the image immediately preceding the image used to generate the later feature map (e.g., the image generated / acquired just before the image corresponding to the feature map to be enhanced). For example, the task-based feature enhancement stage 506 may use the feature map corresponding to the image at time t 0 to enhance the feature map corresponding to the image generated at time t 1 . Similarly, the task-based feature enhancement stage 506 may use the feature maps corresponding to the images at times t 0 , t 1 and t 2 to enhance the feature map corresponding to the image generated at time t 3 and so on. However, it will be understood that any combination of earlier feature maps may be used to enhance the later feature map. For example, the task-based feature enhancement stage 506 may use the feature maps corresponding to the images at times t 0 and / or t 2the feature map corresponding to the image generated at to enhance the one at time t 4 and other places where the feature map corresponding to the generated image is located.

[0105] In addition, the task-based feature enhancement stage 506 can use earlier feature maps to enhance multiple later feature maps. For example, the task-based feature enhancement stage 506 can use the feature map corresponding to the image generated at time t 0 to enhance the feature map corresponding to the image generated at time t 1 , t 2 and / or t 3 and other places where the feature map corresponding to the generated image is located.

[0106] In some cases, the task-based feature enhancement stage 506 can weight the semantic data of the earlier feature map and use the weighted semantic data to modify the semantic data of the later feature map. In certain cases, the task-based feature enhancement stage 506 can weight the semantic data of the earlier feature map according to the time difference between the image corresponding to the earlier feature map and the image corresponding to the later feature map. For example, compared with the semantic data from a feature map with a smaller time distance from the feature map to be enhanced for enhancing the semantic data of a specific feature map, the task-based feature enhancement stage 506 can assign a smaller weight to the semantic data from a feature map with a larger time distance. As another non-limiting example, if the task-based feature enhancement stage 506 uses (corresponding to the images generated at time t 0 , t 1 and t 2 ) feature_map0, feature_map1, and feature_map2 to enhance feature_map3, then the task-based feature enhancement stage 506 can weight the semantic data of feature_map0, feature_map1, and feature_map2 in such a way that the semantic data of feature_map2 has a greater impact on the value of the enhanced feature_map3 compared to feature_map0 or feature_map1. Similarly, the task-based feature enhancement stage 506 can weight feature_map1 in such a way that the semantic data of feature_map1 has a greater impact on the value of the enhanced feature_map3 compared to feature_map0. However, it will be understood that the semantic data of the feature maps can be weighted in various ways (including without weighting relative to each other), etc.

[0107] The task-based feature enhancement stage 506 can enhance a feature map in various ways using one or more earlier feature maps. In some cases, the task-based feature enhancement stage 506 can enhance a later feature map by concatenating semantic data from at least one earlier feature map to the corresponding semantic data of the later feature map and / or performing a cross-attending operation on the semantic data of the later feature map and the semantic data of the earlier feature map.

[0108] In some cases, the task-based feature enhancement stage 506 can identify one or more grid cells in an earlier feature map corresponding to grid cells in a later feature map, and use the semantic data of the (one or more) identified grid cells in the earlier feature map to enhance or modify the semantic data of the corresponding grid cells in the later feature map.

[0109] In some cases, the task-based feature enhancement stage 506 can map grid cells from an earlier feature map to a later feature map based on the positions of the grid cells in the earlier feature map. In some cases, grid cells at the same position in different feature maps can be mapped to each other. For example, the task-based feature enhancement stage 506 can map the grid cell at position (5, 25) in an earlier feature map to (and use it to enhance) the grid cell at position (5, 25) in a later feature map.

[0110] In some cases, the task-based feature enhancement stage 506 can use positioning data associated with the vehicle 200 to identify grid cells in an earlier feature map to map to specific grid cells in a later feature map. For example, the task-based feature enhancement stage 506 can consider the time between an image corresponding to the earlier feature map and an image corresponding to the later feature map, the speed, heading, and location of the vehicle 200, etc., to identify the (one or more) grid cells in the earlier feature map corresponding to (or mapping to) the grid cells in the later feature map.

[0111] In some cases, the task-based feature enhancement stage 506 can use learnable shift values to identify grid cells in an earlier feature map corresponding to grid cells in a later feature map. For example, a neural network can be trained to determine cross-time shift values between pixels based on various parameters of the vehicle and / or images (such as vehicle speed, vehicle direction, time between images, etc.). The learned shift values can be applied to the earlier feature map to determine which grid cells in the earlier feature map correspond to (or should be used to enhance) the grid cells in the later feature map.

[0112] In some cases, such as when enhancing a later feature map by concatenating the semantic data of the later feature map with the semantic data of an earlier feature map, the task-based feature enhancement stage 506 can append the semantic data from the earlier feature map to some or all of the grid cells in the later feature map. For example, the task-based feature enhancement stage 506 can identify the grid cells in the earlier feature map that correspond to the grid cells in the later feature map, and append the semantic data from the identified grid cells in the earlier feature map to the semantic data of the grid cells in the later feature map.

[0113] In some cases, the task-based feature enhancement stage 506 can associate the semantic data of one or more grid cells in the later feature map with one or more grid cells in the earlier feature map. As part of making the grid cells from different feature maps associated, the task-based feature enhancement stage 506 can use one or more linear layers to identify one or more corresponding grid cells in the earlier feature map. For example, the task-based feature enhancement stage 506 can multiply the tensor [1,N] corresponding to the grid cell by the learnable linear layer matrix [N,2] to determine the positions of one or more grid cells in the earlier feature map that correspond to the grid cell in the later feature map.

[0114] The task-based feature enhancement stage 506 can use the features of one or more grid cells (also referred to herein as the mapped grid cells) from the earlier feature map to modify some or all of the features of the grid cells (also referred to herein as the target grid cells) in the later feature map. In some cases, this can include assigning weights to specific features of the target grid cells in the later feature map and to the corresponding features of one or more mapped grid cells in the earlier feature map, and using the result (non-limiting example: sum of products) to modify the specific features of the target grid cells in the later feature map or to assign new values to the specific features of the target grid cells in the later feature map. In certain cases, the task-based feature enhancement stage 506 can use a learnable linear layer matrix to identify multiple grid cells in the earlier feature map and use the identified grid cells to modify the features of the grid cells in the later feature map. In some such cases, the task-based feature enhancement stage 506 can assign different weights to the features of different grid cells and use the weighted features to determine the corresponding features of the target grid cells.

[0115] As described herein, multiple earlier feature maps can be used to enhance a later feature map. In some such cases, the task-based feature enhancement stage 506 can map grid cells (mapped grid cells) from each of the earlier feature maps in the earlier feature maps to target grid cells of the later feature map, and use a combination of the mapped grid cells from different feature maps to enhance or modify the features of the target grid cells. In some such cases, the task-based feature enhancement stage 506 can apply weight values to different earlier feature maps (e.g., based on the time difference from the later feature map), and apply weight values to one or more grid cells of different earlier feature maps (e.g., based on a determined relationship between one or more mapped grid cells and the target grid cells of the later feature map). Additionally, if multiple feature maps are to be enhanced (e.g., because the image feature extractor 504 generates multiple feature maps from an image), the enhancement process can be performed for some or all of the multiple feature maps to be enhanced.

[0116] Detection phase

[0117] The detection stage 510 uses the outputs of the image feature extractor 504 and / or the task-based feature enhancement stage 506 to determine one or more characteristics of the objects in the image 502 and / or generate one or more bounding boxes 512 for the objects in the image 502. The characteristics of the objects can include, but are not limited to, object classification, object location, object depth, object centrality, object size, object rotation, object orientation, and / or object velocity. In some cases, the detection stage 510 can be implemented based on fully convolutional one-stage object detection, a non-limiting example of which is described in Tian et al.'s "FCOS: Fully Convolutional One-Stage Object Detection" (IEEE / CVF International Conference on Computer Vision (ICCV) 2019, October 27, 2019 to November 2, 2019, incorporated herein by reference), and the fully convolutional one-stage object detection can be modified to use the task-based feature enhancement stage 508 to enhance one or more feature maps. However, it will be understood that various detectors can be used as needed.

[0118] In some cases, the detection stage 510 can perform multiple convolutions on a feature map (e.g., output from the image feature extractor 504) or an enhanced feature map (e.g., output by the task-based feature enhancement stage 506) to determine different characteristics of an object in the image. In some cases, the detection stage 510 can use different convolutions to generate different streams. For example, a first set of convolutions can form a classification stream that generates one or more feature maps for classifying an object in the image, and a second set of convolutions can form a regression stream that generates one or more feature maps indicating specific characteristics of an object in the image (e.g., speed, location, depth, centrality, size, rotation, and / or orientation, etc.). Fewer or more streams can be used as needed.

[0119] In the illustrated example, the detection stage 510 includes one or more task-based feature enhancement stages 508. The task-based feature enhancement stage 508 can be similar to the task-based feature enhancement stage 506 in that they can use one or more earlier feature maps to enhance a later feature map. Additionally, the task-based feature enhancement stage 508 can enhance the feature map in a manner similar to the task-based feature enhancement stage 506. For example, the task-based feature enhancement stage 508 can identify one or more grid cells in an earlier feature map that correspond to a specific grid cell (also referred to herein as the target grid cell) in a later feature map and use semantic data associated with the one or more grid cells in the earlier feature map to modify or enhance the semantic data of the target grid cell in the later feature map.

[0120] However, it will be understood that the feature map enhanced by the task-based feature enhancement stage 508 can be different from the feature map enhanced by the task-based feature enhancement stage 506. For example, the feature map enhanced by the task-based feature enhancement stage 508 can have a different size from the feature map enhanced by the task-based feature enhancement stage 506.

[0121] As another non - limiting example, the feature map enhanced by the task - based feature enhancement stage 508 can correspond to the feature map generated from one or more convolutions in the detection stage 510. In some cases, the detection stage 510 can include the task - based feature enhancement stage 508 to enhance the feature map after each convolution, multiple convolutions, or a specific convolution in the detection stage 510. Thus, the detection stage 510 can perform one or more convolutions on the feature map output by the image feature extractor 504 or on the enhanced feature map (also referred to as the convolution output) output by the task - based feature enhancement stage 506, and then use the task - based feature enhancement stage 508 to enhance the convolution output (the feature map obtained from one or more convolutions). In some such cases, the task - based feature enhancement stage 508 can use the earlier feature map it previously enhanced to enhance the convolution output (the feature map obtained from one or more convolutions). As mentioned, the enhancement of the feature map by the task - based feature enhancement stage 508 within the detection stage 510 can occur after each convolution, after a specific convolution, or after multiple convolutions. Additionally, one or more task - based feature enhancement stages 508 can form part of different streams, such as part of a classification stream and / or a regression stream, etc.

[0122] As a non - limiting example, consider the scenario where the detection stage 510 enhances the feature map. Based on the image at time t 0 the perception system 402 can use the image feature extractor 504 to generate a first t 0 feature map and use the detection stage 510 to generate a second t 0 feature map (as a result of one or more convolutions of the first t 0 feature map). Later, based on the image at time t 1 the perception system 402 can use the image feature extractor 504 to generate a first t 1 feature map, use the detection stage 510 to generate a second t 1 feature map (as a result of one or more convolutions of the first t 1 feature map), and use the second t 0 feature map and the second t 1 feature map to generate an enhanced t 1 feature map.

[0123] As another non - limiting example, consider the scenario of enhancing the feature map output by the image feature extractor 504. Based on the image at time t 0 the perception system 402 can use the image feature extractor 504 to generate a first t 0 feature map and use the detection stage 510 (non - limiting example: for example, by performing one or more convolutions on the first t 0An enhanced version of the feature map is subjected to one or more convolutions to generate a second t 0 feature map) based at least in part on the first t 0 feature map. Later, based on the image at time t 1 the perception system 402 can use the image feature extractor 504 to generate the first t 1 feature map, use the first t 0 feature map and the first t 1 feature map to generate an enhanced t 1 feature map, and use the detection stage 510 to generate the second t 1 feature map (e.g., as a result of one or more convolutions of the enhanced t 1 feature map).

[0124] As another non-limiting example, consider the scenario of enhancing the feature map output by the image feature extractor 504 and enhancing the feature map in the detection stage 510 (in the same or different streams). Based on the image at time t 0 the perception system 402 can use the image feature extractor 504 to generate the first t 0 feature map and use the detection stage 510 to generate the second t 0 feature map and the third t 0 feature map (e.g., as a result of one or more convolutions). Later, based on the image at time t 1 the perception system 402 can use the image feature extractor 504 to generate the first t 1 feature map, use the first t 0 feature map and the first t 1 feature map to generate the first enhanced t 1 feature map, use the first enhanced t 1 feature map and the detection stage 510 to generate the second t 1 feature map (e.g., as a result of one or more convolutions of the enhanced t 1 feature map), use the second t 0 feature map and the second t 1 feature map to generate the second enhanced t 1 feature map, use the first enhanced t 1 feature map or the second enhanced t 1 feature map and the detection stage 510 to generate the third t 1 feature map (e.g., as a result of one or more convolutions of the first enhanced t 1 feature map or the second enhanced t 1 feature map, which convolution can be the same as or different from the one or more convolutions used to generate the second t 1 feature map), and use the third t0 Feature map and the third t 1 The feature map generates a third enhanced t 1 Feature map.

[0125] Fewer, more, or different components may be used as part of the perception system 402. For example, in some cases, the perception system 402 may omit the task-based feature enhancement stage 508. In some such cases, the detection stage 510 may determine the characteristics of an object without enhancing the feature map using the task-based feature enhancement stage 508. As another example, the perception system 402 may omit the task-based feature enhancement stage 506. In some such cases, the feature map output by the image feature extractor 504 may be communicated to the detection stage 510, and one or more than one task-based feature enhancement stage 508 may be used to enhance the feature map generated by the detection stage 510 to determine the characteristics of the object. In some cases, the perception system 402 or other components of the vehicle 200 may include a buffer or data storage to store earlier feature maps for enhancing later feature maps.

[0126] Data flow example

[0127] Figure 6A is a data flow diagram illustrating an example of the perception environment 600 in which the perception system 402 generates object characteristics and / or bounding boxes 512 from images 502 (identified as images 502a and 502b, respectively).

[0128] As described herein, the image 502 may correspond to images received from the same image sensor at different times. For example, the image 502a may correspond to an image captured by the image sensor at time t 1 and the image 502b may correspond to an image captured by the image sensor at time t 0 However, it will be understood that fewer or more images may be used. The images 502 may correspond to images taken immediately after each other in an image stream or to images with other images between them (e.g., sampled images). As described herein, the perception system 402 may repeatedly receive images and perform the functions described herein multiple times per second when a new image is received. Thus, it will be understood that the perception system 402 may operate in real time or near real time to determine object characteristics from the images 502 and generate the bounding boxes 512.

[0129] In the illustrated example, the image feature extractor 504 generates a feature map 602a from the image 502a and a feature map 602b from the image 502b. In the illustrated example, the image feature extractor 504 generates one feature map from each image in the images 502. However, it will be understood that the image feature extractor 504 may generate multiple feature maps 602 (e.g., feature maps at multiple levels) from each image in the images 502 and communicate the multiple feature maps 602 to the task-based feature enhancement stage 506.

[0130] As described herein, the image feature extractor 504 may generate the feature map 602b before the feature map 602a, and in some cases, may generate the feature map 602b before the image sensor captures the image 502a. In some such cases, the perception system 402 may store the earlier feature map 602b in a buffer or other data storage before using the earlier feature map 602b to enhance the feature map 602a.

[0131] Each feature map in the feature maps 602 may include an array of grid cells having a specific channel depth. The grid cells may include semantic data (or features) extracted from the pixels in the image(s) 502 from which the feature map 602 is generated. The features may be organized as vectors or some other tensor shape. For example, the features (or semantic data) of the grid cells may indicate the shape, light, texture, reflectivity, edges, object classification, location, etc. of something detected by the image feature extractor 504.

[0132] The task-based feature enhancement stage 506 may combine the feature maps 602 and / or use the earlier feature map 602b to enhance the later feature map 602a to form a first enhanced feature map 604. In some cases, the task-based feature enhancement stage 506 may enhance the feature map 602a by modifying the features of some or all of the grid cells in the feature map 602a using the features from the grid cells in the feature map 602b.

[0133] In some cases, the task-based feature enhancement stage 506 may enhance the grid cells of the feature map 602a by concatenating the features of the corresponding grid cells from the feature map 602b with the features of the grid cells of the feature map 602a. In some such cases, the features of the grid cells at a specific location in the feature map 602b may be concatenated with the features of the grid cells at the same location in the feature map 602a. In certain cases, the task-based feature enhancement stage 506 may use shift values to identify the corresponding grid cells between the feature maps 602 and use the identified corresponding grid cell(s) in the feature map 602b to enhance or modify the features of the target grid cells in the feature map 602a.

[0134] In some cases, the task-based feature enhancement stage 506 may include a cross-attention stage to identify (e.g., of the feature map 602b) the (one or more) grid cells corresponding to the target grid cell of the feature map 602a and use the features of the identified (one or more) grid cells to modify the features of the target grid cell. As described herein, in some cases, this may include: comparing the features of grid cells in different feature maps to determine the shift between the grid cells (e.g., how much an object has moved between the images corresponding to the feature map 602); using the determined shift to identify the grid cells of the feature map 602b that correspond to the target grid cell of the feature map 602a; and using the weighted features of the identified (one or more) grid cells from the feature map 602b to modify the features of the target grid cell from the feature map 602a. For example, the task-based feature enhancement stage 506 may identify the (one or more) grid cells in the feature map 602b that correspond to (or map to) a particular grid cell of the feature map 602a, weight the features, and / or use the (weighted) features of the identified (one or more) grid cells to modify or enhance the features of a particular grid cell of the feature map 602a. In some cases, the task-based feature enhancement stage 506 may use similar techniques to enhance some or all of the grid cells in the feature map 602a.

[0135] By using the grid cells of the feature map 602b to enhance the grid cells of the feature map 602a, the perception system 402 can provide more semantic data for the detection stage 510, through which the detection stage 510 can determine the characteristics (e.g., speed) of the objects in the image 502a and generate the bounding boxes 512. In some cases, the resulting enhanced feature map can result in a more accurate speed calculation for the objects, a more accurate bounding box 512, and / or a more accurate trajectory prediction for these objects.

[0136] The detection stage 510 receives the enhanced feature map 604 from the task-based feature enhancement stage 506 and uses the enhanced feature map 604 to determine object characteristics and / or generate (3D) bounding boxes 512 for some or all of the objects in the image 502a. In some cases, the detection stage 510 generates bounding boxes for certain types of objects (e.g., objects with an object classification of pedestrian, bicycle, vehicle, construction cone, etc.), but not for other types of objects.

[0137] In Figure 6AIn the illustrated example, the detection stage 510 includes a plurality of convolutional streams 601a, 601b, 601c (collectively or generically referred to as convolutional stream 601). Each convolutional stream 601 includes: a set of (unique or shared) convolutions 605 (identified as convolutions 605a - 605j respectively) for generating feature maps (identified as feature maps 606a - 606g respectively) or stream results 616 (identified as stream results 616a, 616b, or 616c respectively); and a task-based feature enhancement stage 508 (identified as task-based feature enhancement stages 508a, 508b, 508c respectively) for enhancing the feature maps 606 of the corresponding convolutional stream 601 using earlier feature maps 612 (identified as earlier feature maps 612a, 612b, 612c respectively) to generate enhanced feature maps 614 (identified as enhanced feature maps 614a, 614b, 614c respectively).

[0138] Each of the earlier feature maps in the earlier feature maps 612 may correspond to the image 502b. For example, each of the earlier feature maps in the earlier feature maps 612 may be generated using the image 502b. In some cases, the earlier feature maps 612 may be generated in a manner similar to the way the perception system 402 generates the feature maps (also referred to herein as target feature maps) to be enhanced by the earlier feature maps 612. For example, the earlier feature map 612a (for enhancing the feature map 606c) may be generated by convolving the feature maps generated from the image 502b (e.g., previously generated enhanced feature maps and / or feature map 602b) with convolutions 605a, 605b, and 605c (the same convolutions used to generate the feature map 606c). Similarly, the earlier feature map 612b (for enhancing the feature map 606f) may be generated by convolving the feature maps generated from the image 502b (e.g., previously generated enhanced feature maps and / or feature map 602b) with convolutions 605e, 605f, and 605g (the same convolutions used to generate the feature map 606f), and the earlier feature map 612c (for enhancing the feature map 606g) may be generated by convolving the feature maps generated from the image 502b (e.g., previously generated enhanced feature maps and / or feature map 602b) with convolutions 605e, 605f, and 605i (the same convolutions used to generate the feature map 606g). Thus, it will be understood that multiple earlier feature maps 612 may be used to enhance multiple later feature maps 606 at different stages within the perception system 402 and / or within the detection stage 510. Additionally, it will be understood that if the target feature map (e.g., the feature map 606c) includes multiple feature levels, the earlier feature map 612a may also include multiple feature levels to enhance the multiple feature levels of the target feature map 606c respectively.

[0139] The convolutions 605 can be different from each other. Accordingly, the content of the feature maps 606 and the flow results 616 can be different from each other. For example, even though the enhanced feature map 604 can be used as the input for both the convolutions 605a and 605e, considering the differences between the convolution 605a and the convolution 605e, the resulting feature maps 606a and 606d may be different respectively. Thus, it will be understood that the detection stage 510 can generate multiple feature maps 606 from the enhanced feature map 604.

[0140] In some cases, the various feature maps 606 can have the same (or different) size (e.g., height and width) and channel depth depending on the convolution 605. For example, the feature maps 606a - 606g can have the same or different heights, widths, and channel depths. Similarly, the flow results 616 can have the same or different sizes depending on the corresponding convolution and the expected result of the corresponding convolution stream 601. For example, the flow results 616 can have the same height and width as the feature maps 606, but can have a channel depth that is different from the feature maps 606 and different from each other. For example, the flow result 616a can have a channel depth corresponding to different classifications of the object, the flow result 616b can have a channel depth of "1" corresponding to the centrality of the object, and the flow result 616c can include multiple flow results (or feature maps) with different channel depths. For example, the flow result 616c can include results with channel depths of "3" (e.g., for the size of the object), "2" (e.g., 1 feature map each for the offset, orientation, and velocity of the object), and "1" (e.g., for the depth of the object). It will be understood that the flow results 616 can include fewer or more features as needed.

[0141] It will be understood that the detection stage 510 can include fewer or more convolution streams 601, convolutions, task - based feature enhancement stages 508, etc. For example, in some cases, each feature map output by the convolution can be enhanced by the task - based feature enhancement stage 508 using the corresponding earlier feature map 612. As another example, in certain cases, only one convolution stream 601 can include the task - based feature enhancement stage 508. As another example, the detection stage 510 can include different convolutions for each feature. For example, if the flow result 616c includes object offset, depth, size, rotation degree, orientation, and velocity, the detection stage 510 can include one or more than one convolution for each of the object offset, depth, size, rotation degree, orientation, and velocity.

[0142] In Figure 6AIn the illustrated example, the first convolutional stream 601a includes convolutions 605a - 605d and a task - based feature enhancement stage 508a. The enhanced feature map 604 is used as the input to convolution 605a, and the output of convolution 605a (feature map 606a), the output of convolution 605b (feature map 606b), and the output of convolution 605c (feature map 606c) are used as the inputs to convolution 605b, convolution 605c, and the task - based feature enhancement stage 508a, respectively. It will be understood that additional convolutions may be used to generate feature maps 606a, 606b, and / or 606c, or additional feature maps may be generated.

[0143] The task - based feature enhancement stage 508a uses an earlier feature map 612a to enhance the feature map 606c to generate an enhanced feature map 614a. As described herein, the task - based feature enhancement stage 508a can use the earlier feature map 612a in various ways (such as, similar to the way the task - based feature enhancement stage 506 enhances the feature map 602a, by concatenating the features of the earlier feature map 612a to the features of the feature map 606c and / or performing a cross - attention operation on the features of the feature map 606c and the features of the earlier feature map 612a, etc.) to enhance the feature map 606c.

[0144] In the illustrated example, the enhanced feature map 614a is used as the input to convolution 605d to generate a stream result 616a. In this example, the stream result 616a is the object classification of the object in the image 502a, which indicates the classification of the object and the probability of correct classification.

[0145] The second convolutional stream 601a includes convolutions 605e - 605h and a task - based feature enhancement stage 508b. The enhanced feature map 604 is used as the input to convolution 605e, and the output of convolution 605e (feature map 606d), the output of convolution 605f (feature map 606e), and the output of convolution 605g (feature map 606f) are used as the inputs to convolution 605f, convolution 605g, and the task - based feature enhancement stage 508b, respectively. It will be understood that additional convolutions may be used to generate feature maps 606d, 606e, and / or 606f, or additional feature maps may be generated.

[0146] The task-based feature enhancement stage 508b uses an earlier feature map 612b to enhance the feature map 606f to generate an enhanced feature map 614b. As described herein, the task-based feature enhancement stage 508b can use the earlier feature map 612b in various ways (such as, similar to the way the task-based feature enhancement stage 506 enhances the feature map 602a, by concatenating the features of the earlier feature map 612b to the features of the feature map 606f and / or performing cross-attention operations on the features of the feature map 606f and the features of the earlier feature map 612b, etc.) to enhance the feature map 606f.

[0147] In the illustrated example, the enhanced feature map 614b is used as the input to the convolution 605h to generate a flow result 616b. In this example, the flow result 616a is the object centrality of the object in the image 502a, which indicates the centrality of the object.

[0148] The third convolution stream 601c includes convolutions 605e, 605f, 605i, and 605j (the third convolution stream 601c shares the convolutions 605e and 605f with the second convolution stream 601b) and a task-based feature enhancement stage 508c. The enhanced feature map 604 is used as the input to the convolution 605e, and the output of the convolution 605e (feature map 606d), the output of the convolution 605f (feature map 606e), and the output of the convolution 605i (feature map 606g) are used as the inputs to the convolution 605f, the convolution 605i, and the task-based feature enhancement stage 508c, respectively. It will be understood that additional convolutions can be used to generate the feature maps 606d, 606e, and / or 606g, or additional feature maps can be generated.

[0149] The task-based feature enhancement stage 508c uses an earlier feature map 612c to enhance the feature map 606g to generate an enhanced feature map 614c. As described herein, the task-based feature enhancement stage 508c can use the earlier feature map 612c in various ways (such as, similar to the way the task-based feature enhancement stage 506 enhances the feature map 602a, by concatenating the features of the earlier feature map 612c to the features of the feature map 606g and / or performing cross-attention operations on the features of the feature map 606g and the features of the earlier feature map 612c, etc.) to enhance the feature map 606g.

[0150] In the illustrated example, the enhanced feature map 614c is used as the input to the convolution 605j to generate a flow result 616c. In this example, the flow result 616c includes the offset, depth, size, rotation degree, orientation, and speed of the object in the image 502a. Therefore, the convolution 605j can represent multiple convolutions (for example, at least one convolution for each of the offset, depth, size, rotation degree, orientation, and speed).

[0151] The flow results 616 from the detection stage 510 can be used to generate bounding boxes 512, determine the path of the vehicle 200, and / or control the vehicle 200. For example, the perception system 402 can generate the bounding boxes 512 and communicate the bounding boxes 512 and / or object characteristics to the planning system 404. The planning system 404 can use the bounding boxes 512 and / or object characteristics to generate a navigation plan or path for the vehicle 200 and communicate the navigation plan to the control system 408. The control system 408 can use the navigation plan to control the vehicle to follow the navigation plan or path.

[0152] Figure 6B is a data flow diagram illustrating another example of the perception environment 650, in which the perception system 402 generates object characteristics and / or bounding boxes 512 from the images 502a and 502b (not shown).

[0153] Figure 6B The illustrated example perception environment 650 is similar to Figure 6A the illustrated example perception environment 600 in some aspects and different in other aspects. For example, similar to the perception environment 600, the perception environment 650 includes an image feature extractor 504 for generating a feature map 602a using the image 502a and a detection stage 510. Additionally, the detection stage 510 includes a plurality of convolutions 605 to generate a feature map 606 and flow results 616.

[0154] However, the perception environment 650 differs from the perception environment 600 in that the perception environment 650 does not include a task-based feature enhancement stage 506 for enhancing the feature map 602a using the feature map 602b. Thus, in the perception environment 650, the feature map 602a is used as an input to the detection stage 510 (e.g., an input to the convolutions 605a and 605e), and the feature map 602a (rather than an enhanced feature map such as the enhanced feature map 604) is convolved.

[0155] Regarding the detection stage 510, in the perception environment 650, the task-based feature enhancement stage 508a, the task-based feature enhancement stage 508b, and the task-based feature enhancement stage 508c are omitted such that the earlier feature maps 612a, 612b, or 612c are not used to generate the enhanced feature maps 614a, 614b, and 614c, respectively. Thus, in Figure 6B the illustrated example, the convolution streams 601a, 601b, and 601c generate the flow results 616a, 616b, and 616c, respectively, without using earlier feature maps to enhance later feature maps.

[0156] In addition, the detection stage 510 of the perception environment 650 includes a fourth convolutional stream 601d. The fourth convolutional stream 601d includes convolutions 605e, 605f, 605k, and 605l (the fourth convolutional stream 601d shares convolutions 605e and 605f with the second convolutional stream 601b and the third convolutional stream 601c), and a task-based feature enhancement stage 508d. The feature map 602a is used as the input to the convolution 605e, and the output of the convolution 605e (feature map 606d), the output of the convolution 605f (feature map 606e), and the output of the convolution 605k (feature map 606h) are respectively used as the inputs to the convolution 605f, the convolution 605k, and the task-based feature enhancement stage 508d. It will be understood that additional convolutions may be used to generate the feature maps 606d, 606e, and / or 606h, or additional feature maps may be generated.

[0157] In the illustrated example, the convolution 605k may obtain a feature map 606h indicating the depth of an object in the image 502a. In some such cases, the feature map 606h may have a channel depth of "1".

[0158] The task-based feature enhancement stage 508d uses an earlier feature map 612d to enhance the feature map 606h to generate an enhanced feature map 614d. As described herein with reference to other earlier feature maps 612, the image 502b may be used to generate the earlier feature map 612d in a manner similar to the generation of the feature map 606h (e.g., by convolving a feature map based on the image 502b using the convolutions 605e, 605f, and 605k). Additionally, as described herein, the earlier feature map 612d may be generated before the feature map 606h, and / or the earlier feature map 612d may be generated before the perception system 402 receives the image 502a. In some such cases, the earlier feature map 612d may be stored by the perception system 402 in a buffer or data storage.

[0159] As described herein, the task-based feature enhancement stage 508d may use the earlier feature map 612d to enhance the feature map 606h in various ways (such as by concatenating the features of the earlier feature map 612d to the features of the feature map 606h and / or performing cross-attention operations on the features of the feature map 606h and the features of the earlier feature map 612d, etc.). In some cases, the task-based feature enhancement stage 508d may also align the objects from the feature maps to account for the speed of the vehicle 200. By considering the speed of the vehicle 200, the task-based feature enhancement stage 508d may enable the perception system 402 to more accurately determine the speed of the object in the image 502a based on the depth of the objects in the images 502a and 502b.

[0160] In the illustrated example, the enhanced feature map 614d is used as the input to the convolution 605l to generate the flow result 616d. In this example, the convolution 605l processes the enhanced feature map 614d (which indicates the depth of the object in the image 502a) such that it can generate the velocity of the object in the image 502a. In this way, the perception system 402 can use the depth determination of the object in multiple images to determine the velocity of the object (in a later image). Thus, the flow result 616d can include the velocity of the object in the image 502a based on the depth of the object in different images. Additionally, the channel depth of the flow result 616d can be different from the channel depth of the feature map 606h. For example, if the feature map 606h (for object depth) has a channel depth of "1", the flow result 616d can have a channel depth of "2" (for object velocity along the x-axis and y-axis).

[0161] Flow example

[0162] Figure 7 is a flow chart illustrating an example of routine 700 implemented by at least one processor to determine object characteristics of an object in an image. Provided for illustrative purposes only Figure 7 the illustrated flow chart. It will be understood that one or more of the steps of the illustrated routine can be removed Figure 7 one or more than one of the steps of the illustrated routine, or the order of the steps can be changed. Additionally, for illustrative clarity, one or more specific system components are described in the context of performing various operations during respective data flow phases in a data flow phase. However, other system arrangements and other distributions of processing steps across system components and / or autonomous vehicle computing 400 can be used.

[0163] At block 702, the perception system 402 receives a first image of the vehicle scene. As described herein, the first image can correspond to an image received from an image sensor or camera located on the vehicle at a particular time.

[0164] At block 704, the perception system 402 generates at least one first feature map based on the image. The (one or more) first feature maps can include an array of grid cells having a particular channel depth (e.g., 256, 512, etc.). As described herein, the grid cells can include features indicating the extracted characteristics of the first image, such as but not limited to color, texture, location, reflectivity, shape, edges, etc.

[0165] In some cases, the perception system 402 may use one or more convolutions to generate one or more first feature maps. In some cases, the perception system 402 may use an image feature extractor (such as but not limited to Resnet and / or Feature Pyramid Network (FPN), etc.) to generate one or more first feature maps, and / or may generate one or more first feature maps as part of a detection stage (such as but not limited to a fully convolutional one-stage object detector, etc.). For example, referring to Figure 6A and Figure 6B , one or more first feature maps may correspond to the feature maps generated by the image feature extractor 504 (e.g., feature map 602a) or the feature maps generated by the detection stage 510 (e.g., feature map 606c, feature map 606f, feature map 606g, and / or feature map 606h).

[0166] At block 706, the perception system 402 obtains one or more second feature maps corresponding to a second image. As described herein, one or more second feature maps may be one or more earlier feature maps stored in a buffer or data storage of the perception system 402, and the one or more second feature maps may be generated before the one or more first feature maps, and / or the one or more second feature maps may be based on an image received before the first image.

[0167] In some cases, one or more second feature maps may have been processed by the perception system 402 in a manner similar to one or more first feature maps. For example, if one or more first feature maps are generated by the image feature extractor 504 using a first image (e.g., feature map 602a), then one or more second feature maps may be generated by the image feature extractor 504 using a second image (e.g., feature map 602b). As another non-limiting example, if one or more first feature maps are the result of three convolutions of the feature maps generated from a first image (e.g., feature map 606c, feature map 606f, feature map 606g, or feature map 606h), then one or more second feature maps may be the result of three convolutions of the feature maps generated from a second image.

[0168] At block 708, the perception system 402 uses one or more second feature maps (or one or more later feature maps) to enhance one or more first feature maps (or one or more earlier feature maps). As described herein, the perception system 402 may use one or more second feature maps to enhance one or more first feature maps in various ways. In some cases, the perception system 402 enhances one or more first feature maps by mapping one or more grid cells of the second feature map (also referred to herein as one or more mapped grid cells) to at least one grid cell of one or more first feature maps (also referred to herein as target grid cells) and using the features of the one or more mapped grid cells of the second feature map to enhance the features of the target grid cells.

[0169] As described herein, grid cells of the second feature map may be mapped to similarly located grid cells of the first feature map (e.g., mapping grid cells at a particular location of the second feature map to grid cells at the same location on the first feature map), or one or more grid cells of the second feature map may be mapped to grid cells of the first feature map based on a shift value or other motion estimate. In some cases, the shift value or other motion estimate may take into account the absolute motion of the object, the absolute motion of the vehicle 200, and / or the relative motion of the object with respect to the vehicle 200.

[0170] The perception system 402 may use the features of the one or more mapped grid cells to enhance the target grid cells in various ways. In some cases, the perception system 402 may concatenate the features (or semantic data) of the one or more mapped grid cells to the features of the target grid cells.

[0171] In some cases, the perception system 402 may perform cross-attention operations on the features of (one or more) mapped grid cells and a target grid cell. As part of performing cross-attention operations on the features, the perception system 402 may determine (e.g., based on a probability relationship between the mapped grid cell and the target grid cell, where the probability relationship may be based on a comparison of the features of the respective mapped grid cell and the target grid cell) the weights between (one or more) mapped grid cells and the target grid cell, weight the features of (one or more) mapped grid cells based on the weights, and use the weighted features of (one or more) mapped grid cells to modify or enhance the features of the target grid cell. In some such cases, as part of the modification / enhancement process, the perception system 402 may also weight the features of the target grid cell and use the weighted features of the target grid cell. As a non-limiting example, the perception system 402 may weight the features f of the target grid cell and the mapped grid cell based on different weight values, and use the combination of the weighted features f of the target grid cell and the mapped grid cell to determine the new value of the features f of the target grid cell. In a similar manner, the perception system 402 may modify some or all of the features in f-f. 0 are weighted, and use the weighted features f of the target grid cell and the mapped grid cell 0 to determine the features f of the target grid cell 0 of the target grid cell. In a similar manner, the perception system 402 may modify the features f of the target grid cell 1 -f n of the target grid cell.

[0172] At block 710, the perception system 402 uses the enhanced feature map to determine object characteristics. As described herein, the perception system 402 may determine the classification, centrality, offset, depth, size, rotation, orientation, and / or velocity of an object based on the enhanced feature map. In some cases, to determine object characteristics, the perception system 402 performs one or more convolutions on the first enhanced feature map. For example, referring Figure 6A to, if the first enhanced feature map corresponds to the enhanced feature map 604, the perception system 402 may perform multiple convolutions on the enhanced feature map 604 to determine object characteristics. In some such cases, the perception system 402 may generate additional enhanced feature maps from one or more convolutions, or may not generate additional enhanced feature maps from one or more convolutions.

[0173] As another example and referring Figure 6A, if the first enhanced feature map corresponds to the enhanced feature map 614c, the perception system 402 may perform multiple convolutions on the enhanced feature map 614c to determine one object characteristic or multiple object characteristics. For example, convolutions may be performed in parallel to generate different object characteristics. In some cases, the perception system 402 (e.g., in parallel) performs different convolutions to determine the offset, depth, size, rotation, orientation, and / or speed of an object.

[0174] As another example and referring to Figure 6B , if the first enhanced feature map corresponds to the enhanced feature map 614d, the perception system 402 may perform one or more convolutions on the enhanced feature map 614d to determine one or more object characteristics. For example, a first one or more convolutions may be used to determine the depth of an object in two images, and a second one or more convolutions (e.g., based on the depth or depth difference of the object between the two images) may be used to determine the speed of the object. Thus, the perception system 402 may perform different convolutions serially to determine the depth and / or speed of an object, and may use one object characteristic (such as the object depth, etc.) (or a feature map indicating one object characteristic) to determine another object characteristic (such as the speed, etc.). In some cases, when determining the speed from a feature map indicating the depth of an object, the perception system 402 may consider the speed of the vehicle 200 to align the objects from the two images.

[0175] Routine 700 may include fewer, more, or different steps. In some cases, the perception system 402 may generate one or more bounding boxes for an object in an image based on the determined characteristics, and control the vehicle based on the one or more bounding boxes. In certain cases, the perception system 402 may generate one or more bounding boxes for an object in an image based on the determined characteristics, estimate a trajectory for the object in the bounding box based on the determined characteristics, determine a path through the vehicle scenario based on the estimated trajectory, and control the vehicle to follow the determined path, and so on. For example, the perception system 402 may determine a bounding box based on the determined characteristics and communicate the bounding box to the planning system 404. The planning system 404 may use the bounding box to determine a path for the vehicle through the vehicle scenario, and the control system 408 may control the vehicle based on the determined path.

[0176] As described herein, the blocks of routine 700 may be implemented by one or more components of the vehicle 200. For example, referring to Figure 6A and Figure 6B, the blocks 704 - 708 can be implemented using the image feature extractor 504 and / or the detection stage 510. For example, the blocks 704 - 708 can be implemented by the image feature extractor 504 and the task-based feature enhancement stage 506 to generate the enhanced feature map 604, or the blocks 704 - 708 can be implemented by the detection stage 510 to generate the enhanced feature map 614a, the enhanced feature map 614b, the enhanced feature map 614c, and / or the enhanced feature map 614d.

[0177] In some cases, some or all of the blocks in the routine 700 can be repeated (e.g., before determining the characteristics of the object at block 710). For example, if the blocks 702 - 708 correspond to generating the enhanced feature map (e.g., the enhanced feature map 604) by the task-based feature enhancement stage 506, the blocks 704 - 708 can be repeated one or more times to enhance the feature map generated by the detection stage 510. In some cases, it can be repeated after each convolution (or any subset of convolutions) in the detection stage 510 (e.g., after each of the convolutions 605a - 605c, 605e - 605g, 605i, and 605k in Figure 6A or Figure 6B ). In certain cases, the blocks 704 - 708 can be repeated once for a specific convolution stream of the detection stage 510, or can be repeated at least once for some or all of the convolution streams in the detection stage 510 (e.g., as exemplified by the task-based feature enhancement stage 508 in Figure 6A ).

[0178] As a non-limiting example and referring to Figure 6A , the perception system 402 can generate a third feature map (e.g., the feature map 606c, the feature map 606f, or the feature map 606g) based on the first enhanced feature map (e.g., the enhanced feature map 604), enhance the third feature map using a fourth feature map (e.g., using the earlier feature map 612a, the earlier feature map 612b, or the earlier feature map 612c respectively) to form a second enhanced feature map (e.g., form the enhanced feature map 614a, the enhanced feature map 614b, or the enhanced feature map 614c respectively), and determine the characteristics of the object based on the second enhanced feature map.

[0179] As described above, blocks 704-708 can be repeated multiple times, and blocks 704-708 can be used to determine different characteristics. For example, if the feature map 606c is the third feature map cited above, the earlier feature map 612 is the fourth feature map, and the enhanced feature map 614a is the second enhanced feature map, the perception system 402 can repeat blocks 704-708 to generate a fifth feature map (e.g., feature map 606f) based on the first enhanced feature map (e.g., enhanced feature map 604), enhance the fifth feature map with a sixth feature map (e.g., earlier feature map 612b) to form a third enhanced feature map (e.g., enhanced feature map 614b), and determine a second (at least one) characteristic of the object based on the third enhanced feature map. Additionally, as a non-limiting example, the perception system 402 can repeat blocks 704-708 to generate a seventh feature map (e.g., feature map 606g) based on the first enhanced feature map (e.g., enhanced feature map 604), enhance the seventh feature map with an eighth feature map (e.g., earlier feature map 612c) to form a fourth enhanced feature map (e.g., enhanced feature map 614c), and determine a third (at least one) characteristic of the object based on the fourth enhanced feature map. In some cases, multiple characteristics can be determined based on the enhanced feature map. For example, the perception system 402 can determine an offset, depth, size, rotation, orientation, and / or velocity of the object based on the enhanced feature map. It will be understood that the vehicle 200 can use any one or any combination of the characteristics determined about the object to generate a bounding box and / or trajectory of the object and navigate the vehicle 200 through the vehicle scenario.

[0180] As described herein, it will be understood that references to a particular feature map (e.g., first feature map, second feature map, third feature map, fourth feature map, etc.) can include references to multiple feature maps at different feature levels. In some such cases, the processing performed on the identified feature maps can occur on each of the feature maps at different feature levels.

[0181] Example

[0182] The various example embodiments of the present disclosure can be described by the following clauses:

[0183] Clause 1. A method, comprising: receiving a first image at a first time; generating a first feature map based on the first image; obtaining a second feature map corresponding to a second image received at a second time, where the second time is before the first time; enhancing the first feature map with the second feature map to form a first enhanced feature map; and determining a characteristic of an object in the first image based on the first enhanced feature map.

[0184] Clause 2. The method according to Clause 1, wherein generating the first feature map based on the first image includes using an image feature extractor including a feature pyramid network to generate the first feature map.

[0185] Clause 3. The method according to Clause 2, further comprising: receiving the second image at the second time; and using the image feature extractor including the feature pyramid network to generate the second feature map based on the second image.

[0186] Clause 4. The method according to any one of Clauses 1 to 3, further comprising: generating at least one bounding box for the object based on the determined characteristics; and causing a vehicle to be controlled based on the at least one bounding box.

[0187] Clause 5. The method according to any one of Clauses 1 to 4, wherein enhancing the first feature map using the second feature map includes concatenating the features of the second feature map with the corresponding features of the first feature map to form the first enhanced feature map.

[0188] Clause 6. The method according to any one of Clauses 1 to 4, wherein enhancing the first feature map and the second feature map includes: identifying a specific grid cell in the first feature map; identifying a set of grid cells in the second feature map associated with the specific grid cell based on a shift value; generating a weight value for each grid cell in the set of grid cells relative to the specific grid cell based on a comparison of the features of the specific grid cell with the features of each grid cell in the set of grid cells; weighting at least one feature of each grid cell in the set of grid cells based on the weight value to provide at least one weighted feature of each grid cell in the set of grid cells; and modifying at least one feature of the specific grid cell based on the at least one weighted feature of each grid cell in the set of grid cells.

[0189] Clause 7. The method according to any one of Clauses 1 to 4, wherein enhancing the first feature map and the second feature map includes: identifying a first grid cell in the first feature map; identifying a set of grid cells in the second feature map associated with the first grid cell based on a shift value; generating a set of weight values for the set of grid cells relative to the first grid cell based on a comparison of the features of the first grid cell with the features of each grid cell in the set of grid cells, wherein the set of weight values includes a second weight value for a second grid cell in the second feature map; weighting at least one feature of the second grid cell based on the weight value to provide at least one weighted feature of the second grid cell; and modifying at least one feature of the first grid cell based on the at least one weighted feature of the second grid cell.

[0190] Clause 8. The method according to any one of Clauses 1 to 7, wherein determining a characteristic of an object in the first image based on the first enhanced feature map includes: generating a third feature map based on the first enhanced feature map; enhancing the third feature map with a fourth feature map to form a second enhanced feature map, the fourth feature map being based on the second feature map; and determining the characteristic of the object in the first image based on the second enhanced feature map.

[0191] Clause 9. The method according to Clause 8, wherein enhancing the third feature map and the fourth feature map includes concatenating the features of the fourth feature map with the corresponding features of the third feature map to form the second enhanced feature map.

[0192] Clause 10. The method according to Clause 8, wherein enhancing the third feature map and the fourth feature map includes: identifying a third grid cell in the third feature map; identifying a second set of grid cells in the fourth feature map associated with the fourth grid cell based on a second shift value; generating a second set of weight values for the second set of grid cells relative to the third grid cell based on a comparison of the features of the third grid cell with the features of each grid cell in the second set of grid cells, wherein the second set of weight values includes a fourth weight value for a fourth grid cell in the fourth feature map; weighting at least one feature of the fourth grid cell based on the weight value to provide at least one weighted feature of the fourth grid cell; and modifying at least one feature of the third grid cell based on the at least one weighted feature of the fourth grid cell.

[0193] Clause 11. The method according to any one of Clauses 8 to 10, wherein the characteristic of the object in the first image includes the depth of the object.

[0194] Clause 12. The method according to any one of Clauses 8 to 10, wherein the characteristics of the object in the first image include the classification of the object.

[0195] Clause 13. The method according to any one of Clauses 8 to 10, wherein the characteristics of the object in the first image include at least one of the centrality, offset, size, rotation degree, direction, and speed of the object.

[0196] Clause 14. The method according to any one of Clauses 8 to 13, wherein the characteristics of the object are the first characteristics, and the method further includes: generating a fifth feature map based on the first enhanced feature map; enhancing the fifth feature map with a sixth feature map to form a third enhanced feature map, the sixth feature map being based on the second feature map; and determining second characteristics of the object in the first image based on the third enhanced feature map.

[0197] Clause 15. The method according to Clause 14, further includes: generating a seventh feature map based on the first enhanced feature map; enhancing the seventh feature map with an eighth feature map to form a fourth enhanced feature map, the eighth feature map being based on the second feature map; and determining third characteristics of the object in the first image based on the fourth enhanced feature map.

[0198] Clause 16. The method according to Clause 15, further includes: generating at least one bounding box for the object based on the first characteristics, the second characteristics, and the third characteristics; and causing a vehicle to be controlled based on the at least one bounding box.

[0199] Clause 17. A system, including: a data storage unit that stores computer-executable instructions; and a processor configured to: receive a first image at a first time; generate a first feature map based on the first image; obtain a second feature map corresponding to a second image received at a second time, wherein the second time is before the first time; enhance the first feature map and the second feature map to form a first enhanced feature map; and determine characteristics of an object in the first image based on the first enhanced feature map.

[0200] Clause 18. The system according to Clause 17, wherein, in order to determine the characteristics of the object in the first image based on the first enhanced feature map, the processor is configured to: generate a third feature map based on the first enhanced feature map; enhance the third feature map with a fourth feature map to form a second enhanced feature map, the fourth feature map being based on the second feature map; and determine the characteristics of the object in the first image based on the second enhanced feature map.

[0201] Clause 19. A non-transitory computer-readable medium comprising computer-executable instructions that, when executed by a computing system, cause the computing system to: receive a first image at a first time; generate a first feature map based on the first image; obtain a second feature map corresponding to a second image received at a second time, where the second time is before the first time; enhance the first feature map and the second feature map to form a first enhanced feature map; and determine a characteristic of an object in the first image based on the first enhanced feature map.

[0202] Clause 20. The non-transitory computer-readable medium according to Clause 19, wherein, to determine a characteristic of an object in the first image based on the first enhanced feature map, the execution of the computer-executable instructions further causes the computing system to: generate a third feature map based on the first enhanced feature map; enhance the third feature map with a fourth feature map to form a second enhanced feature map, the fourth feature map being based on the second feature map; and determine the characteristic of the object in the first image based on the second enhanced feature map.

[0203] Additional example

[0204] All of the methods and tasks described herein can be performed and fully automated by a computer system. The computer system can in some cases include multiple different computers or computing devices (e.g., physical servers, workstations, storage arrays, cloud computing resources, etc.) that communicate and interoperate via a network to perform the described functions. Each such computing device typically includes a processor (or multiple processors) that executes program instructions or modules stored in an execution memory or other non-transitory computer-readable storage medium or device (e.g., solid-state storage device, disk drive, etc.). The various functions disclosed herein can be embodied in such program instructions or can be implemented in dedicated circuitry (e.g., ASIC or FPGA) of the computer system. In cases where the computer system includes multiple computing devices, these devices can but need not be located in the same location. The results of the disclosed methods and tasks can be persistently stored by transforming a physical storage device such as a solid-state memory chip or disk into a different state. In some embodiments, the computer system can be a cloud-based computing system whose processing resources are shared by multiple different business entities or other users.

[0205] The processes described herein or illustrated in the accompanying drawings of the present disclosure may be initiated in response to an event (such as on a pre-determined or dynamically determined schedule, on demand when initiated by a user or system administrator, or in response to some other event). When such a process is initiated, a set of executable program instructions stored on one or more non-transitory computer-readable media (e.g., hard disk drive, flash memory, removable media, etc.) may be loaded into the memory (e.g., RAM) of a server or other computing device. These executable instructions may then be executed by a hardware-based computer processor of the computing device. In some embodiments, such a process or a portion thereof may be implemented serially or in parallel on multiple computing devices and / or multiple processors.

[0206] According to this embodiment, certain actions, events, or functions of any process or algorithm described herein may occur in a different sequence, may be added, combined, or entirely omitted (e.g., not all of the described operations or events are necessary for the practice of the algorithm). Additionally, in certain embodiments, operations or events may occur simultaneously, rather than sequentially, for example, via multi-threading, interrupt handling, or multiple processors or processor cores, or on other parallel architectures.

[0207] The various illustrative logical blocks, modules, routines, and algorithm steps described in connection with the embodiments disclosed herein can be implemented as electronic hardware (e.g., ASIC or FPGA devices), computer software running on computer hardware, or a combination of both. Additionally, the various illustrative logical blocks and modules described in connection with the embodiments disclosed herein can be implemented or performed by: machines, such as a processor device, a digital signal processor (“DSP”), an application specific integrated circuit (“ASIC”), a field programmable gate array (“FPGA”), or other programmable logic device; discrete gate or transistor logic; discrete hardware components; or any combination thereof designed to perform the functions described herein. The processor device can be a microprocessor, but in an alternative, the processor device can be a controller, a microcontroller, or a state machine, or a combination thereof, etc. The processor device can include circuitry configured to process computer-executable instructions. In another embodiment, the processor device includes an FPGA or other programmable device that performs logical operations without processing computer-executable instructions. The processor device can also be implemented as a combination of computing devices (e.g., a combination of a DSP and a microprocessor), multiple microprocessors, one or more than one microprocessor in conjunction with a DSP core, or any other such configuration. Although described herein primarily in terms of digital technology, the processor device can also include primarily analog components. For example, part or all of the rendering techniques described herein can be implemented in analog circuitry or in hybrid analog and digital circuitry. The computing environment can include any type of computer system, including but not limited to a microprocessor-based computer system, a mainframe computer, a digital signal processor, a portable computing device, a device controller, or a computing engine within an appliance (to name just a few examples).

[0208] The elements of a method, process, routine, or algorithm described in connection with the embodiments disclosed herein can be embodied directly in hardware, in a software module executed by a processor device, or in a combination of both. The software module can reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, a hard disk, a removable disk, a CD-ROM, or any other form of non-transitory computer-readable storage medium. The exemplary storage medium can be coupled to the processor device such that the processor device can read information from, and write information to, the storage medium. In an alternative, the storage medium can be integral to the processor device. The processor device and the storage medium can reside in an ASIC. The ASIC can reside in a user terminal. In an alternative, the processor device and the storage medium can reside in the user terminal as discrete components.

[0209] In the foregoing description, aspects and embodiments of the present disclosure have been described with reference to numerous specific details, which may vary according to implementation. Accordingly, the specification and drawings are to be regarded as illustrative rather than restrictive in a limiting sense. The sole and exclusive indication of the scope of the invention, and what the applicant desires to be the scope of the invention, is the literal and equivalent scope of the claims that are granted from this application in the specific form of the issued claims, including any subsequent amendments. Any definition of terms expressly set forth herein for inclusion in such claims shall govern the meaning of such terms as used in the claims. Additionally, when the term "further comprises" is used in the foregoing specification or the appended claims, the text following such phrase may be additional steps or entities, or sub-steps / sub-entities of the previously described steps or entities.

Claims

1. A method, comprising: receiving a first image at a first time; generating a first feature map based on the first image; obtaining a second feature map corresponding to a second image received at a second time, wherein the second time is before the first time; using the second feature map to enhance the first feature map to form a first enhanced feature map; and determining characteristics of an object in the first image based on the first enhanced feature map.

2. The method according to claim 1, wherein generating a first feature map based on the first image comprises: using an image feature extractor including a feature pyramid network to generate the first feature map.

3. The method according to claim 2, further comprising: receiving the second image at the second time; and using the image feature extractor including the feature pyramid network to generate the second feature map based on the second image.

4. The method according to any one of claims 1 to 3, further comprising: generating at least one bounding box for the object based on the determined characteristics; and causing a vehicle to be controlled based on the at least one bounding box.

5. The method according to any one of claims 1 to 4, wherein using the second feature map to enhance the first feature map comprises: concatenating features of the second feature map with corresponding features of the first feature map to form the first enhanced feature map.

6. The method according to any one of claims 1 to 4, wherein enhancing the first feature map and the second feature map comprises: identifying a specific grid cell in the first feature map; identifying a set of grid cells in the second feature map associated with the specific grid cell based on a shift value; generating a weight value for each grid cell in the set of grid cells relative to the specific grid cell based on a comparison of the features of the specific grid cell with the features of each grid cell in the set of grid cells; weighting at least one feature of each grid cell in the set of grid cells based on the weight value to provide at least one weighted feature of each grid cell in the set of grid cells; and modifying at least one feature of the specific grid cell based on the at least one weighted feature of each grid cell in the set of grid cells.

7. The method according to any one of claims 1 to 4, wherein enhancing the first feature map and the second feature map comprises: identifying a first grid cell in the first feature map; identifying a set of grid cells in the second feature map associated with the first grid cell based on a shift value; generating a set of weight values for the set of grid cells relative to the first grid cell based on a comparison of the features of the first grid cell with the features of each grid cell in the set of grid cells, wherein the set of weight values includes a second weight value for a second grid cell in the second feature map; Weight at least one feature of the second grid cell based on the weight value to provide at least one weighted feature of the second grid cell; and Modify at least one feature of the first grid cell based on the at least one weighted feature of the second grid cell.

8. The method according to any one of claims 1 to 7, wherein, Determining characteristics of an object in the first image based on the first enhanced feature map includes: Generating a third feature map based on the first enhanced feature map; Enhancing the third feature map with a fourth feature map to form a second enhanced feature map, the fourth feature map being based on the second feature map; and Determining characteristics of the object in the first image based on the second enhanced feature map.

9. The method according to claim 8, wherein, Enhancing the third feature map and the fourth feature map includes: concatenating features of the fourth feature map with corresponding features of the third feature map to form the second enhanced feature map.

10. The method according to claim 8, wherein, Enhancing the third feature map and the fourth feature map includes: Identifying a third grid cell in the third feature map; Identifying a set of second grid cells in the fourth feature map associated with the fourth grid cell based on a second shift value; Generating a set of second weight values for the set of second grid cells relative to the third grid cell based on a comparison of the features of the third grid cell with the features of each grid cell in the set of second grid cells, wherein the set of second weight values includes a fourth weight value for a fourth grid cell in the fourth feature map; Weight at least one feature of the fourth grid cell based on the weight value to provide at least one weighted feature of the fourth grid cell; and Modify at least one feature of the third grid cell based on the at least one weighted feature of the fourth grid cell.

11. The method according to any one of claims 8 to 10, wherein, The characteristics of the object in the first image include the depth of the object.

12. The method according to any one of claims 8 to 10, wherein, The characteristics of the object in the first image include the classification of the object.

13. The method according to any one of claims 8 to 10, wherein, The characteristics of the object in the first image include at least one of centrality, offset, size, rotation, orientation, and velocity of the object.

14. The method according to any one of claims 8 to 13, wherein, The characteristics of the object are first characteristics, and the method further includes: Generating a fifth feature map based on the first enhanced feature map; Enhancing the fifth feature map with a sixth feature map to form a third enhanced feature map, the sixth feature map being based on the second feature map; and Determining a second characteristic of the object in the first image based on the third enhanced feature map.

15. The method according to claim 14, further including: Generate a seventh feature map based on the first enhanced feature map; Enhance the seventh feature map using an eighth feature map to form a fourth enhanced feature map, where the eighth feature map is based on the second feature map; And Determine a third characteristic of the object in the first image based on the fourth enhanced feature map.

16. The method according to claim 15, further comprising: Generate at least one bounding box for the object based on the first characteristic, the second characteristic, and the third characteristic; And Cause a vehicle to be controlled based on the at least one bounding box.

17. A system, comprising: A data storage unit that stores computer-executable instructions; And A processor configured to: Receive a first image at a first time; Generate a first feature map based on the first image; Obtain a second feature map corresponding to a second image received at a second time, where the second time is before the first time; Enhance the first feature map and the second feature map to form a first enhanced feature map; And Determine a characteristic of an object in the first image based on the first enhanced feature map.

18. The system according to claim 17, wherein, To determine a characteristic of an object in the first image based on the first enhanced feature map, the processor is configured to: Generate a third feature map based on the first enhanced feature map; Enhance the third feature map using a fourth feature map to form a second enhanced feature map, where the fourth feature map is based on the second feature map; And Determine a characteristic of the object in the first image based on the second enhanced feature map.

19. A non-transitory computer-readable medium comprising computer-executable instructions that, when executed by a computing system, cause the computing system to: Receive a first image at a first time; Generate a first feature map based on the first image; Obtain a second feature map corresponding to a second image received at a second time, where the second time is before the first time; Enhance the first feature map and the second feature map to form a first enhanced feature map; And Determine a characteristic of an object in the first image based on the first enhanced feature map.

20. The non-transitory computer-readable medium according to claim 19, wherein, To determine a characteristic of an object in the first image based on the first enhanced feature map, the execution of the computer-executable instructions further causes the computing system to: Generate a third feature map based on the first enhanced feature map; Enhance the third feature map using a fourth feature map to form a second enhanced feature map, where the fourth feature map is based on the second feature map; And Determine a characteristic of the object in the first image based on the second enhanced feature map.