Enhancing feature maps to generate bounding boxes using multiple windows

By grouping and feature enhancing the grid cells of the feature map, the problem of bounding box recognition in autonomous driving vehicles is solved, and the accuracy and speed of identifying objects are improved.

CN120051810APending Publication Date: 2025-05-27MOTIONAL AD LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380073518.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-08-19
Filing Date
2023-08-17
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

In autonomous driving vehicles, it is difficult for neural networks to obtain sufficient semantic and local information, making it challenging to accurately draw 3D bounding boxes in real-time driving environments, and feature maps captured by different cameras lack semantic data, increasing the difficulty of identifying objects in vehicle scenarios.

Method used

By grouping grid cells of feature maps, the grid cell features in the group are combined or correlated, and windows that traverse multiple feature maps are used to enhance grid cell features, reduce processing needs, and improve the speed and efficiency of feature map processing.

Benefits of technology

It improves the accuracy and speed of the autonomous vehicle to identify objects in the image, enhances the ability to determine bounding boxes, and improves the accuracy of object query and bounding box generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120051810A_ABST
    Figure CN120051810A_ABST
Patent Text Reader

Abstract

A perception system may be used to generate bounding boxes for objects in a vehicle scene. A perceptual system may receive an image and a feature map corresponding to the received image. The perceptual system may generate a plurality of windows and use the plurality of windows to enhance semantic data of the feature map. The perceptual system may use the enhanced semantics to generate one or more bounding boxes for objects in the vehicle scene.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to U.S. patent application Ser. No. 17 / 821,154, filed on Aug. 19, 2022, entitled “ENRICHING FEATURE MAPS USING MULTIPLE PLURALITIES OF WINDOWS TO GENERATE BOUNDING BOXES” (Agent Docket No. MOTN.092A1), and U.S. patent application Ser. No. 17 / 821,152, filed on Aug. 19, 2022, entitled “ENRICHING OBJECTQUERIES USING A BIRD'S-EYE VIEW FEATURE MAP TO GENERATE BOUNDING BOXES” (Agent Docket No. MOTN.092A2), which are incorporated herein by reference in their entirety for all purposes. Background Art

[0003] An autonomous vehicle may use images obtained from one or more image sensors to generate bounding boxes for objects in the vehicle scene. BRIEF DESCRIPTION OF THE DRAWINGS

[0004] Figure 1 is an example environment in which a vehicle including one or more components of an autonomous system may be implemented.

[0005] Figure 2 is a diagram of one or more systems of a vehicle including an autonomous system.

[0006] Figure 3 yes Figure 1 and Figure 2 A diagram of one or more devices and / or components of one or more systems.

[0007] Figure 4A is a diagram of some components of an autonomous system.

[0008] Figure 4B It is a diagram of the implementation of a neural network.

[0009] Figure 4C and Figure 4D is a diagram illustrating an example operation of a CNN.

[0010] Figure 5A is a block diagram illustrating an example perception environment in which a perception system receives and processes images to provide one or more (3D) bounding boxes for objects in a vehicle scene.

[0011] Figure 5B is a diagram illustrating an example of rows of windows applied to feature maps corresponding to different images.

[0012] Figure 5C is a diagram illustrating an example of columns of windows applied to feature maps corresponding to different images.

[0013] Figure 5D is a diagram illustrating an example of a grid cell group of a feature map corresponding to different images.

[0014] Figure 6 is a dataflow diagram illustrating an example of a perception environment in which a perception system generates a bounding box from an image.

[0015] Figure 7 is a flow chart illustrating an example of a routine implemented by at least one processor to navigate a vehicle based on at least one bounding box generated from one or more images.

[0016] Figure 8 is a flow chart illustrating an example of a routine implemented by at least one processor to navigate a vehicle based on at least one bounding box generated from one or more images. DETAILED DESCRIPTION

[0017] In the following description, for the purpose of explanation, many specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it will be apparent that the embodiments described in the present disclosure can be implemented without these specific details. In some instances, well-known configurations and devices are illustrated in block diagram form to avoid unnecessarily obscuring aspects of the present disclosure.

[0018] In the accompanying drawings, for ease of description, the specific arrangement or order of schematic elements (such as those representing systems, devices, modules, instruction blocks and / or data elements, etc.) is illustrated. However, those skilled in the art will understand that, unless explicitly described, the specific order or arrangement of schematic elements in the accompanying drawings is not intended to mean that a specific processing order or sequence, or separation of processing is required. In addition, unless explicitly described, the inclusion of schematic elements in the accompanying drawings is not intended to mean that such elements are required in all embodiments, nor is it intended to mean that the features represented by such elements cannot be included in some embodiments or cannot be combined with other elements in some embodiments.

[0019] In addition, in the accompanying drawings, connecting elements (such as solid or dotted lines or arrows, etc.) are used to illustrate the connection, relationship or association between or among two or more other schematic elements, and the absence of any such connecting elements is not intended to mean that there can be no connection, relationship or association. In other words, some connections, relationships or associations between elements are not illustrated in the accompanying drawings so as not to obscure the present disclosure. In addition, for ease of illustration, a single connecting element can be used to represent multiple connections, relationships or associations between elements. For example, if the connecting element represents the communication of a signal, data or instruction (e.g., "software instruction"), it will be understood by those skilled in the art that such an element can represent one or more than one signal path (e.g., bus) that may be needed to affect the communication.

[0020] Although the terms "first", "second" and / or "third", etc. are used to describe various elements, these elements should not be limited by these terms. The terms "first", "second" and / or "third" are only used to distinguish one element from another. For example, a first contact may be referred to as a second contact, and similarly, a second contact may be referred to as a first contact without departing from the scope of the described embodiments. Both the first contact and the second contact are contacts, but they are not the same contacts.

[0021] The terms used in the description of the various embodiments described herein are included only for the purpose of describing a particular embodiment, and are not intended to be limiting. As used in the description of the various embodiments described and in the appended claims, the singular forms "a", "an" and "the" are also intended to include plural forms and can be used interchangeably with "one or more than one" or "at least one", unless the context clearly states otherwise. It will also be understood that the term "and / or" as used herein refers to and includes any and all possible combinations of one or more than one of the associated listed items. In addition, the term "or" is used in its inclusive sense (rather than in its exclusive sense) so that when, for example, a list of elements is used to connect, the term "or" means one, part or all of the elements in the list. Unless otherwise specifically stated, disjunctive language such as phrases "at least one of X, Y and Z" etc. are generally understood in the context of use to represent that an item, term, etc. can be X, Y or Z, or any combination thereof (e.g., X, Y or Z). Thus, such disjunctive language is generally not intended to and should not imply that certain embodiments require at least one of X, at least one of Y and at least one of Z to exist individually. It will also be understood that when the terms "includes," "comprising," "having," and / or "having" are used in this specification, they specifically state the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0022] As used herein, the terms "communication" and "communicating" refer to at least one of receiving, receiving, transmitting, transmitting and / or providing information (or information represented by, for example, data, signals, messages, instructions and / or commands, etc.). For a unit (e.g., a device, a system, a component of a device or system, and / or a combination thereof) to communicate with another unit, this means that the unit is able to directly or indirectly receive information from the other unit and / or send (e.g., transmit) information to the other unit. This can refer to a direct or indirect connection that is wired and / or wireless in nature. In addition, even if the transmitted information can be modified, processed, relayed and / or routed between the first unit and the second unit, the two units can communicate with each other. For example, even if the first unit passively receives information and does not actively transmit information to the second unit, the first unit can communicate with the second unit. As another example, if at least one intermediary unit (e.g., a third unit located between the first unit and the second unit) processes the information received from the first unit and transmits the processed information to the second unit, the first unit can communicate with the second unit. In some embodiments, a message may refer to a network packet (eg, a data packet, etc.) that includes data.

[0023] Conditional language used herein (such as "can", "may", "perhaps", "can", and "for example", etc.), unless otherwise specifically stated or otherwise understood within the context as used, is generally intended to convey that certain embodiments include certain features, elements, or steps, while other embodiments do not include certain features, elements, or steps. Therefore, such conditional language is generally not intended to imply that features, elements, or steps are necessary for one or more embodiments in any way, or that one or more embodiments must include logic for determining whether to include or perform these features, elements, or steps in any particular embodiment in the presence or absence of other inputs or prompts. As used herein, depending on the context, the term "if" is optionally interpreted to mean "when...", "at...", "in response to being determined to be" and / or "in response to being detected", etc. Similarly, depending on the context, the phrase "if it has been determined" or "if [the stated condition or event] is detected" is optionally interpreted to mean "when determining...", "in response to being determined to be" or "when [the stated condition or event] is detected" and / or "in response to being detected [the stated condition or event]", etc. Furthermore, as used herein, the terms "having," "having," or "having," etc. are intended to be open-ended terms. Furthermore, the phrase "based on" is intended to mean "based, at least in part, on" unless expressly stated otherwise.

[0024] Reference will now be made in detail to embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the various embodiments described. However, it will be apparent to one of ordinary skill in the art that the various embodiments described may be implemented without these specific details. In other cases, well-known methods, processes, components, circuits, and networks have not yet been described in detail in order not to unnecessarily obscure aspects of the embodiments.

[0025] Overview

[0026] To effectively navigate various scenes, autonomous vehicles use computer vision to identify objects in the scene and then navigate the scene based on the identified objects. As part of the navigation process, the autonomous vehicle can draw (3D) bounding boxes around objects in the image to understand the spatial relationship of the objects relative to the autonomous vehicle.

[0027] Since neural networks cannot obtain sufficient semantic and local information about objects in real-time driving environments, drawing accurate 3D bounding boxes on objects may be challenging in some cases. In addition, individual cameras on autonomous vehicles may not capture the entire object, which increases the difficulty of identifying objects in the vehicle scene. As such, individual feature maps corresponding to different cameras may not have sufficient semantic data to enable accurate bounding boxes to be drawn.

[0028] Additionally, inter-correlating or cross-correlating grid cells of a feature map may be computationally intensive, and correlating grid cells across an entire feature map or multiple feature maps may not be feasible in a real-time driving environment.

[0029] To address these issues, the autonomous vehicle can group the grid cells of the feature map and combine or correlate the features of the grid cells within the group. In some cases, the grid cells can be grouped based on the contours of the object. In some cases, one or more regions or windows that traverse multiple feature maps can be used to group the grid cells so that the features of the grid cells from different feature maps are enhanced or correlated with each other. Different stages can use different numbers, sizes, or shapes of windows, and multiple layers with shifted windows can be used to further enrich the grid cells.

[0030] By using groups of grid cells (e.g., windows or subsets of images / feature maps) for comparison and enhancement (e.g., for self-attention) instead of entire feature maps or sets of feature maps, autonomous vehicles can reduce processing requirements and increase the speed and efficiency of processing feature maps. These efficiencies can increase the rate at which autonomous vehicles can accurately identify objects in images. For example, in some cases, self-attention may yield 2 Thus, self-attention is performed on the N×M feature maps (which yields (N×M) 2 ) and performing self-attention on Y windows of feature maps (which yields Y*(N×M / Y) 2 can be significantly more computationally intensive (taking more time and processing power) than computations performed on

[0031] Additionally, the enhanced grid cells may improve the autonomous vehicle's ability to identify objects within the autonomous vehicle's scene. For example, the enhanced feature map may improve the autonomous vehicle's ability to determine bounding boxes, and / or the enhanced feature map may be used to enhance object queries and / or to generate a BEV feature map that improves the autonomous vehicle's ability to determine bounding boxes.

[0032] The autonomous vehicle may also generate object queries and enhance these object queries using: (enhanced) feature maps, features from other object queries (similar to how the grid cells of a feature map are enhanced), and / or features from a bird's eye view (BEV) feature map (which in turn may be generated from the original feature map generated by the image feature extractor or from the enhanced feature map). The enhanced object queries may improve the autonomous vehicle's ability to identify objects within the autonomous vehicle's scene.

[0033] General Overview

[0034] By implementing the systems, methods, and computer program products described herein, an autonomous vehicle may more accurately identify objects within an image, more accurately identify the location of an identified object within an image, more accurately predict the trajectory of an identified object within an image, determine additional features of the identified object, and infer additional information related to the scene of the image.

[0035] Reference now Figure 1, illustrates an example environment 100 in which vehicles including autonomous systems and vehicles not including autonomous systems operate. As illustrated, the environment 100 includes vehicles 102a-102n, objects 104a-104n, routes 106a-106n, areas 108, vehicle-to-infrastructure (V2I) devices 110, a network 112, a remote autonomous vehicle (AV) system 114, a fleet management system 116, and a V2I system 118. The vehicles 102a-102n, the vehicle-to-infrastructure (V2I) devices 110, the network 112, the autonomous vehicle (AV) system 114, the fleet management system 116, and the V2I system 118 are interconnected (e.g., establish connections for communication, etc.) via wired connections, wireless connections, or a combination of wired or wireless connections. In some embodiments, objects 104a-104n are interconnected with at least one of vehicles 102a-102n, vehicle-to-infrastructure (V2I) devices 110, networks 112, autonomous vehicle (AV) systems 114, fleet management systems 116, and V2I systems 118 via wired connections, wireless connections, or a combination of wired or wireless connections.

[0036] Vehicles 102a-102n (individually referred to as vehicles 102 and collectively referred to as vehicles 102) include at least one device configured to transport goods and / or people. In some embodiments, vehicles 102 are configured to communicate with V2I devices 110, remote AV systems 114, fleet management systems 116, and / or V2I systems 118 via network 112. In some embodiments, vehicles 102 include cars, buses, trucks, and / or trains, etc. In some embodiments, vehicles 102 are similar to vehicles 200 described herein (see Figure 2 ). In some embodiments, vehicles 200 in the set of vehicles 200 are associated with an autonomous queue manager. In some embodiments, vehicles 102 travel along respective routes 106a-106n (individually referred to as routes 106 and collectively referred to as routes 106) as described herein. In some embodiments, one or more vehicles 102 include an autonomous system (e.g., an autonomous system that is the same as or similar to autonomous system 202).

[0037] Objects 104a-104n (individually referred to as objects 104 and collectively referred to as objects 104) include, for example, at least one vehicle, at least one pedestrian, at least one cyclist, and / or at least one structure (e.g., a building, a sign, a fire hydrant, etc.), etc. Each object 104 is stationary (e.g., located at a fixed location and over a period of time) or moves (e.g., has a speed and is associated with at least one trajectory). In some embodiments, objects 104 are associated with corresponding locations in area 108.

[0038] Routes 106a-106n (individually referred to as routes 106 and collectively referred to as routes 106) are each associated with (e.g., specify a series of actions (also referred to as trajectories) connecting states along which the AV can navigate. Each route 106 begins at an initial state (e.g., a state corresponding to a first spatiotemporal location and / or speed, etc.) and ends at a final target state (e.g., a state corresponding to a second spatiotemporal location different from the first spatiotemporal location) or a target zone (e.g., a subspace of acceptable states (e.g., terminal states)). In some embodiments, the first state includes a location where one or more individuals will board the AV, and the second state or zone includes one or more locations where one or more individuals boarding the AV will disembark. In some embodiments, routes 106 include multiple acceptable state sequences (e.g., multiple spatiotemporal location sequences) that are associated with (e.g., define multiple trajectories). In an example, routes 106 include only high-level actions or imprecise state locations, such as a series of connecting roads indicating a change of direction at a roadway intersection, etc. Additionally or alternatively, the route 106 may include more precise actions or states, such as, for example, specific target lanes or precise locations within lane regions and target speeds at those locations, etc. In an example, the route 106 includes a plurality of precise state sequences along at least one high-level action with a limited look-ahead horizon to an intermediate target, wherein a combination of consecutive iterations of the limited horizon state sequences cumulatively correspond to a plurality of trajectories that collectively form a high-level route terminating at a final target state or region.

[0039] The area 108 includes a physical area (e.g., a geographic region) in which the vehicle 102 can navigate. In an example, the area 108 includes at least one state (e.g., a country, a province, a separate state of a plurality of states included in a country, etc.), at least a portion of a state, at least one city, at least a portion of a city, etc. In some embodiments, the area 108 includes at least one named thoroughfare (referred to herein as a "road"), such as a highway, an interstate highway, a parkway, a city street, etc. Additionally or alternatively, in some examples, the area 108 includes at least one unnamed road, such as a driveway, a section of a parking lot, a section of an open space and / or undeveloped area, a dirt road, etc. In some embodiments, the road includes at least one lane (e.g., a portion of the road that the vehicle 102 can traverse). In an example, the road includes at least one lane associated with (e.g., identified based on) at least one lane marking line.

[0040] The vehicle-to-infrastructure (V2I) device 110 (sometimes referred to as a vehicle-to-everything (V2X) device) includes at least one device configured to communicate with the vehicle 102 and / or the V2I system 118. In some embodiments, the V2I device 110 is configured to communicate with the vehicle 102, the remote AV system 114, the queue management system 116, and / or the V2I system 118 via the network 112. In some embodiments, the V2I device 110 includes a radio frequency identification (RFID) device, a sign, a camera (e.g., a two-dimensional (2D) and / or three-dimensional (3D) camera), lane markings, street lights, parking meters, etc. In some embodiments, the V2I device 110 is configured to communicate directly with the vehicle 102. Additionally or alternatively, in some embodiments, the V2I device 110 is configured to communicate with the vehicle 102, the remote AV system 114, and / or the queue management system 116 via the V2I system 118. In some embodiments, the V2I device 110 is configured to communicate with the V2I system 118 via the network 112 .

[0041] The network 112 includes one or more wired and / or wireless networks. In an example, the network 112 includes a cellular network (e.g., a long-term evolution (LTE) network, a third generation (3G) network, a fourth generation (4G) network, a fifth generation (5G) network, a code division multiple access (CDMA) network, etc.), a public land mobile network (PLMN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), a telephone network (e.g., a public switched telephone network (PSTN)), a private network, an ad hoc network, an intranet, the Internet, a fiber-based network, a cloud computing network, etc., and / or a combination of some or all of these networks, etc.

[0042] The remote AV system 114 includes at least one device configured to communicate with the vehicle 102, the V2I device 110, the network 112, the remote AV system 114, the fleet management system 116, and / or the V2I system 118 via the network 112. In an example, the remote AV system 114 includes a server, a server group, and / or other similar devices. In some embodiments, the remote AV system 114 is co-located with the fleet management system 116. In some embodiments, the remote AV system 114 participates in the installation of some or all of the components of the vehicle (including autonomous systems, autonomous vehicle computing, and / or software implemented by autonomous vehicle computing, etc.). In some embodiments, the remote AV system 114 maintains (e.g., updates and / or replaces) these components and / or software during the life of the vehicle.

[0043] The queue management system 116 includes at least one device configured to communicate with the vehicles 102, the V2I devices 110, the remote AV system 114, and / or the V2I system 118. In an example, the queue management system 116 includes a server, a server group, and / or other similar devices. In some embodiments, the queue management system 116 is associated with a ridesharing company (e.g., an organization for controlling the operation of multiple vehicles (e.g., vehicles including autonomous systems and / or vehicles not including autonomous systems)).

[0044] In some embodiments, the V2I system 118 includes at least one device configured to communicate with the vehicle 102, the V2I device 110, the remote AV system 114, and / or the fleet management system 116 via the network 112. In some examples, the V2I system 118 is configured to communicate with the V2I device 110 via a connection other than the network 112. In some embodiments, the V2I system 118 includes a server, a server group, and / or other similar devices. In some embodiments, the V2I system 118 is associated with a municipality or a private agency (e.g., a private agency for maintaining the V2I device 110, etc.).

[0045] supply Figure 1 The number and arrangement of elements illustrated are examples. Figure 1 There may be additional elements, fewer elements, different elements, and / or differently arranged elements than those illustrated. Additionally or alternatively, at least one element of environment 100 may be described as being Figure 1 Additionally or alternatively, at least one set of elements of environment 100 may perform one or more functions described as being performed by at least one different set of elements of environment 100.

[0046] Reference now Figure 2 , vehicle 200 includes autonomous system 202, powertrain control system 204, steering control system 206, and braking system 208. In some embodiments, vehicle 200 is compatible with vehicle 102 (see Figure 1 ) is the same or similar. In some embodiments, the vehicle 200 has autonomous capabilities (e.g., implementing at least one function, feature and / or device, etc., which enables the vehicle 200 to operate partially or completely without human intervention, including but not limited to fully autonomous vehicles (e.g., vehicles that abandon reliance on human intervention) and / or highly autonomous vehicles (e.g., vehicles that abandon reliance on human intervention in certain situations), etc.). For a detailed description of fully autonomous vehicles and highly autonomous vehicles, reference can be made to SAE International's standard J3016: Taxonomy and Definitions for Terms Related to On-Road Motor Vehicle AutomatedDriving Systems, the entire contents of which are incorporated by reference. In some embodiments, the vehicle 200 is associated with an autonomous queue manager and / or a ride-sharing company.

[0047] Autonomous system 202 includes a sensor suite that includes one or more devices such as camera 202a, LiDAR sensor 202b, Radar sensor 202c, and microphone 202d. In some embodiments, autonomous system 202 may include more or fewer devices and / or different devices (e.g., ultrasonic sensors, inertial sensors, GPS receivers (discussed below), and / or odometer sensors for generating data associated with an indication of the distance that vehicle 200 has traveled, etc.). In some embodiments, autonomous system 202 uses one or more devices included in autonomous system 202 to generate data associated with environment 100 described herein. Data generated by one or more devices of autonomous system 202 can be used by one or more systems described herein to observe the environment (e.g., environment 100) in which vehicle 200 is located. In some embodiments, autonomous system 202 includes communication devices 202e, autonomous vehicle computing 202f, and a drive-by-wire (DBW) system 202h.

[0048] The camera 202a includes a communication device 202e, an autonomous vehicle computer 202f, and / or a safety controller 202g configured to communicate with the communication device 202e via a bus (e.g., Figure 3 The camera 202a includes at least one device for communicating with the autonomous vehicle computing 202f (e.g., an image processing unit 202a, a bus ... Figure 1 In some embodiments, the autonomous vehicle computing 202f determines a depth to one or more objects in a field of view of at least two of the plurality of cameras based on image data from the at least two cameras. In some embodiments, the camera 202a is configured to capture images of objects within a distance relative to the camera 202a (e.g., up to 100 meters and / or up to 1 kilometer, etc.). Thus, the camera 202a includes features such as sensors and lenses that are optimized for sensing objects at one or more distances relative to the camera 202a.

[0049] In an embodiment, the camera 202a includes at least one camera configured to capture one or more images associated with one or more traffic lights, street signs, and / or other physical objects that provide visual navigation information. In some embodiments, the camera 202a generates traffic light data associated with the one or more images. In some examples, the camera 202a generates TLD data associated with one or more images including a format (e.g., RAW, JPEG, and / or PNG, etc.). In some embodiments, the camera 202a that generates TLD data differs from other systems incorporating cameras described herein in that the camera 202a may include one or more cameras with a wide field of view (e.g., a wide-angle lens, a fisheye lens, and / or a lens with a viewing angle of approximately 120 degrees or greater, etc.) to generate images related to as many physical objects as possible.

[0050] The light detection and ranging (LiDAR) sensor 202b includes a communication device 202e, an autonomous vehicle computing device 202f, and / or a safety controller 202g via a bus (e.g., Figure 3 The LiDAR sensor 202b includes at least one device that communicates with a bus (the same or similar bus as the bus 302 of the embodiment of the present invention). The LiDAR sensor 202b includes a system configured to emit light from a light emitter (e.g., a laser emitter). The light emitted by the LiDAR sensor 202b includes light outside the visible spectrum (e.g., infrared light, etc.). In some embodiments, during operation, the light emitted by the LiDAR sensor 202b encounters a physical object (e.g., a vehicle) and is reflected back to the LiDAR sensor 202b. In some embodiments, the light emitted by the LiDAR sensor 202b does not penetrate the physical object encountered by the light. The LiDAR sensor 202b also includes at least one light detector that detects the light emitted from the light emitter after encountering the physical object. In some embodiments, at least one data processing system associated with the LiDAR sensor 202b generates an image (e.g., a point cloud and / or a combined point cloud, etc.) representing objects included in the field of view of the LiDAR sensor 202b. In some examples, at least one data processing system associated with the LiDAR sensor 202b generates an image representing the boundaries of the physical object and / or the surface of the physical object (e.g., the topology of the surface), etc. In such examples, the image is used to determine the boundaries of the physical object in the field of view of the LiDAR sensor 202b.

[0051] The radio detection and ranging (Radar) sensor 202c includes a sensor configured to communicate with the communication device 202e, the autonomous vehicle computing 202f, and / or the safety controller 202g via a bus (e.g., Figure 3At least one device for communicating with a bus (same or similar bus as bus 302 of the embodiment of the present invention). Radar sensor 202c includes a system configured to transmit (pulsed or continuous) radio waves. The radio waves transmitted by Radar sensor 202c include radio waves within a predetermined spectrum. In some embodiments, during operation, the radio waves transmitted by Radar sensor 202c encounter physical objects and are reflected back to Radar sensor 202c. In some embodiments, the radio waves transmitted by Radar sensor 202c are not reflected by some objects. In some embodiments, at least one data processing system associated with Radar sensor 202c generates a signal representing an object included in the field of view of Radar sensor 202c. For example, at least one data processing system associated with Radar sensor 202c generates an image representing the boundary of a physical object and / or the surface of a physical object (e.g., the topology of the surface), etc. In some examples, the image is used to determine the boundary of a physical object in the field of view of Radar sensor 202c.

[0052] The microphone 202d includes a microphone configured to communicate with the communication device 202e, the autonomous vehicle computing device 202f, and / or the safety controller 202g via a bus (e.g., Figure 3 At least one device for communicating with the vehicle 200 (the same or similar bus as bus 302 of FIG. 1 ). Microphone 202d includes one or more microphones (e.g., an array microphone and / or an external microphone, etc.) that capture an audio signal and generate data associated with (e.g., representing) the audio signal. In some examples, microphone 202d includes a transducer device and / or the like. In some embodiments, one or more systems described herein can receive the data generated by microphone 202d and determine the position (e.g., distance, etc.) of an object relative to vehicle 200 based on an audio signal associated with the data.

[0053] The communication device 202e includes at least one device configured to communicate with the camera 202a, the LiDAR sensor 202b, the Radar sensor 202c, the microphone 202d, the autonomous vehicle computing 202f, the safety controller 202g, and / or the drive-by-wire (DBW) system 202h. For example, the communication device 202e may include at least one device configured to communicate with the camera 202a, the LiDAR sensor 202b, the Radar sensor 202c, the microphone 202d, the autonomous vehicle computing 202f, the safety controller 202g, and / or the drive-by-wire (DBW) system 202h. Figure 3 The communication device 202e may be a device that is the same as or similar to the communication interface 314 of the vehicle. In some embodiments, the communication device 202e includes a vehicle-to-vehicle (V2V) communication device (eg, a device for enabling wireless communication of data between vehicles).

[0054] Autonomous vehicle computing 202f includes at least one device configured to communicate with camera 202a, LiDAR sensor 202b, Radar sensor 202c, microphone 202d, communication device 202e, safety controller 202g, and / or DBW system 202h. In some examples, autonomous vehicle computing 202f includes devices such as client devices, mobile devices (e.g., cellular phones and / or tablet computers, etc.), and / or servers (e.g., computing devices including one or more central processing units and / or graphics processing units, etc.). In some embodiments, autonomous vehicle computing 202f is the same or similar to autonomous vehicle computing 400 described herein. Additionally or alternatively, in some embodiments, autonomous vehicle computing 202f is configured to communicate with an autonomous vehicle system (e.g., with Figure 1 remote AV system 114 of the same or similar autonomous vehicle system), a fleet management system (e.g., Figure 1 of the same or similar queue management system as the queue management system 116), V2I devices (e.g., Figure 1 V2I device 110 that is the same as or similar to V2I device 110) and / or V2I system (e.g., Figure 1 The V2I system 118 may communicate with the same or similar V2I system.

[0055] Safety controller 202g includes at least one device configured to communicate with camera 202a, LiDAR sensor 202b, Radar sensor 202c, microphone 202d, communication device 202e, autonomous vehicle computing 202f, and / or DBW system 202h. In some examples, safety controller 202g includes one or more controllers (electrical controllers and / or electromechanical controllers, etc.) configured to generate and / or transmit control signals to operate one or more devices of vehicle 200 (e.g., powertrain control system 204, steering control system 206, and / or braking system 208, etc.). In some embodiments, safety controller 202g is configured to generate control signals that take precedence over (e.g., override) control signals generated and / or transmitted by autonomous vehicle computing 202f.

[0056] The DBW system 202h includes at least one device configured to communicate with the communication device 202e and / or the autonomous vehicle computing 202f. In some examples, the DBW system 202h includes one or more controllers (e.g., electrical controllers and / or electromechanical controllers, etc.) configured to generate and / or transmit control signals to operate one or more devices of the vehicle 200 (e.g., powertrain control system 204, steering control system 206, and / or braking system 208, etc.). Additionally or alternatively, one or more controllers of the DBW system 202h are configured to generate and / or transmit control signals to operate at least one different device of the vehicle 200 (e.g., turn signal lights, headlights, door locks, and / or windshield wipers, etc.).

[0057] The powertrain control system 204 includes at least one device configured to communicate with the DBW system 202h. In some examples, the powertrain control system 204 includes at least one controller and / or actuator, etc. In some embodiments, the powertrain control system 204 receives a control signal from the DBW system 202h, and the powertrain control system 204 causes the vehicle 200 to start moving forward, stop moving forward, start moving backward, stop moving backward, accelerate in a certain direction, decelerate in a certain direction, turn left and / or turn right, etc. In an example, the powertrain control system 204 increases, maintains the same, or decreases the energy (e.g., fuel and / or electricity, etc.) provided to the motor of the vehicle, thereby rotating or not rotating at least one wheel of the vehicle 200.

[0058] The steering control system 206 includes at least one device configured to rotate one or more wheels of the vehicle 200. In some examples, the steering control system 206 includes at least one controller and / or actuator, etc. In some embodiments, the steering control system 206 rotates the two front wheels and / or the two rear wheels of the vehicle 200 to the left or right to turn the vehicle 200 left or right.

[0059] Braking system 208 includes at least one device configured to actuate one or more brakes to decelerate and / or hold vehicle 200 stationary. In some examples, braking system 208 includes at least one controller and / or actuator configured to cause one or more calipers associated with one or more wheels of vehicle 200 to close on corresponding rotors of vehicle 200. Additionally or alternatively, in some examples, braking system 208 includes an automatic emergency braking (AEB) system and / or a regenerative braking system, etc.

[0060] In some embodiments, vehicle 200 includes at least one platform sensor (not explicitly illustrated) for measuring or inferring a property of a state or condition of vehicle 200. In some examples, vehicle 200 includes platform sensors such as a global positioning system (GPS) receiver, an inertial measurement unit (IMU), wheel rate sensors, wheel brake pressure sensors, wheel torque sensors, engine torque sensors, and / or steering angle sensors.

[0061] Reference now Figure 3 , a schematic diagram of an exemplary device 300. As illustrated, device 300 includes a processor 304, a memory 306, a storage component 308, an input interface 310, an output interface 312, a communication interface 314, and a bus 302. In some embodiments, device 300 corresponds to: at least one device of vehicle 102 (e.g., at least one device of a system of vehicle 102); and / or one or more devices of network 112 (e.g., one or more devices of a system of network 112). In some embodiments, one or more devices of vehicle 102 (e.g., one or more devices of a system of vehicle 102), and / or one or more devices of network 112 (e.g., one or more devices of a system of network 112) include at least one device 300 and / or at least one component of device 300. As Figure 3 As shown, apparatus 300 includes a bus 302 , a processor 304 , a memory 306 , a storage component 308 , an input interface 310 , an output interface 312 , and a communication interface 314 .

[0062] The bus 302 includes components that permit communication between components of the device 300. In some cases, the processor 304 includes a processor (e.g., a central processing unit (CPU), a graphics processing unit (GPU), and / or an accelerated processing unit (APU), etc.), a microphone, a digital signal processor (DSP), and / or any processing component that can be programmed to perform at least one function (e.g., a field programmable gate array (FPGA) and / or an application specific integrated circuit (ASIC), etc.). The memory 306 includes a random access memory (RAM), a read-only memory (ROM), and / or another type of dynamic and / or static storage device (e.g., flash memory, magnetic memory, and / or optical memory, etc.) that stores data and / or instructions for use by the processor 304.

[0063] The storage component 308 stores data and / or software related to the operation and use of the device 300. In some examples, the storage component 308 includes a hard disk (e.g., a magnetic disk, an optical disk, a magneto-optical disk, and / or a solid-state disk, etc.), a compact disk (CD), a digital versatile disk (DVD), a floppy disk, a cassette, a tape, a CD-ROM, a RAM, a PROM, an EPROM, a FLASH-EPROM, an NV-RAM, and / or another type of computer-readable medium, and a corresponding drive.

[0064] The input interface 310 includes components that permit the device 300 to receive information, such as via user input (e.g., a touch screen display, a keyboard, a keypad, a mouse, a button, a switch, a microphone, and / or a camera, etc.). Additionally or alternatively, in some embodiments, the input interface 310 includes a sensor for sensing information (e.g., a global positioning system (GPS) receiver, an accelerometer, a gyroscope, and / or an actuator, etc.). The output interface 312 includes components for providing output information from the device 300 (e.g., a display, a speaker, and / or one or more light emitting diodes (LEDs), etc.).

[0065] In some embodiments, communication interface 314 includes a transceiver-like component (e.g., a transceiver and / or a separate receiver and transmitter, etc.) that permits device 300 to communicate with other devices via a wired connection, a wireless connection, or a combination of a wired connection and a wireless connection. In some examples, communication interface 314 permits device 300 to receive information from another device and / or provide information to another device. In some examples, communication interface 314 includes an Ethernet interface, an optical interface, a coaxial interface, an infrared interface, a radio frequency (RF) interface, a universal serial bus (USB) interface, interface and / or cellular network interface, etc.

[0066] In some embodiments, the device 300 performs one or more processes described herein. The device 300 performs these processes based on the processor 304 executing software instructions stored by a computer-readable medium such as a memory 306 and / or a storage component 308. Computer-readable media (e.g., non-transitory computer-readable media) are defined herein as non-transitory memory devices. Non-transitory memory devices include storage space located within a single physical storage device or storage space distributed across multiple physical storage devices.

[0067] In some embodiments, the software instructions are read into the memory 306 and / or storage component 308 from another computer-readable medium or from another device via the communication interface 314. The software instructions stored in the memory 306 and / or storage component 308, when executed, cause the processor 304 to perform one or more processes described herein. Additionally or alternatively, hardwired circuitry is used in place of or in combination with the software instructions to perform one or more processes described herein. Therefore, unless expressly stated otherwise, the embodiments described herein are not limited to any specific combination of hardware circuitry and software.

[0068] The memory 306 and / or the storage component 308 include a data storage unit or at least one data structure (e.g., a database, etc.). The device 300 can receive information from the data storage unit or at least one data structure in the memory 306 or the storage component 308, store information in the data storage unit or at least one data structure, communicate information to the data storage unit or at least one data structure, or search for information stored in the data storage unit or at least one data structure. In some examples, the information includes network data, input data, output data, or any combination thereof.

[0069] In some embodiments, the device 300 is configured to execute software instructions stored in the memory 306 and / or the memory of another device (e.g., another device that is the same as or similar to the device 300). As used herein, the term "module" refers to at least one instruction stored in the memory 306 and / or the memory of another device, which, when executed by the processor 304 and / or the processor of another device (e.g., another device that is the same as or similar to the device 300), causes the device 300 (e.g., at least one component of the device 300) to perform one or more processes described herein. In some embodiments, the module is implemented in software, firmware, and / or hardware, etc.

[0070] supply Figure 3 The number and arrangement of components illustrated are examples. Figure 3 The device 300 may include additional components, fewer components, different components, or differently arranged components than those illustrated. Additionally or alternatively, a set of components (e.g., one or more components) of the device 300 may perform one or more functions described as being performed by another component or set of components of the device 300.

[0071] Reference now Figure 4A, illustrates an example block diagram of an autonomous vehicle computing 400 (sometimes referred to as an "AV stack"). As illustrated, the autonomous vehicle computing 400 includes a perception system 402 (sometimes referred to as a perception module), a planning system 404 (sometimes referred to as a planning module), a positioning system 406 (sometimes referred to as a positioning module), a control system 408 (sometimes referred to as a control module), and a database 410. In some embodiments, the perception system 402, the planning system 404, the positioning system 406, the control system 408, and the database 410 are included in and / or implemented in an automatic navigation system of a vehicle (e.g., the autonomous vehicle computing 202f of the vehicle 200). Additionally or alternatively, in some embodiments, the perception system 402, the planning system 404, the positioning system 406, the control system 408, and the database 410 are included in one or more independent systems (e.g., one or more systems that are the same or similar to the autonomous vehicle computing 400, etc.). In some examples, the perception system 402, planning system 404, positioning system 406, control system 408, and database 410 are included in one or more independent systems located in the vehicle and / or at least one remote system as described herein. In some embodiments, any and / or all of the systems included in the autonomous vehicle computing 400 are implemented in software (e.g., software instructions stored in a memory), computer hardware (e.g., by a microprocessor, microcontroller, application specific integrated circuit (ASIC) and / or field programmable gate array (FPGA), etc.), or a combination of computer software and computer hardware. It will also be understood that in some embodiments, the autonomous vehicle computing 400 is configured to communicate with a remote system (e.g., an autonomous vehicle system that is the same or similar to the remote AV system 114, a fleet management system 116 that is the same or similar to the fleet management system 116, and / or a V2I system that is the same or similar to the V2I system 118, etc.).

[0072] In some embodiments, the perception system 402 receives data associated with at least one physical object in the environment (e.g., data used by the perception system 402 to detect at least one physical object) and classifies the at least one physical object. In some examples, the perception system 402 receives image data captured by at least one camera (e.g., camera 202a), the image being associated with (e.g., representing) one or more physical objects within the field of view of the at least one camera. In such examples, the perception system 402 classifies the at least one physical object based on one or more groups of physical objects (e.g., bicycles, vehicles, traffic signs, and / or pedestrians, etc.). In some embodiments, based on the classification of the physical object by the perception system 402, the perception system 402 transmits data associated with the classification of the physical object to the planning system 404.

[0073] In some embodiments, planning system 404 receives data associated with a destination and generates data associated with at least one route (e.g., route 106) along which a vehicle (e.g., vehicle 102) can travel toward the destination. In some embodiments, planning system 404 periodically or continuously receives data (e.g., data associated with the classification of physical objects described above) from perception system 402, and planning system 404 updates at least one trajectory or generates at least one different trajectory based on the data generated by perception system 402. In some embodiments, planning system 404 receives data associated with an updated position of a vehicle (e.g., vehicle 102) from positioning system 406, and planning system 404 updates at least one trajectory or generates at least one different trajectory based on the data generated by positioning system 406.

[0074] In some embodiments, the positioning system 406 receives data associated with (e.g., representing) a location of a vehicle (e.g., vehicle 102) in an area. In some examples, the positioning system 406 receives LiDAR data associated with at least one point cloud generated by at least one LiDAR sensor (e.g., LiDAR sensor 202b). In some examples, the positioning system 406 receives data associated with at least one point cloud from multiple LiDAR sensors, and the positioning system 406 generates a combined point cloud based on each point cloud. In these examples, the positioning system 406 compares the at least one point cloud or the combined point cloud with a two-dimensional (2D) and / or three-dimensional (3D) map of the area stored in the database 410. Then, based on the positioning system 406 comparing the at least one point cloud or the combined point cloud with the map, the positioning system 406 determines the position of the vehicle in the area. In some embodiments, the map includes a combined point cloud of the area generated before the navigation of the vehicle. In some embodiments, the map includes, but is not limited to, a high-precision map of roadway geometry, a map describing the connectivity of the road network, a map describing the physical properties of the roadway (such as traffic speed, traffic volume, the number of vehicle and bicycle traffic lanes, lane width, lane traffic direction, or the type and location of lane markings, or a combination thereof), and a map describing the spatial location of road features (such as crosswalks, traffic signs, or various types of other driving signals, etc.) In some embodiments, the map is generated in real time based on data received by the perception system.

[0075] In another example, positioning system 406 receives global navigation satellite system (GNSS) data generated by a global positioning system (GPS) receiver. In some examples, positioning system 406 receives GNSS data associated with a location of a vehicle in an area, and positioning system 406 determines the latitude and longitude of the vehicle in the area. In such an example, positioning system 406 determines the position of the vehicle in the area based on the latitude and longitude of the vehicle. In some embodiments, positioning system 406 generates data associated with the position of the vehicle. In some examples, based on the location of the vehicle determined by positioning system 406, positioning system 406 generates data associated with the position of the vehicle. In such an example, the data associated with the position of the vehicle include data associated with one or more semantic properties corresponding to the position of the vehicle.

[0076] In some embodiments, the control system 408 receives data associated with at least one trajectory from the planning system 404, and the control system 408 controls the operation of the vehicle. In some examples, the control system 408 receives data associated with at least one trajectory from the planning system 404, and the control system 408 controls the operation of the vehicle by generating and transmitting control signals to operate a powertrain control system (e.g., DBW system 202h and / or powertrain control system 204, etc.), a steering control system (e.g., steering control system 206), and / or a braking system (e.g., braking system 208). In the example, in the case where the trajectory includes a left turn, the control system 408 transmits a control signal to cause the steering control system 206 to adjust the steering angle of the vehicle 200, thereby causing the vehicle 200 to turn left. Additionally or alternatively, the control system 408 generates and transmits control signals to cause other devices of the vehicle 200 (e.g., headlights, turn signals, door locks, and / or windshield wipers, etc.) to change state.

[0077] In some embodiments, perception system 402, planning system 404, positioning system 406, and / or control system 408 implement at least one machine learning model (e.g., at least one multi-layer perceptron (MLP), at least one convolutional neural network (CNN), at least one recurrent neural network (RNN), at least one autoencoder, and / or at least one transformer, etc.). In some examples, perception system 402, planning system 404, positioning system 406, and / or control system 408 implement at least one machine learning model alone or in combination with one or more of the above systems. In some examples, perception system 402, planning system 404, positioning system 406, and / or control system 408 implement at least one machine learning model as part of a pipeline (e.g., a pipeline for identifying one or more objects located in an environment, etc.). The following is about FIG. 4B to FIG. 4D Includes examples of implementations of machine learning models.

[0078] Database 410 stores data transmitted to, received from, and / or updated by perception system 402, planning system 404, positioning system 406, and / or control system 408. In some examples, database 410 includes a storage component (e.g., a storage component) for storing data and / or software related to operations and using at least one system of autonomous vehicle computing 400. Figure 3In some embodiments, database 410 stores data associated with a 2D and / or 3D map of at least one area. In some examples, database 410 stores data associated with a 2D and / or 3D map of a portion of a city, portions of multiple cities, multiple cities, counties, states, and / or countries (State) (e.g., a country), etc. In such an example, a vehicle (e.g., a vehicle that is the same or similar to vehicle 102 and / or vehicle 200) can be driven along one or more drivable areas (e.g., single-lane roads, multi-lane roads, highways, remote roads, and / or off-road roads, etc.) and cause at least one LiDAR sensor (e.g., a LiDAR sensor that is the same or similar to LiDAR sensor 202b) to generate data associated with an image representing an object included in the field of view of the at least one LiDAR sensor.

[0079] In some embodiments, database 410 can be implemented across multiple devices. In some examples, database 410 includes a vehicle (e.g., a vehicle that is the same or similar to vehicle 102 and / or vehicle 200), an autonomous vehicle system (e.g., an autonomous vehicle system that is the same or similar to remote AV system 114), a fleet management system (e.g., a vehicle that is the same or similar to remote AV system 114), and a fleet management system (e.g., a vehicle that is the same or similar to remote AV system 114). Figure 1 The same or similar queue management system as the queue management system 116 of FIG. 1 and / or the V2I system (e.g., Figure 1 The V2I system 118 is the same as or similar to the V2I system 118).

[0080] Reference now Figure 4B , a diagram illustrating an implementation of a machine learning model. More specifically, a diagram illustrating an implementation of a convolutional neural network (CNN) 420. For purposes of illustration, the following description of CNN 420 will be with respect to implementing CNN 420 by perception system 402. However, it will be understood that in some examples, CNN 420 (e.g., one or more components of CNN 420) is implemented by other systems (such as planning system 404, positioning system 406, and / or control system 408, etc.) other than or in addition to perception system 402. Although CNN 420 includes certain features as described herein, these features are provided for purposes of illustration and are not intended to limit the present disclosure.

[0081] CNN 420 includes a plurality of convolutional layers including a first convolutional layer 422, a second convolutional layer 424, and a convolutional layer 426. In some embodiments, CNN 420 includes a subsampling layer 428 (sometimes referred to as a pooling layer). In some embodiments, subsampling layer 428 and / or other subsampling layers have a dimension that is smaller than the dimension of the upstream system (i.e., the number of nodes). With the subsampling layer 428 having a dimension that is smaller than the dimension of the upstream layer, CNN 420 merges the amount of data associated with the initial input and / or output of the upstream layer, thereby reducing the amount of computation required for CNN 420 to perform downstream convolution operations. Additionally or alternatively, with the subsampling layer 428 being associated with (e.g., configured to perform) at least one subsampling function (as described below with respect to Figure 4C and Figure 4D As described above, CNN 420 incorporates the amount of data associated with the initial input.

[0082] The perception system 402 performs the convolution operation based on the perception system 402 providing respective inputs and / or outputs associated with each of the first convolution layer 422, the second convolution layer 424, and the convolution layer 426 to generate respective outputs. In some examples, the perception system 402 implements the CNN 420 based on the perception system 402 providing data as input to the first convolution layer 422, the second convolution layer 424, and the convolution layer 426. In such examples, the perception system 402 provides data as input to the first convolution layer 422, the second convolution layer 424, and the convolution layer 426 based on the perception system 402 receiving data from one or more different systems (e.g., one or more systems of a vehicle that is the same or similar to the vehicle 102, a remote AV system that is the same or similar to the remote AV system 114, a queue management system that is the same or similar to the queue management system 116, and / or a V2I system that is the same or similar to the V2I system 118, etc.). The following is about Figure 4C Includes a detailed description of the convolution operation.

[0083] In some embodiments, the perception system 402 provides data associated with the input (referred to as the initial input) to the first convolutional layer 422, and the perception system 402 generates data associated with the output using the first convolutional layer 422. In some embodiments, the perception system 402 provides the output generated by the convolutional layer as input to a different convolutional layer. For example, the perception system 402 provides the output of the first convolutional layer 422 as input to the subsampling layer 428, the second convolutional layer 424, and / or the convolutional layer 426. In such an example, the first convolutional layer 422 is referred to as an upstream layer, and the subsampling layer 428, the second convolutional layer 424, and / or the convolutional layer 426 are referred to as downstream layers. Similarly, in some embodiments, the perception system 402 provides the output of the subsampling layer 428 to the second convolutional layer 424 and / or the convolutional layer 426, and in this example, the subsampling layer 428 will be referred to as the upstream layer, and the second convolutional layer 424 and / or the convolutional layer 426 will be referred to as the downstream layer.

[0084] In some embodiments, before the perception system 402 provides the input to the CNN 420, the perception system 402 processes the data associated with the input provided to the CNN 420. For example, the perception system 402 processes the data associated with the input provided to the CNN 420 based on the perception system 402 normalizing the sensor data (e.g., image data, LiDAR data, and / or Radar data, etc.).

[0085] In some embodiments, CNN 420 generates an output based on perception system 402 performing convolution operations associated with each convolution layer. In some examples, CNN 420 generates an output based on perception system 402 performing convolution operations associated with each convolution layer and the initial input. In some embodiments, perception system 402 generates an output and provides the output to fully connected layer 430. In some examples, perception system 402 provides the output of convolution layer 426 to fully connected layer 430, wherein fully connected layer 430 includes data associated with multiple feature values ​​referred to as F1, F2, ..., FN. In this example, the output of convolution layer 426 includes data associated with multiple output feature values ​​representing predictions.

[0086] In some embodiments, perception system 402 identifies a prediction from the plurality of predictions based on perception system 402 identifying a feature value associated with a highest likelihood of being a correct prediction from the plurality of predictions. For example, where fully connected layer 430 includes feature values ​​F1, F2, ..., FN and F1 is the largest feature value, perception system 402 identifies the prediction associated with F1 as the correct prediction from the plurality of predictions. In some embodiments, perception system 402 trains CNN 420 to generate the predictions. In some examples, perception system 402 trains CNN 420 to generate the predictions based on perception system 402 providing training data associated with the predictions to CNN 420.

[0087] Reference now Figure 4C and Figure 4D , a diagram illustrating an example operation of CNN 440 utilizing perception system 402. In some embodiments, CNN 440 (e.g., one or more components of CNN 440) is coupled to CNN 420 (e.g., one or more components of CNN 420) (see Figure 4B ) are the same or similar.

[0088] At step 450, the perception system 402 provides data associated with the image as input to the CNN 440 (step 450). For example, as illustrated, the perception system 402 provides data associated with the image to the CNN 440, where the image is a grayscale image represented as values ​​stored in a two-dimensional (2D) array. In some embodiments, the data associated with the image may include data associated with a color image represented as values ​​stored in a three-dimensional (3D) array. Additionally or alternatively, the data associated with the image may include data associated with an infrared image and / or a Radar image, etc.

[0089] At step 455, CNN 440 performs a first convolution function. For example, CNN 440 performs a first convolution function based on CNN 440 providing a value representing an image as an input to one or more neurons (not explicitly illustrated) included in first convolution layer 442. In this example, the value representing the image may correspond to a value of a region (sometimes referred to as a receptive field) representing the image. In some embodiments, each neuron is associated with a filter (not explicitly illustrated). The filter (sometimes referred to as a kernel) may be represented as an array of values ​​corresponding in size to the value provided as input to the neuron. In one example, the filter may be configured to identify edges (e.g., horizontal lines, vertical lines, and / or straight lines, etc.). In successive convolution layers, the filters associated with the neurons may be configured to continuously identify more complex patterns (e.g., arcs and / or objects, etc.).

[0090] In some embodiments, CNN 440 performs a first convolution function based on CNN 440 multiplying the values ​​of each neuron provided as input to one or more neurons included in the first convolution layer 442 by the values ​​of the filters corresponding to each neuron in the same or more neurons. For example, CNN 440 may multiply the values ​​of each neuron provided as input to one or more neurons included in the first convolution layer 442 by the values ​​of the filters corresponding to each neuron in the one or more neurons to generate a single value or an array of values ​​as output. In some embodiments, the collective output of the neurons of the first convolution layer 442 is referred to as a convolution output. In some embodiments, when each neuron has the same filter, the convolution output is referred to as a feature map.

[0091] In some embodiments, CNN 440 provides the output of each neuron of the first convolutional layer 442 to the neurons of the downstream layer. For clarity, the upstream layer may be a layer that transmits data to a different layer (referred to as the downstream layer). For example, CNN 440 may provide the output of each neuron of the first convolutional layer 442 to the corresponding neurons of the subsampling layer. In the example, CNN 440 provides the output of each neuron of the first convolutional layer 442 to the corresponding neurons of the first subsampling layer 444. In some embodiments, CNN 440 adds a bias value to the aggregate set of all values ​​provided to each neuron of the downstream layer. For example, CNN 440 adds a bias value to the aggregate set of all values ​​provided to each neuron of the first subsampling layer 444. In such an example, CNN 440 determines the final value to be provided to each neuron of the first subsampling layer 444 based on the aggregate set of all values ​​provided to each neuron and the activation function associated with each neuron of the first subsampling layer 444.

[0092] At step 460, CNN 440 performs a first subsampling function. For example, based on CNN 440 providing the values ​​output by first convolutional layer 442 to corresponding neurons of first subsampling layer 444, CNN 440 may perform the first subsampling function. In some embodiments, CNN 440 performs the first subsampling function based on an aggregation function. In an example, CNN 440 performs the first subsampling function based on CNN 440 determining the maximum input (referred to as a max pooling function) among the values ​​provided to a given neuron. In another example, CNN 440 performs the first subsampling function based on CNN 440 determining the average input (referred to as an average pooling function) among the values ​​provided to a given neuron. In some embodiments, based on CNN 440 providing values ​​to individual neurons of first subsampling layer 444, CNN 440 generates an output, which is sometimes referred to as a subsampled convolution output.

[0093] At step 465, CNN 440 performs a second convolution function. In some embodiments, CNN 440 performs the second convolution function in a manner similar to how CNN 440 performs the first convolution function described above. In some embodiments, CNN 440 performs the second convolution function based on CNN 440 providing the value output by first subsampling layer 444 as input to one or more neurons (not explicitly illustrated) included in second convolution layer 446. In some embodiments, as described above, each neuron of second convolution layer 446 is associated with a filter. As described above, the filter (one or more) associated with second convolution layer 446 can be configured to recognize more complex patterns than the filter associated with first convolution layer 442.

[0094] In some embodiments, the CNN 440 performs a second convolution function based on the CNN 440 multiplying the value of each neuron provided as input to the one or more neurons included in the second convolution layer 446 by the value of the filter corresponding to each neuron of the one or more neurons. For example, the CNN 440 may multiply the value of each neuron provided as input to the one or more neurons included in the second convolution layer 446 by the value of the filter corresponding to each neuron of the one or more neurons to generate a single value or a value array as an output.

[0095] In some embodiments, the CNN 440 provides the output of each neuron of the second convolutional layer 446 to the neurons of the downstream layer. For example, the CNN 440 may provide the output of each neuron of the first convolutional layer 442 to the corresponding neurons of the subsampling layer. In an example, the CNN 440 provides the output of each neuron of the first convolutional layer 442 to the corresponding neurons of the second subsampling layer 448. In some embodiments, the CNN 440 adds a bias value to the aggregate set of all values ​​provided to each neuron of the downstream layer. For example, the CNN 440 adds a bias value to the aggregate set of all values ​​provided to each neuron of the second subsampling layer 448. In such an example, the CNN 440 determines the final value provided to each neuron of the second subsampling layer 448 based on the aggregate set of all values ​​provided to each neuron and the activation function associated with each neuron of the second subsampling layer 448.

[0096] At step 470, CNN 440 performs a second subsampling function. For example, based on CNN 440 providing the values ​​output by second convolutional layer 446 to corresponding neurons of second subsampling layer 448, CNN 440 may perform the second subsampling function. In some embodiments, based on CNN 440 using an aggregation function, CNN 440 performs the second subsampling function. In the example, as described above, based on CNN 440 determining the maximum input or average input among the values ​​provided to a given neuron, CNN 440 performs the first subsampling function. In some embodiments, based on CNN 440 providing values ​​to respective neurons of second subsampling layer 448, CNN 440 generates an output.

[0097] At step 475, CNN 440 provides the output of each neuron of the second subsampling layer 448 to the fully connected layer 449. For example, CNN 440 provides the output of each neuron of the second subsampling layer 448 to the fully connected layer 449 so that the fully connected layer 449 generates an output. In some embodiments, the fully connected layer 449 is configured to generate an output associated with a prediction (sometimes referred to as a classification). The prediction may include an indication that the objects included in the image provided as input to CNN 440 include objects and / or sets of objects, etc. In some embodiments, the perception system 402 performs one or more operations and / or provides data associated with the prediction to the various systems described herein.

[0098] Generate bounding boxes for navigation

[0099] As described herein, in order to improve the functionality of an autonomous vehicle and its ability to generate bounding boxes and navigate an environment in real time, the autonomous vehicle can be configured to group the grid cells of a feature map and combine or associate the features of the grid cells within the group. In some cases, the grid cells can be grouped based on the contours of the object. In some cases, one or more regions or windows that traverse multiple feature maps can be used to group the grid cells so that the features of the grid cells from different feature maps are enhanced or correlated with each other. Different stages can use windows of different numbers, sizes, or shapes, and multiple layers with shifted windows can be used to further enhance the grid cells.

[0100] By grouping grid cells into subsets of feature maps, autonomous vehicles can use fewer computational resources to inter-correlate and cross-correlate the features of grid cells, which can increase the speed of self-attention processing and the speed of enhancing feature maps.

[0101] In addition, the enhanced feature map can improve the autonomous vehicle's ability to identify objects within the scene of the autonomous vehicle. For example, the enhanced grid cells can improve the autonomous vehicle's ability to determine bounding boxes for objects in the vehicle scene.

[0102] The autonomous vehicle may also generate object queries and enhance these object queries using: (enhanced) feature maps, features from other object queries (similar to how grid cells of feature maps are enhanced), and / or features from a bird's eye view (BEV) feature map. The enhanced object queries may improve the autonomous vehicle's ability to identify objects within the autonomous vehicle's scene.

[0103] Figure 5A 5 is a block diagram illustrating an example perception environment 500 in which a perception system 402 receives and processes an image 502 to provide one or more (3D) bounding boxes 512 for objects in a vehicle scene (corresponding to the image 502). In the illustrated example, the perception system 402 includes an image feature extractor 504, an attention stage 506, a BEV stage 508, and a detection stage 510. However, it will be understood that the perception system may include fewer or more components. In some cases, the perception system 402 may omit the BEV stage 508. For example, in some cases, the attention stage 506 may output one or more object queries and may not output a feature map. In some such cases, the object will not be associated with the feature map by the BEV stage 508. The various components of the perception system 402 described herein may be implemented using one or more processors and / or as one or more layers or stages of a machine learning model or neural network.

[0104] The image 502 (also referred to herein as the image set 502) for a particular scene may include image data from one or more sensors in the sensor suite. The image 502 may include different types of images corresponding to the sensors or devices used to generate the image. For example, the image 502 may be a camera image generated from one or more cameras (such as camera 202a, etc.), or a LiDAR image generated from one or more LiDAR sensors (such as LiDAR sensor 202b, etc.). Other image types may be used, such as a Radar image generated from one or more Radar sensors (e.g., generated from Radar sensor 202c), etc. Each image may correspond to different image sensors (or cameras) placed at different locations around the self-carrier. In some cases, the combination of images may form a 360-degree view of the scene of the self-carrier from the perspective of the self-carrier. In this way, each image in the image 502 may be adjacent or contiguous to another image in the image 502, and some objects (or parts of objects) may appear in different images in the image 502.

[0105] In addition, images 502 in the image set may be generated at approximately the same time and may form part of a stream of different images. In this way, images 502 may represent the scene of the vehicle at a particular time. Since perception system 402 uses the images to generate bounding box 512 and navigate the vehicle, it will be understood that perception system 402 may process images 502 in real time or near real time to generate bounding box 512.

[0106] The image feature extractor 504 may be implemented using one or more neural networks or layers of neural networks to extract features from the image 502. In some cases, the image feature extractor 504 may be implemented using a backbone having a feature pyramid network (FPN), a residual network (Resnet), or a Swin transformer, a CSWin transformer, etc. The image feature extractor 504 may use the image 502 to generate one or more feature maps. In some cases, the image feature extractor 504 generates at least one feature map for each image 502. For example, if the image feature extractor 504 receives six images corresponding to six cameras placed at different locations around the vehicle and oriented in different ways (e.g., to obtain a 360-degree view of the area around the vehicle), the image feature extractor 504 may generate six feature maps, respectively.

[0107] The feature maps may have the same or different shapes as the images used to generate the feature maps and / or from each other. For example, if the individual images 502 have shapes [900, 1600, 3], then the corresponding feature maps may have shapes [45, 80, 256], however, it will be understood that feature maps may have different shapes from each other.

[0108] Each of the generated feature maps may include an array of grid cells having a particular channel depth. The grid cells may include semantic data (or features) extracted from (pixels in) the (one or more) images from which the feature maps are generated. The features of the grid cells may be organized as vectors or some other tensor shape. For example, the features (or semantic data) of the grid cells may indicate the shape, light, texture, reflectivity, edges, object class, location, etc. of something detected by the image feature extractor 504.

[0109] Object query initialization phase

[0110] The object query formulation phase 505 (also referred to herein as the formulation phase 505) may be used to initialize, seed / modify, and / or enhance the object query. Thus, it will be appreciated that the formulation phase 505 may include one or more sub-phases, including but not limited to an initialization phase, a seed / modification phase, and / or an enhancement phase.

[0111] In some cases, formulation stage 505 may initiate a specific number of object queries. In some cases, formulation stage 505 initiates more object queries than the number of objects expected to be found in image(s) 502. For example, if formulation stage 505 expects that there are no more than 400 objects in the scene of image 502, formulation stage 505 may initiate some number of object queries greater than 400, such as 900 object queries, etc.

[0112] The object query may be organized as a vector or some other tensor shape, and / or may include the same or different number of features. For example, an object query may include 256 dimensional features, or fewer or more dimensional features. Features, alone or in combination, may represent one or more characteristics of an object, such as but not limited to its class, movement, association with other objects, whether it is a foreground or background, location, shape, size, color, texture, reflectivity, etc. In some cases, the formulation stage 505 may initialize the features of the object query randomly and / or pseudo-randomly. For example, the values ​​of the features for the object query may include random numbers or pseudo-random numbers.

[0113] In addition, the formulation stage 505 may initially set or modify the initial (random or pseudo-random) values ​​of the features for the object query. In some cases, the formulation stage 505 may include a positioning network (or receive values ​​from a positioning network) for determining (or helping to determine) the possible locations of the corresponding object query within the vehicle scene. In some cases, the formulation stage 505 may use other data to initially set the object query (or modify the initial value of the object query). For example, the formulation stage 505 may use a heat map for indicating the expected or possible movement or trajectory of the object within the vehicle scene to formulate the object query. In some cases, the formulation stage 505 may include an object cross-attention stage (similar to the object query cross-attention stage 513 described herein), which relates the object query to data from one or more grid cells of one or more feature maps generated by the image feature extractor 504. As part of relating the object query to one or more grid cells, the formulation stage 505 may use one or more linear layers to identify one or more grid cells in one or more feature maps. For example, the formulation stage 505 may multiply the tensor [1, N] corresponding to the object query by the learnable linear layer matrix [N, 2] to determine the location in the feature map corresponding to the grid cell to be associated with the object query. The formulation stage 505 may use the features of the grid cell to modify some or all of the features of the object query. In some cases, this may include: assigning weights to specific features of the object query and assigning weights to corresponding features of the identified grid cells, and using the result (non-limiting example: sum of products) to modify the specific features of the object query or assign new values ​​to the specific features of the object query. In some cases, the formulation stage 505 may use the learnable linear layer matrix to identify multiple grid cells of one or more feature maps, and use the identified grid cells to modify the features of the object query. In some such cases, the formulation stage 505 may assign different weights to the features of different grid cells, and use the weighted features to determine the corresponding features of the object query.

[0114] In some cases, formulation stage 505 may include an object query self-attention stage (e.g., similar to object query self-attention stage 514 described herein) that enables object queries to perform self-attention operations and update themselves. For example, the self-attention stage of formulation stage 505 may modify the values ​​of the object query group based on the features of the object queries in the object query group. As described herein at least with reference to object query self-attention stage 514, the object query self-attention stage of formulation stage 505 may compare the features of a particular object query with some or all of the features in other object queries (or some or all of the object features in the object feature group) to determine the relevance or similarity between the particular object query and the other object queries, use the relevance between the particular object query and the other object queries to weight the features of the various object queries, and use the weighted features to calculate new (or modified) values ​​for the corresponding features of the particular object query. In some cases, object query self-attention stage 514 may update some or all of the features in the object query in this manner. In some cases, the object query self-attention stage 514 can determine a matrix indicating the relationships (or weights) between the features of the various object queries, and use the matrix to update some or all of the features in the object query. The attention stage 506 can be used to enhance the feature map and the object query. In some cases, the attention stage 506 can use self-attention and / or cross-attention techniques to enhance the feature map and the object query.

[0115] In the illustrated example, the attention stage 506 includes one or more layers of an object query cross-attention stage 513, an object query self-attention stage 514, a multi-view self-attention stage 516 (also referred to herein as the multi-view stage 516), and / or a region of interest (ROI) self-attention stage 518 (also referred to herein as the ROI stage 518).

[0116] Attention Stage

[0117] Different layers of the attention stage 506 may include similar components. Figure 6In the illustrated example of , the layers of the attention stage 506 include an object query cross-attention stage 513, an object query self-attention stage 514, a multi-view stage 516, and an ROI stage 518, however, it will be understood that the attention stage 506 and / or different layers of the attention stage 506 may include different components. For example, the attention stage 506 and / or different layers of the attention stage 506 may include different components or different relationships between them (e.g., the ROI stage 518 may be placed in front of the multi-view stage 516, etc.). In some cases, the layers of the attention stage 506 include the same components in the same relationship. In some cases, the layers of the attention stage 506, the components of different layers, may be configured differently or use different parameters. For example, the multi-view stage 516 of the first layer may use different parameters or configurations to process feature maps compared to the multi-view stage 516 of the second layer. Similarly, for the object query cross-attention stage 513, the object query self-attention stage 514, and / or the ROI stage 518, etc., different parameters may be used in different layers.

[0118] The components of the layer of attention stage 506 may process data in parallel or sequentially. In some cases, the output of one stage within a layer may be used as input to another stage within the layer. For example, in Figure 6 In the illustrated example of , the object query self-attention stage 514 processes data output by the object query cross-attention stage 513, the object query cross-attention stage 513 processes data output by the ROI stage 518, and the ROI stage 518 processes data output by the multi-view stage 516. However, it will be appreciated that the components of the attention stage 506 may be aligned in a variety of configurations. For example, the multi-view stage 516 may be configured to process the output of the ROI stage 518.

[0119] The output of one layer of the attention stage 506 may be used as input to a subsequent layer, and the output of the last layer of the attention stage 506 may be provided to the BEV stage 508 and / or the detection stage 510. For example, in an N-layer attention stage 506, the output of the first layer may be used as input to the second layer, and so on, until the output of the N-1th layer is used as input to the Nth layer. In some such cases, the output of the Nth layer may be used as input to the BEV stage 508 and / or the detection stage 510.

[0120] Object query cross-attention stage

[0121] The object query cross-attention stage 513 can be configured to enhance (a set of) object queries (e.g., using cross-attention and / or using data from another stage (such as the multi-view stage 516 and / or the ROI stage 518 described in more detail below).

[0122] In some cases, the object query cross-attention stage 513 can enhance the object query based on data received from the multi-view stage 516 and / or the ROI stage 518. For example, the object query cross-attention stage 513 can modify or edit the object query using semantic data corresponding to one or more grid cells in one or more feature maps output by the multi-view stage 516 and / or the ROI stage 518. In some cases, this can include modifying one or more features of a tensor corresponding to the object query.

[0123] In some cases, the object query criss-cross attention stage 513 may correlate data from one or more grid cells of one or more (enhanced) feature maps enhanced by the multi-view stage 516 and / or the ROI stage 518. As part of correlating the object query with the one or more grid cells, the object query criss-cross attention stage 513 may use one or more linear layers to identify the one or more grid cells in the one or more (enhanced) feature maps. For example, the object query criss-cross attention stage 513 may multiply the tensor [1, N] corresponding to the object query by the learnable linear layer matrix [N, 2] to determine the location (x, y) in the (enhanced) feature map corresponding to the grid cell to be correlated with the object query.

[0124] The object query cross-attention stage 513 may use the features of the identified grid cells to modify some or all of the features of the object query. In some cases, this may include: assigning weights to specific features of the object query and assigning weights to corresponding features of the identified grid cells, and using the result (non-limiting example: sum of products) to modify the specific features of the object query or assign new values ​​to the specific features of the object query. In some cases, the object query cross-attention stage 513 may use a learnable linear layer matrix to identify multiple grid cells of one or more feature maps, and use the identified grid cells to modify the features of the object query. In some such cases, the object query cross-attention stage 513 may assign different weights to corresponding features of different grid cells, and use the weighted features to determine the corresponding features of the object query.

[0125] Object query self-attention stage

[0126] The object query self-attention stage 514 may be configured to generate and / or enhance an object query using features from the object query set (eg, using self-attention).

[0127] As described herein, there may be many object queries, and some or all of these object queries may be modified by the object query cross-attention stage 513 (e.g., using grid cells from the (enhanced) feature map). In some cases, the object query self-attention stage 514 may modify or enhance the object query by comparing features of the object query to each other, determining weighted values ​​based on the comparison, and modifying the features of the object query using the weighted features (weighted based on the determined weighted values). For example, the object query self-attention stage 514 may compare semantic data (or features) of the object query to determine relationships between the object queries, such as the likelihood that different object queries correspond to the same object or different objects, etc. In some cases, this may include comparing features of the object query corresponding to the class of the object, movement, association with other objects, whether the object is foreground or background, color, light reflectivity, edges, texture, shape, etc. Based on the comparison, the object query self-attention stage 514 may update the object query. In some cases, this may include modifying one or more values ​​of a tensor corresponding to the object query.

[0128] In some cases, the object query self-attention stage 514 compares the features of the specific object query with some or all of the features in other object queries (or some or all of the object features in the object feature group) to determine the relevance or similarity between the specific object query and the other object queries. In some cases, the relevance or similarity can be expressed as a probability or weight. Using the relevance between the specific object query and the other object queries, the features of the object query (including the specific object query) can be weighted, and the weighted features can be used to calculate the new (or modified) value of the corresponding feature for the specific object query. For example, the first feature of some or all of the object query can be weighted (relative to the specific object query), and the weighted value is used to determine the value of the first feature for the specific object query. Similarly, other features of the specific object query can be updated (e.g., using the same or different weights). In some cases, the object query self-attention stage 514 can update some or all of the features in the object query in this way. In some cases, the object query self-attention stage 514 can determine a matrix to indicate the relationship (or weight) between the features of various object queries, and use the matrix to update some or all of the features in the object query.

[0129] As a non-limiting example, consider the following three object queries and values ​​for their features: Object query1 [1, 4] = (.2, .2, .4, .7); Object query2 [1, 4] = (.3, .4, .6, .7); Object query3 [1, 4] = (.1, .8, .9, .7).

[0130] After analyzing the features of the three object queries, it is assumed that the object query self-attention stage 514 generates the following relationship (or weighting) matrix between them.

[0131] Object Query 1 Object Query 2 Object Query 3 Object Query 1 .7 .2 .1 Object Query 2 .2 .6 0.2 Object Query 3 .1 .2 .7

[0132] Based on the determined relationships or weightings, the object query self-attention stage 514 may update the values ​​of the features for the object query as follows.

[0133] Object query1[1,4]=(.21,.3,.49,.7) or (.7*.2+.2*.3+.1*.1,.7*.2+.2*.4+.1*.8,.7*.4+.2*.6+.1*.9,.7*.7+.7*.2+.7*.1)

[0134] Object query2[1,4]=(.24,.44,.62,.7) or (.2*.2+.6*.3+.2*.1,.2*.2+.6*.4+.2*.8,.2*.4+.6*.6+.2*.9,.2*.7+.6*.7+.2*.7

[0135] Object query3[1,4]=(.15,.66,.79,.7) or (.1*.2+.2*.3+.7*.1,.1*.2+.2*.4+.7*.8,.1*.4+.2*.6+.7*.9,.1*.7+.2*.7+.7*.7

[0136] Multi-view stage

[0137] The multi-view stage 516 may enhance the feature map by comparing and / or correlating features from different grid cells of the feature map. In some cases, the multi-view stage 516 uses features from grid cells in a grid cell group to update each other (also referred to herein as self-attention). For example, the multi-view stage 516 may use features of a grid cell group in one or more feature maps to enhance or modify features of a particular grid cell in the grid cell group.

[0138] In some cases, the multi-view stage 516 may group the grid cells based on objects (e.g., grouping grid cells that correspond to (or appear to correspond to) the same object or the outline of the same object). In some cases, the multi-view stage 516 may group the grid cells by dividing the feature map into multiple regions (also referred to herein as windows), and / or assigning different grid cells of the feature map to different regions or windows. In some cases, different regions or windows of the feature map may be mutually exclusive (e.g., a grid cell may be assigned to only one region or window). In some cases, the multi-view stage 516 may divide the feature map into multiple rows or columns of regions or windows. Some or all of the regions or windows may have the same (or different) size (e.g., width and height), and one or more of the regions may overlap with multiple feature maps corresponding to different images. The rows of windows may be aligned or offset from each other.

[0139] As described herein, by using groups of grid cells (e.g., windows or subsets of images / feature maps) for comparison and enhancement (e.g., for self-attention) instead of entire feature maps or sets of feature maps, the attention stage 506 can reduce processing requirements and increase the speed and efficiency of processing feature maps using the attention stage 506. These efficiencies can increase the rate at which the perception system 402 can accurately identify objects in the image 502.

[0140] Figure 5B is a diagram illustrating examples of rows of windows 552 (referred to as 552a, 552b, 552c, etc.) applied to feature maps 551 (referred to as 551a to 551f, respectively) corresponding to different images (received from different image sensors). In the illustrated example, the zones are of equal size and the rows are aligned with the rows above and below. However, it will be understood that windows of different sizes may be used and / or the rows may be offset from one another. In some cases, the rows may be offset relative to one another. In some cases, alternating rows may be aligned with an offset in the middle row (e.g., like rows of bricks). Additionally, as Figure 5B As illustrated, window 552 may overlap across multiple feature maps. For example, window 552a includes grid cells from feature map 551a and feature map 551f.

[0141] The multi-view stage 516 may compare semantic data of a group of grid cells (e.g., different grid cells within a particular window or region) with each other. Based on the comparison, the multi-view stage 516 may modify the semantic data of the different grid cells. For example, the multi-view stage 516 may compare certain features of the grid cells (e.g., color, reflectivity, shape, etc.) with corresponding features of different grid cells in the same group (e.g., comparing features of a grid cell within window 552a with corresponding features of a different grid cell within window 552a). Based on the similarity, the multi-view stage 516 may determine a probabilistic relationship between the grid cells in the group (e.g., the probability that the grid cells are part of the same object (such as a vehicle, a bicycle, a pedestrian, a construction cone, etc.)). Based on the determination, the multi-view stage 516 may update one or more features of the grid cells. For example, one grid cell may be updated to indicate that it is the middle of an object, another grid cell may be updated to indicate that it is the beginning of the same object, and so on.

[0142] In some cases, the multi-view stage 516 may correlate some or all of the features of various grid cells within a group (e.g., within a particular region) with each other. In this manner, the multi-view stage 516 may enhance some or all of the grid cells within a particular group. Furthermore, the multi-view stage 516 may repeat the comparison for each of a group (e.g., window) of feature maps and / or across some or all of the feature maps, such that some or all of the grid cells of a feature map are compared / updated based on comparisons with features from other grid cells in the same group (e.g., window or region).

[0143] As described above, with reference to the self-attention of the object query, in some cases, the multi-view stage 516 may generate a matrix including some or all of the grid cells in the group. Then, the multi-view stage 516 may determine the weights or probabilistic relationships between the grid cells and include the weights in the matrix. The multi-view stage 516 may use the weights / relationships in the matrix (indicating the relationship or weights between the grid cells) to calculate the updated values ​​of the features for different grid cells. For example, the multi-view stage 516 may update the specific value of a specific grid cell using the corresponding weighted values ​​of some or all of the other grid cells in the group. Examples of such matrices and calculations (but for object queries) are described herein with reference to the self-attention of the object query. In addition, the process may be repeated across some or all of the grid cell groups of the feature map and across some or all of the feature maps. For example, the image feature extractor 504 may generate multiple feature maps for each image, where each feature map corresponds to one or more detected features of the image. In some such cases, a window (or other form of grouping) may be applied to some or all of the feature maps and grid cells of the feature maps updated as described herein.

[0144] In some cases, the attention stage 506 may include multiple layers of multi-view stages 516. In some such cases, the groups (e.g., windows) in the multi-view stages 516 of different layers may be different. In some cases, the size and / or position of the windows may be different. For example, Figure 5B As illustrated by window 554 (referred to as 554a, 554b, 554c, etc.) on feature map 551, window 554 in the second layer of attention stage 506 can be offset relative to window 552 in the first layer of attention stage 506.

[0145] In some cases where the windows in different layers are offset from each other, alternating layers may use the same or different positions. For example, in an attention stage 506 having six layers, the odd-numbered layers (e.g., layer 1, layer 3, and layer 5) may be similar to Figure 5B The rows of windows 552 shown are aligned, and even-numbered layers (e.g., layer 2, layer 4, and layer 6) may be similar to Figure 5B The rows of windows 554 (individual windows referred to as 554a, 554b, 554c, etc.) are shown aligned.

[0146] However, it will be understood that windows in different layers can be offset from each other as desired. In some cases, each layer can have a different offset than other layers. In some cases, every second, third, fourth (or every other number of) layers in a layer can have the same offset of windows, etc. Figure 5B In the illustrated example, different layers are offset from each other unidirectionally (horizontally), however, it will be understood that different layers may be offset from each other in other directions (e.g., vertically) and / or in multiple directions (e.g., horizontally and vertically).

[0147] By shifting (or changing groups of) windows in different layers, the attention stage 506 can improve the enhancement of (grid cells of) feature maps. For example, when grid cells within window 552a are compared and these network cells are used to enhance each other at one layer, these enhancements can be propagated to grid cells within window 554a in subsequent layers. In this way, enhancements can be propagated across feature maps / images.

[0148] Region of Interest (ROI) Stage

[0149] The ROI stage 518 may enhance the feature map by comparing and / or correlating features from different grid cells of the feature map. In some cases, the ROI stage 518 uses features from grid cells in a grid cell group (such as grid cells in a window, etc.) to update each other (also referred to herein as self-attention). For example, the ROI stage 518 may use features of a grid cell group in one or more feature maps to enhance or modify features of a particular grid cell in the grid cell group.

[0150] In some cases, the ROI stage 518 may group grid cells based on objects (e.g., grouping grid cells that appear to correspond to the same object or the outline of the same object). In some cases, similar to the multi-view stage 516, but using regions or windows of different shapes or sizes, the ROI stage 518 may group grid cells by dividing the feature map into multiple regions or windows and / or assigning different grid cells of the feature map to different regions or windows. In some cases, different regions or windows in a layer of the ROI stage 518 may be mutually exclusive (e.g., grid cells of a feature map may be assigned to only one region or window). In some cases, the ROI stage 518 may divide the feature map into multiple rows or columns of regions or windows. Some or all of these regions or windows may have equal (or different) sizes, and one or more of these regions may overlap with multiple feature maps corresponding to different images. The rows or columns may be aligned or offset from each other.

[0151] As described herein, by using groups, windows, or subsets of grid cells of an image or feature map for comparison and enhancement, the attention stage 506 can reduce processing requirements and increase the speed and efficiency of processing feature maps using the attention stage 506. These efficiencies can increase the rate at which the perception system 402 can accurately identify objects in the image 502.

[0152] Figure 5C 555a to 555f) corresponding to different images (received from different image sensors) is a diagram illustrating an example of columns of windows 556 (referred to as windows 556a, 556b, etc., respectively) applied to feature maps 555 (referred to as individual feature maps 555a to 555f) corresponding to different images (received from different image sensors). In the illustrated example, the windows 556 are of equal size and the columns are aligned with the columns to the left and right. However, it will be understood that windows of different sizes may be used and / or the columns may be offset from each other. In some cases, the columns may be offset relative to one another. In some cases, alternating columns may be aligned with the center column offset (e.g., like rows of bricks).

[0153] The ROI stage 518 may compare semantic data of a group of grid cells (e.g., different grid cells within a particular window or zone) with each other. Based on the comparison, the ROI stage 518 may modify the semantic data of different grid cells. For example, the ROI stage 518 may compare certain features of a grid cell (e.g., color, reflectivity, shape, edge, etc.) with corresponding features of different grid cells in the same group (e.g., comparing features of a grid cell within window 556a with corresponding features of different grid cells within window 556a). Based on the similarity, the ROI stage 518 may determine a probabilistic relationship between the grid cells in the group (e.g., the probability that the grid cell is part of the same object (such as a vehicle, a bicycle, a pedestrian, a construction cone, etc.)). For example, one grid cell may be updated to indicate that it is the middle of an object, and another grid cell may be updated to indicate that it is the beginning of the same object. As another non-limiting example, one grid cell may be updated to indicate that it is moving at 60 m / s, and another grid cell may be updated to indicate that it is moving at 10 m / s, and so on.

[0154] In some cases, the ROI stage 518 may correlate some or all of the features of various grid cells within a group (e.g., within a particular region) with each other. In this manner, the ROI stage 518 may enhance some or all of the grid cells within a particular group. Furthermore, the ROI stage 518 may repeat the comparison for each of a group (e.g., window) of feature maps and / or across some or all of the feature maps, such that some or all of the grid cells of a feature map are compared / updated based on comparisons with features from other grid cells in the same group (e.g., window or region).

[0155] As described above, with reference to the self-attention of object queries, in some cases, the ROI stage 518 may generate a matrix including some or all of the grid cells in the group. Then, the ROI stage 518 may determine the weights or probabilistic relationships between the grid cells and include the weights in the matrix. The ROI stage 518 may use the weights / relationships in the matrix (indicating the relationships or weights between the grid cells) to calculate updated values ​​for features of different grid cells. For example, the ROI stage 518 may update a specific value of a specific grid cell using the corresponding weighted values ​​of some or all of the other grid cells in the group. Examples of such matrices and calculations (but for object queries) are described herein with reference to the self-attention of object queries. In addition, the process may be repeated across some or all of the grid cell groups of the feature map and across some or all of the feature maps. For example, the image feature extractor 504 may generate multiple feature maps for each image, wherein each feature map corresponds to one or more detected features of the image. In some such cases, a window (or other form of grouping) may be applied to some or all of the feature maps and grid cells of the feature maps updated as described herein.

[0156] In some cases, the attention stage 506 may include multiple layers of ROI stages 518. In some such cases, the groups (e.g., windows) in the ROI stages 518 of different layers may be different. In some cases, the size and / or position of the windows may be different. For example, Figure 5C As illustrated by window 558 (the individual windows referred to as 558a, 554b, etc.) on the feature map 555 of , window 558 in the second layer of the attention stage 506 can be offset relative to window 556 in the first layer of the attention stage 506.

[0157] In some cases where the windows in different layers are offset from each other, alternating layers may use the same or different positions. For example, in an attention stage 506 having six layers, the odd-numbered layers (e.g., layer 1, layer 3, and layer 5) may be similar to Figure 5C 556, and even-numbered layers (e.g., layers 2, 4, and 6) may be similar to Figure 5C 558 of the attention stage. However, it will be appreciated that windows in different layers may be offset from each other as desired. In some cases, each layer may have a different offset than other layers. For example, in an attention stage 506 having six layers, the layers may be offset from each other by 1 / 6 of the image size so that the windows advance across the image.

[0158] In some cases, the borders of a window may overlap and / or wrap around multiple feature maps 555. For example, Figure 5CAs illustrated, window 558b is on the lower right side of first feature map 555f and on the upper left side of second feature map 555a. Figure 5C In the illustrated example, the windows 556, 558 at different layers are offset from each other in multiple directions (horizontally and vertically), however, it will be understood that the windows 556, 558 at different layers can be offset from each other in other ways (e.g., unidirectionally).

[0159] By offsetting the windows in different layers, the attention stage 506 can improve the enhancement of the feature map. For example, when the grid cells within window 556a are compared and these network cells are used to enhance each other at one layer, these enhancements can be propagated to the grid cells within windows 558a and 558b in subsequent layers. In this way, the enhancements can be propagated across the grid cells of the feature map.

[0160] As described above, in some cases, the ROI stage 518 can enhance the feature map using points of interest. For example, the feature map generated by the image feature extractor 504 can include indications of objects in the scene, such as, but not limited to, pedestrians, bicycles, motorcycles, vehicles, construction cones, etc. In addition, the grid cells of the feature map can indicate the outline of the object. In some such cases, the ROI stage 518 can use grid cells corresponding to the outline of the object (e.g., rather than grid cells within the window) as a group of grid cells for self-attention. For example, as described herein, the ROI stage 518 can compare the features of grid cells corresponding to the outline of the object to determine a relationship (e.g., a correlation or weight) between the grid cells, and then update the features of the grid cells based on the determined relationship. By using grid cells associated with the outline of the potential object (rather than grid cells within the window), the ROI stage 518 can further reduce the computational time and resources used to determine a bounding box for the object. For example, the ROI stage 518 can use grid cells that form the outer edge or outline of the object, rather than a group of grid cells that include grid cells that form an area of ​​the window. Examples of different groups of grid cells within different feature maps indicating points of interest (e.g., the outline of the object) are shown by Figure 5D The grid cell groups 560, 562, 564, 566, 568, 570 are shown. Figure 5D In the illustrated examples of , each grid cell group 560, 562, 564, 566, 568, 570 includes grid cells that may include portions of the same object (e.g., grid cells on the outline of an object). Although the illustrated examples are shown with three points of interest for each grid cell group, it will be appreciated that fewer or more points of interest may be included in a grid cell group. For example, a grid cell group may include twenty, thirty, fifty, or more grid cells that form at least a portion of the outline of an object.

[0161] Bird's-Eye View Stage

[0162] The BEV stage 508 may combine outputs from different stages of the attention stage 506. For example, the BEV stage 508 may combine or correlate (or perform a cross-attention operation on) one or more enhanced object queries output from the object query self-attention stage 514 with one or more enhanced feature maps (or BEV feature maps generated from the enhanced feature maps) output from the multi-view stage 516 and / or the ROI stage 518.

[0163] exist Figure 5A In the illustrated example of , the BEV stage 508 includes a BEV generator 520 and a cross-attention stage 522. However, it will be understood that the BEV stage 508 may include fewer or more components.

[0164] The BEV generator 520 may be implemented using an image to BEV encoder such as a Lift, Splat, Shoot encoder, examples of which are described in "Lift, Splat, Shoot: Encoding Images from Arbitrary Camera Rigs by Implicitly Unprojecting to 3D" by Philion et al., August 13, 2020 (arXiv:2008.05711v1) (incorporated herein by reference). In some cases, the BEV generator 520 converts the enhanced feature map output by the attention stage 506 into a BEV feature map. When converting the enhanced feature map into the BEV feature map, the BEV generator 520 may correlate or group grid cells from the enhanced feature map that are mapped to the same grid cell of the BEV feature map. Thus, in some cases, multiple grid cells from the enhanced feature map may be mapped to the same grid cell of the BEV feature map.

[0165] The cross-attention stage 522 may use the BEV feature map to enhance the enhanced object query received from the attention stage 506, and vice versa. In some cases, as part of performing a cross-attention operation on the enhanced object query, the cross-attention stage 522 identifies a grid cell of the BEV feature map corresponding to the enhanced object query, and uses the value of the identified BEV grid cell to modify the value of the object query. As described herein, in some cases, the cross-attention stage 522 may determine (e.g., using a linear layer) a location associated with the enhanced object query, and identify relevant BEV grid cells based on the determined location. For example, the cross-attention stage 522 may multiply a tensor [1, N] corresponding to the enhanced object query by a learnable linear layer matrix [N, 3] to determine a location in the BEV feature map corresponding to a grid cell to be associated with the enhanced object query. In some cases, the cross-attention stage 522 may use multiple BEV grid cells of the BEV feature map to modify or enhance the value of the object query.

[0166] In the same manner, the cross-attention stage 522 can enhance the BEV feature map with information from the object query. For example, in some cases, each grid cell of the BEV feature map can be considered as a BEV object query. Thus, the enhanced object query corresponding to the grid cell of the BEV feature map can be used to modify / update the corresponding BEV grid cell. For example, the features of the enhanced object query mapped to a particular BEV grid cell can be used to modify / update the features of the particular BEV grid cell (e.g., using a weighted value).

[0167] Detection Phase

[0168] The detection stage 510 uses the output of the cross-attention stage 522 to determine a bounding box for an object query and can be implemented using a detector such as a center point detector, an example of which is described in “Center-based 3D Object Detection and Tracking” by Yin et al., published on January 6, 2021 (arXiv:2006.11275v2).

[0169] Fewer, more, or different components may be used as part of the perception system 402. For example, in some cases, the perception system 402 may omit the BEV stage 508. In some such cases, the enhanced object query output by the attention stage 506 may be communicated to the detection stage 510 to detect the bounding box 512. As another example, in some cases, the object query cross-attention stage 513, the multi-view stage 516, and / or the ROI stage 518 may be omitted. In some such cases, the feature map generated by the image feature extractor 504 may be communicated to the BEV stage 508, and the BEV stage 508 may generate a BEV feature map (e.g., using the BEV generator 520) and perform a cross-attention operation on the object query (e.g., received from the attention stage 506) and the BEV feature map. As another example, in some cases, the multi-view stage 516 and / or the ROI stage 518 may be omitted or combined. For example, one or a combination of the multi-view stage 516 and / or the ROI stage 518 may be used to enhance the feature map generated by the image feature extractor 504 in one or more layers.

[0170] Figure 6 6 is a data flow diagram illustrating an example of a perception environment 600 in which a perception system 402 generates a bounding box 512 from an image 502. In the illustrated example, two general data paths are shown: an object query data path and a feature map data path. Although there is a crossover of data between the two general data paths, for simplicity, the feature map data path (including the image 502, the image feature extractor 504, the feature map 602, the multi-view stage 516, the ROI stage 518, and the enhanced feature map 606) will be described first, followed by the object query data path (including the formulation stage 505, the object query 604, the object query cross-attention stage 513, the object query self-attention stage 514, and the enhanced object query 608).

[0171] As described herein, image 502 can correspond to images received from different cameras (or other image sensors) around the self-vehicle. In the illustrated example, six images are shown, however, it will be understood that fewer or more images can be used. Image 502 can correspond to images taken at the same (or approximately the same) time (e.g., within milliseconds of each other). In this way, the images can correspond to the same scene for the vehicle. In addition, perception system 402 can repeatedly receive images and perform the functions described herein multiple times per second when new images are received. Therefore, it will be understood that perception system 402 can operate in real time or near real time to generate bounding box 512 from image 502.

[0172] As described herein, image feature extractor 504 generates feature maps 602 from images 502. In the illustrated example, image feature extractor 504 generates one feature map from each of images 502, however, it will be appreciated that image feature extractor 504 may generate multiple feature maps 602 from each of images 502 and communicate the multiple feature maps 602 to attention stage 506.

[0173] Each feature map in feature map 602 may include an array of grid cells with a particular channel depth. The grid cells may include semantic data (or features) extracted from (pixels in) the (one or more) images from which the feature map is generated. Features may be organized as vectors or some other tensor shape. For example, the features (or semantic data) of a grid cell may indicate the shape, light, texture, reflectivity, edge, object class, location, etc. of something detected by image feature extractor 504.

[0174] The multi-view stage 516 and the ROI stage 518 enhance the feature map to form the enhanced feature map 606. In some cases, the multi-view stage 516 and the ROI stage 518 may enhance the feature map by modifying features in a grid cell using features from another grid cell (and vice versa). As described herein, the multi-view stage 516 and / or the ROI stage 518 may cross-correlate features in grid cells within different groups (e.g., windows) of the enhanced feature map 606. In some cases, the group of grid cells (e.g., windows) used by the multi-view stage 516 may differ (e.g., in shape, number, or placement) from the group of grid cells (e.g., windows) used by the ROI stage 518.

[0175] In addition, one or more of the groups of grid cells may overlap with multiple feature maps such that a grid cell in one feature map is cross-correlated with a grid cell from another feature map. In some such cases, the multi-view stage 516 and / or the ROI stage 518 may use windows that overlap feature maps 602 corresponding to images from cameras that are close to each other (e.g., one window may overlap a feature map corresponding to a front view image of a vehicle and a feature map corresponding to a left front view of the vehicle). Thus, features from a grid cell in one feature map may propagate to (or be correlated with) features from a grid cell in another (e.g., adjacent, contiguous, or abutting) feature map.

[0176] As described herein, there may be multiple layers of multi-view stages 516 and / or ROI stages 518, such that the enhanced feature map 606 generated by a first layer is communicated to a second layer (as illustrated by dashed line 610A), etc. In some cases, there may be 2, 4, 6, or more layers.

[0177] The parameters or configurations of the multi-view stage 516 and / or ROI stage 518 in different layers may be the same or different. For example, the weights or nodes in the multi-view stage 516 and / or ROI stage 518 of the first layer may be different from the weights or nodes in the multi-view stage 516 and / or ROI stage 518 of the second layer.

[0178] In addition, the group of grid cells (e.g., windows) used to perform self-attention operations on grid cells in one layer may be different from the group of grid cells in another layer. For example, the group of grid cells used by the multi-view stage 516 in the first layer may be different from the group of grid cells used in the second layer. In some cases, the window used by the multi-view stage 516 in the second layer may be shifted in one or more directions relative to the window used by the multi-view stage 516 in the first layer. Subsequent layers may include additional shifts, or oscillate between the placement of windows in the first layer and the placement of windows in the second layer. In some cases, windows in the multi-view stage 516 of different layers may be shifted in a different manner than windows in the ROI stage 518 of different layers. For example, windows in the multi-view stage 516 of different layers may be shifted in one direction (e.g., horizontally), and windows in the ROI stage 518 of different layers may be shifted in multiple directions (e.g., vertically and horizontally). In addition, depending on how the windows are aligned, a portion of the window may include an opposite end or corner of the feature map. For example, a window may include the lower right corner of one feature map and the upper left corner of another (eg, adjacent or neighboring) feature map.

[0179] Referring to the object query data path, the formulation stage 505 generates an object query 604. As described herein, concurrently with the image feature extractor 504 generating the feature map 602, the formulation stage 505 can generate and / or initialize the object query 604. As described herein, when generating the object query 604, the formulation stage 505 can randomly or pseudo-randomly initialize the features of the object query.

[0180] The formulation stage 505 may also generate an object query 604 using features from the feature map 602, other feature maps (e.g., from a localization network), or other data (e.g., heat map data associated with a heat map), etc. In some cases, the formulation stage 505 may include a cross-attention stage to identify grid cells (e.g., of the feature map 602) corresponding to the object query, and use the features of the identified (one or more) grid cells to modify the features of the object query. As described herein, in some cases, this may include: using a linear layer matrix to identify the grid cells, and using weighted features of the grid cells to perform a cross-attention operation on the features of the identified grid cells and the features of the object query. In addition, in some cases, the formulation stage 505 may include a self-attention stage to enable the grid cells to correlate or associate features and update themselves.

[0181] The formulation stage 505 communicates the object query 604 to the attention stage 506 for further processing (e.g., the object query cross-attention stage 513 of the attention stage 506). As described herein, the object query cross-attention stage 513 can use grid cells from the enhanced feature map 606 to enhance the object query 604. In some cases, the object query cross-attention stage 513 enhances the object query 604 in a manner similar to the manner in which the formulation stage 505 uses the feature map 602 to generate the object query 604. For example, the object query cross-attention stage 513 can identify (one or more) grid cells in the enhanced feature map 606 that correspond to a particular object query in the object queries 604, weight the features, and / or use the (weighted) features of the identified (one or more) grid cells to modify or enhance the features of the particular object query. In some cases, the object query cross-attention stage 513 can use similar techniques to enhance some or all of the object queries 604.

[0182] The object query self-attention stage 514 may further enhance the (enhanced) object query 604 received from the ROI stage 518 to provide an enhanced object query 608. In some cases, the object query self-attention stage 514 enhances the object query 604 by modifying or correlating (or performing a self-attention operation) the features of the object query 604, similar to the way the multi-view stage 516 and the ROI stage 518 correlate features of grid cells within a group of grid cells (e.g., a window). For example, the object query self-attention stage 514 may compare the features of the object queries 604 to determine a probabilistic relationship between them, generate weighted values ​​based on the probabilistic relationship, and modify the features of one object query based on the weighted features from other object queries (after being weighted using the weighted values). In some cases, the object query self-attention stage 514 may enhance some or all of the object queries 604 using similar techniques to provide the enhanced object query 608, such that the semantic data from some or all of the object queries 604 is updated or enhanced.

[0183] As described herein, there may be multiple layers for the object query cross-attention stage 513 and the object query self-attention stage 514 in the attention stage 506, such that as illustrated by dashed line 610B, etc., the enhanced object query 608 generated in the first layer of the attention stage 506 is communicated to the second layer (e.g., the second object query cross-attention stage 513 and / or the second object query self-attention stage 514). In some cases, there may be 2, 4, 6, or more layers.

[0184] Once processed by one or more layers of the attention stage 506, the enhanced feature map 606 and the enhanced object query 608 are communicated to the BEV stage 508 for further processing, as described herein. As described herein, the BEV generator 520 can generate a BEV feature map 612 from the enhanced feature map 606. As described herein, as part of the transformation, multiple grid cells from the same or different enhanced feature maps 606 can be mapped to one BEV grid cell of the BEV feature map 612. In this way, the BEV grid cell can include more features or more enhanced features than the grid cells of the enhanced feature map 606.

[0185] The cross-attention stage 522 uses the BEV feature map 612 to enhance the enhanced object query 608 (and / or vice versa) to provide a multi-representation (MR) query 614 (also referred to herein as a further enhanced query). The multi-representation query 614 may include some or all of the enhanced object query 608 and / or some or all of the enhanced BEV query from the BEV feature maps 612, 614. Similar to the way the object query cross-attention stage 513 uses the enhanced feature map 606 to enhance the object query 604, the cross-attention stage 522 may use the BEV feature map 612 to enhance the enhanced object query 608. For example, the cross-attention stage 522 may identify (e.g., using a linear layer matrix) one or more BEV grid cells corresponding to a particular object query, and use the features of the identified (one or more) BEV grid cells to modify the features of the object query. In some such cases, the cross-attention stage 522 may weight features of the identified BEV grid cell(s) and use the weighted features to determine features for the corresponding object query.

[0186] In some cases, the cross-attention stage 522 may treat the grid cells (or other portions) of the BEV feature map 612 as object queries (also referred to herein as BEV object queries). In some such cases, the cross-attention stage 522 may enhance the BEV object queries using features from the enhanced object queries 608. For example, the cross-attention stage 522 may identify (e.g., using a linear layer matrix) one or more enhanced object queries 608 corresponding to a particular BEV grid cell (or BEV object query) and use the features of the identified object queries to modify the features of the BEV object queries. In some such cases, the cross-attention stage 522 may weight the features of the identified object queries and use the weighted features to determine features for the corresponding BEV object queries.

[0187] Thus, the multi-representation query 614 may include different types of object queries or different types of representation queries. For example, the multi-representation query 614 may include an enhanced object query (e.g., corresponding to the enhanced object query 608) from the object query data path, and an enhanced BEV object query (corresponding to one or more grid cells of the BEV feature map 612). In some cases, the enhanced object query 608 (or object query 604) may also be referred to as a floating query (or a first type of representation query) because the locations they determine may change. For example, when the feature values ​​of the object query (or corresponding tensor) change due to cross-attention and / or self-attention (or other modifications), the combination of the modified object query with the (same) linear layer matrix may result in different locations being determined. In addition, since different linear layer matrices may be used to determine the corresponding locations of the object query (e.g., at different layers of the attention stage 506), the combination of different linear layer matrices with the (same or modified) object query may result in different locations being associated with the object query. As such, the location of the object query 604 or the enhanced object query 608 may vary or “float” to different grid cells of the feature map 602 and / or the enhanced feature map 606. In some cases, a BEV object query from the BEV feature map 612 may also be referred to as a fixed query (or a second type of representation query) because the BEV object query may correspond to (or be the same as) a particular grid cell of the BEV feature map 612. As such, the location associated with a particular BEV object query from the BEV feature map 612 may not change, and therefore may be considered “fixed.”

[0188] By enhancing and / or including a BEV object query in the multi-representation query 614 communicated to the detection stage 510, the detection stage 510 may have more data with which it may generate the bounding box 512. The (BEV) object query and / or more enhanced object queries may result in a more accurate bounding box 512.

[0189] The detection stage 510 receives the multi-representation query 614 from the BEV stage 508 and uses the multi-representation query 614 to generate a (3D) bounding box 512 for some or all of the object query. In some cases, the detection stage 510 generates bounding boxes for certain types of object queries (e.g., object queries with object classes of pedestrians, bicycles, vehicles, construction cones, etc.).

[0190] Perception system 402 may use fewer, more, or different components to generate bounding box 512. In some cases, attention stage 506 may include one layer. In some cases, BEV stage 508 may be omitted, and detection stage 510 uses enhanced object query 608 to generate bounding box 512. In some cases, multi-view stage 516 and / or ROI stage 518 may be combined or omitted. In some such cases, feature map 602 may be used to generate BEV feature map 612, and / or object query cross-attention stage 513 may be omitted.

[0191] Figure 7 is a flow chart illustrating an example of a routine 700 implemented by at least one processor to navigate a vehicle based on at least one bounding box generated from one or more images. Figure 7 The illustrated flowchart is provided for illustration purposes only. It will be understood that the Figure 7 One or more of the steps of the illustrated routine, or the order of the steps may be changed. In addition, for purposes of illustrating clear examples, one or more specific system components are described in the context of performing various operations during various data flow stages. However, other system arrangements and other distributions of processing steps across system components and / or autonomous vehicle computing 400 may be used.

[0192] At box 702, perception system 402 receives images of a vehicle scene. As described herein, these images may correspond to different image sensors or cameras located at different locations around the vehicle. Combined, these images may represent a 360-degree view of the environment from the perspective of the vehicle.

[0193] At box 704, the perception system 402 generates a feature map based on the image. As described herein, in some cases, the perception system 402 generates at least one feature map for each received image. In some cases, the feature map includes a location relationship corresponding to the image based on which the feature map is generated. For example, an adjacent or neighboring image can correspond to an adjacent or neighboring feature map. In some cases, the perception system 402 generates the feature map using a feature pyramid network (such as, but not limited to, Resnet or a feature pyramid network (FPN)). The feature map may include an array of grid cells with a specific channel depth (e.g., 256, 512, etc.). As described herein, the grid cell may include features indicating extracted characteristics of the image (such as, but not limited to, color, texture, location, reflectivity, shape, edges, etc.).

[0194] At box 706, the perception system 402 determines a first plurality of windows for the feature map. As described herein, the first plurality of windows can have a specific size and shape and can cover the entire set of feature maps. In some cases, there can be multiple rows and / or columns of windows, where the rows are aligned or offset from each other. In some cases, one or more windows can cross the boundary of the feature map, so that the grid cells from the first feature map and the grid cells from the different feature maps are included in (one or more than one) corresponding windows.

[0195] At box 708, the perception system 402 enhances the semantic data set associated with the feature map based on the first plurality of windows to provide a first enhanced semantic data set. As described herein, for some or all of the first plurality of windows, the perception system 402 can (e.g., based on the placement of the grid cells within the corresponding windows) inter-correlate or correlate features (semantic data) of grid cells within the corresponding windows. For windows that include grid cells that span multiple feature maps, this can include: using semantic data (e.g., features) from grid cells in the first feature map to modify or enhance semantic data (e.g., features) of grid cells in the second feature map. Thus, semantic data from a first grid cell within a first window in the first plurality of windows can be used to modify, enhance, or determine semantic data of a second grid cell (in the same or different feature map) also within the first window in the first plurality of windows, and vice versa.

[0196] When correlating or correlating the features, the perception system 402 may weight the features of different grid cells relative to each other (e.g., based on a probabilistic relationship between the grid cells based on a comparison of the features of the grid cells) and use the weighted features to determine a new or modified value for the feature of a particular grid cell. In some cases, when determining weights between grid cells, the perception system 402 may generate a matrix that includes the grid cells in rows and columns and weight values ​​indicating the relationship between two grid cells (and the weight to be applied to a feature from one grid cell to another grid cell).

[0197] At box 710, the perception system 402 determines a second plurality of windows for the feature map. As described herein, the second plurality of windows can have a specific size and shape and can cover the entire feature map set. In some cases, there can be multiple rows and / or columns of windows, where the rows are aligned or offset from each other. In some cases, a window in the second plurality of windows can cross the boundary of the feature map so that grid cells from two different feature maps are included in the window.

[0198] As a non-limiting example, because the second plurality of windows is different from the first plurality of windows, in some cases, the first grid cell may be placed in the first window of the first plurality of windows, while the second grid cell (from the same or different feature map) and the third grid cell (from the same or different feature map as the first grid cell or the second grid cell) may be placed in the second window of the first plurality of windows, in which case the semantic data of the first grid cell may be used to determine or update the semantic data of the second grid cell, and vice versa. The first grid cell may then be placed in a third window of the second plurality of windows together with the third grid cell, and the second grid cell may be placed in a fourth window of the second plurality of windows together with the fourth grid cell that may have been in the first window or the second window of the first plurality of windows.

[0199] As described herein, the second plurality of windows may correspond to another stage of the same layer of the attention stage 506 and / or may correspond to the same stage at a different layer of the attention stage 506. For example, if the first plurality of windows corresponds to the multi-view stage 516 of the first layer of the attention stage 506, the second plurality of windows may correspond to the ROI stage 518 of the first layer of the attention stage 506 or the multi-view stage 516 of the second layer of the attention stage 506.

[0200] If the second plurality of windows corresponds to another stage in the same layer, the second plurality of windows may be a different size and / or shape than the first plurality of windows. Additionally, the second plurality of windows may be located at a different position than the first plurality of windows. For example, if the first plurality of windows corresponds to the multi-view stage 516 in the first layer and the second plurality of windows corresponds to the ROI stage 518 in the first layer, the second plurality of windows may include fewer or more windows, fewer or more columns or rows, and / or the size of the windows in the second plurality of windows may be different than the first plurality of windows.

[0201] If the second plurality of windows corresponds to the same stage of a different layer, the second plurality of windows may be the same size and shape as the first plurality of windows (and include the same number of windows, the same number of columns, and / or the same number of rows), but in a different (shifted) location. For example, if the first plurality of windows corresponds to the multi-view stage 516 in the first layer and the second plurality of windows corresponds to the multi-view stage 516 in the second layer, the second plurality of windows may be horizontally (or vertically, as the case may be) shifted relative to the first plurality of windows. As another example, if the first plurality of windows corresponds to the ROI stage 518 in the first layer and the second plurality of windows corresponds to the ROI stage 518 in the second layer, the second plurality of windows may be vertically shifted relative to the first plurality of windows.

[0202] At block 712, the perception system 402 enhances the first enhanced semantic data set associated with the feature map based on the second plurality of windows to provide a second enhanced semantic data set. Similar to the first plurality of windows as described herein at least with reference to block 708, for some or all of the second plurality of windows, the perception system 402 may intercorrelate or correlate (also referred to herein as self-attention) features of grid cells within the respective windows. Because the second plurality of windows is different from the first plurality of windows, enhancements to one grid cell may propagate to additional grid cells.

[0203] Continuing to refer to the above examples of the first grid cell, the second grid cell, the third grid cell, and the fourth grid cell, it will be understood that by placing the first grid cell and the third grid cell in a third window in the second plurality of windows, the semantic data of the first grid cell can be used to modify, enhance, and / or determine the semantic data of the third grid cell, and vice versa. Similarly, by placing the second grid cell and the fourth grid cell in a fourth window in the second plurality of windows, the semantic data of the second grid cell can be used to modify, enhance, and / or determine the semantic data of the fourth grid cell, and vice versa.

[0204] At block 714, the perception system 402 generates at least one bounding box based on the second enhanced semantic data set. As described herein, the perception system 402 can generate a bounding box based on the second enhanced semantic data set in various ways. For example, the perception system 402 can generate a bird's eye view (BEV) feature map and / or an enhanced object query using, at least in part, the second enhanced semantic data set (e.g., after one or more additional rounds of enhancement of the second enhanced semantic data set). The BEV feature map and / or the enhanced object query can be used to determine at least one bounding box. In some cases, the BEV feature map can be used to further enhance the enhanced object query (and / or vice versa), and the further enhanced object query and / or the BEV feature map enhanced by the enhanced object query can be used to generate at least one bounding box.

[0205] As described herein, the perception system 402 may generate a bird's eye view (BEV) feature map using, at least in part, the second enhanced semantic data set. For example, the perception system 402 may further enhance the second enhanced semantic data set. As part of generating the BEV feature map, the perception system 402 may assign multiple grid cells from the second enhanced feature map (or further enhanced grid cells corresponding to the second enhanced feature map after passing through additional enhancement stages or layers of the perception system 402) to specific grid cells of the BEV feature map. In some cases, one or more of the grid cells of the BEV feature map may be considered or deemed to be an object query or a BEV object query. In some such cases, the BEV object query may be used by the perception system 402 to generate a bounding box.

[0206] As described herein, the perception system 402 can use a second enhanced semantic data set to enhance the object query. For example, using the features of the object query, the perception system 402 can (e.g., using a linear layer matrix) identify grid cells corresponding to the object query from the enhanced feature map. The identified grid cells can be used to enhance or modify the features of the object query. In addition, the perception system 402 can perform a self-attention function on the object query it has generated to make the features of the object query interrelated and / or intercorrelated. Similar to the enhancement of the feature map, the perception system 402 can use multiple layers of enhanced feature maps to (further) enhance the object query. The perception system 402 can use the resulting enhanced object query to generate a bounding box.

[0207] As described herein, in some cases, the perception system 402 may use the second enhanced semantic data set to enhance the object query and generate a BEV feature map. The enhanced object query and the BEV feature map may be used alone or in combination to generate a bounding box. In some cases, the enhanced object query may be used to enhance the BEV feature map, and vice versa. The perception system 402 may use the resulting (further) enhanced object query and the enhanced feature map (which may include the BEV object query) to generate a bounding box.

[0208] At block 716, perception system 402 causes navigation of the vehicle based on the at least one bounding box. In some cases, perception system 402 may communicate the bounding box to planning system 404. Planning system 404 may use the bounding box to determine how to navigate the vehicle scene.

[0209] Fewer, more, or different steps may be included in routine 700. For example, reference is made to the use of a window and correlating features of grid cells within the window, however, it will be understood that other grid cell groups may be used, such as but not limited to those described herein at least with reference to Figure 5DThe outline of the object (or an approximation of the outline) corresponds to the grid cells, etc.

[0210] In some cases, the perception system 402 may enhance the first enhanced semantic data using a plurality of points or grid cells of interest (rather than using a second plurality of windows). In some such cases, block 710 may be replaced with determining a group of points of interest or a group of grid cells, and block 712 may be replaced with enhancing the first enhanced semantic data based on the group of points of interest or the group of grid cells.

[0211] Similarly, groups of grid cells or interest points may be identified and used to enhance semantic data associated with the feature map (e.g., replacing blocks 706 and 708). Figure 5D As described herein, in some cases, the interest points can be used to identify objects in an image or feature map. In some such cases, a self-attention operation can be performed on a group of interest points corresponding to the same object (rather than grid cells in the same window) to enhance the first enhanced semantic data. As described herein, the group of interest points can form an outline or rough outline of an object, in which case the first enhanced semantic data can be enhanced using interest points corresponding to the outer edge of the object (rather than points corresponding to a region of the object or a region of the window).

[0212] In addition, the order of the steps can be changed and / or some steps can be repeated. For example, if boxes 710 and 712 correspond to different stages of the same layer of the attention stage 506, the perception system 402 can repeat boxes 706 to 712 corresponding to different layers of the attention stage 506. In some such cases, the third (fifth, seventh, ninth, eleventh, etc.) plurality of windows can have the same size, shape, and number as the first plurality of windows, and can have the same or different positions as the first plurality of windows. For example, the third, seventh, eleventh, etc. plurality of windows can have different positions (e.g., vertically and / or horizontally shifted) from the first plurality of windows, and the fifth, ninth, etc. plurality of windows can have the same (or different) positions as the first plurality of windows.

[0213] For multiple windows having the same size, shape, number, and position as the first multiple windows, the grid units of the first multiple windows can further enhance each other. For example, referring to the first grid unit, the second grid unit, the third grid unit, and the fourth grid unit example, if the subsequent multiple windows have the same size, shape, number, and position as the first multiple windows, the semantic data of the first grid unit and the second grid unit can be used to determine the semantic data for each other. Similarly, if the subsequent multiple windows have the same size, shape, number, and position as the second multiple windows, the semantic data of the first grid unit and the third grid unit and the semantic data of the second grid unit and the fourth grid unit can be used respectively to determine the semantic data for each other.

[0214] Similarly, the fourth (sixth, eighth, tenth, etc.) plurality of windows may have the same size, shape, and number as the second plurality of windows, and may have the same or different positions as the second plurality of windows. For example, the fourth, eighth, etc. plurality of windows may have different positions (e.g., vertically and / or horizontally shifted) than the second plurality of windows, and the sixth, tenth, etc. plurality of windows may have the same (or different) positions as the second plurality of windows.

[0215] As another non-limiting example, if blocks 710 and 712 correspond to the same stage at different layers of attention stage 506, perception system 402 may repeat blocks 706 to 708 and / or blocks 710 to 712 corresponding to different layers of attention stage 506. In some such cases, the additional multiple windows may have the same size, shape, and number as the first multiple windows and the second multiple windows, and may have the same (or different) position as the first multiple windows or the second multiple windows. For example, the third, fifth, seventh, ninth, etc. multiple windows may have the same (or different) position as the first multiple windows (e.g., vertically and / or horizontally shifted), and the fourth, sixth, eighth, tenth, etc. multiple windows may have the same (or different) position as the second multiple windows.

[0216] Figure 8 is a flow chart illustrating an example of a routine 800 implemented by at least one processor to navigate a vehicle based on at least one bounding box generated from one or more images. Figure 8 The illustrated flowchart is provided for illustration purposes only. It will be understood that the Figure 8 One or more of the steps of the illustrated routine, or the order of the steps may be changed. In addition, for the purpose of illustrating clear examples, one or more specific system components are described in the context of performing various operations during various data flow stages. However, other system arrangements and other distributions of processing steps across system components and / or autonomous vehicle computing 400 may be used.

[0217] At box 802, perception system 402 receives images of a vehicle scene. As described herein, these images may correspond to different image sensors or cameras located at different locations around the vehicle. Combined, these images may represent a 360-degree view of the environment from the perspective of the vehicle.

[0218] At block 804, as at least referenced herein Figure 7 As described in block 704 of , the perception system 402 generates a feature map based on the image.

[0219] At block 806, the perception system 402 generates an object query based on the feature map. As described herein, the perception system 402 may use the features of the particular object query to identify one or more grid cells from the feature map corresponding to the particular object query. The perception system 402 may use the identified grid cells to modify the features of the particular object query. In some cases, the perception system 402 may weight the features of the identified one or more grid cells and use the weighted features to modify the features of the particular object query. In a similar manner, the perception system 402 may modify some or all of the features in the object query. In some cases, as part of generating the object query, the perception system 402 may use different feature maps (e.g., generated from a localization network different from the image feature extractor 504) and / or different data (e.g., heat map data associated with a heat map). In addition, in some cases, as part of the generation process, the perception system 402 may perform a cross-attention operation on the object query (e.g., using the feature map generated at block 804) or a self-attention operation on the object query.

[0220] At box 808, the perception system 402 enhances the object query. As described herein, the perception system 402 can enhance the object query in various ways. In some cases, the perception system 402 can enhance the object query based on any one or any combination of the following: (enhanced) feature maps (described herein at least with reference to the non-limiting example of the object query cross-attention stage 513), features from other object queries (described herein at least with reference to the non-limiting example of the object query self-attention stage 514), and / or one or more BEV feature maps (described herein at least with reference to the non-limiting example of the cross-attention stage 522). In some cases, the perception system 402 can enhance the object query based on features from other object queries and / or BEV feature maps, rather than based on an enhanced feature map (e.g., enhanced feature map 606). In some such cases, the BEV feature map can be generated from a feature map generated by an image feature network (e.g., image feature extractor 504).

[0221] As described herein, in order to enhance an object query based on an (enhanced) feature map, the perception system 402 may identify (one or more than one) grid cells from the (enhanced) feature map corresponding to different object queries, and use features of the identified grid cells to enhance or modify features of the corresponding object query.

[0222] As described herein, to enhance an object query based on features of other object queries, the perception system 402 may perform a self-attention function to determine a relationship between object queries (e.g., by comparing features of different object queries). Based on the determined relationship, the perception system 402 may weight features from the object queries relative to each other, and use a portion or all of the weighted features from the object queries to modify features of a particular object query. For example, to enhance a first object query, the perception system 402 may compare features of the first object query with features of a set of at least one second object query, and determine a relationship based on the comparison. The perception system 402 may further determine a weighted value to be applied to features from a set of at least one second object query to generate a weighted feature from the set of at least one second object query. The perception system 402 may then use the weighted values ​​from the set of at least one object query to determine or modify one or more features of the first object query. For example, as described herein, the perception system 402 may multiply a weighted value associated with the second object query by a particular feature of the second object query, and use the weighted feature to determine or modify a corresponding feature of the first object query.

[0223] As described herein, to enhance an object query based on a BEV feature map, perception system 402 may generate a BEV feature map from one or more (enhanced) feature maps (e.g., feature map 602 and / or enhanced feature map 606). Perception system 402 may use features of the object query to identify one or more BEV grid cells (or BEV object queries) in the BEV feature map that correspond to the object query, and use the features of the BEV grid cells (or BEV object queries) to modify features of the corresponding object query. In some cases, features from the BEV grid cells / BEV object queries may be weighted based on a determined relationship between the BEV grid cells / BEV object queries and the corresponding object query being modified.

[0224] In some cases, the perception system 402 may enhance the object query based on the (enhanced) feature map, features from other object queries, and one or more BEV feature maps. Figure 6 Non-limiting examples of enhancing object queries based on (enhanced) feature maps, features from other object queries, and one or more BEV feature maps are illustrated.

[0225] At block 810, the perception system 402 generates at least one bounding box based on the enhanced object query. As described herein, the perception system 402 may use one or more encoders to identify bounding boxes for objects in the image based on the enhanced object query. In some cases, the perception system 402 may use the enhanced object query and a BEV object query corresponding to one or more BEV grid cells of the BEV feature map. In some cases, more object queries used to generate the bounding box may result in improved accuracy of the bounding box.

[0226] At block 812, as at least referenced herein Figure 7 As described in block 716 of , perception system 402 enables navigation of the vehicle based on at least one bounding box.

[0227] Fewer, more, or different blocks may be used in routine 800. In some cases, any one or any combination of blocks from routine 700 may be combined with blocks from routine 800, or vice versa.

[0228] In some cases, routine 800 may include: using object query to enhance the BEV feature map, and / or generating at least one bounding box based on the (enhanced) BEV feature map. In some cases, these functions may be included in routine 800 and / or replace blocks 808 and 810 of routine 800.

[0229] Example

[0230] Various example embodiments of the present disclosure can be described by the following clauses.

[0231] Clause 1. A method comprising: receiving a plurality of images from a plurality of image sensors, the plurality of images corresponding to a plurality of views of a scene of a vehicle; generating a plurality of feature maps based on the plurality of images; determining a first plurality of windows for the plurality of feature maps, wherein a first window in the first plurality of windows includes a first grid cell from a first feature map in the plurality of feature maps and a second grid cell from a second feature map in the plurality of feature maps; enhancing a set of semantic data associated with the plurality of feature maps based on the first plurality of windows to provide a first enhanced set of semantic data, wherein enhancing the set of semantic data associated with the plurality of feature maps based on the first plurality of windows includes: determining first semantic data for the first grid cell using second semantic data associated with the second grid cell based on the first grid cell and the second grid cell being included in the first window; determining first semantic data for the first grid cell based on the first grid cell and the second grid cell being included in the first window; Figure determines a second plurality of windows, wherein a third window in the second plurality of windows includes the first grid cell and the third grid cell, and a fourth window in the second plurality of windows includes the second grid cell; based on the second plurality of windows, a first enhanced semantic data set associated with the plurality of feature maps is enhanced to provide a second enhanced semantic data set associated with the plurality of feature maps, wherein enhancing the first enhanced semantic data set associated with the plurality of feature maps based on the second plurality of windows includes: based on the first grid cell and the third grid cell being included in the third window, using fourth semantic data associated with the third grid cell to determine third semantic data for the first grid cell; based on the second enhanced semantic data set, at least one bounding box is generated for an object in a scene of the vehicle; and the vehicle is controlled based on the at least one bounding box.

[0232] Clause 2. A method according to clause 1, wherein generating at least one bounding box for an object in the scene of the vehicle based on the second enhanced semantic data set includes: enhancing an object query set based on the second enhanced semantic data set to provide an enhanced object query set; and generating the at least one bounding box based on the enhanced object query set.

[0233] Clause 3. A method according to clause 1 or 2, wherein generating at least one bounding box for an object in the scene of the vehicle based on the second enhanced semantic data set includes: generating a bird's-eye view feature map based on the second enhanced semantic data set; and generating the at least one bounding box based on the bird's-eye view feature map.

[0234] Clause 4. A method according to any one of clauses 1 to 3, wherein generating at least one bounding box for an object in the scene of the vehicle based on the second enhanced semantic data set includes: enhancing an object query set based on the second enhanced semantic data set to provide a first enhanced object query set; generating a bird's-eye view feature map based on the second enhanced semantic data set; enhancing the first enhanced object query set based on the bird's-eye view feature map to provide a second enhanced object query set; and generating the at least one bounding box based on the second enhanced object query set.

[0235] Clause 5. The method of any one of clauses 1 to 4, wherein the plurality of image sensors are placed around the vehicle in different orientations.

[0236] Clause 6. The method of any one of clauses 1 to 5, wherein the plurality of images provide a 360 degree view around the vehicle.

[0237] Clause 7. The method of any one of clauses 1 to 6, wherein the first plurality of windows forms at least one row of windows, and each window in the first plurality of windows is the same width and height.

[0238] Clause 8. The method of any one of clauses 1 to 7, wherein the first plurality of windows forms a plurality of rows of windows.

[0239] Clause 9. A method according to any one of clauses 1 to 8, wherein, with respect to the plurality of feature maps, the second plurality of windows are horizontally shifted relative to the first plurality of windows.

[0240] Clause 10. A method according to any one of clauses 1 to 8, wherein, with respect to the plurality of feature maps, the second plurality of windows are vertically and horizontally shifted relative to the first plurality of windows.

[0241] Clause 11. The method of any one of Clauses 1 to 8, wherein the second plurality of windows are a different size than the first plurality of windows.

[0242] Item 12. A system comprising: a data storage unit storing computer executable instructions; and a processor configured to execute the computer executable instructions, wherein execution of the computer executable instructions causes the system to: receive a plurality of images from a plurality of image sensors, the plurality of images corresponding to a plurality of views of a scene of a vehicle; generate a plurality of feature maps based on the plurality of images; determine a first plurality of windows for the plurality of feature maps, wherein a first window in the first plurality of windows includes a first grid cell from a first feature map in the plurality of feature maps and a second grid cell from a second feature map in the plurality of feature maps; enhance a semantic data set associated with the plurality of feature maps based on the first plurality of windows to provide a first enhanced semantic data set, wherein enhancing the semantic data set associated with the plurality of feature maps based on the first plurality of windows includes: based on the first grid cell and the second grid cell being included in the first window, using a second grid cell associated with the second grid cell The method comprises the steps of: determining first semantic data for the first grid cell based on semantic data of the first grid cell; determining a second plurality of windows for the plurality of feature maps, wherein a third window in the second plurality of windows includes the first grid cell and the third grid cell, and a fourth window in the second plurality of windows includes the second grid cell; enhancing a first enhanced semantic data set associated with the plurality of feature maps based on the second plurality of windows to provide a second enhanced semantic data set associated with the plurality of feature maps, wherein enhancing the first enhanced semantic data set associated with the plurality of feature maps based on the second plurality of windows comprises: determining third semantic data for the first grid cell using fourth semantic data associated with the third grid cell based on the first grid cell and the third grid cell being included in the third window; generating at least one bounding box for an object in a scene of the vehicle based on the second enhanced semantic data set; and controlling the vehicle based on the at least one bounding box.

[0243] Clause 13. A system according to clause 12, wherein, in order to generate at least one bounding box for an object in the scene of the vehicle based on the second enhanced semantic data set, the processor is configured to: enhance an object query set based on the second enhanced semantic data set to provide an enhanced object query set; and generate the at least one bounding box based on the enhanced object query set.

[0244] Clause 14. A system according to clause 12 or 13, wherein, in order to generate at least one bounding box for an object in the scene of the vehicle based on the second enhanced semantic data set, the processor is configured to: generate a bird's-eye view feature map based on the second enhanced semantic data set; and generate the at least one bounding box based on the bird's-eye view feature map.

[0245] Clause 15. A system according to any one of clauses 12 to 14, wherein, in order to generate at least one bounding box for an object in the scene of the vehicle based on the second enhanced semantic data set, the processor is configured to: enhance an object query set based on the second enhanced semantic data set to provide a first enhanced object query set; generate a bird's-eye view feature map based on the second enhanced semantic data set; enhance the first enhanced object query set based on the bird's-eye view feature map to provide a second enhanced object query set; and generate the at least one bounding box based on the second enhanced object query set.

[0246] Clause 16. The system of any of Clauses 12 to 15, wherein the second plurality of windows are a different size than the first plurality of windows.

[0247] Item 17. A non-transitory computer-readable medium comprising computer-executable instructions which, when executed by a computing system, cause the computing system to: receive a plurality of images from a plurality of image sensors, the plurality of images corresponding to a plurality of views of a scene of a vehicle; generate a plurality of feature maps based on the plurality of images; determine a first plurality of windows for the plurality of feature maps, wherein a first window in the first plurality of windows comprises a first grid cell from a first feature map of the plurality of feature maps and a second grid cell from a second feature map of the plurality of feature maps; enhance a set of semantic data associated with the plurality of feature maps based on the first plurality of windows to provide a first enhanced set of semantic data, wherein enhancing the set of semantic data associated with the plurality of feature maps based on the first plurality of windows comprises: determining a set of semantic data for the plurality of feature maps using a second semantic data associated with the second grid cell based on the first grid cell and the second grid cell being included in the first window The method comprises: providing a first semantic data of the first grid cell; determining a second plurality of windows for the plurality of feature maps, wherein a third window in the second plurality of windows includes the first grid cell and the third grid cell, and a fourth window in the second plurality of windows includes the second grid cell; enhancing a first enhanced semantic data set associated with the plurality of feature maps based on the second plurality of windows to provide a second enhanced semantic data set associated with the plurality of feature maps, wherein enhancing the first enhanced semantic data set associated with the plurality of feature maps based on the second plurality of windows comprises: determining third semantic data for the first grid cell using fourth semantic data associated with the third grid cell based on the first grid cell and the third grid cell being included in the third window; generating at least one bounding box for an object in a scene of the vehicle based on the second enhanced semantic data set; and controlling the vehicle based on the at least one bounding box.

[0248] Clause 18. A non-transitory computer-readable medium according to clause 17, wherein, in order to generate at least one bounding box for an object in the scene of the vehicle based on the second enhanced semantic data set, execution of the computer executable instructions further causes the computing system to: enhance an object query set based on the second enhanced semantic data set to provide an enhanced object query set; and generate the at least one bounding box based on the enhanced object query set.

[0249] Clause 19. A non-transitory computer-readable medium according to clause 17 or 18, wherein, in order to generate at least one bounding box for an object in the scene of the vehicle based on the second enhanced semantic data set, execution of the computer executable instructions further causes the computing system to: generate a bird's-eye view feature map based on the second enhanced semantic data set; and generate the at least one bounding box based on the bird's-eye view feature map.

[0250] Clause 20. A non-transitory computer-readable medium according to any one of clauses 17 to 19, wherein, in order to generate at least one bounding box for an object in the scene of the vehicle based on the second enhanced semantic data set, execution of the computer-executable instructions further causes the computing system to: enhance an object query set based on the second enhanced semantic data set to provide a first enhanced object query set; generate a bird's-eye view feature map based on the second enhanced semantic data set; enhance the first enhanced object query set based on the bird's-eye view feature map to provide a second enhanced object query set; and generate the at least one bounding box based on the second enhanced object query set.

[0251] Item 21. A method comprising: receiving a plurality of images from a plurality of image sensors, the plurality of images corresponding to a plurality of views of a scene of a vehicle; generating a plurality of feature maps based on the plurality of images; generating a plurality of object queries based on the plurality of feature maps; enhancing the plurality of object queries using a bird's-eye view feature map to provide a plurality of enhanced object queries; generating at least one bounding box for an object in the scene of the vehicle based on the plurality of enhanced object queries; and controlling the vehicle based on the at least one bounding box.

[0252] Clause 22. A method according to Clause 21, wherein using a bird's-eye view feature map to enhance the multiple object queries to provide a plurality of enhanced object queries comprises: generating the bird's-eye view feature map based on the multiple feature maps; and enhancing the multiple object queries based on the bird's-eye view feature map to provide the multiple enhanced object queries.

[0253] Clause 23. A method according to Clause 22, wherein enhancing the multiple object queries based on the bird's-eye view feature map to provide the multiple enhanced object queries includes: using a linear layer matrix to identify at least one grid cell in the bird's-eye view feature map corresponding to a first object query in the multiple object queries; and modifying at least one feature of the first object query based on at least one feature of the at least one grid cell.

[0254] Clause 24. A method according to any one of clauses 21 to 23, wherein using a bird's-eye view feature map to enhance the multiple object queries to provide multiple enhanced object queries includes: determining relationships between the multiple object queries based on comparison of features of the multiple object queries; generating weighted values ​​for the multiple object queries relative to each other based on the determined relationships between the multiple object queries; and modifying at least one feature of each object query in the multiple object queries based on the weighted values.

[0255] Clause 25. A method according to any one of clauses 21 to 23, wherein using a bird's-eye view feature map to enhance the multiple object queries to provide multiple enhanced object queries includes: determining a relationship between a first object query in the multiple object queries and a second object query in the multiple object queries based on a comparison of features of the first object query in the multiple object queries relative to features of the second object query in the multiple object queries; generating a weighted value for the first object query relative to the second object query based on the determined relationship between the first object query and the second object query; weighting at least one feature of the first object query based on the weighted value to provide at least one weighted feature of the first object query; and modifying at least one feature of the second object query based on the at least one weighted feature of the first object query.

[0256] Clause 26. A method according to any one of clauses 21 to 25, wherein using a bird's-eye view feature map to enhance the multiple object queries to provide multiple enhanced object queries comprises: generating multiple enhanced feature maps based on a semantic data set of the multiple feature maps; and enhancing the multiple object queries based on the multiple enhanced feature maps.

[0257] Clause 27. A method according to any one of Clauses 21, 24 and 25, wherein the multiple enhanced object queries are multiple second enhanced object queries, and wherein using a bird's-eye view feature map to enhance the multiple object queries to provide multiple second enhanced object queries includes: generating multiple enhanced feature maps based on a semantic data set of the multiple feature maps; enhancing the multiple object queries based on the multiple enhanced feature maps to provide multiple first enhanced object queries; generating the bird's-eye view feature map based on the multiple enhanced feature maps; and enhancing the multiple first enhanced object queries based on the bird's-eye view feature map to provide the multiple second enhanced object queries.

[0258] Clause 28. The method of any one of clauses 21 to 27, wherein the plurality of image sensors are placed around the vehicle in different orientations.

[0259] Clause 29. The method of any one of clauses 21 to 28, wherein the plurality of images provide a 360 degree view around the vehicle.

[0260] Item 30. A system comprising: a data storage unit storing computer executable instructions; and a processor configured to execute the computer executable instructions, wherein execution of the computer executable instructions causes the system to: receive multiple images from multiple image sensors, the multiple images corresponding to multiple views of a scene of a vehicle; generate multiple feature maps based on the multiple images; generate multiple object queries based on the multiple feature maps; enhance the multiple object queries using a bird's-eye view feature map to provide multiple enhanced object queries; generate at least one bounding box for objects in the scene of the vehicle based on the multiple enhanced object queries; and control the vehicle based on the at least one bounding box.

[0261] Clause 31. A system according to Clause 30, wherein, in order to enhance the multiple object queries using a bird's-eye view feature map to provide a plurality of enhanced object queries, the processor is configured to: generate the bird's-eye view feature map based on the multiple feature maps; and enhance the multiple object queries based on the bird's-eye view feature map to provide the multiple enhanced object queries.

[0262] Clause 32. A system according to clause 31, wherein, in order to enhance the multiple object queries based on the bird's-eye view feature map to provide the multiple enhanced object queries, the processor is configured to: use a linear layer matrix to identify at least one grid cell in the bird's-eye view feature map corresponding to a first object query among the multiple object queries; and modify at least one feature of the first object query based on at least one feature of the at least one grid cell.

[0263] Clause 33. A system according to any one of clauses 30 to 32, wherein, in order to enhance the multiple object queries using a bird's-eye view feature map to provide a plurality of enhanced object queries, the processor is configured to: determine relationships between the multiple object queries based on a comparison of features of the multiple object queries; generate weighted values ​​for the multiple object queries relative to each other based on the determined relationships between the multiple object queries; and modify at least one feature of each object query in the multiple object queries based on the weighted values.

[0264] Clause 34. A system according to any one of clauses 30 to 32, wherein, in order to enhance the multiple object queries using a bird's-eye view feature map to provide multiple enhanced object queries, the processor is configured to: determine a relationship between a first object query in the multiple object queries and a second object query in the multiple object queries based on a comparison of features of the first object query relative to features of the second object query in the multiple object queries; generate a weighted value for the first object query relative to the second object query based on the determined relationship between the first object query and the second object query; weight at least one feature of the first object query based on the weighted value to provide at least one weighted feature of the first object query; and modify at least one feature of the second object query based on the at least one weighted feature of the first object query.

[0265] Clause 35. A system according to any one of clauses 30 to 34, wherein using a bird's-eye view feature map to enhance the multiple object queries to provide multiple enhanced object queries comprises: generating multiple enhanced feature maps based on a semantic data set of the multiple feature maps; and enhancing the multiple object queries based on the multiple enhanced feature maps.

[0266] Clause 36. A system according to any one of clauses 30, 33 and 34, wherein the multiple enhanced object queries are multiple second enhanced object queries, and wherein, in order to enhance the multiple object queries using a bird's-eye view feature map to provide multiple enhanced object queries, the processor is configured to: generate a multiple enhanced feature map based on a semantic data set of the multiple feature maps; enhance the multiple object queries based on the multiple enhanced feature maps to provide a multiple first enhanced object queries; generate the bird's-eye view feature map based on the multiple enhanced feature maps; and enhance the multiple first enhanced object queries based on the bird's-eye view feature map to provide the multiple second enhanced object queries.

[0267] Item 37. A non-transitory computer-readable medium comprising computer-executable instructions which, when executed by a computing system, cause the computing system to: receive a plurality of images from a plurality of image sensors, the plurality of images corresponding to a plurality of views of a scene of a vehicle; generate a plurality of feature maps based on the plurality of images; generate a plurality of object queries based on the plurality of feature maps; enhance the plurality of object queries using a bird's-eye view feature map to provide a plurality of enhanced object queries; generate at least one bounding box for an object in the scene of the vehicle based on the plurality of enhanced object queries; and cause the vehicle to be controlled based on the at least one bounding box.

[0268] Clause 38. A non-transitory computer-readable medium according to Clause 37, wherein, in order to enhance the multiple object queries using the bird's-eye view feature map to provide multiple enhanced object queries, execution of the computer-executable instructions also causes the computing system to: generate the bird's-eye view feature map based on the multiple feature maps; and enhance the multiple object queries based on the bird's-eye view feature map to provide the multiple enhanced object queries.

[0269] Clause 39. A non-transitory computer-readable medium according to clause 38, wherein, in order to enhance the multiple object queries based on the bird's-eye view feature map to provide the multiple enhanced object queries, execution of the computer-executable instructions also causes the computing system to: use a linear layer matrix to identify at least one grid cell in the bird's-eye view feature map corresponding to a first object query in the multiple object queries; and modify at least one feature of the first object query based on at least one feature of the at least one grid cell.

[0270] Clause 40. A non-transitory computer-readable medium according to any one of clauses 37 to 39, wherein, in order to enhance the multiple object queries using a bird's-eye view feature map to provide a plurality of enhanced object queries, execution of the computer-executable instructions also causes the computing system to: determine relationships between the multiple object queries based on a comparison of features of the multiple object queries; generate weighted values ​​for the multiple object queries relative to each other based on the determined relationships between the multiple object queries; and modify at least one feature of each of the multiple object queries based on the weighted values.

[0271] Additional Examples

[0272] All methods and tasks described herein can be performed and fully automated by a computer system. The computer system may include a plurality of different computers or computing devices (e.g., physical servers, workstations, storage arrays, cloud computing resources, etc.) in some cases, which communicate and interoperate through a network to perform the described functions. Each such computing device typically includes a processor (or multiple processors) that executes program instructions or modules stored in a memory or other non-transient computer-readable storage medium or device (e.g., a solid-state storage device, a disk drive, etc.). The various functions disclosed herein may be embodied in such program instructions, or may be implemented in a dedicated circuit (e.g., ASIC or FPGA) of a computer system. In the case where a computer system includes a plurality of computing devices, these devices may be, but need not be, located in the same location. The results of the disclosed methods and tasks may be persistently stored by transforming a physical storage device such as a solid-state memory chip or a disk into different states. In some embodiments, the computer system may be a cloud-based computing system whose processing resources are shared by a plurality of different business entities or other users.

[0273] The processing described herein or illustrated in the drawings of the present disclosure may be initiated in response to an event (such as on a predetermined or dynamically determined schedule, on demand when initiated by a user or system administrator, or in response to some other event). When such processing is initiated, a set of executable program instructions stored on one or more non-transitory computer-readable media (e.g., hard drive, flash memory, removable media, etc.) may be loaded into a memory (e.g., RAM) of a server or other computing device. These executable instructions may then be executed by a hardware-based computer processor of the computing device. In some embodiments, such processing or a portion thereof may be implemented serially or in parallel on multiple computing devices and / or multiple processors.

[0274] According to the present embodiment, certain actions, events or functions of any process or algorithm described herein may be performed in a different sequence, may be added, combined or omitted entirely (e.g., not all described operations or events are necessary for the practice of the algorithm). In addition, in some embodiments, operations or events may be performed simultaneously, for example, through multithreading, interrupt processing, or multiple processors or processor cores or on other parallel architectures, rather than sequentially.

[0275] The various illustrative logic blocks, modules, routines, and algorithm steps described in conjunction with the embodiments disclosed herein may be implemented as electronic hardware (e.g., ASIC or FPGA devices), computer software running on computer hardware, or a combination of the two. In addition, the various illustrative logic blocks and modules described in conjunction with the embodiments disclosed herein may be implemented or performed by: a machine, such as a processor device, a digital signal processor ("DSP"), an application-specific integrated circuit ("ASIC"), a field programmable gate array ("FPGA"), or other programmable logic device; discrete gate or transistor logic; discrete hardware components; or any combination thereof designed to perform the functions described herein. The processor device may be a microprocessor, but in an alternative, the processor device may be a controller, a microcontroller, or a state machine, or a combination thereof, etc. The processor device may include a circuit configured to process computer executable instructions. In another embodiment, the processor device includes an FPGA or other programmable device that performs logical operations without processing computer executable instructions. The processor device may also be implemented as a combination of computing devices (e.g., a combination of a DSP and a microprocessor), multiple microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration. Although the present invention is primarily described with respect to digital techniques, the processor device may also include primarily analog components. For example, some or all of the rendering techniques described herein may be implemented in analog circuits or mixed analog and digital circuits. The computing environment may include any type of computer system, including but not limited to a microprocessor-based computer system, a mainframe computer, a digital signal processor, a portable computing device, a device controller, or a computing engine within an appliance (to name a few examples).

[0276] The elements of the methods, processes, routines or algorithms described in conjunction with the embodiments disclosed herein may be directly embodied in hardware, software modules executed by a processor device, or a combination of the two. The software module may reside in a RAM memory, a flash memory, a ROM memory, an EPROM memory, an EEPROM memory, a register, a hard disk, a removable disk, a CD-ROM, or any other form of non-transient computer-readable storage medium. An exemplary storage medium may be coupled to a processor device so that the processor device may read information from the storage medium and write information to the storage medium. In an alternative, the storage medium may be integrated with the processor device. The processor device and the storage medium may reside in an ASIC. The ASIC may reside in a user terminal. In an alternative, the processor device and the storage medium may reside in a user terminal as discrete components.

[0277] In the previous description, aspects and embodiments of the present disclosure have been described with reference to many specific details, which may vary depending on the implementation. Therefore, the specification and drawings should be regarded as illustrative, not restrictive. The only and exclusive indication of the scope of the invention, and the applicant's expectation that the scope of the invention is the literal and equivalent scope of the claims authorized from this application in the specific form of the claims of the authorization announcement, including any subsequent amendments. Any definition of the terms used to be included in such claims explicitly set forth herein should be based on the meaning of such terms as used in the claims. In addition, when the term "also includes" is used in the previous specification or the attached claims, the phrase may be followed by additional steps or entities, or sub-steps / sub-entities of the previously described steps or entities.

Claims

1. A method, include: receiving a plurality of images from a plurality of image sensors, the plurality of images corresponding to a plurality of views of a scene of the vehicle; generating a plurality of feature maps based on the plurality of images; Determining a first plurality of windows for the plurality of feature maps, wherein a first window in the first plurality of windows includes a first grid cell of a first feature map from the plurality of feature maps and a second grid cell of a second feature map from the plurality of feature maps; enhancing the semantic data sets associated with the plurality of feature maps based on the first plurality of windows to provide a first enhanced semantic data set, wherein enhancing the semantic data sets associated with the plurality of feature maps based on the first plurality of windows comprises: determining first semantic data for the first grid cell using second semantic data associated with the second grid cell based on the first grid cell and the second grid cell being included in the first window; Determining a second plurality of windows for the plurality of feature maps, wherein a third window in the second plurality of windows includes the first grid unit and the third grid unit, and a fourth window in the second plurality of windows includes the second grid unit; enhancing the first enhanced semantic data set associated with the plurality of feature maps based on the second plurality of windows to provide a second enhanced semantic data set associated with the plurality of feature maps, wherein enhancing the first enhanced semantic data set associated with the plurality of feature maps based on the second plurality of windows comprises: determining third semantic data for the first grid cell using fourth semantic data associated with the third grid cell based on the first grid cell and the third grid cell being included in the third window; generating at least one bounding box for an object in a scene of the vehicle based on the second set of enhanced semantic data; and The vehicle is caused to be controlled based on the at least one bounding box.

2. The method according to claim 1, in, Generating at least one bounding box for an object in the scene of the vehicle based on the second enhanced semantic data set includes: enhancing a set of object queries based on the second enhanced semantic data set to provide an enhanced set of object queries; and The at least one bounding box is generated based on the enhanced object query set.

3. The method according to claim 1 or 2, in, Generating at least one bounding box for an object in the scene of the vehicle based on the second enhanced semantic data set includes: generating a bird's-eye view feature map based on the second enhanced semantic data set; and The at least one bounding box is generated based on the bird's eye view feature map.

4. The method according to any one of claims 1 to 3, in, Generating at least one bounding box for an object in the scene of the vehicle based on the second enhanced semantic data set includes: enhancing a set of object queries based on the second enhanced set of semantic data to provide a first enhanced set of object queries; generating a bird's-eye view feature map based on the second enhanced semantic data set; enhancing the first enhanced object query set based on the bird's eye view feature map to provide a second enhanced object query set; and The at least one bounding box is generated based on the second enhanced set of object queries.

5. The method according to any one of claims 1 to 4, in, The plurality of image sensors are positioned around the vehicle in different orientations.

6. The method according to any one of claims 1 to 5, in, The plurality of images provides a 360 degree view around the vehicle.

7. The method according to any one of claims 1 to 6, in, The first plurality of windows form at least one row of windows, and each window in the first plurality of windows is of the same width and height.

8. The method according to any one of claims 1 to 7, in, The first plurality of windows form a plurality of rows of windows.

9. The method according to any one of claims 1 to 8, in, With respect to the plurality of feature maps, the second plurality of windows are horizontally shifted relative to the first plurality of windows.

10. The method according to any one of claims 1 to 8, in, With respect to the plurality of feature maps, the second plurality of windows are vertically and horizontally shifted relative to the first plurality of windows.

11. The method according to any one of claims 1 to 8, in, The second plurality of windows are a different size than the first plurality of windows.

12. A system, include: A data storage unit storing computer executable instructions; as well as a processor configured to execute the computer executable instructions, wherein execution of the computer executable instructions causes the system to: receiving a plurality of images from a plurality of image sensors, the plurality of images corresponding to a plurality of views of a scene of the vehicle; generating a plurality of feature maps based on the plurality of images; Determining a first plurality of windows for the plurality of feature maps, wherein a first window in the first plurality of windows includes a first grid cell of a first feature map from the plurality of feature maps and a second grid cell of a second feature map from the plurality of feature maps; enhancing the semantic data sets associated with the plurality of feature maps based on the first plurality of windows to provide a first enhanced semantic data set, wherein enhancing the semantic data sets associated with the plurality of feature maps based on the first plurality of windows comprises: determining first semantic data for the first grid cell using second semantic data associated with the second grid cell based on the first grid cell and the second grid cell being included in the first window; Determining a second plurality of windows for the plurality of feature maps, wherein a third window in the second plurality of windows includes the first grid unit and the third grid unit, and a fourth window in the second plurality of windows includes the second grid unit; enhancing the first enhanced semantic data set associated with the plurality of feature maps based on the second plurality of windows to provide a second enhanced semantic data set associated with the plurality of feature maps, wherein enhancing the first enhanced semantic data set associated with the plurality of feature maps based on the second plurality of windows comprises: determining third semantic data for the first grid cell using fourth semantic data associated with the third grid cell based on the first grid cell and the third grid cell being included in the third window; generating at least one bounding box for an object in a scene of the vehicle based on the second set of enhanced semantic data; and The vehicle is caused to be controlled based on the at least one bounding box.

13. The system according to claim 12, in, To generate at least one bounding box for an object in the scene of the vehicle based on the second enhanced semantic data set, the processor is configured to: enhancing a set of object queries based on the second enhanced semantic data set to provide an enhanced set of object queries; as well as The at least one bounding box is generated based on the enhanced object query set.

14. The system according to claim 12 or 13, in, To generate at least one bounding box for an object in the scene of the vehicle based on the second enhanced semantic data set, the processor is configured to: generating a bird's-eye view feature map based on the second enhanced semantic data set; as well as The at least one bounding box is generated based on the bird's eye view feature map.

15. A system according to any one of claims 12 to 14, in, To generate at least one bounding box for an object in the scene of the vehicle based on the second enhanced semantic data set, the processor is configured to: enhancing a set of object queries based on the second enhanced set of semantic data to provide a first enhanced set of object queries; generating a bird's-eye view feature map based on the second enhanced semantic data set; enhancing the first enhanced object query set based on the bird's eye view feature map to provide a second enhanced object query set; as well as The at least one bounding box is generated based on the second enhanced set of object queries.

16. A system according to any one of claims 12 to 15, in, The second plurality of windows are a different size than the first plurality of windows.

17. A non-transitory computer readable medium comprising computer executable instructions which, when executed by a computing system, cause the computing system to: receiving a plurality of images from a plurality of image sensors, the plurality of images corresponding to a plurality of views of a scene of the vehicle; generating a plurality of feature maps based on the plurality of images; A first plurality of windows is determined for the plurality of feature maps, wherein: A first window in the first plurality of windows comprises a first grid cell of a first feature map from the plurality of feature maps and a second grid cell of a second feature map from the plurality of feature maps; enhancing the semantic data sets associated with the plurality of feature maps based on the first plurality of windows to provide a first enhanced semantic data set, wherein enhancing the semantic data sets associated with the plurality of feature maps based on the first plurality of windows comprises: determining first semantic data for the first grid cell using second semantic data associated with the second grid cell based on the first grid cell and the second grid cell being included in the first window; Determining a second plurality of windows for the plurality of feature maps, wherein a third window in the second plurality of windows includes the first grid unit and the third grid unit, and a fourth window in the second plurality of windows includes the second grid unit; enhancing the first enhanced semantic data set associated with the plurality of feature maps based on the second plurality of windows to provide a second enhanced semantic data set associated with the plurality of feature maps, wherein enhancing the first enhanced semantic data set associated with the plurality of feature maps based on the second plurality of windows comprises: determining third semantic data for the first grid cell using fourth semantic data associated with the third grid cell based on the first grid cell and the third grid cell being included in the third window; generating at least one bounding box for an object in a scene of the vehicle based on the second set of enhanced semantic data; and The vehicle is caused to be controlled based on the at least one bounding box.

18. The non-transitory computer readable medium of claim 17, in, To generate at least one bounding box for an object in the scene of the vehicle based on the second set of enhanced semantic data, execution of the computer executable instructions further causes the computing system to: enhancing a set of object queries based on the second enhanced semantic data set to provide an enhanced set of object queries; as well as The at least one bounding box is generated based on the enhanced object query set.

19. The non-transitory computer readable medium according to claim 17 or 18, in, To generate at least one bounding box for an object in the scene of the vehicle based on the second set of enhanced semantic data, execution of the computer executable instructions further causes the computing system to: generating a bird's-eye view feature map based on the second enhanced semantic data set; as well as The at least one bounding box is generated based on the bird's eye view feature map.

20. The non-transitory computer readable medium according to any one of claims 17 to 19, in, To generate at least one bounding box for an object in the scene of the vehicle based on the second set of enhanced semantic data, execution of the computer executable instructions further causes the computing system to: enhancing a set of object queries based on the second enhanced set of semantic data to provide a first enhanced set of object queries; generating a bird's-eye view feature map based on the second enhanced semantic data set; enhancing the first enhanced object query set based on the bird's eye view feature map to provide a second enhanced object query set; as well as The at least one bounding box is generated based on the second enhanced set of object queries.