Calculating feature correlations to estimate depth information for stereo images
Patent Information
- Application Number
- DE102025115307
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-19
- Filing Date
- 2025-04-17
- Publication Date
- 2025-10-23
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
BACKGROUND
[0001] A computer vision system used for stereovision (referred to as a computer stereovision system) perceives objects in a real-world scene, typically based on a pair of two-dimensional (2D) images of the scene, captured horizontally from each other using two cameras (e.g., stereo cameras). Such a pair of digital 2D images is often referred to as the stereo image pair, the left and right stereo image, or simply the left and right image. The system perceives the scene in three dimensions (3D) by extracting depth information from the stereo image pair. The depth information is typically calculated based on the distance between two corresponding pixels in the stereo image pair (typically referred to as the disparity between two points).Such a system can be used in an autonomous mobile robot (AMR) as part of a perception system configured to perceive the depth of objects and structures in its environment in real or near real time. Such a system can also be used in a robotic arm as part of a perception system configured to perceive the depth of itself and the objects it manipulates in real or near real time.
[0002] A computer stereo vision system can predict a depth map for the left stereo image or the right stereo image. For example, a depth map for the left stereo image shows a depth value for each pixel of the stereo image. Such a depth map is often represented as a 2D array / vector of depth values. A depth map is frequently calculated based on a disparity map predicted for the same stereo image and the intrinsic values of the stereo cameras.
[0003] To compute a depth map for a given stereo image, a set of feature maps is extracted from each of the left and right stereo images. Each feature map represents an extracted feature from its respective stereo image. Examples of features include edges, textures, shapes, and higher-level features such as object semantics and categories. Each feature map is often represented as a 2D data structure (e.g., a 2D vector or tensor), where each position in the map contains a value (e.g., a pixel value) indicating the presence, absence, or level of a given feature. Each feature map is often referred to as a feature channel or channel. All channels are often collectively referred to as a channel dimension. In this way, the extracted features from each of the stereo images are represented in a 3D structure.For example, if there is a C number of channels in a set of feature maps and each feature map has a width of W and a height of H (e.g., dimensions of W times H), then the set will have dimensions of C times W times H.
[0004] Using the two sets of feature maps extracted from the left and right stereo images, a feature correlation engine (also known as cost volume calculation) calculates the correlations between the two sets of feature maps, which are used to compute the depth map. Specifically, for each channel, the feature correlation engine calculates correlations between the left feature map for the left stereo image and the right feature map for the right stereo image.
[0005] In some existing approaches, for each position in a row of the left feature map, a correlation value is calculated relative to each position within a search window of the corresponding row in the right feature map. Intuitively, the search window represents a set of possible corresponding positions in the right feature map for a given position in the left feature map. The search window typically has a predefined width (often referred to as the maximum disparity) that is less than the full width of the right feature map. The calculation of correlation values can be performed using a feature-shifting technique. With the left and right feature maps aligned, such a technique shifts the right feature map across the left feature map to the rightmost column over time, up to the maximum disparity.In this way, with each shift, an overlapping area is formed between the two feature maps, and correlation scores are then calculated based on positions in the two feature maps that are part of the overlapping area.
[0006] However, feature shifting requires a significant amount of processing and computation time. Furthermore, the maximum disparity may not be large enough to include the corresponding position in the right-hand feature map for a given position in the left-hand feature map. Consequently, the disparity predicted for a given pixel may be incorrect. Additionally, the computation time of a feature shift increases linearly and is therefore not scalable, even if the maximum disparity increases (e.g., to avoid missing the corresponding position in the right-hand feature map).
[0007] As such, there is a need for more efficient and effective techniques for calculating correlations between feature maps extracted from a pair of stereo images. OVERVIEW
[0008] The invention is defined by the claims. To illustrate the invention, aspects and embodiments that may or may not be within the scope of the claims are described herein.
[0009] Embodiments of the present disclosure relate to techniques for calculating feature correlations for estimating depth information for stereo images. The techniques described herein include, for each channel in a set of feature channels, receiving a first feature map for a first image in a stereo image pair and a second feature map for a second image in the stereo image pair, and calculating a corresponding set of correlation maps representing correlations between each position in a row in the first feature map and each position in a corresponding row in the second feature map. The techniques further include generating a set of compressed correlation maps based on compressing the sets of correlation maps corresponding to the set of feature channels over a dimension associated with the set of feature channels.The techniques further include: masking one or more sections of each of the set of compressed correlation maps based on a respective correlation filter to generate a corresponding set of masked correlation maps. The techniques further include: generating a depth map associated with the stereo image pair based on the set of masked correlation maps.
[0010] The disclosed technique offers several technical advantages over previous approaches. Because the disclosed technique computes correlations between a given column in the left feature map invariably relative to all columns in the right feature map, it avoids, in particular, the need for the dynamic computation performed in previous approaches, thereby improving computational efficiency. Furthermore, the disclosed technique is designed to utilize the parallel processing capabilities of a processor(s) (e.g., GPU(s), programmable vision accelerators (PVAs), deep learning accelerators, optical flow accelerators (OFAs), etc.), further improving computational efficiency.
[0011] The revelation extends to any novel aspects or features described and / or illustrated herein.
[0012] Further features of the disclosure are characterized by the independent and dependent claims.
[0013] Any feature in one aspect of the disclosure can be applied to other aspects of the disclosure in any suitable combination. In particular, procedural aspects can be applied to device or system aspects, and vice versa.
[0014] Furthermore, features implemented in hardware can also be implemented in software, and vice versa. Any reference to software and hardware features in this description should be interpreted accordingly.
[0015] Each system or device feature as described herein can also be provided as a process feature, and vice versa. Functionally described system and / or device aspects (including means plus functional features) can alternatively be expressed in terms of their corresponding structure, such as a suitably programmed processor and associated memory.
[0016] It should also be understandable that certain combinations of the various features described and defined in any aspect of the revelation can be implemented and / or provided and / or used independently of one another.
[0017] This disclosure also provides computer programs and computer program products comprising software code adapted, when executed on a data processing device, to perform any of the procedures described herein and / or to embody any of the device and system features described herein, including any or all of the partial steps of a procedure.
[0018] The disclosure also provides a computer or computing system (including networked or distributed systems) that includes an operating system supporting a computer program for performing any of the procedures described herein and / or embodying any of the device or system features described herein.
[0019] The revelation also provides a computer-readable medium on which one or more of the aforementioned computer programs are stored.
[0020] The revelation also provides a signal that carries one or more of the aforementioned computer programs.
[0021] The disclosure extends to processes and / or devices and / or systems as described herein with reference to the accompanying drawings.
[0022] Aspects and embodiments of the disclosure will now be described exclusively by way of example with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The systems and methods presented here for calculating feature correlations to estimate depth information for stereo images are described in detail below with reference to the attached drawings, wherein: Fig. 1 illustrates a computing device configured to implement one or more aspects of different embodiments; Fig. 2 a more detailed illustration of a depth estimation engine by Fig. 1 according to different embodiments; Fig. 3 a more detailed illustration of a feature correlation engine of Fig. 2 according to different embodiments; Fig. 4A-4B Illustrate operations of a feature correlation engine according to different embodiments; Fig. Figure 5 illustrates a flowchart of a feature correlation method according to different embodiments; Fig. 6 is a block representation of an exemplary computing device suitable for use in implementing some embodiments of the present disclosure; Fig. 7 is a block representation of an exemplary data center suitable for use in implementing some embodiments of the present disclosure; Fig. Figure 8A is an illustration of an exemplary autonomous vehicle according to some embodiments of the present disclosure; Fig. 8B is an example of camera locations and fields of view for the exemplary autonomous vehicle of Fig. 8A according to some embodiments of the present disclosure; Fig. 8C a block diagram of an exemplary system architecture for the exemplary autonomous vehicle of Fig. 8A according to some embodiments of the present disclosure; and Fig. 8D a system representation according to some embodiments of the present disclosure for communication between one or more cloud-based server(s) and the exemplary autonomous vehicle of Fig. 8A is. DETAILED DESCRIPTION
[0024] Techniques for a feature correlation engine for calculating correlations between feature maps generated from a pair of stereo images (e.g., two or more images acquired using two or more image sensors with at least partially overlapping fields of view) are disclosed. The feature correlation engine can be part of a depth estimation engine (e.g., a neutral deep learning network (DNN)) configured to predict a depth map (or synonymously, a disparity map) for a given stereo image (e.g., a left stereo image). The depth estimation engine extracts a set of feature maps from each of the stereo images. Each feature map in the set of feature maps represents a distinct feature and is typically referred to as a feature channel or channel. These channels are collectively referred to as a channel dimension.
[0025] Using the two sets of feature maps within each channel, the feature correlation engine calculates correlations between all pixels in a given row (e.g., the first row) of one feature map (e.g., a left feature map) and all pixels in a corresponding row (e.g., the first row) of the other feature map (e.g., a right feature map) using multiplication. Since the pixels in a column are located in different rows of a feature map, the feature correlation engine, in at least one embodiment, calculates correlations between all columns in the left feature map and all columns in the right feature map. Specifically, within each channel, the feature correlation engine calculates correlations between a given column in the left feature map and all columns in the right feature map to generate a correlation map for that given column in the left map.Since the number of columns in the right-hand feature map corresponds to the width of the right-hand feature map, such a generated correlation map has the dimensions of the right-hand feature map, which correspond to W times H. Furthermore, since the number of columns in the left-hand feature map corresponds to the width of the left-hand feature map, a W-number of correlation maps is formed within each channel. As a result, all correlation maps across all channels form a structure with the dimensions C times W times W times H.
[0026] As a further optimization, the feature correlation engine compresses the correlation maps to improve computational efficiency. Specifically, the feature correlation engine compresses the correlation maps across the channel dimension. In particular, the correlation maps created using the corresponding columns in the left-hand feature maps are summed across all channels. Since there is a W number of columns in the left-hand feature maps, a W number of compressed correlation maps are generated, each with the dimensions of the original correlation maps equal to W times H. This compression operation consolidates all channels into a single channel and efficiently eliminates the channel dimension of the original correlation maps. As a result, the compressed correlation maps have the dimension W times W times H.
[0027] As a further optimization, the feature correlation engine performs a masking operation to further improve computational efficiency and accuracy. Specifically, since all columns of the right feature map are used in the feature correlation calculation, certain positions in the right feature map that might not be necessary for disparity prediction are also included in the calculation. In particular, according to the intrinsic properties of a stereo image pair, a position in the right feature map that corresponds to a given position in the left feature map is always to the left of the given position. That is, positions in the right feature map that are to the right of the given position cannot correspond to the given position in the left feature map and are therefore not used for disparity prediction.Rather, a disparity predicted based on such positions in the right feature map and the given position in the left feature map would yield an unrealistic negative disparity value. Accordingly, the disclosed feature correlation engine applies a masking operation (also known as bit masking or positional masking) to each of the compressed correlation maps to remove compressed correlations that are not used in downstream processing to predict a disparity map.
[0028] During the masking operation, a different mask (also called a correlation filter) is generated for each of the compressed correlation maps based on the spatial position of the column in the left feature map for which the respective correlation map was generated. Because the masks are calculated based on the spatial position of the columns and not the pixel values in those columns, the computational cost of calculating the masks remains constant and negligible, regardless of the size of the left and right stereo images and the pixel values. Once the masking operation is complete, the resulting masked correlation maps are provided for downstream processing to predict a disparity map for the left stereo image.
[0029] The disclosed technique offers several technical advantages over previous approaches. Because the disclosed technique calculates correlations between a given column in the left feature map invariably relative to all columns in the right feature map, it avoids, in particular, the need for the dynamic calculation performed in previous approaches, thereby improving computational efficiency. Furthermore, the disclosed technique employs matrix multiplication, which can better utilize the parallel processing capabilities of the processor(s) (e.g., GPU(s)) and further improves computational efficiency.
[0030] Fig. Figure 1 illustrates a computing device 100 configured to implement one or more aspects of various embodiments. In at least one embodiment, the computing device 100 includes: a desktop computer, a laptop computer, a smartphone, a personal digital assistant (PDA), a tablet computer, a server, one or more virtual machines, an embedded system, a system-on-a-chip (SoC) system(s), an in-vehicle computing device, and / or any other type of computing device configured to receive input, process data, and optionally display images, and suitable for implementing one or more embodiments. The computing device 100 is configured to operate a depth estimation engine 122, which resides in a memory 116.It should be noted that the computing device described herein is illustrative and that other technically feasible configurations fall within the scope of this disclosure. For example, multiple instances of the depth estimation engine 122 can be run on a set of nodes in a distributed and / or cloud computing system to implement the functionality of the computing device 100.
[0031] In one embodiment, the computing device 100 includes, among other things: an intermediate connection (bus) 112 connecting one or more processors 102, an input / output (I / O) device interface 104 coupled to one or more input / output (I / O) devices 108, a memory 116, a storage 114, and / or a network interface 106. The processor(s) 102 may include any suitable processor implemented as follows: as a central processing unit (CPU), as a graphics processing unit (GPU), as an application-specific integrated circuit (ASIC), as a field-programmable gate array (FPGA), as an artificial intelligence accelerator (AI accelerator), or as a parallel processing unit (PPU).as a data processing unit (DPU), as a programmable vision accelerator (PVA), which may include one or more vector processing units (VPUs), one or more pixel processing engines (PPEs), and / or one or more direct memory access (DMA) systems, any other type of processing unit, or a combination of different processing units, such as a CPU(s) configured to operate in conjunction with a GPU(s). In general, the processor(s) 102 may include any feasible hardware unit capable of processing data and / or executing software applications. Furthermore, in the context of this disclosure,The computing elements shown in computing device 100 correspond to a physical computing system (e.g., a system in a data center) and / or can correspond to a virtual computing instance running within a computing cloud.
[0032] In at least one embodiment, the I / O devices 108 include devices capable of receiving input, such as a keyboard, mouse, touch screen, touchpad, VR / MR / AR headset, gesture recognition system, and / or microphone, as well as devices capable of providing output, such as a display device(s), a haptic device(s), and / or one or more loudspeakers. Furthermore, I / O devices 108 may include devices capable of both receiving input and providing output, such as a touch device, a universal serial bus (USB) port, etc. I / O devices 108 may be configured to receive various types of input from an end user (e.g.,to receive data from the computer device 100 (e.g., from a designer) and also to provide the end user of the computer device 100 with various types of output, such as displayed digital images, digital videos, or text. In some embodiments, one or more I / O devices 108 are configured to connect the computer device 100 to a network 110.
[0033] In one embodiment, the network 110 is any technically feasible type of communication network that allows data to be exchanged between the computing device 100 and internal, local, remote, or external entities or devices, such as a web server and / or another networked computing device. For example, the network 110 may include, among others, a wide area network (WAN), a local area network (LAN), a wireless network (e.g., a WiFi network), a cellular network, and / or the Internet.
[0034] In at least one embodiment, the storage 114 includes non-volatile storage for applications and data and may include fixed or removable disk drives, flash memory devices, and CD-ROM, DVD-ROM, Blu-ray, HD-DVD, or other magnetic, optical, or solid-state storage devices. The depth estimation engine 122 may be stored in the storage 114 and loaded into memory 116 when it is executed.
[0035] In one embodiment, the memory 116 includes a random-access memory (RAM) module, a flash memory unit, and / or any other type of memory unit or combination thereof. The processor(s) 102, the I / O device interface 104, and the network interface 106 can be configured to read data from and write data to the memory 116. The memory 116 can include various software programs executed by the processor(s) 102 and application data associated with the software programs, including the depth estimation engine 122.
[0036] Depth Estimation Engine 122 includes functionality to estimate a depth map based on a stereo image pair. Using a stereo image pair, Depth Estimation Engine 122 can extract various features from the stereo image pair. Examples of features include edges, textures, shapes, and higher-level features such as object semantics and categories. In at least one embodiment, Depth Estimation Engine 122 extracts a set of left feature maps and a set of right feature maps from the left stereo image and the right stereo image, respectively. Each feature map in the set of left or right feature maps represents a distinct textured feature and is often referred to as a feature channel or channel. Two feature maps in the set of feature maps are considered correlated (e.g., at least for feature correlation purposes) if they correspond to the same channel.
[0037] To calculate a depth map for one of the stereo images (e.g., the left stereo image), the depth estimation engine calculates 122 correlations between the two sets of feature maps. The operation of such a correlation calculation is described in Fig. 2 and Fig. 3. Described in more detail. Based on the calculated correlations, the Depth Estimation Engine 122 determines, for each pixel of the left stereo image, a corresponding pixel in the right stereo image. Intuitively, the correspondence between the two pixels indicates that they represent the same point in space in a real-world scene captured by the stereo image pair. For each pair of corresponding pixels, the Depth Estimation Engine 122 determines a disparity value between the two pixels. Such determined disparity values for the left stereo image form a disparity map for the left stereo image. The Depth Estimation Engine 122 can calculate a depth map for the left stereo image based on the disparity map using the stereo camera configurations and / or intrinsic values (e.g., focal length and baseline) according to the following formula: Depth = (Focal length * Baseline) Disparity
[0038] Fig. Figure 2 is a more detailed illustration of a depth estimation engine 122 by Fig. 1 according to various embodiments. As shown, the depth estimation engine 122 includes a first feature extractor 212, a second feature extractor 214, a feature correlation engine 232 (also referred to as cost volume calculation), and a depth estimator 242. Using a left stereo image 202 and a right stereo image 204, the depth estimation engine 122 can generate a depth map 252 that represents the depth information for the pixels in one of the stereo images. In some embodiments, the depth estimation engine 122 is implemented as a neutral network (NN) (e.g., as a neutral deep learning network (DNN)).
[0039] In at least one embodiment, the depth estimation engine 122 is configured to generate a depth map 252 for the left stereo image 202. In such an embodiment, the first feature extractor 212 and the second feature extractor 214 can receive the left stereo image 202 and the right stereo image 204, respectively, as input and generate a left feature map set 222 and a right feature map set 224, respectively. Each of the first feature extractor 212 and the second feature extractor 214 can use an identical set of one or more convolutional layers to generate their respective set of feature maps. For example, with the left stereo image 202, the first feature extractor 212 can use a convolutional layer to extract features from the left stereo image 202 using a set of convolutional filters (also called filters or kernels).Each of the set of convolutional filters is used to extract a different feature from the left stereo image 202 and output a corresponding left feature map. In this way, a set of left feature maps, such as a left feature map set 222, can be extracted from the left stereo image 202 using the convolutional filters. Similarly, the second feature extractor 214 can extract the right feature map set 224 from the right stereo image 204 using the left stereo image 202. In some embodiments, the first feature extractor 212 and the second feature extractor 214 are implemented using CNN(s) or a variant thereof. In such cases, the two feature extractors can share network weights, ensuring that the feature maps are extracted from them in a consistent manner.
[0040] Continuing with the embodiment above, the feature correlation engine 232 calculates correlations between the left feature map set 222 and the right feature map set 224 to generate the depth map 252 for the left stereo image 202. Specifically, for each channel, correlations are calculated between each position in a row of the left feature map and each position in a corresponding row of the right feature map. The feature correlation engine 232 performs a compression operation on the calculated correlations from each channel, for example, to improve processing efficiency. In particular, the feature correlation engine 232 compresses the calculated correlations across all channels so that the compressed correlations are consolidated into a single channel instead of the multiple channels associated with the original calculated correlations.
[0041] Feature correlation engine 232 performs a masking operation on the compressed correlations. Specifically, feature correlation engine 232 masks those compressed correlations that are compressed based on calculated correlations between a given position in a left feature map and the position(s) in a corresponding right feature map that cannot be a possible corresponding position for the given position. In particular, feature correlation engine 232 masks those compressed correlations that are compressed based on calculated correlations between a given position in the left feature map and the positions in the right feature map that are to the right of the given position (e.g., because such positions would result in a negative disparity value, as described above).Once the masking is complete, the feature correlation engine 232 outputs the masked correlations as feature correlations 234. This avoids unnecessary downstream processing of the compressed correlations, which cannot contribute to calculating the depth map 252. The operation(s) of the feature correlation engine 232 is / are described in . Fig. 3 described in more detail.
[0042] Continuing with the embodiment described above, the depth estimator 242 receives the feature correlations 252 and the left feature map set 222 as input to compute the depth map 252 for the left stereo image 202. In at least one embodiment, the depth estimator 242 combines the feature correlations 252 with the extracted features in the left feature map set 222 to create the depth map 252. The depth estimator 242 can be implemented as a unet architecture that includes an encoder and a decoder. Using the left feature map set 222, the encoder can include convolutional layers, which, in some embodiments, are followed by max-pool operations, for example, to reduce the spatial resolution of the feature maps in the left feature map set 222. Such an encoder can be implemented using a rest network (e.g., a Res-net 18).With the output of such an encoder, the decoder can increase the spatial resolution of the feature maps through upsampling, e.g., to generate a dense segmentation map (e.g., the Depth Map 252). The Depth Map 252 can be used by a perception system (e.g., a computer stereo vision system) to perceive objects and / or structures in the environment in a real-world scene in real time.
[0043] Fig. Figure 3 is a more detailed illustration of the feature correlation engine 232 by Fig. 2 according to different embodiments. As shown, with the left feature card set 222 and the right feature card set 224, as in Fig. 2 described, the feature correlation engine 232 performs various operations, including generating correlation maps 312, compressing correlation maps 322 and masking correlation maps 332, as well as outputting masked correlation maps 334.
[0044] For example, the feature correlation engine 232 calculates a depth map 252 for the left stereo image 202 (in Fig. 2 shown), with the left feature card set 222 and the right feature card set 224, in the generate-of-correlation-cards-312 operation, for each channel, correlations between each position in a row of the left feature card and each position in a corresponding row of the right feature card. Fig. Figure 4A illustrates an exemplary implementation of such a generate-of-correlation-maps operation. As shown, an exemplary left feature map set includes N channels, ch1-chN. Each of the N channels has a corresponding feature map with dimensions W × H. Each feature map includes positions organized into three columns. Columns Col1L, Col2L, Col3L include positions L11, L21, and L31; positions L12, L22, and L32; and positions L13, L23, and L33, respectively. Each position can include an associated value (e.g., a pixel value) indicating the presence or absence of the feature associated with the feature map. In other words, each feature map includes positions organized into three rows: L1, L12, and L13; L21, L22, and L23; and L31, L32, and L33. As shown, an exemplary right-hand feature card set 224 is similarly structured and organized.It should be understood that the number of channels and dimensions shown in these feature maps are for illustrative purposes only. Any other suitable number of channels and / or any other suitable dimensions can be implemented for these feature maps.
[0045] In operation 312, the feature correlation engine 232 performs a multiplication on each channel between each column in the left feature map and all columns in the right feature map to generate a corresponding correlation map. For example, as shown, the multiplication for channel ch1, between Col1L and Col1R, Col2R, Col3R, produces the leftmost correlation map of the correlation maps 314Ch1. The multiplication is performed similarly for the other two columns in the left feature map to produce the other two correlation maps in channel ch1. In this way, three correlation maps are produced based on the three columns in the left feature map. Correlation maps 314Ch2 through 314ChN are also produced similarly for channels ch2 through chN.The disclosed technique of performing multiplications between each column in the left feature map and all columns in the right feature map inherently calculates correlations between each position in a row of the left feature map and each position in a corresponding row of the right feature map.
[0046] Returning to Fig. 3, continuing with the example above, the feature correlation engine 232, with the correlation maps 314, generates compressed correlation maps 324 during the compress-of-correlation-maps-322 operation. Fig. Figure 4A illustrates an exemplary implementation of such a compression-of-correlation-maps operation. In operation 322, the correlation maps generated from the same column of the left feature map are summed across all channels (e.g., using matrix addition). For example, as shown, the correlation maps generated from Col1L of the left feature map are summed across all channels. As a result, three compressed correlation maps 324 are produced. In this way, the channels ch1-chN for the correlation maps 314Ch1-314ChN are consolidated into a single channel. In other words, the channel dimension of the correlation maps 314Ch1-314ChN is effectively eliminated.
[0047] As shown, each of the compressed correlation maps 324 maintains the spatial arrangement of the positions in the corresponding correlation maps. For example, the positions with the values ΣL11×R11, ΣL11×R12 and ΣL11×R13 in the leftmost compressed correlation map correspond to the positions with the values L11×R11, L11×R12 and L11×R13 in the leftmost correlation map in each of the channels ch1-chN.
[0048] Returning to Fig. 3, continuing with the example above, the feature correlation engine 232, with the compressed correlation maps 324, masks certain cooperations in the compressed correlation maps 324 during the masking-of-correlation-maps-332 operation, as in Fig. 2 described. Fig. Figure 4B illustrates an exemplary implementation of such a masking-of-correlation-maps-322 operation. In operation 332, correlation filters 333 are applied to respective compressed correlation maps 324 to produce respective masked correlation maps 334.
[0049] Each of the correlation filters 333 is generated specifically for a given compressed correlation map based on how those correlation maps are calculated. For example, the far left correlation filter is generated for the far left compressed correlation map based on how the far left correlation map is calculated in each of the channels ch1-chN. Each far left correlation map is generated based on calculating correlations between Col1L of a left feature map and Col1R, Col2R, and Col3R in a corresponding right feature map. As described above, a position in the right feature map that is to the right of a given position in the left feature map cannot correspond to that given position, and the correlations calculated based on such two positions cannot therefore be used for disparity prediction.Therefore, in this example, since R12 lies to the right of L11, R12 cannot correspond to L11. Accordingly, as shown, the leftmost correlation filter is generated to include a value of 0 at the same spatial position as the position with the value ∑L11×R12 in the leftmost correlation map. Such a value of 0 indicates that the value at the same spatial position in the leftmost compressed correlation map is to be masked during Operation 332. As another example, a positive value is generated in the leftmost correlation filter at the same spatial position as the position with the value ΣL11×R11 in the leftmost compressed correlation map. Such a positive value indicates that the value at the same spatial position in the leftmost compressed correlation map is not to be masked during Operation 332.In at least one embodiment, in order to calculate a correlation filter for a given compressed correlation map, the feature correlation engine 232 first generates a negative bound value for position(s) in the row to be masked for each row in the compressed correlation map, and then distributes the collection of the negative bound value(s) and the original values in that row to positive values and 0s (e.g. using a SoftMax function).
[0050] As shown, when the feature correlation engine 232 performs operation 332, the position with the value ∑L11×R12 in the leftmost compressed correlation map is updated to include a value of 0. The position with the value ΣL11×R11 in the leftmost compressed coalition map retains its original value. In this way, the feature correlation engine 232 masks the correlations in the compressed correlation maps 324 that are calculated based on positions that cannot correspond to each other.
[0051] Now, referring to Fig. 5. Each block of the Method 500 described herein comprises a computational process that can be performed using any combination of hardware, firmware, and / or software. For example, various functions can be performed by one or more processors executing instructions stored in memory. The methods can also be embodied as computationally usable instructions stored in computer storage media. The methods can be provided by a standalone application, a service, or a hosted service (alone or in combination with another hosted service), or a plug-in to another product, to name only a few. Furthermore, Method 500 is exemplified with respect to the system of Fig. 1-4 described. However, these procedures can additionally or alternatively be carried out by any system or any combination of systems, including those described herein.
[0052] Fig. Figure 5 illustrates a flowchart representing a process 500 for Fig. Figure 3 shows various embodiments. The method calculates 500 feature correlations for a stereo image pair to generate a depth map using the stereo image pair. As shown in Fig. As shown in Figure 5, procedure 500 begins with operation 502, in which a depth estimation engine (e.g., depth estimation engine 122) performs operations 504 and 506 for at least one channel in a set of feature channels. Specifically, in operation 504, the depth estimation engine receives a first feature map for a first image in a stereo image pair and a second feature map for a second image in a stereo image pair. In operation 506, the depth estimation engine computes a corresponding set of correlation maps representing correlations between one or more positions in a row in the first feature map and one or more positions in a corresponding row in the second feature map.In some embodiments, the set of correlation maps is calculated at least on the basis of multiplying at least one column of the first feature map with one or more columns of the right feature map using matrix multiplication.
[0053] In Operation 508, the depth estimation engine generates a set of compressed feature maps based at least on compressing the sets of correlation maps corresponding to the feature channel set over a dimension associated with the feature channel set. In some embodiments, the generation of the set of compressed correlation maps includes summing corresponding correlation maps in one or more feature channels of the feature channel set.
[0054] In Operation 510, the depth estimation engine masks one or more sections of individual compressed correlation maps from the set of compressed correlation maps, based on at least one correlation filter, to generate a corresponding set of masked correlation maps. In some embodiments, the masked sections of individual compressed correlation maps from the set of compressed correlation maps include compressed correlations that are not used in the depth map generation.
[0055] In Operation 512, the depth estimation engine generates a depth map associated with the stereo image pair, based at least on the set of masked correlation maps. In some embodiments, the generated depth map is used by a perception system (e.g., a computer stereo vision system) to perceive objects and / or structures in the environment in a real-world scene in real time or near real time.
[0056] The systems and procedures described herein may be used by, among others, non-autonomous vehicles, semi-autonomous vehicles (e.g., in one or more adaptive driver assistance systems (ADAS)), guided and unguided robots or robotic platforms, warehouse vehicles, off-road vehicles, vehicles coupled to one or more trailers, flying vehicles, boats, shuttle vehicles, emergency vehicles, motorcycles, electric or motorized bicycles, aircraft, construction vehicles, underwater vehicles, drones, and / or other types of vehicles.Furthermore, the systems and methods described herein can be used for a number of purposes, including but not limited to machine control, machine locomotion, machine propulsion, synthetic data generation, model training, perception, augmented reality, virtual reality, mixed reality, robotics, security and monitoring, simulation and digital twinning, autonomous or semi-autonomous machine applications, deep learning, environmental simulation, object or actuator simulation and / or digital twinning, data center processing, conversational AI, light transport simulation (e.g., ray tracing, path tracing, etc.), collaborative content creation for 3D assets, cloud computing, and / or other suitable applications.
[0057] Disclosed embodiments may be included in a range of different systems, such as automotive systems (e.g., a control system for an autonomous or semi-autonomous machine, a perception system for an autonomous or semi-autonomous machine), systems implemented using a robot, antenna systems, media systems, boat systems, intelligent area monitoring systems, systems for performing deep learning operations, systems for performing simulation operations, systems for performing digital twinning operations, systems implemented using an edge device, systems involving one or more virtual machines (VMs), systems for performing synthetic data generation operations, systems implemented at least partially in a data center, and systems for performing conversational AI operations.Systems for performing light transport simulation, systems for performing collaborative content creation for 3D assets, systems that are implemented at least partially using cloud computing resources, and / or other types of systems. EXAMPLE CALCULATION DEVICE
[0058] Fig. Figure 6 is a block diagram of an exemplary computing device(s) 600 suitable for use in implementing some embodiments of the present disclosure. The computing device 600 may include a connection system 602 that directly or indirectly couples the following devices: memory 604, one or more central processing units (CPUs) 606, one or more graphics processing units (GPUs) 608, a communication interface 610, input / output (I / O) ports 612, input / output components 614, a power supply 616, one or more presentation components 618 (e.g., display(s)), and one or more logic units 620. In at least one embodiment, the computing device(s) 600 may include one or more virtual machines (VMs) and / or any of its components may include virtual components (e.g., virtual hardware components).For non-restrictive examples, one or more of the GPUs 608 may comprise one or more vGPUs, one or more of the CPUs 606 may comprise one or more vCPUs, and / or one or more of the logic units 620 may comprise one or more virtual logic units. As such, a computing device(s) 600 may include discrete components (e.g., a complete GPU dedicated to the computing device 600), virtual components (e.g., a portion of a GPU dedicated to the computing device 600), or a combination thereof.
[0059] Although the various blocks of Fig. Where components 6 are shown to be connected via the connection system 602 by lines, this is not intended to be restrictive and is merely for clarity. For example, in some embodiments, a presentation component 618, such as a display device, can be considered an I / O component 614 (e.g., if the display is a touchscreen). As another example, the CPUs 606 and / or GPUs 608 can include memory (e.g., the memory 604 can be representative of a storage device, in addition to the memory of the GPUs 608, the CPUs 606, and / or other components). In other words, the computing device of Fig. Figure 6 is for illustrative purposes only. No distinction is made between categories such as "workstation", "server", "laptop", "desktop", "tablet", "client device", "mobile device", "handheld device", "game console", "electronic control unit (ECU)", "virtual reality system" and / or other device or system types, as all fall under the category of computing device. Fig. 6 are considered.
[0060] The 602 interconnect system can represent one or more links or buses, such as an address bus, a data bus, a control bus, or a combination thereof. The 602 interconnect system can include one or more bus or link types, such as an Industry Standard Architecture (ISA) bus, an Extended Industry Standard Architecture (EISA) bus, a Video Electronics Standard Association (VESA) bus, a Peripheral Component Connection (PCI) bus, a Peripheral Component Connection Express (PCIe) bus, and / or other bus or link types. In some embodiments, there are direct connections between components. For example, the CPU 606 can be directly connected to the memory 604. Furthermore, the CPU 606 can be directly connected to the GPU 608. In the case of a direct or point-to-point connection between components, the 602 interconnect system can include a PCIe link to perform the connection.In these examples, the computing device 600 does not need to include a PCI bus.
[0061] Memory 604 can include any of a range of computer-readable media. The computer-readable media can be any available media accessible to the computing device 600. The computer-readable media can include both volatile and non-volatile media, and both removable and non-removable media. As an example, and not as a limitation, the computer-readable media can include computer storage media and communication media.
[0062] Computer storage media can include both volatile and non-volatile media and / or removable and non-removable media, implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, and / or other data types. For example, memory 604 can store computer-readable instructions (e.g., representing a program and / or program element, such as an operating system).Computer storage media may include, but are not limited to, RAM, ROM, EEPROM, flash memory or other storage technology, CD-ROM, Digital Versatile Discs (DVDs) or other optical disk storage, magnetic cartridges, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and that can be accessed by the Computing Device 600. As used herein, computer storage media do not, per se, include signals.
[0063] Computer storage media can embody computer-readable instructions, data structures, program modules, and / or other data types in a modulated data signal, such as a carrier wave or other transport mechanism, and include information delivery media. The term "modulated data signal" can refer to a signal for which one or more of its characteristics are set or modified in a way that encodes information in the signal. By way of example, and not limited to this, computer storage media can include wired media, such as a wired network or a directly wired connection, and wireless media, such as acoustic, RF, infrared, and other wireless media. Combinations of any of the foregoing should also be included in the scope of computer-readable media.
[0064] The CPU(s) 606 can be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 600 to perform one or more of the procedures and / or processes described herein. The CPU(s) 606 can include one or more cores (e.g., one, two, four, eight, twenty-eight, seventy-two, etc.) capable of handling a plurality of software threads concurrently. The CPU(s) 606 can include any type of processor and may include different types of processors depending on the type of computing device 600 implemented (e.g., processors with fewer cores for mobile devices and processors with more cores for servers).For example, depending on the type of computing device 600, the processor can be an Advanced RISC Machines (ARM) processor implemented using Reduced Instruction Set Computing (RISC), or an x86 processor implemented using Complex Instruction Set Computing (CISC). The computing device 600 can include one or more CPUs 606 in addition to one or more microprocessors or supplementary coprocessors, such as math coprocessors.
[0065] In addition to or as an alternative to the CPU(s) 606, the GPU(s) 608 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 600 to perform one or more of the methods and / or processes described herein. One or more of the GPU(s) 608 may be an integrated GPU (e.g., in one or more of the CPU(s) 606) and / or one or more of the GPU(s) 608 may be a discrete GPU. In embodiments, one or more of the GPU(s) 608 may be a coprocessor of one or more of the CPU(s) 606. The GPU(s) 608 may be used by the computing device 600 to render graphics (e.g., 3D graphics) or to perform general-purpose computing. For example, the GPU(s) 608 can be used for general-purpose computing on GPUs (GPGPU).The GPU(s) 608 can include hundreds or thousands of cores capable of handling hundreds or thousands of software threads simultaneously. The GPU(s) 608 can generate pixel data for output images in response to rendering commands (e.g., rendering commands from the CPU(s) 606 received via a host interface). The GPU(s) 608 can include graphics memory, such as display memory, for storing pixel data or any other suitable data, such as GPGPU data. The display memory can be included as part of the memory 604. The GPU(s) 608 can include two or more GPUs operating in parallel (e.g., via a link). The link can connect the GPUs directly (e.g., using an NVLINK) or connect them via a switch (e.g., using an NVSwitch).When combined, each GPU can generate 608 pixel data or GPGPU data for different dividers, or output for different outputs (e.g., a first GPU for a first image and a second GPU for a second image). Each GPU can include its own memory or share memory with other GPUs.
[0066] In addition to or as an alternative to the CPU(s) 606 and / or the GPU(s) 608, the logic unit(s) 620 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 600 to perform one or more of the methods and / or processes described herein. In embodiments, the CPU(s) 606, the GPU(s) 608, and / or the logic unit(s) 620 may discretely or jointly perform any combination of the methods, processes, and / or parts thereof. One or more of the logic units 620 may be part of and / or integrated into one or more of the CPU(s) 606 and / or the GPU(s) 608, and / or one or more of the logic units 620 may be discrete components or otherwise external to the CPU(s) 606 and / or the GPU(s) 608.In embodiments, one or more of the logic units 620 can be a coprocessor of one or more of the CPU(s) 606 and / or one of the GPU(s) 608.
[0067] Examples of the Logic Unit(s) 620 include: one or more processing cores and / or components thereof, such as Data Processing Units (DPUs), Tensor Cores (TCs), Tensor Processing Units (TPUs), Pixel Visual Cores (PVCs), Vision Processing Units (VPUs), Graphics Processing Clusters (GPCs), Texture Processing Clusters (TPCs), Streaming Multiprocessors (SMs), Tree Traversal Units (TTUs), Artificial Intelligence Accelerators (AIAs), Deep Learning Accelerators (DLAs), Arithmetic Logic Units (ALUs), Application-Specific Integrated Circuits (ASICs), Floating-Point Units (FPUs), Input / Output (I / O) elements, Peripheral Component Connector (PCI) or Peripheral Component Connector Express (PCIe) components and / or the like.
[0068] In various embodiments, one or more CPU(s) 606, GPU(s) 608 and / or logic unit(s) 1020 are configured to execute one or more distances of the depth estimation engine 122 and / or feature correlation engine 232.
[0069] The communication interface 610 can include one or more receivers, transmitters, and / or transceivers that enable the computing device 600 to communicate with other computing devices over an electronic communication network, including wired and / or wireless communication. The communication interface 610 can include components and functionality to enable communication over any number of different networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., communicating over Ethernet or InfiniBand), low-power wide-area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet.In one or more embodiments, the logic unit(s) 620 and / or communication interface 610 may include one or more data processing units (DPUs) to transfer data received via a network and / or the connection system 602 directly to (e.g., a memory of) one or more GPU(s) 608.
[0070] The I / O ports 612 enable the computing device 600 to be logically coupled with other devices, including the I / O components 614, the presentation component(s) 618, and / or other components, some of which may be built into (e.g., integrated with) the computing device 600. Illustrative I / O components 614 include a microphone, mouse, keyboard, joystick, gamepad, game controller, satellite table, scanner, printer, wireless device, etc. The I / O components 614 can provide a natural user interface (NUI) that processes air gestures, speech, or other physical input generated by a user. In some cases, input can be transmitted to an appropriate network element for further processing.A NUI can implement any combination of speech recognition, pen recognition, facial recognition, biometric recognition, gesture recognition both on-screen and adjacent to the screen, air gestures, head and eye tracking, and touch recognition (as described in more detail below), associated with a display of the Computing Device 600. The Computing Device 600 can include depth cameras, such as stereoscopic camera systems, infrared camera systems, RGB camera systems, touchscreen technology, and combinations thereof for gesture detection and recognition. Furthermore, the Computing Device 600 can include accelerometers or gyroscopes (e.g., as part of an inertia measurement unit, IMU) that enable motion detection.In some examples, the output from the accelerometers or gyroscopes can be used by the Computing Device 600 to render immersive augmented reality or virtual reality.
[0071] The power supply 616 can include a hardwired power supply, a battery power supply, or a combination of both. The power supply 616 can provide power to the computing device 600 to enable the components of the computing device 600 to operate.
[0072] The presentation component(s) 618 can include a display (e.g., a monitor, a touchscreen, a television screen, a heads-up display (HUD), other display types, or a combination thereof), speakers, and / or other presentation components. The presentation component(s) 618 can receive data from other components (e.g., the GPU(s) 608, the CPU(s) 606, DPUs, etc.) and output the data (e.g., as an image, video, sound, etc.). EXEMPLARY DATA CENTER
[0073] Fig. Figure 7 illustrates an exemplary data center 700 that can be used in at least one embodiment of the present disclosure. The data center 700 can include a data center infrastructure layer 710, a framework layer 720, a software layer 730, and / or an application layer 740.
[0074] As in Fig. As shown in Figure 7, the data center infrastructure layer 710 can include a resource orchestrator 712, clustered compute resources 714, and node compute resources (“node CRs”) 716(1)-716(N), where “N” is a positive integer. In at least one embodiment, node CRs 716(1)-716(N) can include, among other things, any number of central processing units (CPUs) or other processors (including DPUs, accelerators, field-programmable gate arrays (FPGAs), graphics processing units or graphics processing units (GPUs), etc.), memory devices (e.g., dynamic read-only memories), storage devices (e.g., solid-state or disk drives), network input / output (NW I / O) devices, network switches, virtual machines (VMs), power modules, and / or cooling modules, etc. In some embodiments, one or more node CRs can be controlled by the node CRs.s 716(1)-716(N) correspond to a server that has one or more of the computing resources mentioned above. Furthermore, in some embodiments, the node CRs 716(1)-716(N) may include one or more virtual components, such as vGPUs, vCPUs, and / or the like, and / or may correspond to one or more of the node CRs 716(1)-716(N) of a virtual machine (VM).
[0075] In at least one embodiment, grouped compute resources 714 can include separate groupings of node CRs 716 located in one or more racks (not shown), or many racks located in data centers at different geographic locations (also not shown). Separate groupings of node CRs 716 within grouped compute resources 714 can include grouped compute, network, storage, or memory resources that can be configured or allocated to support one or more workloads. In at least one embodiment, multiple node CRs 716, including CPUs, GPUs, DPUs, and / or other processors, can be grouped in one or more racks to provide compute resources to support one or more workloads.The one or more racks can also include any number of power modules, cooling modules and / or network switches in any combination.
[0076] The resource orchestrator 712 can configure or otherwise control one or more node CRs 716(1)-716(N) and / or grouped compute resources 714. In at least one embodiment, the resource orchestrator 712 can include a software design infrastructure (SDI) management entity for the data center 700. The resource orchestrator 712 can include hardware, software, or a combination thereof.
[0077] In at least one embodiment, as in Fig. As shown in Figure 7, a framework layer 720 can include a job scheduler 733, a configuration manager 734, a resource manager 736, and / or a distributed file system 738. The framework layer 720 can include a framework to support software 732 of software layer 730 and / or one or more applications 742 of application layer 740. The software 732 or application 742 can each include web-based service software or applications, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. The framework layer 720 can be, among other things, a type of free and open-source software web application framework, such as Apache Spark™ (hereinafter "Spark"), which can use the distributed file system 738 for large-scale data processing (e.g., "big data").In at least one embodiment, the job scheduler 733 can include a Spark driver to facilitate the scheduling of workloads supported by various layers of the data center 700. The configuration manager 734 can be capable of configuring various layers, such as the software layer 730 and the framework layer 720, including Spark and the distributed file system 738, to support high-volume data processing. The resource manager 736 can be capable of managing clustered or grouped compute resources allocated or assigned to support the distributed file system 738 and the job scheduler 733. In at least one embodiment, the clustered or grouped compute resources can include a grouped compute resource 714 at the data center infrastructure layer 710.The resource manager 736 can coordinate with the resource orchestrator 712 to manage these allocated or assigned computing resources.
[0078] In at least one embodiment, software 732, which is enclosed in software layer 730, can include software used by at least parts of the node CRs 716(1)-716(N), grouped computing resources 714, and / or distributed file system 738 of framework layer 720. One or more types of software can include, among others, internet website search software, email virus scanning software, database software, and streaming video content software.
[0079] In at least one embodiment, one or more applications 742 enclosed in the application layer 740 may include one or more types of applications used by at least parts of the node CRs 716(1)-716(N), grouped compute resources 714, and / or distributed file system 738 of the framework layer 720. One or more types of applications may include, among others, any number of a genomics application, a cognitive computing application, and a machine learning application, including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), and / or other machine learning applications used in conjunction with one or more embodiments.
[0080] In at least one embodiment, any configuration manager 734, resource manager 736, and resource orchestrator 712 can implement any number and any type of self-modifying operations based on any set and any type of data captured in any technically feasible manner. Self-modifying operations can relieve a data center operator of the data center 700 of potentially making poor configuration decisions and potentially avoiding underutilized and / or poorly performing parts of a data center.
[0081] The Data Center 700 may include tools, services, software, or other resources to train one or more machine learning models or to predict or infer information using one or more machine learning models according to one or more embodiments described herein. For example, one or more machine learning models may be trained by calculating weight parameters according to a neural network architecture using software and / or computational resources as described above in relation to the Data Center 700.In at least one embodiment, trained or deployed machine learning models corresponding to one or more neural networks can be used to infer and predict information regarding the data center 700 using the resources described above, by using weight parameters calculated via one or more training techniques, such as, but not limited to, those described above.
[0082] In at least one embodiment, the data center can use 700 CPUs, application-specific integrated circuits (ASICs), GPUs, FPGAs, and / or other hardware (or equivalent virtual computing resources) to perform training and / or inference using the resources described above. Additionally, one or more of the software and / or hardware resources described above can be configured as a service to allow users to train or perform information inference, such as image recognition, speech recognition, or other artificial intelligence services. EXEMPLARY NETWORK ENVIRONMENTS
[0083] Network environments suitable for use in implementing embodiments of the disclosure may include one or more client devices, servers, network-attached storage (NAS), other backend devices, and / or other device types. The client devices, servers, and / or other device types (e.g., each device) may be connected to one or more instances of the computing device(s) 600 of Fig. 6. For example, each device can include similar components, features, and / or functionality to the computing device(s) 600. Furthermore, if backend devices (e.g., servers, NAS, etc.) are implemented, the backend devices can be implemented as part of a data center 700; an example of this is described in more detail herein with reference to Fig. 7 described.
[0084] Components in a network environment can communicate with each other over a network, which can be wired, wireless, or both. The network can include multiple networks or a network of networks. For example, the network can include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks, such as the internet, and / or a public switched telephone network (PSTN), and / or one or more private networks. If the network includes a wireless telecommunications network, components such as a base station, a communications tower, or even access points (as well as other components) can provide wireless connectivity.
[0085] Compatible network environments can include one or more peer-to-peer network environments—in which case no server may be included in a network environment—and one or more client-server network environments—in which case one or more servers may be included in a network environment. In peer-to-peer network environments, functionality described herein with respect to one or more servers can be implemented on any number of client devices.
[0086] In at least one embodiment, a network environment can include one or more cloud-based network environments, a distributed computing environment, a combination thereof, etc. A cloud-based network environment can include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more servers, which may include one or more core network servers and / or edge servers. A framework layer can include a framework to support software of a software layer and / or one or more application(s) of an application layer. The software or application(s) may each include web-based service software or applications. In embodiments, one or more of the client devices can use the web-based service software or applications (e.g.,by accessing the service software and / or applications via one or more application programming interfaces (APIs). The framework layer can be, among other things, a type of free and open-source software web application framework that can use a distributed file system for large-volume data processing (e.g., "big data").
[0087] A cloud-based network environment can provide cloud computing and / or cloud storage, thereby performing any combination of computing and / or data storage functions as described herein (or one or more parts thereof). Any of these various functions can be distributed across multiple locations of central or core servers (e.g., one or more data centers that may be distributed across a state, region, country, the Earth, etc.). If a connection to a user (e.g., a client device) is relatively close to an edge server, a core server may designate at least some of the functionality for that edge server(s). A cloud-based network environment can be private (e.g., restricted to a single organization), public (e.g., available to many organizations), and / or a combination thereof (e.g., a hybrid cloud environment).
[0088] The client device(s) may include at least some of the components, features, and functionality described herein with reference to Fig. The 6 exemplary computing device(s) described include 600. By way of example, and not limited to, a client device may be a personal computer (PC), laptop computer, mobile device, smartphone, tablet computer, smartwatch, wearable computer, personal digital assistant (PDA), MP3 player, virtual reality headset, global positioning system (GPS) or global positioning device, video player, video camera, surveillance device or system, vehicle, boat, flying vehicle, virtual machine, drone, robot, handheld communication device, hospital device, gaming device or system, entertainment system, vehicle computer system, embedded system controller, remote control, device, consumer electronics device, workstation, edge device, any combination of these described devices, or any other suitable device. EXEMPLARY AUTONOMOUS VEHICLE
[0089] Fig. Figure 8A is an illustration of an exemplary autonomous vehicle 800 according to some embodiments of the present disclosure. The autonomous vehicle 800 (alternatively referred to herein as the “vehicle 800”) may include, among others, a passenger vehicle such as a car, truck, bus, rescue vehicle, shuttle, electric or motorized bicycle, motorcycle, fire engine, police vehicle, ambulance, boat, construction vehicle, underwater vehicle, robotic vehicle, drone, aircraft, a vehicle coupled to a trailer (e.g., a semi-trailer truck used for carrying cargo), and / or another type of vehicle (e.g., one that is unmanned and / or that carries one or more passengers).Autonomous vehicles are generally described in terms of automation levels, as defined by the National Highway Traffic Safety Administration (NHTSA), a division of the U.S. Department of Transportation, and the Society of Automotive Engineers (SAE) in their "Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles" (Standard No. J3016-201806, published June 15, 2018; Standard No. J3016-201609, published September 30, 2016; and earlier and future versions of this standard). The Vehicle 800 may be capable of functionality corresponding to one or more of Levels 3 through 5 of autonomous driving levels. The Vehicle 800 may be capable of functionality corresponding to one or more of Levels 1 through 5 of autonomous driving levels.The Vehicle 800, for example, may be capable of driver assistance (Level 1), partial automation (Level 2), conditional automation (Level 3), high automation (Level 4), and / or full automation (Level 5), depending on the configuration. The term "autonomous," as used herein, may include any and / or all types of autonomy for the Vehicle 800 or any other machine, such as fully autonomous, highly autonomous, conditionally autonomous, partially autonomous, providing assistive autonomy, semi-autonomous, mainly autonomous, or any other designation.
[0090] The vehicle 800 can include components such as a chassis, body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other vehicle components. The vehicle 800 can include a propulsion system 850, such as an internal combustion engine, a hybrid-electric drive unit, a fully electric motor, and / or another type of propulsion system. The propulsion system 850 can be connected to a drivetrain of the vehicle 800, which may include a transmission to enable the propulsion of the vehicle 800. The propulsion system 850 can be controlled in response to receiving signals from the throttle / accelerator device 852.
[0091] A steering system 854, which may include a steering wheel, can be used to steer the vehicle 800 (e.g., along a desired path or route) when the drive system 850 is in operation (e.g., when the vehicle is in motion). The steering system 854 can receive signals from a steering actuator 856. The steering wheel may be optional for full automation functionality (Level 5).
[0092] The brake sensor system 846 can be used to operate the vehicle brakes in response to receiving signals from the brake actuators 848 and / or brake sensors.
[0093] One or more Controller 836, which contains one or more CPU(s), System-on-Chips (“SoCs”) 804 ( Fig. The controller(s) 836 may include one or more GPUs (e.g., 8C) and / or one or more GPUs, and may provide signals (e.g., representing commands) for one or more components and / or systems of the vehicle 800. For example, the controller(s) may send signals to operate the vehicle brakes via one or more brake actuators 848, to operate the steering system 854 via one or more steering actuators 856, and / or to operate the propulsion system 850 via throttle / accelerator device 852. The controller(s) 836 may include one or more (e.g., integrated) onboard computing devices (e.g., supercomputers) that process sensor signals and issue operational commands (e.g., signals representing commands) to enable autonomous driving and / or to assist a human driver in driving the vehicle 800.The controller(s) 836 may include a first controller 836 for autonomous driving functions, a second controller 836 for functional safety functions, a third controller 836 for artificial intelligence functionality (e.g., computer vision), a fourth controller 836 for infotainment functionality, a fifth controller 836 for redundancy in emergency situations, and / or other controllers. In some embodiments, a single controller 836 may handle two or more of the above functionalities, two or more controllers 836 may handle a single functionality, and / or any combination thereof.
[0094] The controller(s) 836 can provide signals to control one or more components and / or one or more systems of the vehicle 800 in response to sensor data received from one or more sensors (e.g. sensor inputs). The sensor data can be received, for example, by the following: Global Navigation Satellite System (GNSS) sensors 858 (e.g., global positioning system sensors), radar sensors 860, ultrasonic sensors 862, lidar sensors 864, inertial measurement unit (IMU) sensors 866 (e.g., accelerometers, gyroscopes, magnetic compasses, magnetometers, etc.), microphones 896, stereo cameras 868, wide-angle cameras 870 (e.g., fisheye cameras), infrared cameras 872, surround-view cameras 874 (e.g., 360-degree cameras), long-range and / or medium-range cameras 898, velocity sensors 844 (e.g.,for measuring the speed of the vehicle 800), vibration sensors 842, steering sensors 840, brake sensors (e.g. as part of the brake sensor system 846) and / or other sensor types.
[0095] One or more of the controllers 836 can receive inputs (e.g., represented by input data) from an instrument cluster 832 of the vehicle 800 and provide outputs (e.g., represented by output data, display data, etc.) via a human-machine interface (HMI) display 834, an acoustic alarm, a loudspeaker, and / or via other components of the vehicle 800. The outputs can include information such as vehicle speed, velocity, time, map data (e.g., the high-definition ("HD") map 822 of Fig. 8C), location data (e.g., the location data of vehicle 800, as shown on a map), direction, location of other vehicles (e.g., an occupancy grid), information about objects and the status of objects as perceived by the controller(s) 836, etc. The HMI display 834 can, for example, show information about the presence of one or more objects (e.g., a road sign, caution sign, traffic light sequence, etc.) and / or information about driving maneuvers that the vehicle has made, is making, or will make (e.g., changing lanes now, exit 34B in two miles, etc.).
[0096] The vehicle 800 may further include a network interface 824, which can use one or more wireless antenna(s) 826 and / or one or more modem(s) to communicate over one or more networks. For example, the network interface 824 may be capable of communication over Long-Term Evolution (“LTE”), Wideband Code Division Multiple Access (“WCDMA”), Universal Mobile Telecommunications System (“UMTS”), Global System for Mobile Communications (“GSM”), IMT CDMA Multi-Carrier (“CDMA2000”), etc. The wireless antenna(s) 826 can also enable communication between objects in the environment (e.g., vehicles, mobile devices, etc.) using local area network(s), such as Bluetooth, Bluetooth Low Energy (“LE”), Z-Wave, ZigBee, etc., and / or for a low-power wide-area network (“LPWANs”), such as LoRaWAN, SigFox, etc.
[0097] Fig. 8B is an example of camera locations and fields of view for the exemplary autonomous vehicle 800 from Fig. 8A according to some embodiments of the present disclosure. The cameras and their respective fields of view are an exemplary embodiment and are not intended to be limiting. For example, additional and / or alternative cameras may be included and / or the cameras may be located at different locations on the vehicle 800.
[0098] The camera types may include, but are not limited to, digital cameras adapted for use with the components and / or systems of the Vehicle 800. The camera(s) may operate at Automotive Safety Integrity Level (ASIL) B and / or another ASIL. Depending on the specific configuration, the camera types may be capable of any frame rate, such as 60 frames per second (fps), 120 fps, 240 fps, etc. The cameras may be capable of using rolling shutters, global shutters, another type of shutter, or a combination thereof.In some embodiments, the color filter array may include a red-clear-clear-clear (RCCC) color filter array, a red-clear-clear-blue (RCCB) color filter array, a red-blue-green clear (RBGC) color filter array, a Foveon X3 color filter array, a Bayer sensor (RGGB) color filter array, a monochrome sensor color filter array, and / or another type of color filter array. In some embodiments, clear-pixel cameras, such as cameras with an RCCC, RCCB, and / or RBGC color filter array, may be used in an effort to increase light sensitivity.
[0099] In some examples, one or more of the cameras can be used to perform advanced driver assistance system (ADAS) functions (e.g., as part of a redundant or fail-safe system). For example, a multi-function monocular camera can be installed to provide functions such as lane departure warning, traffic sign recognition, and intelligent headlight control. One or more of the cameras (e.g., all of the cameras) can simultaneously record and provide image data (e.g., video).
[0100] One or more of the cameras can be mounted in a custom-designed (three-dimensionally printed) assembly to eliminate stray light and reflections from inside the car (e.g., reflections off the dashboard in the windshield mirrors) that could impair the camera's image capture capabilities. Regarding side mirror mounting assemblies, the assemblies can be custom 3D printed so that the camera mounting plate matches the shape of the side mirror. In some examples, the camera(s) can be integrated into the side mirror. For side-view cameras, the camera(s) can also be integrated within four pillars at each corner of the cabin.
[0101] Cameras with a field of view that includes parts of the environment in front of the vehicle (e.g., forward-facing cameras) can be used for surround view to identify ahead routes and obstacles and, with the aid of one or more controllers and / or control SoCs, to assist in providing information critical for generating an occupancy grid and / or determining preferred vehicle routes. Forward-facing cameras can be used to perform many of the same ADAS functions as LiDAR, including emergency braking, pedestrian detection, and collision avoidance. Forward-facing cameras can also be used for ADAS functions and systems, including lane departure warnings (LDW), autonomous cruise control (ACC), and / or other functions such as traffic sign recognition.
[0102] Various cameras can be used in a forward-facing configuration, for example, including a monocular camera platform that incorporates a CMOS (Complementary Metal Oxide Semiconductor) color imager. Another example could be one or more wide-angle cameras (870) that can be used to detect objects entering the field of view from the periphery (e.g., pedestrians, crossing traffic, or bicycles). Although in Fig. Since only one wide-angle camera is illustrated in Figure 8B, there can be any number (including zero) of wide-angle cameras 870 on the vehicle 800. Furthermore, any number of remote cameras 898 (e.g., a pair of long-range stereo cameras) can be used for depth-based object detection, especially for objects for which a neural network has not yet been trained. The remote camera(s) 898 can also be used for object detection and classification, as well as simple object tracking.
[0103] Any number of stereo cameras 868 can also be included in a forward-facing configuration. In at least one embodiment, one or more stereo camera(s) 868 can include an integrated control unit comprising a scalable processing unit that can provide programmable logic (“FPGA”) and a multi-core microprocessor with an integrated controller area network (“CAN”) or Ethernet interface on a single chip. Such a unit can be used to generate a 3D map of the vehicle's surroundings, including a distance estimate for all points in the image.An alternative stereo camera(s) 868 may include a compact stereo vision sensor, which may include two camera lenses (one on the left and one on the right) and an image processing chip, with which the distance from the vehicle to the target object can be measured and the generated information (e.g., metadata) can be used to activate the autonomous emergency braking and lane departure warning functions. Other types of stereo camera(s) 868 may be used in addition to or as an alternative to those described herein.
[0104] Cameras with a field of view that includes parts of the surroundings at the side of the vehicle 800 (e.g., side-view cameras) can be used for surround view, providing information used to generate and update the occupancy grid and to generate side-impact collision warnings. For example, one or more surround camera(s) 874 (e.g., four surround cameras 874 as in Fig. (Figure 8B illustrates) is positioned on the vehicle 800. The surround view camera(s) 874 can include one or more wide-angle cameras 870, fisheye cameras, 360-degree cameras, and / or the like. For example, four fisheye cameras can be positioned at the front, rear, and sides of the vehicle. In an alternative arrangement, the vehicle can use three surround view cameras 874 (e.g., left, right, and rear) and can use one or more other cameras (e.g., a forward-facing camera) as a fourth surround view camera.
[0105] Cameras with a field of view that includes parts of the surroundings at the rear of the vehicle 800 (e.g., rear-view cameras) can be used for parking assistance, surround view, rear collision warnings, and generating and updating the occupancy grid. A wide range of cameras can be used, including, but not limited to, cameras that are also suitable as forward-facing cameras (e.g., long-range and / or medium-range cameras 898, stereo cameras 868, infrared cameras 872, etc.), as described herein.
[0106] Fig. 8C is a block diagram of an exemplary system architecture for the exemplary autonomous vehicle 800 from Fig. 8A according to some embodiments of the present disclosure. It should be understood that these and other arrangements described herein are presented only as examples. Other arrangements and elements (e.g., machines, interfaces, functions, orders, groupings of functions, etc.) may be used in addition to or instead of those shown, and other elements may be omitted entirely. Furthermore, many of the elements described herein are functional entities that may be implemented as discrete or distributed components, or in conjunction with other components, and in any suitable combination and position. Various functions described herein as being performed by entities may be executed by hardware, firmware, and / or software. For example, various functions may be performed by a processor executing instructions stored in memory.
[0107] The components, features and systems of the 800 vehicle in Fig. 8C are each illustrated as connected via bus 802. Bus 802 can include a Controller Area Network (CAN) data interface (alternatively referred to herein as a CAN bus). A CAN can be a network within the vehicle 800 that is used to support the control of various features and functionality of the vehicle 800, such as actuation of the brakes, acceleration, braking, steering, windshield wipers, etc. A CAN bus can be configured to have dozens or even hundreds of nodes, each with its own unique identifier (e.g., a CAN ID). The CAN bus can be read to obtain steering wheel angle, ground speed, engine revolutions per minute (RPM), button positions, and / or other vehicle status indicators. The CAN bus can be ASIL-B compliant.
[0108] Although described herein as a CAN bus, this is not intended to be restrictive. For example, FlexRay and / or Ethernet may be used in addition to or as an alternative to the CAN bus. Furthermore, although a single line is used to represent the 802 bus, this is not intended to be restrictive. For example, there may be any number of 802 buses, which may include one or more CAN buses, one or more FlexRay buses, one or more Ethernet buses, and / or one or more other types of buses using a different protocol. In some examples, two or more 802 buses may be used to perform different functions and / or for redundancy. For example, a first 802 bus may be used for collision avoidance functionality, and a second 802 bus may be used for actuation control.In any given example, the 802 bus can communicate with any of the components of the vehicle 800, and two or more 802 buses can communicate with the same components. In some examples, each SoC 804, each controller 836, and / or each computer within the vehicle can have access to the same input data (e.g., inputs from sensors of the vehicle 800) and be connected to a common bus, such as the CAN bus.
[0109] The vehicle 800 can include one or more controllers 836, such as those mentioned herein in relation to Fig. 8A. The 836 controller(s) can be used for a variety of functions. The 836 controller(s) can be coupled with any of the various other components and systems of the 800 vehicle and used for controlling the 800 vehicle, the 800 vehicle's artificial intelligence, the 800 vehicle's infotainment system, and / or the like.
[0110] The Vehicle 800 can include one or more System-on-a-Chip (SoC) 804. The SoC 804 can include one or more CPUs 806, GPUs 808, one or more processors 810, caches 812, accelerators 814, data storage 816, and / or other components and features not illustrated. The SoC(s) 804 can be used to control the Vehicle 800 in a variety of platforms and systems. For example, the SoC(s) 804 in a system (e.g., the Vehicle 800 system) can be combined with an HD card 822, allowing map refreshes and / or updates to be received via a network interface 824 from one or more servers (e.g., the server(s) 878). Fig. 8D) can be obtained.
[0111] The CPU(s) 806 may include a CPU cluster or CPU complex (alternatively referred to herein as "CCPLEX"). The CPU(s) 806 may include multiple cores and / or L2 caches. For example, in some embodiments, the CPU(s) 806 may include eight cores in a coherent multiprocessor configuration. In some embodiments, the CPU(s) 806 may include four dual-core clusters, each cluster having a dedicated L2 cache (e.g., a 2 MB L2 cache). The CPU(s) 806 (e.g., the CCPLEX) may be configured to support concurrent cluster operation, allowing a combination of CPU(s) 806 clusters to be active at any given time.
[0112] The 806 CPU(s) can implement power management capabilities that include one or more of the following features: individual hardware blocks can be automatically clock-gated (automatically disconnected from the clock signal in a gate) when idle to save dynamic power; each core clock can be disconnected when the core is not actively executing instructions due to the execution of WFI / WFE instructions; each core can be independently power-gated; each core cluster can be independently clock-gated when all cores are clock-gated or power-gated; and / or each core cluster can be independently power-gated when all cores are power-gated.The CPU(s) 806 can further implement an improved power state management algorithm, specifying the permissible power states and expected wake-up times, and the hardware / microcode for the core, cluster, and CCPLEX determining the best power state to enter. The processing cores can support simplified power state entry sequences in software, offloading the work to the microcode.
[0113] The GPU(s) 808 may include an integrated GPU (alternatively referred to herein as the "iGPU"). The GPU(s) 808 may be programmable and efficient for parallel workloads. The GPU(s) 808 may, in some examples, use an enhanced tensor instruction set. The GPU(s) 808 may include one or more streaming microprocessors, each of which may include an L1 cache (for example, an L1 cache with a minimum storage capacity of 96 KB), and two or more of the streaming microprocessors may share an L2 cache (for example, an L2 cache with a storage capacity of 512 KB). In some embodiments, the GPU(s) 808 may include at least eight streaming microprocessors. The GPU(s) 808 can use Compute Application Programming Interface(s) (API(s)). Furthermore, the GPU(s) 808 can use one or more parallel computing platforms and / or programming models (e.g.,NVIDIA's CUDA).
[0114] The GPU(s) 808 can be performance-optimized to achieve the best performance in automotive and embedded applications. For example, the GPU(s) 808 can be manufactured using a FinFET field-effect transistor. However, this is not intended to be a limitation, and the GPU(s) 808 can be manufactured using other semiconductor fabrication processes. Each streaming microprocessor can contain a number of mixed-precision processing cores, which are divided into multiple blocks. For example, 64 PF32 cores and 32 PF64 cores can be divided into four processing blocks, among others. In such an example, each processing block could be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two NVIDIA TENSOR CORES for mixed-precision deep learning matrix arithmetic, an L0 instruction cache, a warp scheduler, a dispatch unit and / or a 64 KB register file.Furthermore, streaming microprocessors can include independent parallel integer and floating-point data paths to enable efficient execution of workloads with a mix of computational and addressing operations. Streaming microprocessors can include independent thread scheduling functions to allow for finer-grained synchronization and collaboration between parallel threads. Streaming microprocessors can also include a combined L1 data cache and shared memory to improve performance while simplifying programming.
[0115] The GPU(s) 808 can include high-bandwidth memory (HBM) and / or a 16 GB HBM2 memory subsystem to provide a peak memory bandwidth of approximately 900 GB / second in some examples. In some examples, synchronous graphics random access memory (SGRAM), such as double-data-rate type five synchronous graphics random access memory (GDDR5), can be used in addition to or as an alternative to HBM memory.
[0116] The GPU(s) 808 can include unified memory technology, including access counters, to allow more accurate migration of memory pages to the processor that accesses them most frequently, thereby improving efficiency for memory areas shared by processors. In some examples, support for Address Translation Services (ATS) can be used to allow the GPU(s) 808 to directly access the page tables of the CPU(s) 806. In such examples, an address translation request can be sent to the CPU(s) 806 if the memory management unit (MMU) of the GPU(s) 808 experiences an erroneous access. In response, the CPU(s) 806 can search its page tables for the virtual-to-physical mapping for the address and send the translation back to the GPU(s) 808.As such, Unified Memory technology can allow a single, unified virtual address space for the memory of both the CPU(s) 806 and the GPU(s) 808, thereby simplifying the programming of the GPU(s) 808 and the porting of applications to the GPU(s) 808.
[0117] Furthermore, GPU(s) 808 can include an access counter that tracks the frequency of GPU(s) accessing the memory of other processors. This access counter can help ensure that memory pages are moved to the physical memory of the processor that accesses them most frequently.
[0118] The SoC(s) 804 can include any number of Cache(s) 812, including those described herein. For example, the Cache(s) 812 can include an L3 cache available to both the CPU(s) 806 and the GPU(s) 808 (e.g., connected to both the CPU(s) 806 and the GPU(s) 808). The Cache(s) 812 can include a write-back cache capable of tracking row states, e.g., using a cache coherence protocol (e.g., MEI, MESI, MSI, etc.). The L3 cache can include 4 MB or more, although smaller cache sizes can be used.
[0119] The SoC(s) 804 can include one or more ALUs (Arithmetic Logic Units) that can be used to perform processing with respect to any of the tasks or operations of the Vehicle 800, such as processing DNNs. In addition, the SoC(s) 804 can include one or more Floating-Point Units (FPUs) – or other mathematical or numeric coprocessor types – for performing mathematical operations within the system. For example, the SoC(s) 104 can include one or more FPUs that are integrated as execution units within one or more CPU(s) 806 and / or GPU(s) 808.
[0120] The SoC(s) 804 can include one or more Accelerators 814 (e.g., hardware accelerators, software accelerators, or a combination thereof). For example, the SoC(s) 804 can include a hardware acceleration cluster, which may include optimized hardware accelerators and / or a large amount of on-chip memory. The large on-chip memory (e.g., 4 MB SRAM) can enable the hardware acceleration cluster to accelerate neural networks and other computations. The hardware acceleration cluster can be used to complement the GPU(s) 808 and offload some tasks from the GPU(s) 808 (e.g., to free up more GPU cycles for other tasks). As an example, the Accelerator 814 can be used for targeted workloads (e.g. perception, convolutional neural networks (CNNs) etc.).), which are stable enough to be amenable to acceleration. The term “CNN”, as used herein, can include all types of CNNs, including region-based or regional convolutional neural networks (RCNNs) and fast RCNNs (e.g., such as those used for object detection).
[0121] The Accelerator 814 (e.g., the hardware acceleration cluster) can include one or more deep learning accelerators (DLAs). The DLA(s) can include one or more tensor processing units (TPUs), which can be configured to provide an additional ten trillion operations per second for deep learning applications and inference. The TPUs can be accelerators configured and optimized to perform image processing functions (e.g., for CNNs, RCNNs, etc.). The DLA(s) can further be optimized for a specific set of neural network types and floating-point operations, as well as for inference. The design of the DLA(s) can deliver more performance per millimeter than a general-purpose GPU and far surpasses the performance of a CPU.The TPU(s) can perform several functions, including a single-instance folding function that supports, for example, INT8, INT16 and FP16 data types for both features and weights, as well as post-processor functions.
[0122] The DLA(s) can quickly and efficiently execute neural networks, especially CNNs, on processed or unprocessed data for any variety of functions, including, but not limited to: a CNN for object identification and detection using camera sensor data; a CNN for distance estimation using camera sensor data; a CNN for emergency vehicle detection and identification using microphone data; a CNN for facial recognition and vehicle owner identification using camera sensor data; and / or a CNN for security-related events.
[0123] The DLA(s) can perform any function of the GPU(s) 808, and using an inference accelerator, a designer can, for example, use either the DLA(s) or the GPU(s) 808 for any given function. A designer can, for instance, focus the processing of CNNs and floating-point operations on the DLA(s) and leave other functions to the GPU(s) 808 and / or another accelerator 814.
[0124] The Accelerator 814 (e.g., the hardware acceleration cluster) can include one or more programmable vision accelerators (PVAs), which may alternatively be referred to herein as computer vision accelerators. The PVA(s) can be designed and configured to accelerate computer vision algorithms for advanced driver assistance systems (ADAS), autonomous driving applications, augmented reality (AR), and / or virtual reality (VR). The PVA(s) can provide a balance between performance and flexibility. For example, each PVA can include, among other things, any number of reduced instruction set compute cores (RISC), direct memory access (DMA), and / or any number of vector processors.
[0125] The RISC cores can interact with image sensors (e.g., the image sensors of any of the cameras described herein), image signal processor(s), and / or the like. Each RISC core can include any memory location. The RISC cores can use any of a number of protocols, depending on the implementation. In some examples, the RISC cores can run a real-time operating system (RTOS). The RISC cores can be implemented using one or more integrated circuit devices, application-specific integrated circuits (ASICs), and / or memory devices. For example, the RISC cores can include an instruction cache and / or tightly coupled RAM.
[0126] The DMA can enable components of the PVA(s) to access system memory independently of the CPU(s) 806. The DMA can support any number of features provided to optimize the PVA, including, but not limited to, support for multidimensional addressing and / or circular addressing. In some examples, the DMA can support up to six or more addressing dimensions, which may include block width, block height, block depth, horizontal block gradation, vertical block gradation, and / or depth gradation.
[0127] The vector processors can be programmable processors designed to efficiently and flexibly execute computer vision algorithms and provide signal processing capabilities. In some examples, the PVA may include a PVA core and two vector processing subsystem partitions. The PVA core may include a processor subsystem, a DMA engine (e.g., two DMA engines), and / or other peripherals. The vector processing subsystem may act as the primary processing engine of the PVA and may include a vector processing unit (VPU), an instruction cache, and / or vector memory (e.g., VMEM). A VPU core may include a digital signal processor, such as a single-instruction multiple data (SIMD) or very-long instruction word (VLIW) signal processor. The combination of SIMD and VLIW can improve throughput and speed.
[0128] Each vector processor can include an instruction cache and be coupled to dedicated memory. Consequently, in some examples, each vector processor can be configured to operate independently of the others. In other examples, vector processors enclosed in a given PVA can be configured to utilize data parallelism. For example, in some embodiments, the multitude of vector processors enclosed in a single PVA can execute the same computer vision algorithm, but in different regions of an image. In other examples, the vector processors enclosed in a given PVA can simultaneously execute different computer vision algorithms on the same image, or even different algorithms on sequential images or portions of an image.Among other things, any number of PVAs can be included in a hardware acceleration cluster, and any number of vector processors can be included in each of the PVAs. Furthermore, the PVA(s) can include additional ECC (Error Correction Code) memory to improve overall system security.
[0129] The Accelerator 814 (e.g., the hardware acceleration cluster) can include a computer vision network on a single chip and SRAM to provide high-bandwidth, low-latency SRAM for the Accelerator 814. In some examples, the on-chip memory can include at least 4 MB of SRAM, consisting, for example, of eight field-configurable memory blocks accessible by both the PVA and the DLA. Each pair of memory blocks can include an Advanced Peripheral Bus (APB) interface, a configuration circuit, a controller, and a multiplexer. Any type of memory can be used. The PVA and DLA can access the memory via a backbone, providing high-speed memory access for both the PVA and DLA.The backbone can include a computer vision network on a chip that connects the PVA and DLA to the memory (e.g., using the APB).
[0130] The computer vision network can include an interface on a single chip that, prior to the transmission of control signals / addresses / data, ensures that both the PVA and the DLA provide ready-to-use and valid signals. Such an interface can provide separate phases and channels for the transmission of control signals / addresses / data, as well as burst communication for continuous data transfer. This type of interface can conform to ISO 26262 or IEC 61508 standards, although other standards and protocols can also be used.
[0131] In some examples, the SoC(s) 804 may include a real-time ray-tracing hardware accelerator, as described in US Patent Application No. 16 / 101,232, filed on August 10, 2018. The real-time ray-tracing hardware accelerator can be used for fast and efficient determination of the positions and extents of objects (e.g., within a world model), for generating real-time visualization simulations, for interpreting radar signals, for synthesizing and / or analyzing sound propagation, for simulating sonar systems, for general wave propagation simulation, for comparison with lidar data for localization purposes, and / or for other functions and / or other purposes. In some embodiments, one or more Tree Traversal Units (TTUs) can be used to perform one or more ray-tracing-related operations.
[0132] The Accelerator 814 (e.g., the hardware accelerator cluster) has a wide range of applications in autonomous driving. The PVA can be a programmable vision accelerator used for key processing stages in ADAS and autonomous vehicles. The PVA's capabilities are well-suited for algorithmic domains requiring predictable processing with low power consumption and low latency. In other words, the PVA performs well with semi-dense or dense regular computations, even with small datasets, that require predictable runtimes with low latency and low power consumption. In the context of autonomous vehicle platforms, PVAs are therefore designed to execute classic computer vision algorithms, as they are efficient at object detection and operate with integer mathematics.
[0133] For example, according to one embodiment of the technology, the PVA is used to perform computer stereo vision. A semi-global matching-based algorithm can be used in some examples, although this is not intended to be a limiting factor. Many Level 3-5 autonomous driving applications require motion estimation / stereo matching during operation (on-the-fly) (e.g., structure from motion, pedestrian detection, lane detection, etc.). The PVA can perform computer stereo vision functions with input from two monocular cameras.
[0134] In some examples, the PVA can be used to perform dense optical flow processing, such as processing raw radar data (e.g., using 4D Fast Fourier Transform) to produce processed radar data. In other examples, the PVA is used for time-of-flight depth processing, by processing the raw time-of-flight data to produce, for example, processed time-of-flight data.
[0135] The DLA can be used to operate any type of network to improve steering and driving safety, for example, a neural network that outputs a confidence score for each object detection. Such a confidence score can be interpreted as a probability or as providing a relative "weight" to each detection compared to other detections. This confidence score allows the system to make further decisions about which detections should be considered true positives and not false positives. For example, the system can set a confidence threshold and only consider detections that exceed this threshold as true positives.In an automatic emergency braking (AEB) system, false positive detections would cause the vehicle to automatically initiate emergency braking, which is obviously undesirable. Therefore, only the most confident detections should be considered as triggers for AEB. The DLA can operate a neural network to regress the confidence level. The neural network can use as input at least a subset of parameters, such as the bounding frame dimensions, the ground plane estimate obtained (e.g. from another subsystem), the output of an inertial measurement unit (IMU) sensor 866 correlated with the vehicle orientation 800, the distance, 3D position estimates of the object obtained from the neural network and / or other sensors (e.g., LIDAR sensor(s) 864 or RADAR sensor(s) 860).
[0136] The SoC(s) 804 may include one or more data stores 816 (e.g., memory). The data store 816 may be on-chip memory of the SoC(s) 804 capable of storing neural networks to be executed on the GPU and / or the DLA. In some examples, the data store 816 may be large enough to store multiple instances of neural networks for redundancy and safety. The data store 812 may include L2 or L3 cache(s) 812. Reference to the data store 816 may include reference to the memory associated with the PVA, DLA, and / or other accelerator(s) 814, as described herein.
[0137] The SoC(s) 804 can include one or more Processor(s) 810 (e.g., embedded processors). The Processor(s) 810 can include a Boot and Power Management Processor, which may be a dedicated processor and subsystem that handles boot power and management functions and associated security enforcement. The Boot and Power Management Processor can be part of the SoC(s) 804 boot sequence and provide runtime power management services. The Boot and Power Management Processor can provide clock and voltage programming, support for low-power system state transitions, management of the SoC(s) 804's thermal and temperature sensors, and / or management of the SoC(s) 804's power supply states.Each temperature sensor can be implemented as a ring oscillator whose output frequency is proportional to the temperature, and the SoC(s) 804 can use ring oscillators to detect temperatures of the CPU(s) 806, GPU(s) 808, and / or accelerator(s) 814. If it is determined that the temperatures exceed a threshold, the boot and power management processor can enter a temperature fault routine and put the SoC(s) 804 into a lower power consumption state and / or put the vehicle 800 into a chauffeur-to-safe-stop mode (e.g., bring the vehicle 800 to a safe stop).
[0138] The 810 processor(s) may also include a set of embedded processors that can serve as an audio processing engine. The audio processing engine may be an audio subsystem providing full hardware support for multi-channel audio across multiple interfaces and a wide and flexible range of audio I / O interfaces. In some examples, the audio processing engine is a dedicated processor core with a digital signal processor and dedicated RAM.
[0139] The 810 processor(s) may further include an always-on processor engine that can provide the necessary hardware features to support the management of low-power sensors and wake-up use cases. The always-on processor engine may include a processor core, tightly coupled RAM, supporting peripherals (such as timers and interrupt controllers), various I / O controller peripherals, and routing logic.
[0140] The 810 processor(s) can further include a security cluster engine, which comprises a dedicated processor subsystem for handling security management for automotive applications. The security cluster engine can include two or more processor cores, tightly coupled RAM, supporting peripherals (such as timers, an interrupt controller, etc.), and / or routing logic. In a security mode, the two or more cores can operate in lockstep mode, functioning as a single core with comparison logic to detect differences between their operations.
[0141] The 810 processor(s) may also include a real-time camera engine, which may include a dedicated processor subsystem for handling real-time camera management.
[0142] The 810 processor(s) may also include a high dynamic range signal processor, which may include an image signal processor, which is a hardware engine that is part of the camera processing pipeline.
[0143] The 810 processor(s) can include a video image compositor, which may be a processing block (e.g., implemented on a microprocessor) that implements video post-processing functions required by a video playback application to produce the final image for the playback window. The video image compositor can perform lens distortion correction on the 870 wide-angle camera(s), the 874 surround-view camera(s), and / or the in-cabin surveillance camera sensors. An in-cabin surveillance camera sensor is preferably monitored by a neural network running on a separate instance of the Advanced SoC, configured to identify and respond to in-cabin events.An in-cabin system can perform lip reading to activate cellular service and make phone calls, dictate emails, change the vehicle's destination, activate or change the infotainment system and vehicle settings, or provide voice-activated web browsing. Certain functions are only available to the driver when the vehicle is operating in autonomous mode and are otherwise disabled.
[0144] The video image compositor can include enhanced temporal noise reduction for both spatial and temporal noise reduction. For example, if motion occurs in a video, the noise reduction can weight spatial information accordingly, thereby reducing the impact of information provided by adjacent frames. If an image or part of an image does not contain motion, the temporal noise reduction performed by the video image compositor can use information from the previous image to reduce noise in the current image.
[0145] The video image compositor can also be configured to perform stereo equalization on the input stereo lens frames. Furthermore, the video image compositor can be used for user interface composition when the operating system desktop is in use and the GPU(s) 808 are not required to continuously render new surfaces. The video image compositor can also be used to offload the GPU(s) 808, even when the GPU(s) 808 are powered on and actively performing 3D rendering, to improve performance and responsiveness.
[0146] The SoC(s) 804 may further include a Mobile Industry Processor Interface (MIPI) camera serial interface for receiving video and camera input, a high-speed interface, and / or a video input block that can be used for camera and associated pixel input functions. The SoC(s) 804 may further include one or more input / output controllers that can be software-controlled and used to receive I / O signals not designated for a specific role.
[0147] The SoC(s) 804 can also include a wide selection of peripheral interfaces to enable communication with peripheral devices, audio codecs, power management, and / or other devices. The SoC(s) 804 can be used to process data from cameras (e.g., connected via Gigabit Multimedia Serial Link and Ethernet), sensors (e.g., LiDAR sensor(s) 864, RADAR sensor(s) 860, etc., which may be connected via Ethernet), data from bus 802 (e.g., vehicle speed 800, steering wheel position, etc.), data from GNSS sensor(s) 858 (e.g., connected via an Ethernet bus or a CAN bus), etc. The SoC(s) 804 can also include dedicated high-performance mass storage controllers, which can include their own DMA engines and can be used to offload routine data management tasks from the CPU(s) 806.
[0148] The 804 SoC(s) can be an end-to-end platform with a flexible architecture that encompasses automation levels 3-5, providing a comprehensive functional safety architecture that leverages and efficiently utilizes computer vision and ADAS techniques for diversity and redundancy, thus providing a platform for a flexible, reliable propulsion software stack, along with deep learning tools. The 804 SoC(s) can be faster, more reliable, and even more energy- and space-efficient than conventional systems. For example, the 814 accelerator(s), in combination with the 806 CPU(s), 808 GPU(s), and 816 data storage(s), can provide a fast, efficient platform for Level 3-5 autonomous vehicles.
[0149] This technology thus provides capabilities and functionality that cannot be achieved by conventional systems. For example, computer vision algorithms can be run on CPUs that can be configured using a higher-level programming language, such as C, to execute a wide range of processing algorithms on a wide range of visual data. However, CPUs are often unable to meet the performance requirements of many computer vision applications, such as those related to execution time and power consumption. In particular, many CPUs are unable to execute complex object detection algorithms in real time, which is a requirement for in-vehicle ADAS applications and for practical Level 3-5 autonomous vehicles.
[0150] Unlike conventional systems, the technology described herein, by providing a CPU complex, a GPU complex, and a hardware acceleration cluster, allows multiple neural networks to run simultaneously and / or sequentially and to combine the results to enable Level 3-5 autonomous driving functionality. For example, a CNN running on the DLA or the dGPU (e.g., the GPU(s) 820) can include text and word recognition, allowing the supercomputer to read and understand traffic signs, including those for which the neural network has not been specifically trained. The DLA can further include a neural network capable of identifying and interpreting a sign, providing a semantic understanding of the sign, and passing this semantic understanding to the route planning modules running on the CPU complex.
[0151] As another example, multiple networks can operate simultaneously, as required for Level 3, 4, or 5 driving. For instance, a warning sign consisting of "Caution: Flashing lights indicate black ice" along with an electric light can be interpreted independently or jointly by several neural networks. The sign itself can be identified as a traffic sign by a first neural network (e.g., a trained neural network), while the text "Flashing lights indicate black ice" can be interpreted by a second neural network, which informs the vehicle's route planning software (preferably running on the CPU complex) that black ice is present when flashing lights are detected.The flashing light can be identified by a third, deployed neural network over several frames, thus informing the vehicle's route planning software of the presence (or absence) of flashing lights. All three neural networks can operate simultaneously, for example, within the DLA and / or on the GPU(s) 808.
[0152] In some examples, a CNN for facial recognition and vehicle owner identification can use data from camera sensors to identify the presence of an authorized driver and / or vehicle owner of the Vehicle 800. The always-on sensor processing engine can be used to unlock the vehicle and turn on the lights when the owner approaches the driver's door, and to disable the vehicle in security mode when the owner leaves. In this way, the SoC(s) 804 provides security against theft and / or forced vehicle removal.
[0153] In another example, a CNN for emergency vehicle detection and identification can use data from microphones 896 to detect and identify emergency vehicle sirens. Unlike conventional systems that use general classifiers to detect sirens and manually extract features, the SoC(s) 804 can use the CNN to classify ambient and urban noise as well as visual data. In a preferred embodiment, the CNN, operating on the DLA, is trained to identify the relative approach speed of the emergency vehicle (e.g., using the Doppler effect). The CNN can also be trained to identify emergency vehicles specific to the local area in which the vehicle operates, as identified by the GNSS sensor(s) 858.For example, when operating in Europe, the CNN will attempt to detect European sirens, and when operating in the United States, the CNN will attempt to identify only North American sirens. A control program, with the assistance of 862 ultrasonic sensors, can be used, once an emergency vehicle is detected, to execute a safety routine for the emergency vehicle, such as slowing down, pulling over to the side of the road, parking the vehicle, and / or holding the vehicle in neutral until the emergency vehicle(s) has / have passed.
[0154] The vehicle can include one or more CPU(s) 818 (e.g., discrete CPU(s) or dCPU(s)) that can be coupled to the SoC(s) 804 via a high-speed connection (e.g., PCIe). The CPU(s) 818 can include, for example, an x86 processor. The CPU(s) 818 can be used to perform any of a number of functions, including, for example, arbitrating potentially inconsistent results between ADAS sensors and the SoC(s) 804 and / or monitoring the status and state of the Controller(s) 836 and / or Infotainment SoC 830.
[0155] The Vehicle 800 can include one or more GPU(s) 820 (e.g., discrete GPU(s) or dGPU(s)) that can be coupled to the SoC(s) 804 via a high-speed connection (e.g., NVIDIA's NVLINK). The GPU(s) 820 can provide additional artificial intelligence functionality, such as running redundant and / or different neural networks, and can be used to train and / or update neural networks based on input (e.g., sensor data) from sensors in the Vehicle 800.
[0156] The vehicle 800 can further include the network interface 824, which can include one or more wireless antennas 826 (e.g., one or more wireless antennas for different communication protocols, such as a cellular antenna, a Bluetooth antenna, etc.). The network interface 824 can be used to enable wireless connectivity over the internet to the cloud (e.g., to the server(s) 878 and / or other network devices), to other vehicles, and / or to computing devices (e.g., passenger client devices). Communication with other vehicles can be established via a direct connection between the two vehicles and / or an indirect connection (e.g., via networks and the internet). Direct connections can be provided using a vehicle-to-vehicle communication link.The vehicle-to-vehicle communication link can provide the Vehicle 800 with information about vehicles in its vicinity (e.g., vehicles in front of, to the side of, and / or behind the Vehicle 800). This functionality can be part of a cooperative adaptive cruise control feature of the Vehicle 800.
[0157] The 824 network interface can include a system-on-a-chip (SoC) that provides modulation and demodulation functionality, enabling the 836 controller(s) to communicate over wireless networks. The 824 network interface can include a radio frequency front end for upconverting baseband to radio frequency and downconverting radio frequency to baseband. Frequency conversions can be performed using known processes and / or superheterodyne processes. In some examples, the radio frequency front-end functionality can be provided by a separate chip. The network interface can include wireless functionality for communication over LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols.
[0158] The vehicle 800 may further include one or more data storage devices 828, which may include off-chip (e.g., off-SoC(s) 804) storage. The data storage device(s) 828 may include one or more storage elements, including RAM, SRAM, DRAM, VRAM, flash, hard disks, and / or other components and / or devices capable of storing at least one data bit.
[0159] The Vehicle 800 can also include one or more GNSS Sensor(s) 858. The GNSS Sensor(s) 858 (e.g., GPS, supported GPS sensors, differential GPS sensors (DGPS), etc.) support mapping, perception, occupancy grid generation, and / or route planning functions. Any number of GNSS Sensor(s) 858 can be used, including, for example, a single GPS using a USB connection with an Ethernet-to-serial (e.g., RS-232) bridge.
[0160] The vehicle 800 can further include one or more RADAR sensor(s) 860. The RADAR sensor(s) 860 can be used by the vehicle 800 for long-range vehicle detection, even in darkness and / or adverse weather conditions. The functional RADAR safety levels can be ASIL B. In some examples, the RADAR sensor(s) 860 can use CAN and / or bus 802 (e.g., to transmit data generated by the RADAR sensor(s) 860) for control and access to object tracking data, with access to Ethernet for raw data. A wide range of RADAR sensor types can be used. For example, the RADAR sensor(s) 860 can be suitable for front, rear, and side RADAR applications, among others. In some examples, a pulse Doppler radar sensor(s) is / are used.
[0161] The 860 RADAR sensor(s) can include various configurations, such as long-range with a narrow field of view, short-range with a wide field of view, side coverage with short-range, etc. In some examples, long-range RADAR can be used for adaptive cruise control functionality. Long-range RADAR systems can provide a wide field of view, achieved through two or more independent scans, for example, within a range of 250 m. The 860 RADAR sensor(s) can assist in distinguishing between stationary and moving objects and can be used by ADAS systems for emergency braking assistance and forward collision warning. Long-range radar sensors can include monostatic multimodal radar with multiple (e.g., six or more) fixed radar antennas and a high-speed CAN and FlexRay interface.In an example with six antennas, the central four antennas can generate a focused beam pattern designed to record the area around vehicle 800 at higher speeds with minimal interference from traffic in adjacent lanes. The other two antennas can expand the field of view, allowing vehicles entering or leaving vehicle 800's lane to be detected quickly.
[0162] Mid-range radar systems, for example, can have a range of up to 860 m (front) or 80 m (rear) and a field of view of up to 42 degrees (front) or 850 degrees (rear). Short-range radar systems can include, among other things, radar sensors designed for installation at both ends of the rear bumper. When such radar sensor systems are installed at both ends of the rear bumper, they can generate two beams that constantly monitor the blind spot in one direction, both rearward and close to the vehicle.
[0163] Short-range radar systems can be used in the ADAS system to detect a blind spot and / or to assist with a lane change.
[0164] The vehicle 800 can further include one or more ultrasonic sensor(s) 862. The ultrasonic sensor(s) 862 can be positioned on the front, rear, and / or sides of the vehicle 800 and can be used for parking assistance and / or for generating and updating an occupancy grid. A wide range of ultrasonic sensor(s) 862 can be used, and different ultrasonic sensor(s) 862 can be used for different detection ranges (e.g., 2.5 m, 4 m). The ultrasonic sensor(s) 862 can operate according to the functional safety levels of ASIL B.
[0165] The vehicle 800 can include one or more LiDAR sensor(s) 864. The LiDAR sensor(s) 864 can be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. The LiDAR sensor(s) 864 can operate at functional safety level ASIL B. In some examples, the vehicle 800 can include multiple LiDAR sensors 864 (e.g., two, four, six, etc.) that can use Ethernet (e.g., to provide data to a Gigabit Ethernet switch).
[0166] In some examples, the LiDAR sensor(s) 864 can provide a list of objects and their distances for a 360-degree field of view. A commercially available LiDAR sensor(s) 864, for example, may have a specified range of approximately 800 m with an accuracy of 2 cm to 3 cm and support for an 800 Mbps Ethernet connection. In some examples, one or more non-protruding LiDAR sensors 864 may be used. In such examples, the LiDAR sensor(s) 864 may be implemented as a small device that can be embedded in the front, rear, sides, and / or corners of the vehicle 800. The LIDAR sensor(s) 864 can provide a horizontal field of view of up to 120 degrees and a vertical field of view of up to 35 degrees in such examples, with a range of 200 m, even for objects with low reflectivity.One or more front-mounted LIDAR sensor(s) 864 can be configured for a horizontal field of view between 45 degrees and 135 degrees.
[0167] In some examples, LiDAR technologies, such as 3D flash LiDAR, can also be used. 3D flash LiDAR uses a laser flash as a transmission source to illuminate the vehicle's surroundings up to approximately 200 m. A flash LiDAR unit includes a sensor that records the laser pulse time-of-flight and the reflected light at each pixel, which corresponds to the vehicle's range to objects. With flash LiDAR, highly accurate and distortion-free images of the surroundings can be generated with each laser flash. In some examples, four flash LiDAR sensors can be used, one on each side of the vehicle. Available 3D flash LiDAR systems include a solid-state 3D LiDAR camera with a rigid array and no moving parts except for a fan (e.g., a non-scanning LiDAR device).The Blitz LIDAR device can use a 5-nanosecond Class I (eye-safe) laser pulse per frame and capture reflected laser light in the form of 3D distance point clouds and co-recorded intensity data. By using Blitz LIDAR, and because Blitz LIDAR is a solid-state device with no moving parts, the LIDAR sensor(s) 864 may be less susceptible to motion blur, vibration, and / or shock.
[0168] The vehicle may further include one or more IMU sensor(s) 866. The IMU sensor(s) 866 may be located in the center of the rear axle of the vehicle 800 in some examples. The IMU sensor(s) 866 may include, for example, one or more accelerometers, one or more magnetometers, one or more gyroscopes, one or more magnetic compasses, and / or other sensor types. In some examples, such as in six-axis applications, the IMU sensor(s) 866 may include, among others, accelerometers and gyroscopes, while in nine-axis applications, the IMU sensor(s) 866 may include accelerometers, gyroscopes, and magnetometers.
[0169] In some embodiments, the IMU sensor(s) 866 can be implemented as a miniature, high-performance GPS-aided inertial navigation system (GPS / INS) that combines microelectromechanical system (MEMS) inertial sensors, a high-sensitivity GPS receiver, and advanced Kalman filter algorithms to provide estimates of position, velocity, and orientation. As such, in some examples, the IMU sensor(s) 866 can enable the vehicle 800 to estimate its course without requiring input from a magnetic sensor by directly observing and correlating velocity changes from a GPS to the IMU sensor(s) 866. In some examples, the IMU sensor(s) 866 and the GNSS sensor(s) 858 can be combined in a single integrated unit.
[0170] The vehicle can include one or more microphone(s) 896, which are placed in and / or around the vehicle 800. The microphone(s) 896 can be used, among other things, for emergency vehicle detection and identification.
[0171] The vehicle may also include any number of camera types, including one or more stereo camera(s) 868, wide-angle camera(s) 870, infrared camera(s) 872, surround-view camera(s) 874, long-range camera(s), and / or medium-range camera(s) 898, and / or other camera types. The cameras may be used to capture image data around the entire periphery of the vehicle 800. The types of cameras used depend on the embodiment and requirements for the vehicle 800, and any combination of camera types may be used to provide the necessary coverage around the vehicle 800. Furthermore, the number of cameras may vary depending on the embodiment. For example, the vehicle may include six cameras, seven cameras, ten cameras, twelve cameras, and / or any other number of cameras.The cameras, for example, can support Gigabit Multimedia Serial Link (GMSL) and / or Gigabit Ethernet communication, among other things. Each camera is described in more detail below. Fig. 8A and Fig. 8B described.
[0172] The vehicle 800 may further include one or more vibration sensor(s) 842. The vibration sensor(s) 842 may measure vibrations of vehicle components, such as the axle(s). For example, changes in vibration may indicate a change in the road surface. In another example, if two or more vibration sensors 842 are used, the differences between the vibrations may be used to determine friction or slipperiness of the road surface (e.g., if the difference in vibration is between a driven axle and a freely rotating axle).
[0173] The vehicle 800 may include an ADAS system 838. The ADAS system 838 may include a SoC in some examples. The ADAS system 838 may include systems for autonomous / adaptive / automatic cruise control (ACC), cooperative adaptive cruise control (CACC), forward crash warning (FCW), automatic emergency braking (AEB), lane departure warnings (LDW), lane keep assist (LKA), blind spot warning (BSW), rear cross-traffic warning (RCTW), collision warning systems (CWS), lane centering (LC), and / or other features and functions.
[0174] The ACC systems can use one or more radar sensors (860), lidar sensors (864), and / or one or more cameras. The ACC systems can include longitudinal ACC and / or lateral ACC. Longitudinal ACC monitors and controls the distance to the vehicle immediately in front of the vehicle (800) and automatically adjusts the vehicle's speed to maintain a safe distance from vehicles ahead. Lateral ACC maintains the distance and advises the vehicle (800) to change lanes if necessary. Lateral ACC is related to other ADAS applications, such as LCA and CWS.
[0175] CACC uses information from other vehicles, which can be received via the network interface 824 and / or the wireless antenna(s) 826 from other vehicles either wirelessly or indirectly via a network connection (e.g., the internet). Direct connections can be provided through a vehicle-to-vehicle (V2V) communication link, while indirect connections can be an infrastructure-to-vehicle (I2V) communication link. Generally, the V2V communication concept provides information about vehicles immediately ahead (e.g., vehicles directly in front of and in the same lane as vehicle 800), while the I2V communication concept provides information about traffic at a greater distance. CACC systems can incorporate one or both of the I2V and V2V information sources.Thanks to the information about the vehicles in front of the vehicle 800, the CACC can be more reliable and has the potential to improve traffic flow and reduce congestion on the road.
[0176] FCW systems are designed to alert the driver to a hazard, allowing the driver to take corrective action. FCW systems utilize a forward-facing camera and / or RADAR 860 sensors, coupled with a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback, such as a display, speaker, and / or vibrating component. FCW systems can provide a warning, for example, in the form of an audible signal, a visual warning, a vibration, and / or a rapid braking pulse.
[0177] AEB systems detect an impending forward collision with another vehicle or object and can automatically apply the brakes if the driver does not take corrective action within a specified time or distance parameter. AEB systems use one or more forward-facing cameras and / or one or more radar sensors coupled to a dedicated processor, DSP, FPGA, and / or ASIC. When the AEB system detects a hazard, it will typically first warn the driver to take corrective action to avoid the collision. If the driver does not take corrective action, the AEB system will automatically apply the brakes in an effort to avoid or at least mitigate the impact of the predicted collision. AEB systems may include techniques such as dynamic brake assist and / or anticipatory braking.
[0178] Lane Departure Warning (LDW) systems provide visual, audible, and / or tactile warnings, such as steering wheel or seat vibrations, to alert the driver if the vehicle crosses lane markings. An LDW system will not activate if the driver indicates an intention to leave the lane, for example, by using a turn signal. LDW systems may use forward-facing cameras coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback, such as a display, speaker, and / or vibration component.
[0179] LKA systems are a variant of LDW systems. LKA systems provide steering or braking input to correct vehicle 800 if it begins to leave its lane.
[0180] Blind Spot Warning (BSW) systems detect vehicles in a car's blind spot and warn the driver. BSW systems can provide a visual, audible, and / or tactile alert to indicate that merging or changing lanes is unsafe. The system can provide an additional warning if the driver uses a turn signal. BSW systems can use rear-facing camera(s) and / or radar sensor(s) coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback, such as a display, speaker, and / or vibration component.
[0181] RCTW systems can provide visual, audible, and / or tactile alerts when an object outside the reversing camera's field of view is detected while the vehicle is reversing. Some RCTW systems include AEB to ensure the vehicle's brakes are applied to avoid a collision. RCTW systems can use one or more rear-facing radar sensors coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback, such as a display, speaker, and / or vibrating component.
[0182] Conventional ADAS systems can be prone to false positives, which can be annoying and distracting for a driver, but are generally not catastrophic because the ADAS systems warn the driver and allow them to decide whether a safety condition actually exists and act accordingly. In an autonomous vehicle 800, however, the vehicle 800 itself must decide, in the case of conflicting results, whether the result should be considered by a primary computer or a secondary computer (e.g., a first controller 836 or a second controller 836). For example, the ADAS system 838 in some embodiments may be a backup and / or secondary computer that provides perceptual information to a rationality module of the backup computer. The rationality monitor of the backup computer may run redundant software on hardware components to detect perceptual errors and dynamic driving tasks.Output from the ADAS system 838 can be provided to a higher-level MCU. If output from the primary and secondary computers conflict, the higher-level MCU must decide how to resolve the conflict to ensure safe operation.
[0183] In some examples, the primary computer can be configured to provide the higher-level MCU with a confidence score indicating its level of trust in the chosen result. If the confidence score exceeds a certain threshold, the higher-level MCU can follow the primary computer's guidance, regardless of whether the secondary computer provides a conflicting or inconsistent result. If the confidence score does not reach a certain threshold, and if the primary and secondary computers display different results (e.g., a conflict), the higher-level MCU can arbitrate between the computers to determine the appropriate result.
[0184] The higher-level MCU can be configured to run a neural network(s) trained and configured to determine, based on output from the primary and secondary computers, the conditions under which the secondary computer will generate false alarms. Thus, the neural network(s) in the higher-level MCU can learn when the secondary computer's output can be trusted and when it cannot. For example, if the secondary computer is a radar-based FCW system, the neural network(s) in the higher-level MCU can learn when the FCW system identifies metallic objects that are not actually hazards, such as a drainage grate or a manhole cover, triggering an alarm.Similarly, if the secondary computer is a camera-based lane departure warning (LDW) system, a neural network in the higher-level MCU can learn to override the LDW when cyclists or pedestrians are present and leaving the lane is indeed the safest maneuver. In embodiments that include one or more neural networks running on the higher-level MCU, the higher-level MCU can include at least one DLA or GPU suitable for running the neural network(s) with associated memory. In preferred embodiments, the higher-level MCU can include a component of the SoC(s) 804 and / or be included as such.
[0185] In other examples, the ADAS System 838 can include a secondary computer that performs ADAS functionality using traditional computer vision rules. As such, the secondary computer can use classic computer vision rules (if-then), and the presence of one or more neural networks in the higher-level MCU can improve reliability, safety, and performance. For example, the overall system becomes more fault-tolerant due to the different implementation and intentional non-identity, particularly with regard to errors caused by software functionality (or the software-hardware interface).For example, if there is a software bug or error in the software running on the primary computer, and the non-identical software code running on the secondary computer provides the same overall result, the higher-level MCU can have greater confidence that the overall result is correct and that the bug in the software or hardware on the primary computer does not cause a significant error.
[0186] In some examples, the output of the ADAS system 838 can be fed into the perception block of the primary computer and / or into the dynamic driving task block of the primary computer. For example, if the ADAS system 838 displays a forward collision warning due to an object immediately in front of it, the perception block can use this information in object identification. In other examples, the secondary computer may have its own neural network that is trained, thus reducing the risk of false positives, as described herein.
[0187] The Vehicle 800 may further include the Infotainment SoC 830 (e.g., an In-Vehicle Infotainment System (IVI)). Although illustrated and described as an SoC, the Infotainment System may not be an SoC and may include two or more discrete components. The Infotainment SoC 830 may include a combination of hardware and software that can be used to provide the Vehicle 800 with audio (e.g., music, a personal digital assistant, navigation instructions, news, radio, etc.), video (e.g., television, movies, streaming, etc.), telephone (e.g., hands-free calling), network connectivity (e.g., LTE, Wi-Fi, etc.), and / or information services (e.g., navigation systems, rear parking assist, a radio data system, vehicle-related information such as fuel level, total distance traveled, brake fluid level, oil level, door open / closed status, air filter information, etc.). The Infotainment SoC 830 may include, for example, a driver information system ... B.This includes radios, record players, navigation systems, video players, USB and Bluetooth connectivity, car computers, in-car entertainment, WiFi, steering wheel audio controls, hands-free voice control, a heads-up display (HUD), an HMI display 834, a telematics device, a control panel (e.g., for controlling and / or interacting with various components, functions, and / or systems), and / or other components. The Infotainment SoC 830 can also be used to provide information (e.g., visual and / or audible) to the vehicle's user(s), such as information from the ADAS system 838, autonomous driving information like planned vehicle maneuvers, trajectories, environmental information (e.g., intersection information, vehicle information, road information, etc.), and / or other information.
[0188] The Infotainment SoC 830 can include GPU functionality. The Infotainment SoC 830 can communicate with other devices, systems, and / or components of the Vehicle 800 via the 802 bus (e.g., CAN bus, Ethernet, etc.). In some examples, the Infotainment SoC 830 can be coupled with a higher-level MCU so that the Infotainment System's GPU can perform some self-regulating functions if the primary controller(s) 836 (e.g., the Vehicle 800's primary and / or backup computer) fails. In such an example, the Infotainment SoC 830 can put the Vehicle 800 into a chauffeur-to-safe-stop mode, as described herein.
[0189] The Vehicle 800 may further include an Instrument Cluster 832 (e.g., a digital instrument cluster, an electronic instrument cluster, a digital instrument panel, etc.). The Instrument Cluster 832 may include a controller and / or a supercomputer (e.g., a discrete controller or a discrete supercomputer). The Instrument Cluster 832 may include a set of instruments such as a speedometer, fuel gauge, oil pressure gauge, tachometer, odometer, turn signals, gear shift indicator, seatbelt warning light(s), parking brake warning light(s), engine malfunction light(s), airbag (SRS) system information, lighting controls, safety system controls, navigation information, etc. In some examples, information can be displayed and / or shared between the Infotainment SoC 830 and the Instrument Cluster 832.In other words, the Instrument Cluster 832 can be included as part of the Infotainment SoC 830, or vice versa.
[0190] Fig. 8D is a system representation according to some embodiments of the present disclosure for communication between one or more cloud-based server(s) and the exemplary autonomous vehicle 800 of Fig.8A is the system. The 876 system can include one or more 878 servers, one or more 890 networks, and vehicles, including the 800 vehicle. The 878 server(s) can include a variety of 884(A)-884(H) GPUs (collectively referred to hereafter as GPUs 884), 882(A)-882(H) PCIe switches (collectively referred to hereafter as PCIe switches 882), and / or 880(A)-880(B) CPUs (collectively referred to hereafter as CPUs 880). The 884 GPUs, 880 CPUs, and PCIe switches can be interconnected using high-speed links, such as, but not limited to, NVIDIA's NVLink 888 interfaces and / or PCIe 886 links. In some examples, the GPUs 884 are connected via an NVLink and / or NVSwitch SoC, and the GPUs 884 and PCIe switches 882 are connected via PCIe links. Although eight GPUs 884, two CPUs 880, and two PCIe switches are illustrated, this is not intended to be a limiting factor.Depending on the configuration, each Server 878 can include any number of GPUs 884, CPUs 880, and / or PCIe switches. For example, the Server 878 can include eight, sixteen, thirty-two, and / or more GPUs 884.
[0191] The server(s) 878 can receive image data from the network(s) 890 and the vehicles, representing images that show unexpected or changed road conditions, such as recently started roadworks. The server(s) 878 can transmit neural networks 892, updated neural networks 892, and / or map information 894, including information regarding traffic and road conditions, to the vehicles via the network(s) 890 and the vehicles. The map information updates 894 can include updates to the HD map 822, such as information about construction sites, potholes, detours, floods, and / or other obstacles.In some examples, the neural networks 892, the updated neural networks 892 and / or the map information 894 may result from new training and / or new experiences represented in data received from any number of vehicles in the vicinity, and / or based on training performed at a data center (e.g. using server(s) 878 and / or other servers).
[0192] Server 878 can be used to train machine learning models (e.g., neural networks) based on training data. The training data can be generated by the vehicles and / or in a simulation (e.g., using a game engine). In some examples, the training data is tagged (e.g., if the neural network benefits from supervised learning) and / or preprocessed, while in other examples, the training data is untagged and / or preprocessed (e.g., if the neural network does not require supervised learning).Training can be performed according to one or more classes of machine learning techniques, including, but not limited to, classes such as: supervised training, semi-supervised training, unsupervised training, self-learning, reinforcement learning, federated learning, transfer learning, feature learning (including principal component and cluster analysis), multilinear subspace learning, manifold learning, representational learning (including sparse dictionary learning), rule-based machine learning, anomaly detection, and any other variants or combinations thereof. After the machine learning models have been trained, they can be used by the vehicles (e.g., transmitted to the vehicles via network(s) 890) and / or the machine learning models can be used by server(s) 878 for remote monitoring of the vehicles.
[0193] In other examples, the Server 878 can receive data from the vehicles and apply that data to current real-time neural networks for intelligent real-time inference. The Server 878 can include supercomputers for deep learning and / or dedicated AI computers powered by GPU 884, such as NVIDIA's DGX and DGX Station machines. However, in some examples, the Server 878 can also include deep learning infrastructure that uses only CPU-powered data centers.
[0194] The deep learning infrastructure of the server(s) 878 can be capable of fast, real-time inference and use this capability to assess and verify the state of the processors, software, and / or associated hardware in the vehicle 800. For example, the deep learning infrastructure can receive periodic updates from the vehicle 800, such as a sequence of images and / or objects that the vehicle 800 has located within that sequence of images (e.g., through computer vision and / or other machine learning object classification techniques).The deep learning infrastructure can operate its own neural network to identify the objects and compare them with the objects identified by the vehicle 800, and if the results do not match and the infrastructure concludes that the AI in the vehicle 800 is faulty, the server(s) 878 can send a signal to the vehicle 800 and instruct a fail-safe computer of the vehicle 800 to take over control, notify the passengers and perform a safe parking maneuver.
[0195] The Server 878 can include GPU(s) 884 and one or more programmable inference accelerators (e.g., NVIDIA's TensorRT). The combination of GPU-powered servers and inference acceleration can enable real-time responsiveness. In other examples, such as when performance is less critical, servers powered by CPUs, FPGAs, and other processors can be used for inference purposes.
[0196] In some embodiments of the systems described herein, procedures can be performed in a simulation environment (e.g., NVIDIA's DriveSIM) using simulated data (e.g., simulated sensor data from a virtual or simulated machine). For example, simulated sensor data and / or map data can be used to identify regions of interest (e.g., parking spaces) and subregions of interest (e.g., subregions of a parking space that include a curb, wheel stop, etc.) within the simulation environment, and this information can be used to perform operations (e.g., parking) associated with the virtual machine within the environment. These simulated operations can be used to test the performance of the underlying algorithms, systems, and / or processes before their deployment in the real world.In some cases, the simulation can be used to generate synthetic training data—for example, training data that includes regions of interest and / or subregions of interest within the simulation. This synthetic training data can then be processed (in addition to or as an alternative to real-world data) to determine geometry and / or other information related to the regions of interest, such as parking spaces or pallet drop-off points within a warehouse. For example, when a simulation environment is used for testing, validation, training, etc., the simulation environment and / or associated training data can be rendered or otherwise generated using one or more light transport algorithms—such as ray tracing and / or path tracing algorithms.In some embodiments, the simulation environment and / or one or more objects, features, or components thereof can be generated or managed within a three-dimensional content collaboration platform (3D content collaboration platform) (e.g., NVIDIA's OMNIVERSE) for industrial digitization, generative physical AI, and / or other use cases, applications, or services. For example, the content collaboration platform or system may include a system for using or employing Universal Scene Descriptor (USD) data (e.g., OpenUSD) to manage objects, features, scenes, etc., within a simulated environment, digital environment, etc. The platform may include real-world physics simulation, such as using NVIDIA's PhysX SDK, to simulate real-world physics and physical interactions with simulations hosted by the platform.The platform can integrate OpenUSD together with ray tracing / path tracing / light transport simulation (e.g. NVIDIA's RTX rendering technologies) into software tools and simulation workflows for building, training, deploying or testing AI systems - such as systems for testing, validating, training (e.g. machine learning models, neural networks etc.) and / or other tasks related to automotive, robotics, machinery or other applications.
[0197] The disclosure of this application also includes the following numbered clauses: Clause 1: In some embodiments, a method for at least one channel in a set of feature channels comprises: receiving a first feature map for a first image in a stereo image pair and a second feature map for a second image in the stereo image pair, and calculating a corresponding set of correlation maps representing correlations between one or more positions in a row in the first feature map and one or more positions in a corresponding row in the second feature map; generating a set of compressed correlation maps at least based on compressing the sets of correlation maps corresponding to the set of feature channels over a dimension associated with the set of feature channels;Masking one or more sections of individual compressed correlation maps from the set of compressed correlation maps, at least based on a respective correlation filter, to generate a corresponding set of masked correlation maps; and generating a depth map associated with the stereo image pair, at least based on the set of masked correlation maps. Clause 2. Procedure according to Clause 1, wherein, for at least one channel in the set of feature channels, the set of correlation maps is calculated at least on the basis of multiplying one or more columns of the first feature map with one or more columns of the second feature map using matrix multiplication. Clause 3. Procedure according to Clause 1 or 2, wherein the compression of the set of correlation maps corresponding to the set of feature channels includes a dimension associated with the set of feature channels, and the summation of corresponding correlation maps in one or more channels of the set of feature channels. Clause 4. Procedure according to any of Clauses 1-3, wherein the masked sections of individual compressed correlation maps of the set of compressed correlation maps include compressed correlations that are not used in the generation of the depth map. Clause 5. Procedure according to one of Clauses 1-4, wherein generating the depth map associated with the stereo image pair includes converting a disparity map generated at least on the basis of the set of masked correlation maps. Clause 6. Procedure according to any of Clauses 1-5, which further includes downscaling the set of masked correlation maps using one or more convolutional layers. Clause 7. Procedure according to one of Clauses 1-6, wherein the first feature map and the second feature map are generated by a first feature extractor and a second feature extractor of a neural network, respectively. Clause 8. Procedure according to one of Clauses 1-7, wherein the first feature map and the second feature map are associated with a feature channel and are represented as a matrix with a width dimension and a height dimension. Clause 9. Procedure according to any one of Clauses 1-8, wherein the first feature map, the second feature map, each correlation map of the set of correlation maps, each compressed correlation map of the set of compressed correlation maps, and each compressed correlation map of the compressed correlation maps have the same width dimension and the same height dimension. Clause 10. Procedure according to one of Clauses 1-9, wherein, in at least one channel in the set of feature channels, the number of coalition maps corresponds to the width of the first feature map. Clause 11. At least one processor comprising: one or more circuits for: receiving, for at least one channel in a set of feature channels, a first feature map for a first image in a stereo image pair and a second feature map for a second image in the stereo image pair, and calculating a corresponding set of correlation maps representing correlations between one or more positions in a row in the first feature map and one or more positions in a corresponding row in the second feature map; generating a set of compressed correlation maps at least based on compressing the sets of correlation maps corresponding to the set of feature channels over a dimension associated with the set of feature channels;Masking one or more sections of individual compressed correlation maps from the set of compressed correlation maps, at least based on a respective correlation filter, to generate a corresponding set of masked correlation maps; and generating a depth map associated with the stereo image pair, at least based on the set of masked correlation maps. Clause 12. The at least one processor according to Clause 11, wherein the at least one processor consists of at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing digital twinning operations; a system for performing light transport simulation; a system for performing collaborative content creation for 3D assets; a system for performing deep learning operations; a system implemented using an edge device; a system for generating or presenting at least one of virtual reality content, augmented reality content, or mixed reality content; a system implemented using a robot; a system for performing conversational AI operations;a system for implementing one or more large language models (LLMs); a system for implementing one or more vision language models (VLMs); a system for implementing one or more multimodal language models; a system for generating synthetic data; a system that involves one or more virtual machines (VMs); a system that is implemented at least partially in a data center; or a system that is implemented at least partially using cloud computing resources. Clause 13. The at least one processor according to Clause 11 or 12, wherein, for at least one channel in the set of feature channels, the set of correlation maps is calculated at least on the basis of multiplying one or more columns of the first feature map with one or more columns of the second feature map using matrix multiplication. Clause 14. The at least one processor according to Clauses 11-13, wherein the compression of the set of correlation maps corresponding to the set of feature channels includes a dimension associated with the set of feature channels, and the summation of corresponding correlation maps in one or more channels of the set of feature channels. Clause 15. The at least one processor according to Clauses 11-14, wherein the masked sections of individual compressed correlation maps of the set of compressed correlation maps include compressed correlations that are not used in the generation of the depth map. Clause 16. The at least one processor according to Clause 11-15, wherein generating the depth map associated with the stereo image pair includes converting a disparity map generated at least on the basis of the set of masked correlation maps. Clause 17. The at least one processor according to Clauses 11-16, wherein the one or more circuits further serve to downscale the set of masked coalition parties using one or more convolutional layers. Clause 18. The at least one processor according to Clause 11-17, wherein the first feature map and the second feature map are generated by a first feature extractor and a second feature extractor of a neural network, respectively. Clause 19. A system comprising: one or more processors for: receiving, for at least one channel in a set of feature channels, a first feature map for a first image in a stereo image pair and a second feature map for a second image in the stereo image pair, and calculating a corresponding set of correlation maps representing correlations between one or more positions in a row in the first feature map and one or more positions in a corresponding row in the second feature map; generating a set of compressed correlation maps at least based on compressing the sets of correlation maps corresponding to the set of feature channels over a dimension associated with the set of feature channels;Masking one or more sections of individual compressed correlation maps from the set of compressed correlation maps, at least based on a respective correlation filter, to generate a corresponding set of masked correlation maps; and generating a depth map associated with the stereo image pair, at least based on the set of masked correlation maps. Clause 20. System according to Clause 19, wherein the system consists of at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing digital twinning operations; a system for performing light transport simulation; a system for performing collaborative content creation for 3D assets; a system for performing deep learning operations; a system implemented using an edge device; a system for generating or presenting at least one of virtual reality content, augmented reality content, or mixed reality content; a system implemented using a robot; a system for performing conversational AI operations;a system for implementing one or more large language models (LLMs); a system for implementing one or more vision language models (VLMs); a system for implementing one or more multimodal language models; a system for generating synthetic data; a system that involves one or more virtual machines (VMs); a system that is implemented at least partially in a data center; or a system that is implemented at least partially using cloud computing resources.
[0198] Any and all combinations of any of the claim elements according to any of the claims and / or of elements or clauses described in this application fall in any way within the considered scope of the present disclosure and protection.
[0199] The disclosure can generally be described in the context of computer code or machine-usable instructions, including computer-executable instructions such as program modules that are executed by a computer or other machine, such as a personal data assistant or a handheld device. Generally, program modules, including routines, programs, objects, components, data structures, etc., refer to code that performs specific tasks or implements specific abstract data types. The disclosure can be exercised in various system configurations, including handheld devices, consumer electronics, general-purpose computers, other specialized computing devices, etc. The disclosure can also be exercised in distributed computing environments where tasks are performed by remote processing devices linked via a communication network.
[0200] As used herein, any mention of "and / or" with reference to two or more elements shall be interpreted as referring to only one element or a combination of elements. For example, "element A, element B, and / or element C" may include only element A, only element B, only element C, element A and element B, element A and element C, element B and element C, or elements A, B, and C. Furthermore, "at least one of element A or element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B. Additionally, "at least one of element A and element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B.
[0201] The subject matter of this disclosure is described herein with specificity to satisfy legal requirements. However, the description itself is not intended to limit the scope of this disclosure. Rather, the inventors have provided that the claimed subject matter may also be embodied in other ways to include different steps or combinations of steps similar to those described in this document, in conjunction with other current or future technologies. Furthermore, although the terms "step" and / or "block" may be used herein to denote different elements of methods employed, they should not be interpreted as implying a particular sequence of or between different steps disclosed herein, except where the sequence of individual steps is explicitly described.
[0202] It is understood that the aspects and embodiments described above are merely examples and that modifications in detail may be made within the scope of the claims.
[0203] Each device, method and feature disclosed in the description and (if applicable) in the claims and drawings can be provided independently or in any suitable combination.
[0204] Reference numerals appearing in the claims serve only for illustration and are not intended to have any limiting effect on the scope of the claims. QUOTES INCLUDED IN THE DESCRIPTION
[0000] This list of documents cited by the applicant was automatically generated and is included solely for the reader's convenience. The list is not part of the German patent or utility model application. The DPMA accepts no liability for any errors or omissions. Cited patent literature
[0000] US 16 / 101,232
[0131] Cited non-patent literature
[0000] Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles” (Standard No. J3016-201806, published on June 15, 2018, Standard No. J3016-201609, published on September 30, 2016
[0089]
Claims
[1] A procedure comprising the following: for at least one channel in a set of feature channels: Receiving a first feature card for a first image in a stereo image pair and a second feature card for a second image in the stereo image pair, and Calculating a corresponding set of correlation maps that represent correlations between one or more positions in a row in the first feature map and one or more positions in a corresponding row in the second feature map; Generating a set of compressed correlation maps at least based on compressing the sets of correlation maps corresponding to the set of feature channels, across a dimension associated with the set of feature channels; Masking one or more sections of individual compressed correlation maps from the set of compressed correlation maps, based at least on a respective correlation filter, to generate a corresponding set of masked correlation maps; and Generating a depth map associated with the stereo image pair, at least based on the set of masked correlation maps. [2] Method according to claim 1, wherein, for at least one channel in the set of feature channels, the set of correlation maps is calculated at least on the basis of multiplying one or more columns of the first feature map with one or more columns of the second feature map using matrix multiplication. [3] Method according to a preceding claim, wherein the compression of the set of correlation maps corresponding to the set of feature channels over a dimension associated with the set of feature channels comprises summing corresponding correlation maps in one or more channels of the set of feature channels. [4] Method according to a preceding claim, wherein the masked sections of individual compressed correlation maps of the set of compressed correlation maps comprise compressed correlations that are not used in the generation of the depth map. [5] Method according to a preceding claim, wherein generating the depth map associated with the stereo image pair comprises converting a disparity map generated at least on the basis of the set of masked correlation maps. [6] Method according to a preceding claim, further comprising downscaling the set of masked correlation maps using one or more convolutional layers. [7] Method according to a preceding claim, wherein the first feature map and the second feature map are generated by a first feature extractor and a second feature extractor of a neural network, respectively. [8] Method according to a preceding claim, wherein the first feature map and the second feature map are associated with a feature channel and are represented as a matrix with a width dimension and a height dimension. [9] Method according to a preceding claim, wherein the first feature map, the second feature map, each correlation map of the set of correlation maps, each compressed correlation map of the set of compressed correlation maps and each compressed correlation map of the compressed correlation maps have the same width dimension and the same height dimension. [10] Method according to a preceding claim, wherein, in at least one channel in the set of feature channels, the number of coalition maps corresponds to the width of the first feature map. [11] At least one processor comprising the following: one or more circuits for: for at least one channel in a set of feature channels, Receiving a first feature card for a first image in a stereo image pair and a second feature card for a second image in the stereo image pair, and Calculating a corresponding set of correlation maps that represent correlations between one or more positions in a row in the first feature map and one or more positions in a corresponding row in the second feature map; Generating a set of compressed correlation maps at least based on compressing the sets of correlation maps corresponding to the set of feature channels, across a dimension associated with the set of feature channels; Masking one or more sections of individual compressed correlation maps from the set of compressed correlation maps, based at least on a respective correlation filter, to generate a corresponding set of masked correlation maps; and Generating a depth map associated with the stereo image pair, at least based on the set of masked correlation maps. [12] The at least one processor according to claim 11, wherein the at least one processor consists of at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for conducting digital twinning operations; a system for performing a light transport simulation; a system for conducting collaborative content creation for 3D assets; a system for performing deep learning operations; a system implemented using an edge device; a system for generating or presenting at least one type of virtual reality content, augmented reality content, or mixed reality content; a system that is implemented using a robot; a system for performing conversational AI operations; a system for implementing one or more large language models (LLMs); a system for implementing one or more Visual Language Models (VLMs); a system for implementing one or more multimodal language models; a system for generating synthetic data; a system that includes one or more virtual machines (VMs); a system that is at least partially implemented in a data center; or a system that is implemented at least partially using cloud computing resources. [13] The at least one processor according to one of claims 11 or 12, wherein, for at least one channel in the set of feature channels, the set of correlation maps is calculated at least on the basis of multiplying one or more columns of the first feature map with one or more columns of the second feature map using matrix multiplication. [14] The at least one processor according to one of claims 11-13, wherein the compression of the set of correlation maps corresponding to the set of feature channels comprises summing corresponding correlation maps in one or more channels of the set of feature channels over a dimension associated with the set of feature channels. [15] The at least one processor according to any one of claims 11-14, wherein the masked sections of individual compressed correlation maps of the set of compressed correlation maps comprise compressed correlations that are not used in the generation of the depth map. [16] The at least one processor according to one of claims 11-15, wherein generating the depth map associated with the stereo image pair comprises converting a disparity map generated at least on the basis of the set of masked correlation maps. [17] The at least one processor according to one of claims 11-16, wherein the one or more circuits further serve to downscale the set of masked coalition parties using one or more convolutional layers. [18] The at least one processor according to one of claims 11-17, wherein the first feature map and the second feature map are generated by a first feature extractor and a second feature extractor of a neural network, respectively. [19] A system that includes the following: one or more processors for: for at least one channel in a set of feature channels, Receiving a first feature card for a first image in a stereo image pair and a second feature card for a second image in the stereo image pair, and Calculating a corresponding set of correlation maps that represent correlations between one or more positions in a row in the first feature map and one or more positions in a corresponding row in the second feature map; Generating a set of compressed correlation maps at least based on compressing the sets of correlation maps corresponding to the set of feature channels, across a dimension associated with the set of feature channels; Masking one or more sections of individual compressed correlation maps from the set of compressed correlation maps, based at least on a respective correlation filter, to generate a corresponding set of masked correlation maps; and Generating a depth map associated with the stereo image pair, at least based on the set of masked correlation maps. [20] System according to claim 19, wherein the system consists of at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for conducting digital twinning operations; a system for performing a light transport simulation; a system for conducting collaborative content creation for 3D assets; a system for performing deep learning operations; a system implemented using an edge device; a system for generating or presenting at least one type of virtual reality content, augmented reality content, or mixed reality content; a system that is implemented using a robot; a system for performing conversational AI operations; a system for implementing one or more large language models (LLMs); a system for implementing one or more Visual Language Models (VLMs); a system for implementing one or more multimodal language models; a system for generating synthetic data; a system that includes one or more virtual machines (VMs); a system that is at least partially implemented in a data center; or a system that is implemented at least partially using cloud computing resources.
Citation Information
Patent Citations
US-PATENTANMELDUNGNR.16/101,232