Computing optical flow using semi-global matching

By modifying the SGM method and using parallel processing units, the problems of insufficient efficiency and accuracy in optical flow calculation are solved, achieving efficient and accurate optical flow calculation and supporting real-time processing for applications such as autonomous devices and medical imaging.

CN116703965BActive Publication Date: 2026-04-28NVIDIA CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NVIDIA CORP
Filing Date
2022-09-27
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing optical flow calculation methods are insufficient in terms of computational efficiency and accuracy, making it difficult to meet the requirements of real-time processing and high precision, especially in applications such as autonomous devices and medical imaging.

Method used

A modified semi-global matching (SGM) method is adopted, which combines noise reduction, edge detection, object detection and decision logic to generate disparity maps and optical flow maps, and achieves efficient computation through parallel processing units (such as GPUs).

Benefits of technology

It improves the accuracy and efficiency of optical flow calculation, enabling it to meet the real-time processing needs of applications such as autonomous devices and medical imaging, and supporting complex vision tasks such as motion estimation and object detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116703965B_ABST
    Figure CN116703965B_ABST
Patent Text Reader

Abstract

The present disclosure relates to computing optical flow using semi-global matching, and in particular to apparatuses, systems, and techniques for determining optical flow. In at least one embodiment, a set of disparity values is used to determine optical flow between an input image and a reference image. For each of a plurality of image regions of the input image, the set of disparity values can include disparity values for a plurality of directions that intersect the image region.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] At least one embodiment relates to a method for determining optical flow for at least one pair of images. For example, at least one embodiment relates to a processor or computing system for obtaining a disparity map for determining optical flow according to various new techniques described herein. Background Technology

[0002] Optical flow is a measure of the apparent motion of a subject (e.g., an object) occurring from a first image to a second image (e.g., a video frame). For example, optical flow can be calculated based on the apparent motion of an image region (e.g., a pixel) occurring from the first image to the second image. Optical flow can be used for motion estimation, object detection, object tracking, image principal plane extraction, motion detection, robot navigation, visual odometry, camera motion detection, and video compression. Attached Figure Description

[0003] Figure 1 An example system for determining optical flow according to at least one embodiment is shown;

[0004] Figure 2 The description illustrates a method that can be derived from at least one embodiment. Figure 1 A block diagram of a method for generating optical flow graphs executed by the optical flow hardware of the system;

[0005] Figure 3A An example input image according to at least one embodiment is shown, which includes a... Figure 1 The system's optical flow hardware selects a set of image regions and is composed of... Figure 1 The system's optical flow hardware is configured for a specific direction in these image regions;

[0006] Figure 3B The diagram illustrates a method according to at least one embodiment. Figure 1 The system's optical flow hardware generates an example object graph for the input image;

[0007] Figure 3C The diagram illustrates a method according to at least one embodiment. Figure 1 The system's optical flow hardware generates an example edge map for the input image;

[0008] Figure 4 The diagram illustrates a method according to at least one embodiment. Figure 1 The system's optical flow hardware generates an example disparity map for the input image;

[0009] Figure 5 An example reference image is shown side-by-side with an example input image according to at least one embodiment;

[0010] Figure 6Example values ​​of one or more metrics assigned to each image region in a reference image and an input image according to at least one embodiment are shown;

[0011] Figure 7 The illustration shows a first set of image regions determined for a reference image according to at least one embodiment, example values ​​of one or more metrics assigned to each image region in the first set, a second set of image regions determined for an input image, and example values ​​of one or more metrics assigned to each image region in the second set.

[0012] Figure 8 An example disparity map according to at least one embodiment is shown;

[0013] Figure 9 The diagram illustrates a method according to at least one embodiment. Figure 1 The system's optical flow hardware determines the set of image regions and orientations of reference images for specific image regions in these image regions;

[0014] Figure 10 It is shown that, according to at least one embodiment, it can be made by Figure 1 A flowchart of the optical flow hardware execution method;

[0015] Figure 11A The inference and / or training logic according to at least one embodiment is illustrated;

[0016] Figure 11B The inference and / or training logic according to at least one embodiment is illustrated;

[0017] Figure 12 The training and deployment of a neural network according to at least one embodiment are illustrated;

[0018] Figure 13 An example data center system according to at least one embodiment is shown;

[0019] Figure 14A A chip-level supercomputer according to at least one embodiment is illustrated;

[0020] Figure 14B A rack-mounted supercomputer according to at least one embodiment is illustrated;

[0021] Figure 14C A rack-mounted supercomputer according to at least one embodiment is shown;

[0022] Figure 14D A supercomputer at the entire system level according to at least one embodiment is shown;

[0023] Figure 15 This is a block diagram illustrating a computer system according to at least one embodiment;

[0024] Figure 16 This is a block diagram illustrating a computer system according to at least one embodiment;

[0025] Figure 17 A computer system according to at least one embodiment is shown;

[0026] Figure 18 A computer system according to at least one embodiment is shown;

[0027] Figure 19A A computer system according to at least one embodiment is shown;

[0028] Figure 19B A computer system according to at least one embodiment is shown;

[0029] Figure 19C A computer system according to at least one embodiment is shown;

[0030] Figure 19D A computer system according to at least one embodiment is shown;

[0031] Figure 19E and Figure 19F A shared programming model according to at least one embodiment is shown;

[0032] Figure 20 An exemplary integrated circuit and a related graphics processor according to at least one embodiment are shown.

[0033] Figures 21A-21B An exemplary integrated circuit and an associated graphics processor according to at least one embodiment are shown.

[0034] Figures 22A-22B Additional exemplary graphics processor logic according to at least one embodiment is shown;

[0035] Figure 23 A computer system according to at least one embodiment is shown;

[0036] Figure 24A A parallel processor according to at least one embodiment is shown;

[0037] Figure 24B A partitioning unit according to at least one embodiment is shown;

[0038] Figure 24C A processing cluster according to at least one embodiment is shown;

[0039] Figure 24D A graphics multiprocessor according to at least one embodiment is shown;

[0040] Figure 25 A multi-graphics processing unit (GPU) system according to at least one embodiment is illustrated;

[0041] Figure 26 A graphics processor according to at least one embodiment is shown;

[0042] Figure 27 It is a block diagram illustrating a processor microarchitecture for a processor according to at least one embodiment;

[0043] Figure 28 A deep learning application processor according to at least one embodiment is shown;

[0044] Figure 29 A block diagram of an example neuromorphic processor is shown according to at least one embodiment;

[0045] Figure 30 At least a portion of a graphics processor according to one or more embodiments is shown;

[0046] Figure 31 At least a portion of a graphics processor according to one or more embodiments is shown;

[0047] Figure 32 At least a portion of a graphics processor according to one or more embodiments is shown;

[0048] Figure 33 A block diagram of a graphics processing engine of a graphics processor is shown according to at least one embodiment;

[0049] Figure 34 It is a block diagram of at least a portion of a graphics processor core according to at least one embodiment;

[0050] Figures 35A-35B The diagram illustrates thread execution logic according to at least one embodiment, which includes an array of processing elements of a graphics processor core.

[0051] Figure 36 A parallel processing unit (PPU) according to at least one embodiment is shown;

[0052] Figure 37 A general-purpose processing cluster (“GPC”) according to at least one embodiment is shown;

[0053] Figure 38 A memory partitioning unit of a parallel processing unit (PPU) according to at least one embodiment is shown;

[0054] Figure 39 A streaming multiprocessor according to at least one embodiment is illustrated;

[0055] Figure 40 This is an example data flow diagram of an advanced computing pipeline according to at least one embodiment;

[0056] Figure 41 This is a system diagram of an example system for training, adapting, instantiating, and deploying machine learning models in an advanced computing pipeline, according to at least one embodiment;

[0057] Figure 42 Example illustrations include an advanced computing pipeline 4110A for processing imaging data according to at least one embodiment;

[0058] Figure 43A Includes example data flow diagrams of virtual instruments supporting ultrasound equipment according to at least one embodiment;

[0059] Figure 43B Includes example data flow diagrams of virtual instruments supporting CT scanners according to at least one embodiment;

[0060] Figure 44A A data flow diagram of a process for training a machine learning model according to at least one embodiment is shown;

[0061] Figure 44B This is an example illustration of a client-server architecture that utilizes a pre-trained annotation model to enhance an annotation tool according to at least one embodiment;

[0062] Figure 45 A software stack of a programming platform according to at least one embodiment is shown;

[0063] Figure 46 The illustration shows an embodiment according to at least one of the embodiments. Figure 45 The CUDA implementation of the software stack;

[0064] Figure 47 The illustration shows an embodiment according to at least one of the embodiments. Figure 45 The ROCm implementation method of the software stack;

[0065] Figure 48 The illustration shows an embodiment according to at least one of the embodiments. Figure 45 The OpenCL implementation of the software stack;

[0066] Figure 49 Software supported by a programming platform according to at least one embodiment is shown;

[0067] Figure 50 A method for using at least one embodiment is shown. Figures 45-48 Compiled code executed on the programming platform;

[0068] Figure 51 A multimedia system according to at least one embodiment is shown;

[0069] Figure 52 A distributed system according to at least one embodiment is shown;

[0070] Figure 53 An oversampling neural network according to at least one embodiment is shown;

[0071] Figure 54 An architecture of an oversampling neural network according to at least one embodiment is shown;

[0072] Figure 55 An example of streaming using an oversampling neural network according to at least one embodiment is shown;

[0073] Figure 56 Examples of simulations using an oversampled neural network according to at least one embodiment are shown; and

[0074] Figure 57 An example of a device using an oversampling neural network according to at least one embodiment is shown. Detailed Implementation

[0075] Figure 1 An example system 100 for determining optical flow according to at least one embodiment is shown. System 100 includes optical flow hardware 102 capable of implementing a modified semi-global matching (“SGM”) method. As described below, the modified SGM method differs significantly from the conventional SGM algorithm described in “Accurate and Efficient Stereo Processing by Semi-Global Matching and Mutual Information”, IEEE Conference on Computer Vision and Pattern Recognition (“CVPR”), (June 20-26, 2005), by Heiko Hirschmüller, San Diego, USA, the entire contents of which are incorporated herein by reference.

[0076] refer to Figure 1 Upstream hardware 104 provides reference image 106 and input image 108 (e.g., a pair of stereo images, a pair of video frames, a pair of consecutive images, a pair of simultaneously captured images, etc.) to optical flow hardware 102. Upstream hardware 104 may include at least one data storage device, at least one camera, at least one video camera, a computing device, at least one microcontroller, at least one microprocessor, at least one controller, at least one central processing unit (“CPU”), at least one parallel processing unit (e.g., at least one graphics processing unit (“GPU”)), one or more hardware state machines, etc.

[0077] Reference image 106 and input image 108 may at least partially depict the same subject. For example, reference image 106 and input image 108 may depict one or more objects at different times, from different viewpoints, or from different camera angles. As a non-limiting example, reference image 106 may depict one or more objects at a first location, and input image 108 may depict at least one of one or more objects at the same first location or different second locations. In such embodiments, one or more objects may have moved after reference image 106 is captured but before input image 108 is captured. As another non-limiting example, reference image 106 and input image 108 may depict the same scene captured from different viewpoints or from different camera angles.

[0078] Optical flow hardware 102 maps image regions in reference image 106 to corresponding image regions in input image 108 and outputs a disparity map or optical flow map 110. The optical flow map 110 shows the amount of positional offset (or motion) occurring in at least a portion of the image regions between reference image 106 and input image 108. Optical flow hardware 102 can provide the optical flow map 110 to downstream hardware 112.

[0079] Figure 1 The optical flow hardware 102 is shown to include at least one processor 114, at least one interface 115, a memory 116, and one or more buses 117. Interface 115 is connected to upstream hardware 104 and downstream hardware 112 via connections 119A and 119B, respectively. Each of connections 119A and 119B can be implemented using one or more buses, one or more conductors (e.g., at least one wire, at least one signal trace, and / or the like), one or more switches, and / or the like. One or more interfaces 115 receive a reference image 106 and an input image 108 from upstream hardware 104 via connection 119A and provide the reference image 106 and the input image 108 to one or more processors 114 and / or memory 116 via one or more buses 117. One or more interfaces 115 receive optical flow maps 110 from one or more processors 114 and / or memory 116 via one or more buses 117 and provide the optical flow maps 110 to downstream hardware 112 via connection 119B.

[0080] Optical flow hardware 102 can implement denoising process 120, edge detection process 122, object detection process 124, optional thresholding process 126, decision logic 128, extraction process 130, and SGM process 132. Memory 116 can store processor-executable instructions 118 that, when executed by processor 114, implement denoising process 120, edge detection process 122, object detection process 124, optional thresholding process 126, decision logic 128, extraction process 130, and / or SGM process 132. As a non-limiting example, one or more processors 114 may include at least one microcontroller, at least one microprocessor, at least one controller, at least one CPU, at least one parallel processing unit (e.g., at least one GPU), one or more hardware state machines, etc.

[0081] Instruction 118 can be incorporated into an optical flow software development kit (“SDK”) to be used with one or more parallel processing units (e.g., graphics processing units (“GPUs”, Turing GPUs, Ampere GPUs, etc.) capable of calculating the relative motion of image regions (e.g., pixels) between a reference image 106 and an input image 108. The optical flow SDK can be used to implement computer games, medical imaging software, computer animation, virtual reality, augmented reality, video editing, computer vision, etc. Optical flow hardware 102 can be integrated into computer vision systems that include such parallel processing units. Optical flow hardware 102 can be integrated into autonomous devices (e.g., autonomous vehicles), medical imaging devices, etc. Optical flow graph 110 can be used by optical flow hardware 102 and / or downstream hardware 112 to perform intelligent video analysis. As a non-limiting example, a device incorporating optical flow hardware 102 may include upstream hardware 104 that provides reference image 106 and input image 108 to optical flow hardware 102 and / or downstream hardware 112 that receives optical flow graph 110 from optical flow hardware 102. The device (e.g., an autonomous device) may use optical flow graph 110 in further processing. For example, the device may use optical flow graph 110 to perform motion estimation, object detection, frame generation (e.g., using deep learning or deep neural networks), frame extrapolation, frame interpolation, object tracking, image principal plane extraction, motion detection, robot navigation, visual odometry, camera motion detection, image (e.g., video) compression, image (e.g., video) decompression, etc.

[0082] Figure 2 The illustration depicts optical flow hardware 102 according to at least one embodiment (see [link]). Figure 1 A block diagram of a method for generating an optical flow map 110, executed by [the method described above]. As described above, the optical flow hardware 102 obtains a reference image 106 and an input image 108 (e.g., from [the image source]). Figure 1(Showing upstream hardware 104). Optical flow hardware 102 (e.g., processor 114) can then execute instructions 118 that implement the denoising process 120 to perform the denoising process 120 on the input image 108 and obtain a denoised image 202. The denoising process 120 can forward the denoised image 202 to an edge detection process 122, an object detection process 124, an optional thresholding process 126 (if present), and / or decision logic 128. As a non-limiting example, the denoising process 120 can use spatial filtering, Gaussian filters, median filters, mean filters, frequency domain filters (e.g., notch filters), machine learning techniques, deep convolutional neural networks (“CNN”), denoising autoencoders, bilateral filters, Kuwahara filters, anisotropic diffusion techniques, weighted least squares, edge-avoiding wavelet techniques, edge-preserving filters, geodesic editing, guided filtering, iterative guided filtering, one or more domain transforms, and / or similar methods to generate the denoised image 202.

[0083] Edge detection process 122 receives a denoised image 202 (e.g., from denoising process 120) created by denoising process 120, and optical flow hardware 102 (e.g., processor 114) can execute instructions 118 that implement edge detection process 122 to perform edge detection process 122 on denoised image 202 and obtain edge map 204. Edge detection process 122 can forward denoised image 202 and / or edge map 204 to object detection process 124, optional thresholding process 126 (if present), and / or decision logic 128. As a non-limiting example, edge detection process 122 can use Canny edge detector, Sobel edge detector, etc. to generate edge map 204.

[0084] Object detection process 124 receives an edge map 204 created by edge detection process 122 (e.g., from edge detection process 122) and may optionally receive a denoised image 202 (e.g., from denoising process 120 and / or edge detection process 122). Optical flow hardware 102 (e.g., processor 114) may then execute instructions 118 that implement object detection process 124 to perform object detection process 124 on the denoised image 202 and / or edge map 204 and obtain object map 206. Object detection process 124 may forward the denoised image 202, edge map 204, and / or object map 206 to an optional thresholding process 126 (if present) and / or decision logic 128. As a non-limiting example, object detection process 124 may use neural networks (e.g., cellular neural networks (“CNN”)), scale-invariant feature transform (“SIFT”), etc., to generate object map 206.

[0085] Optionally, the optional thresholding process 126 may receive an edge map 204 created by the edge detection process 122 (e.g., from the edge detection process 122). The optional thresholding process 126 may receive an object map 206 from the object detection process 124, and / or a denoised image 202 from the denoising process 120, the edge detection process 122, and / or the object detection process 124. The optical flow hardware 102 may execute instructions 118 that implement the optional thresholding process 126 to perform the optional thresholding process 126 on the edge map 204 and obtain an optional thresholded edge map 208. The optional thresholding process 126 may forward the denoised image 202, the edge map 204, the object map 206, and / or the optional thresholded edge map 208 to the decision logic 128. As a non-limiting example, the optional thresholding process 126 may generate the optional thresholded edge map 208 by removing any edges from the edge map 204 that are not thicker than a threshold.

[0086] Optical flow hardware 102 (e.g., processor 114) can execute instructions 118 that implement the extraction process 130 to extract a set 210 of image regions from the input image 108. The set 210 of image regions may include feature points and / or pixels. For example, refer to... Figure 3A The extraction process 130 can select every nth pixel (e.g., every fourth pixel) along the rows and columns of the input image 108. In other words, the extraction process 130 can downsample the input image 108. In such an embodiment, the selected pixel can be characterized as being located at the center of a pixel block or as being surrounded by a pixel neighborhood in the input image 108. For ease of illustration, set 210 (composed of array P) i The image region 210 is described as including image regions P1-P9, which will be described as pixels. However, each image region in set 210 can be any part of the image, including feature points. Optionally, the extraction process 130 may skip or otherwise not select image regions along the boundary of the input image 108. However, this is not required, and in at least one embodiment, the extraction process 130 may select one or more image regions along the boundary of the input image 108.

[0087] Optical flow hardware 102 (e.g., processor 114) can execute instructions 118 that implement decision logic 128 to generate one or more disparity maps (e.g., penalty maps PM1 and PM2). Reference Figure 2 The input image 108 can be denoised as an image 202, an edge map 204, an object map 206, an optional thresholded edge map 208 (if present), and a set 210 of image regions P1-P9 (see [link to image map]). Figures 3A-3CThe code is forwarded to decision logic 128, which generates two disparity (penalty) maps, PM1 and PM2, for the set 210 of image regions P1-P9. (See reference.) Figure 3A Decision Logic 128 (see Figure 1 and Figure 2 For each image region P1-P9 in set 210 (in Figure 2 The diagram is shown and is composed of array P. i (representation) determines the direction (by array r) j (Representation). For each image region P1-P9 in set 210, each direction determined for the image region either passes through or intersects with that image region. For example, in Figure 3A In this process, decision logic 128 can determine eight directions for each image region P1-P9 in set 210. In the example shown, directions R1-R8 are shown for image region P5. Directions similar to R1-R8 can be determined for each of the other image regions P1-P4 and P6-P9. Along the boundaries of input image 108, fewer than eight directions can be considered. For example, only three directions can be considered, namely the corners of input image 108 (e.g., image regions P1, P3, P7, and P9), and only five directions along the boundaries of input image 108 between the corners (e.g., in image regions P2, P4, P6, and P8). Alternatively, as described above, extraction process 130 can skip or otherwise not select image regions along the boundaries of input image 108.

[0088] refer to Figure 4 The penalty graphs PM1 and PM2 each include a storage location corresponding to each image region P1-P9 in set 210 (by array P). i (represented) and each direction (by array r) j (Represented). Each storage location stores the disparity (penalty) value for one of the directions and one of the image regions P1-P9. For example, each of the penalty maps PM1 and PM2 can be implemented as a two-dimensional array that stores the disparity (penalty) value for each direction (by array r). j The penalty values ​​are represented by the image regions P1-P9 and the image regions P1-P9. Therefore, when the penalty images PM1 and PM2 are implemented as two-dimensional arrays, one dimension (e.g., rows) corresponds to the image regions P1-P9, and the other dimension (e.g., columns) corresponds to the orientation (determined by the array r). j (Represented by) Correspondingly. As a non-limiting example, penalty maps PM1 and PM2 may store only the disparity values ​​of image regions P1-P9 in set 210. By another non-limiting example, penalty maps PM1 and PM2 may store the disparity values ​​of all image regions (e.g., pixels) in input image 108.

[0089] Decision logic 128 assigns disparity values ​​to each storage location in each penalty map PM1 and PM2. As a non-limiting example, see [reference needed]. Figure 4 The penalty maps PM1 and PM2 can each include eight disparity values ​​of image region P5, where one of the eight disparity values ​​is used for eight directions (by array r). j Each of them in (representation) is in Figures 3A-3C The directions are shown as R1-R8. Similarly, refer to... Figure 4 The penalty maps PM1 and PM2 will each include eight disparity values ​​for each of the image regions P1-P4 and P6-P9, where one of the eight disparity values ​​is used for eight directions similar to those of directions R1-R8 (by array r). j Each of the terms in the expression.

[0090] Decision Logic 128 (see Figure 1 and Figure 2 Based on object graph 206, for each image region P1-P9 in set 210 (by array P), i Each direction (represented by array r) j The decision logic 128 assigns disparity values ​​to the penalty map PM1 based on object map 206. For example, decision logic 128 can assign one of two disparity values, V1a and V1b, to each storage location in the penalty map PM1. Value V1a can be greater than value V1b. For a specific direction and a specific image region in image regions P1-P9, if object map 206 indicates that the specific image region is part of an object identical to at least one adjacent image region in set 210 along the specific direction, decision logic 128 can assign a larger value V1a. Otherwise, the minimum value V1b can be assigned to the penalty map PM1 for the specific image and the specific direction. For example, if object map 206 indicates that the specific image region is part of an object identical to the nearest neighbor image region in set 210 along the specific direction, decision logic 128 can assign a larger value V1a. Otherwise, the minimum value V1b can be assigned to the penalty map PM1 for the specific image region and the specific direction.

[0091] For example, refer to Figure 3B Object detection process 124 (see Figure 1 and Figure 2 The input image 108 (see) can be determined. Figure 1 , Figure 2 , Figure 5 and Figure 6This includes three objects 320, 322, and 324. In this example, image region P1 is shown inside object 320, image regions P2, P3, P6, and P9 are shown inside object 322, and image regions P4, P5, P7, and P8 are shown inside object 324. Along each of directions R6-R8, image regions P4, P7, and P8 are located inside the same object (i.e., object 324) as image region P5. Therefore, as... Figure 4 As shown, decision logic 128 can assign a larger value V1a to the penalty map PM1 of image region P5 for each of directions R6-R8. Furthermore, refer to... Figure 3B Along each of directions R1-R5 and R9, image regions P1-P3, P6, and P9 are located inside objects different from image region P5. Specifically, image region P1 is located inside object 320, while image regions P2, P3, P6, and P9 are located inside object 322. Therefore, as Figure 4 As shown, decision logic 128 can assign a smaller value V1b to the penalty map PM1 for each of the image regions P5 in directions R1-R5. In this example, the penalty map PM1 will store values ​​V1b, V1b, V1b, V1b, V1b, V1a, V1a, and V1a for the image regions P5 in directions R1-R8, respectively.

[0092] Decision Logic 128 (see Figure 1 and Figure 2 Based on edge graph 204 (see...) Figure 2 and Figure 3C ), and optionally, based on an optional thresholded edge map 208 (see Figure 2 For each image region P1-P9 in set 210 (by array P...), i Each direction (represented by array r) j(This indicates that) a penalty value is assigned to the penalty map PM2. For example, decision logic 128 may assign one of three penalty values ​​V2a-V2c to the penalty map PM2 for each image region P1-P9 in the set 210 for each direction. Value V2a can be the largest and value V2c can be the smallest of the three values ​​V2a-V2c. The values ​​V2a-V2c assigned to the penalty map PM2 are greater than the values ​​V1a and V1b assigned to the penalty map PM1. Therefore, value V2a can be the largest penalty and value V1b can be the smallest penalty. For a specific direction and a specific image region in the set 210, if edge map 204 does not indicate that an edge lies between the specific image region and at least one adjacent image region along the specific direction (in the set 210 selected by extraction process 130), decision logic 128 may assign the maximum value V2a. On the other hand, if the optional thresholded edge map 208 indicates that an edge (with a thickness greater than the threshold) lies between a specific image region and at least one adjacent image region (in the set 210 selected by the extraction process 130) along a specific direction, then the decision logic 128 may assign a minimum value V2c. If the edge map 204 indicates that an edge lies between a specific image region and at least one adjacent image region (in the set 210 selected by the extraction process 130) along a specific direction, and the edge thickness is less than or equal to the threshold, then the decision logic 128 may assign an intermediate value V2b.

[0093] For example, Figure 3C An example of edge graph 204 including edges E1-E7 is shown. In this example, only edge E1 is thicker than the threshold and will therefore be included in the optional thresholded edge graph 208 (see [link to example]). Figure 2 In this example, there are no edges between image region P5 and its neighboring image regions P7 and P8 along directions R7 and R6, respectively. Therefore, decision logic 128 can assign the maximum value V2a to image region P5 for directions R6 and R7. Edges E1 with a thickness greater than the threshold are located between image region P5 and its neighboring image regions P1-P3 along directions R1-R3, respectively. Therefore, decision logic 128 can assign the minimum value V2c to image region P5 for directions R1-R3. Edge E3 is located between image region P5 and image region P6 along direction R4, edge E4 is located between image region P5 and image region P9 along direction R5, and edge E5 is located between image region P5 and image region P4 along direction R8. Edges E3-E5 are not thicker than the threshold and therefore will not be included in the optional thresholded edge graph 208. Therefore, decision logic 128 can assign the intermediate value V2b to image region P5 for directions R4, R5, and R8. In this example, as Figure 4As shown, the penalty image PM2 will store the values ​​V2c, V2c, V2c, V2b, V2b, V2a, V2a and V2b for the image region P5 in directions R1-R8 respectively.

[0094] As described above, each penalty map PM1 and PM2 can be implemented as a two-dimensional array. The first dimension (e.g., rows) can be associated with image regions P1-P9 in set 210 (e.g., by array P). i The second dimension (represented by pixels) corresponds to the first dimension (e.g., column) and can be associated with the direction (determined by array r). j (Representation) Correspondingly. For example, penalty images PM1 and PM2 can be represented by arrays PM1(i,j) and PM2(i,j), respectively, where the variable "i" identifies one of the image regions in set 210 and the variable "j" identifies one of the directions.

[0095] Optical flow hardware 102 (e.g., processor 114) can execute instructions 118 that implement the SGM process 132 to generate an optical flow map 110. A set 210 of reference image 106, input image 108, penalty maps PM1 and PM2, and image regions P1-P9 is forwarded to the SGM process 132. Figure 2 The SGM process 132 uses penalty maps PM1 and PM2 to generate an optical flow map 110, which encodes motion from the reference image 106 to the input image 108. The SGM process 132 implements an SGM method that modifies the output optical flow map 110, as described above, which can then be forwarded to downstream hardware 112 (see [link to SGM process 132]). Figure 1 ).

[0096] Figure 5 This is a block diagram showing example content of reference image 106 and example content of input image 108 side by side. In this example, reference image 106 includes image regions R-1 to R-49 arranged in rows MR1-MR7 and columns NR1-NR7. Similarly, input image 108 includes image regions I-1 to I-49 arranged in rows MI1-MI7 and columns NI1-NI7. In this example, image regions P1-P9 (see...) Figures 3A-3C and Figure 7 These correspond to image regions I-1, I-4, I-7, I-22, I-25, I-28, I-43, I-46, and I-49, respectively. The image regions R-17 to R-20, R-24 to R-27, and R-31 to R-34 of reference image 106 depict object 502, which is shown as a rectangle. The same object 502 is also depicted in input image 108 by image regions I-18 to I-21, I-25 to I-28, and I-32 to I-35. Therefore, object 502 may appear to have shifted one column to the right from reference image 106 into input image 108.

[0097] Figure 6 Example values ​​of one or more metrics assigned to each image region R-1 to R-49 in reference image 106 and each image region I-1 to I-49 in input image 108 according to at least one embodiment are shown. These values ​​can be generated by optical flow hardware 102 (see Figure 1 The assignment can be an attribute of the reference image 106 and the input image 108 themselves. The metric can include any parameter or feature of the image region. For example, the metric can include intensity, color, mutual information, etc. For ease of illustration, the values ​​of the metric are already shown. Figure 6 It is described as ranging from zero to ten.

[0098] Figure 7 The illustration shows a set 710 of image regions R-1 to R-49 determined for reference image 106 according to at least one embodiment, a two-dimensional array 706 depicting example values ​​of metrics in image regions PR1-PR9 of set 710, a set 210 determined for input image 108, and a two-dimensional array 708 depicting example values ​​of metrics in image regions P1-P9 of set 210. SGM process 132 (see...) Figure 1 and Figure 2 The location of each image region in the set 710 of image regions R-1 to R-49 in reference image 106 is determined in input image 108. For ease of illustration, reference... Figure 7 Set 710 will be described as including image regions PR1-PR9. Image regions PR1-PR9 correspond to image regions R-1, R-4, R-7, R-22, R-25, R-28, R-43, R-46, and R-49, respectively. Figure 5 As shown. SGM process 132 can use extraction process 130 (see...) Figure 1 and Figure 2 The SGM process 132 may select set 710 from image regions R-1 to R-49, or it may include a separate extraction process (not shown) that is substantially similar to the extraction process 130 for selecting set 710. For example, the SGM process 132 may downsample the reference image 106 to obtain set 710. Alternatively, set 710 may include all image regions R-1 to R-49, in which case the SGM process 132 may determine where the contents of all image regions R-1 to R-49 appear in the input image 108.

[0099] exist Figure 5In the example shown, the content of image region R-25 in reference image 106 appears in image region I-26 of input image 108. Parallax is the distance between a first point in reference image 106 (e.g., image region R-25) and a second point in input image 108 (e.g., image region I-26). For example, if reference image 106 and input image 108 are regularized, rows MR1-MR7 of reference image 106 should correspond to rows MI1-MI7 of input image 108. In other words, input image 108 can be shifted only a few columns relative to reference image 106. Therefore, Figure 1 , Figure 2 , Figure 5 and Figure 6 The example reference image 106 and input image 108 shown have thirteen possible parallaxes (e.g., negative six to six). However, the reference... Figure 7 If image regions PR1-PR9 of set 710 are compared with image regions P1-P9 of set 210, there are only five possible disparities (e.g., negative two to two). On the other hand, if reference image 106 and input image 108 are not regularized, input image 108 may be shifted by several rows and / or columns relative to reference image 106.

[0100] The values ​​of the metric can be used to generate disparity maps. For example, Figure 8 Example disparity maps 802-810 according to at least one embodiment are shown. SGM process 132 (see Figure 1 and Figure 2 Disparity maps 802-810 can be generated by comparing the metric values ​​of image regions PR1-PR9 in set 710 with the metric values ​​of image regions P1-P9 in set 210. In the example shown, disparity maps 802-810 are calculated for disparities from -2 to 2. The disparity maps 802-810 shown store the disparity metric value for each of the image regions PR1-PR9 in set 710 at a disparity of -2 to 2. For example, each of the disparity maps 802-810 includes multiple locations corresponding to the image regions PR1-PR9 in set 710. Within each of the multiple locations, each of the disparity maps 802-810 stores the disparity metric value of the corresponding image region of reference image 106. For ease of illustration, when set 710 is offset from set 210 by disparity, the disparity metric in each disparity map 802-810 shown in the figure is the absolute value of the difference between the metric values ​​of image regions PR1-PR9 and the metric values ​​of image regions P1-P9.

[0101] In other words, disparity map 806 depicts a situation where the disparity is zero, where the SGM process 132 evaluates whether image regions PR1-PR9 correspond to image regions P1-P9. Therefore, disparity map 806 includes a metric for each of image regions PR1-PR9, indicating the difference in values ​​of metrics (e.g., intensity, color, mutual information, etc.) between image regions PR1-PR9 and image regions P1-P9. Similarly, if the disparity is one, the SGM process 132 compares image regions PR1, PR2, PR4, PR5, PR7, and PR8 with image regions P2, P3, P5, P6, P8, and P9. On the other hand, if the disparity is negative one, the SGM process 132 compares image regions PR2, PR3, PR5, PR6, PR8, and PR9 with image regions P1, P2, P4, P5, P7, and P8. Different disparity maps can be calculated in this way for each available disparity.

[0102] SGM process 132 (see also) Figure 1 and Figure 2 Optical flow is determined by calculating the cumulative cost (represented by the expression S(p,d)) of each image region PR1-PR9 in set 710 (represented by the variable p) at each possible disparity (represented by the variable d) with respect to the input image 108. For example, as described above, if the image is regularized, rows MR1-MR7 of the reference image 106 should correspond to rows MI1-MI7 of the input image 108, and the input image 108 can be shifted only by a few columns relative to the reference image 106. Figure 7 In the example shown, SGM procedure 132 can calculate three cumulative costs (each represented by the expression S(p,d)) for each image region PR1-PR9 in set 710. For example, image region PR5 can be compared with image regions P4, P5, and P6 at disparities of -1, zero, and one, respectively. Therefore, in this example, SGM procedure 132 can calculate the cumulative cost (represented by the expression S(p,d)) for image region PR5 for each of the disparities of -1, zero, and one (each represented by the variable d in the expression S(p,d)).

[0103] In SGM process 132 (see...) Figure 1 and Figure 2 After calculating the cumulative cost for each image region PR1-PR9 (represented by variable p) in set 710 at each possible disparity (represented by variable d) with respect to input image 108, SGM process 132 selects the one with the minimum cumulative cost for each image region PR1-PR9 in set 710 (e.g., by the expression min). dS(p,d) represents the selected cumulative cost. For each image region PR1-PR9 in set 710, SGM process 132 includes the selected disparity (or a value determined at least in part based on the selected disparity) in the optical flow map 110 at the location corresponding to the image region.

[0104] The cumulative cost (represented by the expression S(p,d)) is the cost (represented by the expression L) r (p,d) represents the sum of costs, where each cost is calculated along one of a predetermined number of directions traversing a specific image region (represented by variable p) and for a specific disparity (represented by variable d). Therefore, the cumulative cost can be calculated using the following Equation 1:

[0105] S(p,d))=∑ r L r (p,d) Equation 1

[0106] For ease of explanation, the direction of the predetermined quantity will be described as relative to the direction of the predetermined quantity (represented by variable j) used to generate the penalty maps PM1 and PM2 (represented by array r). j (This is the same as the previous statement.) However, this is not necessary, and the SGM process 132 can use different predetermined numbers of directions. Figure 9 The diagram illustrates a method according to at least one embodiment. Figure 1 The system's optical flow hardware 102 is for a specific set 710 of image regions R-1 to R-49 and directions D1-D8. For Figure 9 The example shown can have three disparity values ​​for each image region PR1-PR9, if eight directions D1-D8 are used (which are similar to...). Figures 3A-3C The directions shown are R1-R8). The SGM process 132 can calculate 24 costs for each image region PR1-PR9 (as shown by expression L). r (p,d) represents). In Figure 9 In the image, region PR5 is shown as having eight directions D1-D8 (which is similar to...). Figures 3A-3C The directions shown are R1-R8). SGM process 132 will calculate the cost (by expression L) for the same disparity and the same image region. r (p,d) represents the summation of values ​​to produce the cumulative cost for each image region. As mentioned above, in Figure 8 In the example shown, SGM process 132 will calculate three cumulative costs for each image region PR1-PR9.

[0107] According to Equation 2 below, by combining the matching term (represented by the expression C(p,d)) and the regularization term (represented by the expression R(d)) p ,dq The cost is calculated by adding the terms (as expressed by expression L). r (p,d) represents).

[0108] L r (p,d))=C(p,d))+R(d p ,d q Equation 2

[0109] In Equation 2 above, the matching term (represented by the expression C(p,d)) is a measure of how well a specific image region (e.g., image region PR5) among image regions PR1-PR9 matches one of image regions P1-P9 (e.g., image region P5) at a specific disparity (e.g., disparity zero). Regularization (represented by the expression R(d)) p ,d q The term "(indicated)" improves smoothness by penalizing disparity variations assigned to adjacent image regions.

[0110] SGM procedure 132 can determine a match (represented by the expression C(p,d)) for a specific image region within image regions PR1-PR9 (represented by the variable p) based at least in part on the value of its metric (e.g., intensity) and the value of the metric of one of image regions P1-P9 located at a specific disparity (represented by the variable d). For example, SGM procedure 132 can determine the match using any method, including any method suitable for use by conventional SGM algorithms. As a non-limiting example, a method for computing the match can include the sampling-insensitive measurement described in the paper entitled “Depth discontinuities by pixel-to-pixel stereo”, published on pages 1073-1080 of the proceedings of the 6th IEEE International Conference on Computer Vision in Mumbai, India (January 1998), authored by Birchfield and Tomasi, S. Birchfield and C. Tomasi, which is incorporated herein by reference in its entirety. As another non-limiting example, the method for computing matches can include the mutual information method described by Hirschmüller Heiko in his paper entitled "Accurate and Efficient Stereo Processing by Semi-Global Matching and Mutual Information" presented at the IEEE Conference on Computer Vision and Pattern Recognition (“CVPR”) (June 20-26, 2005) in San Diego, California, USA.

[0111] Now turn to the regularization term (by the expression R(d)). p ,d q The modified SGM method implemented by SGM procedure 132 differs from the traditional SGM algorithm in several ways. For example, the traditional SGM algorithm uses Equation 3 (below) to determine the value of the regularization term (from the expression R(d)). p ,d q )express):

[0112]

[0113] In equation 3 above, the variable d p This represents the computational cost (determined by expression L) of an image region (e.g., image region PR5). r (p,d) represents the disparity metric at the disparity (represented by variable d). Variable d qThis indicates that the computational cost (determined by expression L) of adjacent image regions (e.g., image region PR4) along a direction (e.g., direction D8) is less than that of adjacent image regions (e.g., image region PR4). r (p,d) represents the disparity metric at the disparity (represented by variable d). (See reference) Figure 8 variable d p and d q The value of d can be obtained from the disparity map created for a specific disparity. In other words, the SGM procedure 132 can simply look up the variable d from the disparity map created for a specific disparity. p and d q The value of .

[0114] In Equation 3 above, variables "P1" and "P2" represent penalties. Variables "P1" and "P2" can be two constant parameters, with the value of "P1" being less than the value of "P2". The regularization term is set to zero when the disparity metric remains unchanged. When the disparity metric changes slightly (|d... p -d q When |=1), the regularization term is set to the value of the variable "P1". On the other hand, when the disparity metric varies greatly (|d p -d q |>1), the regularization term is set to be equal to the larger value of variable "P2". A smaller value of variable "P1" penalizes smaller variations and allows one or more image regions PR1-PR9 depicting inclined or curved surfaces to map more accurately to corresponding image regions in input image 108 depicting the same surface. A larger value of variable "P2" helps to maintain discontinuities. To further maintain discontinuities, the larger value of variable "P2" can be adjusted or modified, at least in part, based on comparing the metric of each of image regions PR1-PR9 with the metric of its neighboring image regions. For example, Equation 4 below can be used to determine the value of variable "P2":

[0115]

[0116] In equation 4 above, variable P2′ represents a constant, and variable I bp The variable I represents a metric (e.g., intensity) of an image region. bq Indicates the computational cost (as expressed by expression L). r (p,d) represents the metric (e.g., intensity) of the adjacent image region in the direction that the target is.

[0117] The cost (expressed by expression L) of a specific image region (represented by variable p) at a specific disparity (represented by variable d) along a specific direction (represented by variable r). r (p,d) is typically implemented recursively using Equation 5 (as follows):

[0118] Lr (p,d)=D(p,d) Equation 5

[0119] +min{L r (pr,d),

[0120] L r (pr,d-1)+P1,

[0121] L r (pr,d+1)+P1,

[0122] min i L r (pr,i)+P2}

[0123] -min k L r (pr,k)

[0124] In Equation 5 (above), the expression "pr" represents the previous image region along a specific direction (represented by the variable r) preceding a specific image region (represented by the variable p). For example, refer to... Figure 9 If a specific image region is image region PR5, then the previous image region along direction D1 is image region PR1. In Equation 5 (above), the expression min k L r (pr,k) represents the minimum cost of the previous image region.

[0125] When set 710 includes fewer image regions than the reference image 106 (e.g., image regions PR1-PR9), using a constant value for variable "P1" across the entire reference image 106 and adjusting a larger value for variable "P2" using Equation 4 (above) will not work as expected because the image regions (e.g., pixels) are spaced apart from each other. For example, if these objects are located between image regions for which optical flow is being determined, the conventional SGM algorithm will miss thin or sharp objects. To correct this problem, the modified SGM method implemented by SGM procedure 132 uses Equation 5 below to determine the value of the regularization term (by the expression R(d)). p ,d q )express):

[0126]

[0127] As shown above, the variables “P1” and “P2” present in equations 3 and 5 above are replaced by the expressions P1(p,q) and P2(p,q) in equation 6 above. Expressions P1(p,q) and P2(p,q) refer to penalty maps PM1 and PM2. Specifically, SGM process 132 can look up the penalty value used by the modified SGM method in penalty maps PM1 and PM2. Therefore, in the recursive formula (equation 5 above), expression P1(p,q) replaces the value “P1”, and expression P2(p,q) replaces the value “P2”. For example, if penalty maps PM1 and PM2 are represented by arrays P1(i,j) and P2(i,j), then expression P1(p,q) will return the value at P1(2,3), and expression P2(p,q) will return the value at P2(2,3) for the second image region (e.g., image region PR2) and the third direction (e.g., direction D3). In the modified SGM method, penalty maps PM1 and PM2 adapt the flow generated in a specific image region to the flow in its neighboring image regions. The modified SGM method can be used to compute optical flow between images that are not stereo images because the modified SGM algorithm uses a two-dimensional search window instead of a one-dimensional search for the disparity range.

[0128] Penalized maps PM1 and PM2 may be more helpful in preserving discontinuities than using constant values ​​for variables "P1" and "P2" or using a constant value for variable "P1" and adjusting the value of variable "P2" using Equation 4 above (e.g., based on intensity differences). In particular, when set 710 includes fewer image regions than reference image 106 (e.g., sparse feature points, single pixels per block, etc.), penalized maps PM1 and PM2 can be used well for calculating optical flow.

[0129] Using fewer image regions than all of the reference image 106 (e.g., sparse feature points, each block of a single pixel, etc.) when determining optical flow can help accelerate the computation of optical flow that does not require per-image-region (e.g., per-pixel) optical flow. The modified SGM method allows computation only for those image regions (e.g., pixels) that are geographically separated from each other. For example, optical flow can be computed for every nth (e.g., fourth) image region (e.g., pixel) in both the horizontal and vertical dimensions of the reference image 106.

[0130] Figure 10 It is shown that optical flow hardware 102 (see at least one embodiment) can be used according to at least one embodiment. Figure 1 The flowchart of method 1000 is shown in the first box 1002. Figure 1 The optical flow hardware 102 acquires the reference image 106 and the input image 108. Then, in box 1004 (see...) Figure 10 In the process, optical flow hardware 102 performs a noise reduction process 120 on the input image 108 to obtain a noise-reduced image 202 (see...). Figure 2 ). refer to Figure 2 The noise reduction process 120 forwards the noise reduction image 202 to the edge detection process 122 and / or stores the noise reduction image 202 in the memory 116 (see...). Figure 1 The edge detection process 122 can access the location in the noise reduction process 120. The noise reduction process 120 may optionally forward the noise reduction image 202 to the decision logic 128 and / or store the noise reduction image 202 in the memory 116 (see [link]). Figure 1 The decision logic 128 is accessible at the location.

[0131] Next, in box 1006 (see...) Figure 10 In the optical flow hardware 102, an edge detection process 122 is performed on the denoised image 202 to obtain an edge map 204. The edge detection process 122 may forward the edge map 204 to the object detection process 124 and / or store the edge map 204 in memory 116 (see [link to image 104]). Figure 1 The edge detection process 122 can forward the edge map 204 to an optional thresholding process 126 and / or store the edge map 204 in memory 116 at a location accessible by the optional thresholding process 126. The edge detection process 122 can forward the edge map 204 to decision logic 128 and / or store the edge map 204 in memory 116 at a location accessible by decision logic 128. The edge detection process 122 can forward the denoised image 202 to the object detection process 124, the optional thresholding process 126, and / or decision logic 128. The edge detection process 122 can store the denoised image 202 in memory 116 at a location accessible by the object detection process 124, the optional thresholding process 126, and / or decision logic 128.

[0132] Then, in box 1008 (see Figure 10 In the optical flow hardware 102, an object detection process 124 is performed on the denoised image 202 and / or edge map 204 to obtain an object map 206. The object detection process 124 forwards the object map 206 to the decision logic 128 and / or stores the object map 206 in memory 116 (see [link to relevant documentation]). Figure 1 The object detection process 124 can forward the object image 206 to an optional thresholding process 126 and / or store the object image 206 in memory 116 at a location accessible by the optional thresholding process 126. The object detection process 124 can also forward the edge image 204 and / or the denoised image 202 to an optional thresholding process 126 and / or the decision logic 128. The object detection process 124 can also store the edge image 204 and / or the denoised image 202 in memory 116 at a location accessible by the optional thresholding process 126 and / or the decision logic 128.

[0133] Next, in optional box 1010 (see...) Figure 10 In the optical flow hardware 102, an optional thresholding process 126 can be performed on edge map 204 to obtain an optional thresholded edge map 208. The optional thresholding process 126 forwards the optional thresholded edge map 208 to decision logic 128 and / or stores the optional thresholded edge map 208 in memory 116 (see [link to relevant documentation]). Figure 1 The decision logic 128 can access the decision logic 128. An optional thresholding process 126 may forward the object map 206, edge map 204, and / or denoised image 202 to the decision logic 128. The optional thresholding process 126 may store the object map 206, edge map 204, and / or denoised image 202 in memory 116 at a location accessible to the decision logic 128. In embodiments where optional box 1010 is omitted, the optical flow hardware 102 may execute box 1012 or box 1014 after box 1008.

[0134] In box 1012 (see Figure 10 In the optical flow hardware 102, an extraction process 130 is performed on the input image 108 to obtain a set 210. Although box 1012 is shown in... Figure 10 The optical flow hardware 102 can execute box 1014 at any time before box 1014, but box 1012 can occur after optional box 1010 (or in box 1008 in embodiments where optional box 1010 is omitted). In embodiments where optional box 1010 is present and box 1012 is executed before optional box 1010, the optical flow hardware 102 can execute box 1014 after optional box 1010. Similarly, in embodiments where optional box 1010 is absent and box 1012 is executed before box 1008, the optical flow hardware 102 can execute box 1014 after box 1008.

[0135] In box 1014, refer to Figure 2 The optical flow hardware 102 uses decision logic 128 to determine penalty maps PM1 and PM2 for the input image 108. Decision logic 128 forwards penalty maps PM1 and PM2 to the SGM process 132 and / or stores penalty maps PM1 and PM2 in memory 116 (see [link to SGM process]). Figure 1 The location accessible in the SGM procedure 132.

[0136] In box 1016 (see Figure 10 In the process, optical flow hardware 102 executes SGM procedure 132, which obtains reference image 106 and a set 210 of image regions in input image 108 (see [link]). Figure 2 and Figure 7 Select a set 710 of image regions in reference image 106 (see...) Figure 7 and Figure 9The method then uses penalty maps PM1 and PM2 to determine optical flow map 110. Optical flow map 110 includes metrics for each image region in set 710, indicating the direction and magnitude of motion occurring in the image region from reference image 106 to input image 108. Method 1000 then terminates.

[0137] Reasoning and training logic

[0138] Figure 11A Inference and / or training logic 1115 for performing inference and / or training operations associated with one or more embodiments is shown. The following is in conjunction with... Figure 11A and / or Figure 11B Provide details about reasoning and / or training logic 1115.

[0139] In at least one embodiment, the inference and / or training logic 1115 may include, but is not limited to, code and / or data storage 1101 for storing forward and / or output weights and / or input / output data, and / or other parameters configuring neurons or layers of a neural network trained for and / or used for inference in one or more embodiments. In at least one embodiment, the training logic 1115 may include or be coupled to code and / or data storage 1101 for storing graph code or other software to control timing and / or sequence, wherein weight and / or other parameter information is loaded to configure logic, including integer and / or floating-point units (collectively, arithmetic logic units (ALUs)). In at least one embodiment, code (such as graph code) loads weight or other parameter information into the processor ALU based on the architecture of the neural network to which the code corresponds. In at least one embodiment, the code and / or data storage 1101 stores the weight parameters and / or input / output data of each layer of a neural network trained or used in one or more embodiments during forward propagation of input / output data and / or weight parameters during training and / or inference using one or more embodiments. In at least one embodiment, any portion of the code and / or data storage 1101 may be included within other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory.

[0140] In at least one embodiment, any portion of the code and / or data storage 1101 may be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, the code and / or data storage 1101 may be a cache memory, dynamic random-addressable memory (DRAM), static random-addressable memory (SRAM), non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the choice of whether the code and / or data storage 1101 is internal or external to the processor, for example, or composed of DRAM, SRAM, flash memory, or some other storage type, may depend on the available on-chip or off-chip storage space, the latency requirements of the training and / or inference functions being performed, the batch size of the data used in the inference and / or training of the neural network, or some combination of these factors.

[0141] In at least one embodiment, the inference and / or training logic 1115 may include, but is not limited to, code and / or data storage 1105 for storing backpropagation and / or output weights and / or input / output data neural networks corresponding to neurons or layers of a neural network trained and / or used for inference in one or more embodiments. In at least one embodiment, during training and / or inference using one or more embodiments, the code and / or data storage 1105 stores weight parameters and / or input / output data for each layer of a neural network trained or used in one or more embodiments during backpropagation of input / output data and / or weight parameters. In at least one embodiment, the training logic 1115 may include or be coupled to code and / or data storage 1105 for storing graph code or other software to control timing and / or sequence, wherein weight and / or other parameter information is loaded to configure logic including integer and / or floating-point units (collectively, arithmetic logic units (ALUs)).

[0142] In at least one embodiment, code (such as graph code) causes the architecture of the neural network corresponding to that code to load weights or other parameter information into the processor ALU. In at least one embodiment, any portion of the code and / or data storage 1105 may be included together with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory. In at least one embodiment, any portion of the code and / or data storage 1105 may be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, the code and / or data storage 1105 may be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the choice between the code and / or data storage 1105 being internal or external to the processor, for example, whether it consists of DRAM, SRAM, flash memory, or some other type of storage, depends on whether the available storage is on-chip or off-chip, the latency requirements of the training and / or inference functions being performed, the data batch size used in the inference and / or training of the neural network, or some combination of these factors.

[0143] In at least one embodiment, code and / or data storage 1101 and code and / or data storage 1105 may be separate storage structures. In at least one embodiment, code and / or data storage 1101 and code and / or data storage 1105 may be the same storage structure. In at least one embodiment, code and / or data storage 1101 and code and / or data storage 1105 may be partially combined and partially separated. In at least one embodiment, any portion of code and / or data storage 1101 and code and / or data storage 1105 may be included with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory.

[0144] In at least one embodiment, the inference and / or training logic 1115 may include, but is not limited to, one or more arithmetic logic units (“ALUs”) 1110 (including integer and / or floating-point units) for performing logical and / or mathematical operations at least in part based on or instructed by training and / or inference code (e.g., graph code), the results of which may produce activations (e.g., output values ​​from layers or neurons within a neural network) stored in activation storage 1120, which are functions of input / output and / or weight parameter data stored in code and / or data storage 1101 and / or code and / or data storage 1105. In at least one embodiment, activation is activated in response to execution instructions or other code, and linear algebraic and / or matrix-based mathematical generation performed by ALU 1110 is stored in activation storage 1120, wherein weight values ​​stored in code and / or data storage 1105 and / or code and / or data storage 1101 are used as operands with other values, such as bias values, gradient information, momentum values, or other parameters or hyperparameters, and any or all of these can be stored in code and / or data storage 1105 or code and / or data storage 1101 or other on-chip or off-chip storage.

[0145] In at least one embodiment, one or more processors or other hardware logic devices or circuits include one or more ALUs 1110, while in another embodiment, one or more ALUs 1110 may be located outside the processor or other hardware logic device or the circuitry that uses them (e.g., a coprocessor). In at least one embodiment, one or more ALUs 1110 may be included within an execution unit of a processor, or otherwise included in a group of ALUs accessible by the execution unit of the processor, which may be within the same processor or distributed among different processors of different types (e.g., a central processing unit, a graphics processing unit, a fixed-function unit, etc.). In at least one embodiment, code and / or data storage 1101, code and / or data storage 1105, and activation storage 1120 may share a processor or other hardware logic device or circuitry, while in another embodiment, they may be located in different processors or other hardware logic devices or circuitry, or in some combination of the same and different processors or other hardware logic devices or circuitry. In at least one embodiment, any portion of activation storage 1120 may be included together with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory. Furthermore, inference and / or training code may be stored together with other code accessible to the processor or other hardware logic or circuitry, and may be retrieved and / or processed using the processor’s fetch, decode, schedule, execute, exit, and / or other logic circuitry.

[0146] In at least one embodiment, the active memory 1120 may be a cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other memory. In at least one embodiment, the active memory 1120 may be wholly or partially located inside or outside one or more processors or other logic circuits. In at least one embodiment, the choice of whether the active memory 1120 is internal to or external to the processor may depend on the available on-chip or off-chip storage, the latency requirements for training and / or inference functions, the batch size of data used in inference and / or training the neural network, or some combination of these factors. For example, it may include DRAM, SRAM, flash memory, or other memory types.

[0147] In at least one embodiment, Figure 11A The inference and / or training logic 1115 shown can be used in conjunction with an application-specific integrated circuit (ASIC), such as those from Google. Processing unit, from Graphcore TM Inference processing units (IPUs) or from Intel Corp. (e.g., a Lake Crest processor). In at least one embodiment, Figure 11A The inference and / or training logic 1115 shown can be used in conjunction with central processing unit (CPU) hardware, graphics processing unit (GPU) hardware, or other hardware (such as field programmable gate array (FPGA)).

[0148] Figure 11B An inference and / or training logic 1115 according to at least one embodiment is illustrated. In at least one embodiment, the inference and / or training logic 1115 may include, but is not limited to, hardware logic, wherein computational resources are dedicated or otherwise uniquely used in conjunction with weight values ​​or other information corresponding to one or more layers of neurons within a neural network. In at least one embodiment, Figure 11B The inference and / or training logic 1115 shown can be used in conjunction with an application-specific integrated circuit (ASIC), such as those from Google. Processing unit, from Graphcore TM Inference processing units (IPUs) or from Intel Corp. (e.g., a Lake Crest processor). In at least one embodiment, Figure 11BThe inference and / or training logic 1115 shown can be used in conjunction with central processing unit (CPU) hardware, graphics processing unit (GPU) hardware, or other hardware (e.g., field-programmable gate array (FPGA)). In at least one embodiment, the inference and / or training logic 1115 includes, but is not limited to, code and / or data storage 1101 and code and / or data storage 1105, which can be used to store code (e.g., graph code), weight values, and / or other information, including bias values, gradient information, momentum values, and / or other parameter or hyperparameter information. Figure 11B In at least one embodiment shown, each of code and / or data storage 1101 and code and / or data storage 1105 is associated with dedicated computing resources (e.g., computing hardware 1102 and computing hardware 1106), respectively. In at least one embodiment, each of computing hardware 1102 and computing hardware 1106 includes one or more ALUs that perform mathematical functions (e.g., linear algebraic functions) only on the information stored in code and / or data storage 1101 and code and / or data storage 1105, respectively, and the results of the function execution are stored in activation memory 1120.

[0149] In at least one embodiment, each of the code and / or data storage 1101 and 1105 and the corresponding computing hardware 1102 and 1106 corresponds to a different layer of the neural network, such that an activation obtained from a storage / computation pair 1101 / 1102 of the code and / or data storage 1101 and computing hardware 1102 provides input as input to the next storage / computation pair 1105 / 1106 of the code and / or data storage 1105 and computing hardware 1106, in order to reflect the conceptual organization of the neural network. In at least one embodiment, each storage / computation pair 1101 / 1102 and 1105 / 1106 may correspond to more than one neural network layer. In at least one embodiment, additional storage / computation pairs (not shown) may be included in the inference and / or training logic 1115 following or paralleling the storage / computation pairs 1101 / 1102 and 1105 / 1106.

[0150] System 100 (see) Figure 1 The upstream hardware 104 can be implemented by at least a portion of the inference and / or training logic 1115. For example, computing hardware 1102 and / or computing hardware 1106 can implement the upstream hardware 104 (see...). Figure 1 Optical flow hardware 102 (see) Figure 1 ) and / or downstream hardware 112 (see Figure 1 By way of another non-limiting example, code and / or data storage 1105 can implement memory 116 (see...). Figure 1 ). refer to Figure 1The inference and / or training logic 1115 (see Figure 11) can perform inference and / or training operations for any one of the noise reduction process 120, edge detection process 122, object detection process 124, optional thresholding process 126, decision logic 128, extraction process 130 and SGM process 132.

[0151] Neural network training and deployment

[0152] Figure 12 Training and deployment of a deep neural network according to at least one embodiment are illustrated. In at least one embodiment, an untrained neural network 1206 is trained using a training dataset 1202. In at least one embodiment, the training framework 1204 is the PyTorch framework, while in other embodiments, the training framework 1204 is TensorFlow, Boost, Caffe, Microsoft Cognitive Toolkit / CNTK, MXNet, Chainer, Keras, Deeplearning4j, or other training frameworks. In at least one embodiment, the training framework 1204 trains the untrained neural network 1206 and enables it to be trained using the processing resources described herein to generate a trained neural network 1208. In at least one embodiment, the weights may be randomly selected or pre-trained using a deep belief network. In at least one embodiment, training may be performed in a supervised, partially supervised, or unsupervised manner.

[0153] In at least one embodiment, supervised learning is used to train an untrained neural network 1206, wherein the training dataset 1202 includes inputs paired with desired outputs for input, or wherein the training dataset 1202 includes inputs with known outputs and the neural network 1206 is manually graded output. In at least one embodiment, the untrained neural network 1206 is trained in a supervised manner, and inputs from the training dataset 1202 are processed, and the resulting outputs are compared with a set of expected or desired outputs. In at least one embodiment, errors are then propagated back through the untrained neural network 1206. In at least one embodiment, a training framework 1204 adjusts the weights controlling the untrained neural network 1206. In at least one embodiment, the training framework 1204 includes tools for monitoring the degree to which the untrained neural network 1206 converges to a model (e.g., a trained neural network 1208) adapted to generate the correct answer (e.g., result 1214) based on input data (e.g., a new dataset 1212). In at least one embodiment, the training framework 1204 repeatedly trains the untrained neural network 1206 while adjusting the weights to improve the output of the untrained neural network 1206 using a loss function and tuning algorithm (e.g., stochastic gradient descent). In at least one embodiment, the training framework 1204 trains the untrained neural network 1206 until the untrained neural network 1206 reaches the desired accuracy. In at least one embodiment, the trained neural network 1208 can then be deployed to implement any number of machine learning operations.

[0154] In at least one embodiment, unsupervised learning is used to train an untrained neural network 1206, wherein the untrained neural network 1206 attempts to train itself using unlabeled data. In at least one embodiment, the unsupervised learning training dataset 1202 will include input data without any associated output data or "ground truth" data. In at least one embodiment, the untrained neural network 1206 can learn groupings within the training dataset 1202 and can determine how each input relates to the untrained dataset 1202. In at least one embodiment, unsupervised training can be used to generate a self-organizing graph in the trained neural network 1208, which is capable of performing operations useful for reducing the dimensionality of the new dataset 1212. In at least one embodiment, unsupervised training can also be used to perform anomaly detection, which allows the identification of data points in the new dataset 1212 that deviate from the normal patterns of the new dataset 1212.

[0155] In at least one embodiment, semi-supervised learning can be used, a technique in which a mixture of labeled and unlabeled data is included in the training dataset 1202. In at least one embodiment, the training framework 1204 can be used to perform incremental learning, for example, through transfer learning techniques. In at least one embodiment, incremental learning enables the trained neural network 1208 to adapt to a new dataset 1212 without forgetting the knowledge injected into the trained neural network 1208 during initial training.

[0156] Figure 12 The training and deployment of the deep neural network shown can be used to train and / or deploy one or more deep neural networks used or incorporated by any of the denoising process 120, edge detection process 122, object detection process 124, optional thresholding process 126, decision logic 128, extraction process 130, and SGM process 132. Furthermore, the optical flow map 110, or at least partially derived from the optical flow map 110, can be used as... Figure 12 The input to the deep neural network shown.

[0157] Data Center

[0158] Figure 13 An example data center 1300 that can be used with at least one embodiment is shown. In at least one embodiment, the data center 1300 includes a data center infrastructure layer 1310, a framework layer 1320, a software layer 1330, and an application layer 1340.

[0159] In at least one embodiment, such as Figure 13 As shown, the data center infrastructure layer 1310 may include a resource coordinator 1312, packet computing resources 1314, and node computing resources ("nodes CR") 1316(1)-1316(N), where "N" represents a positive integer (which may be an integer "N" different from the integers used in other diagrams). In at least one embodiment, nodes CR 1316(1)-1316(N) may include, but are not limited to, any number of central processing units ("CPUs") or other processors (including accelerators, field-programmable gate arrays (FPGAs), graphics processors, etc.), memory storage devices 1318(1)-1318(N) (e.g., dynamic read-only memory, solid-state drives, or disk drives), network input / output ("NW I / O") devices, network switches, virtual machines ("VMs"), power modules, and cooling modules, etc. In at least one embodiment, one or more nodes CR 1316(1)-1316(N) may be servers having one or more of the aforementioned computing resources.

[0160] In at least one embodiment, the grouped computing resource 1314 may include individual groups (not shown) of node CRs housed within one or more racks, or a plurality of racks (also not shown) housed within data centers in various geographical locations. In at least one embodiment, the individual groups of node CRs within the grouped computing resource 1314 may include computing, networking, memory, or storage resources that can be configured or allocated to support groups of one or more workloads. In at least one embodiment, several node CRs, including CPUs or processors, may be grouped within one or more racks to provide computing resources to support one or more workloads. In at least one embodiment, the one or more racks may also include any number of power modules, cooling modules, and network switches, in any combination.

[0161] In at least one embodiment, resource coordinator 1312 may be configured or otherwise control one or more nodes CR1316(1)-1316(N) and / or grouped computing resources 1314. In at least one embodiment, resource coordinator 1312 may include a Software Design Infrastructure (“SDI”) management entity for data center 1300. In at least one embodiment, resource coordinator 1112 may include hardware, software, or some combination thereof.

[0162] In at least one embodiment, such as Figure 13 As shown, framework layer 1320 includes a job scheduler 1322, a configuration manager 1324, a resource manager 1326, and a distributed file system 1328. In at least one embodiment, framework layer 1320 may include a framework of software 1332 supporting software layer 1330 and / or one or more applications 1342 supporting application layer 1340. In at least one embodiment, software 1332 or application 1342 may respectively include web-based service software or applications, such as services or applications provided by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, framework layer 1320 may be, but is not limited to, a free and open-source software web application framework, such as Apache Spark, which can utilize distributed file system 1328 for large-scale data processing (e.g., "big data"). TM(Hereinafter referred to as "Spark"). In at least one embodiment, the job scheduler 1322 may include a Spark driver to facilitate the scheduling of workloads supported by various layers of data center 1300. In at least one embodiment, the configuration manager 1324 may be able to configure different layers, such as software layer 1330 and framework layer 1320 including Spark and a distributed file system 1328 for supporting large-scale data processing. In at least one embodiment, the resource manager 1326 is able to manage cluster or group computing resources mapped to or allocated to support distributed file system 1328 and job scheduler 1322. In at least one embodiment, cluster or group computing resources may include group computing resources 1314 on data center infrastructure layer 1310. In at least one embodiment, the resource manager 1326 may coordinate with resource coordinator 1312 to manage these mapped or allocated computing resources.

[0163] In at least one embodiment, the software 1332 included in software layer 1330 may include software used by at least a portion of nodes CR1316(1)-1316(N), grouped computing resources 1314, and / or the distributed file system 1328 of framework layer 1320. In at least one embodiment, one or more types of software may include, but are not limited to, Internet web page search software, email virus scanning software, database software, and streaming video content software.

[0164] In at least one embodiment, one or more applications 1342 included in application layer 1340 may include one or more types of applications used by at least a portion of nodes CR1316(1)-1316(N), grouped computing resources 1314, and / or the distributed file system 1328 of framework layer 1320. In at least one embodiment, one or more types of applications may include, but are not limited to, any number of genomics applications, cognitive computing, applications, and machine learning applications, including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), or other machine learning applications used in conjunction with one or more embodiments.

[0165] In at least one embodiment, any of the configuration manager 1324, resource manager 1326, and resource coordinator 1312 can implement any number and type of self-modification actions based on any amount and type of data acquired in any technically feasible manner. In at least one embodiment, self-modification actions can mitigate potentially poor configuration decisions by data center operators of data center 1300 and can prevent underutilization and / or poor performance of the data center.

[0166] In at least one embodiment, data center 1300 may include tools, services, software, or other resources to train one or more machine learning models or to use one or more machine learning models to predict or infer information according to one or more embodiments described herein. For example, in at least one embodiment, a machine learning model can be trained by calculating weight parameters based on a neural network architecture using the software and computing resources described above with respect to data center 1300. In at least one embodiment, information can be inferred or predicted using trained machine learning models corresponding to one or more neural networks using the resources described above with respect to data center 1300, by using weight parameters calculated through one or more training techniques described herein.

[0167] In at least one embodiment, the data center may use a CPU, application-specific integrated circuit (ASIC), GPU, FPGA, or other hardware to utilize the aforementioned resources to perform training and / or inference. Furthermore, one or more of the aforementioned software and / or hardware resources may be configured as a service to allow a user to train or perform information inference, such as image recognition, speech recognition, or other artificial intelligence services.

[0168] Inference and / or training logic 1115 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 11A and / or Figure 11B Details regarding the inference and / or training logic 1115 are provided. In at least one embodiment, the inference and / or training logic 1115 can be implemented in the system. Figure 13 Used in this context for inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0169] In at least one embodiment, system 100 (see...) Figure 1 This can be implemented by data center 1300. For example, instruction 118 can be executed by one or more of the packet computing resources 814 and / or one or more of CR816(1)-816(N) and used to obtain optical flow map 110 (see Figure 1 and Figure 2 As described above, system 100 (see...) Figure 1 This can be implemented by at least a portion of the reasoning and / or training logic 1115. (See reference...) Figure 1 Data Center 1300 (see Figure 13It can perform inference and / or training operations (e.g., at least a portion of inference and / or training logic 1115) for any of the noise reduction process 120, edge detection process 122, object detection process 124, optional thresholding process 126, decision logic 128, extraction process 130 and SGM process 132.

[0170] Supercomputing

[0171] The following figures illustrate, but are not limited to, exemplary supercomputer-based systems that can be used to implement at least one embodiment.

[0172] In at least one embodiment, a supercomputer can refer to a hardware system exhibiting substantial parallelism and comprising at least one chip, wherein the chips in the system are interconnected via a network and housed in a hierarchically organized enclosure. In at least one embodiment, a large hardware system filling a machine room with several racks is a specific example of a supercomputer, each rack containing several board / rack modules, each board / rack module containing several chips all interconnected via a scalable network. In at least one embodiment, a single rack of such a large hardware system is another example of a supercomputer. In at least one embodiment, a single chip exhibiting substantial parallelism and comprising several hardware components can also be considered a supercomputer because as feature size can be reduced, the amount of hardware that can be incorporated into a single chip can also increase.

[0173] Figure 14A A chip-level supercomputer 1400 according to at least one embodiment is illustrated. In at least one embodiment, within an FPGA or ASIC chip, main computation is performed within a finite state machine (1404) called a thread unit. In at least one embodiment, a task and synchronization network (1402) connects to the finite state machine and is used to schedule threads and perform operations in the correct order. In at least one embodiment, a memory network (1406, 1410) is used to access a multi-level cache hierarchy (1408, 1412) partitioned on the chip. In at least one embodiment, a memory controller (1416) and an off-chip memory network (1414) are used to access off-chip memory. In at least one embodiment, an I / O controller (1418) is used for cross-chip communication when the design is not suitable for a single logic chip.

[0174] Figure 14BA supercomputer at the rack module level is illustrated according to at least one embodiment. In at least one embodiment, within the rack module, there are multiple FPGA or ASIC chips (1420) connected to one or more DRAM cells (1422) constituting the main accelerator memory. In at least one embodiment, each FPGA / ASIC chip is connected to its adjacent FPGA / ASIC chip using a wide on-board bus with differential high-speed signaling (1424). In at least one embodiment, each FPGA / ASIC chip is also connected to at least one high-speed serial communication cable.

[0175] Figure 14C A rack-mounted supercomputer according to at least one embodiment is shown. Figure 14D A supercomputer at the entire system level is illustrated according to at least one embodiment. In at least one embodiment, reference is made to... Figure 14C and Figure 14DHigh-speed serial optical or copper cables (1426, 1428) are used to implement a scalable, potentially incomplete, hypercube network between rack modules within the rack and across racks throughout the system. In at least one embodiment, one of the FPGA / ASIC chips in the accelerator is connected to the host system via a PCI-Express connection (1430). In at least one embodiment, the host system includes a host microprocessor (1434) running the software portion of an application and a memory consisting of one or more host memory DRAM cells (1432) aligned with the memory on the accelerator. In at least one embodiment, the host system may be a standalone module on one of the racks or may be integrated with one of the modules of the supercomputer. In at least one embodiment, a cubic-connected loop topology provides communication links to create a hypercube network for a large supercomputer. In at least one embodiment, a group of FPGA / ASIC chips on a rack module may act as a single hypercube node, increasing the total number of external links per group compared to a single chip. In at least one embodiment, a group comprises chips A, B, C, and D on a rack module having an internal wide differential bus connecting A, B, C, and D in a toroidal organization. In at least one embodiment, there are 12 serial communication cables connecting the rack module to the outside world. In at least one embodiment, chip A on the rack module is connected to serial communication cables 0, 1, and 2. In at least one embodiment, chip B is connected to cables 3, 4, and 5. In at least one embodiment, chip C is connected to cables 6, 7, and 8. In at least one embodiment, chip D is connected to cables 9, 10, and 11. In at least one embodiment, the entire group {A, B, C, D} constituting the rack module can form a hypercube node within a supercomputer system, with up to 2^12 = 4096 rack modules (16384 FPGA / ASIC chips). In at least one embodiment, for chip A to send a message on link 4 of group {A, B, C, D}, the message must first be routed to chip B, which has an onboard differential wide bus connection. In at least one embodiment, messages arriving at group {A, B, C, D} (i.e., arriving at B) on link 4, destined for chip A, must also first be routed to the correct destination chip (A) within group {A, B, C, D}. In at least one embodiment, parallel supercomputer systems of other sizes can also be implemented.

[0176] In at least one embodiment, system 100 (see...) Figure 1 It can be implemented, at least in part, by a supercomputer, for example. Figures 14A-14D One or more of the supercomputers shown. For example, instruction 118 (see...) Figure 1 This can be executed by a supercomputer and used to obtain optical flow diagram 110 (see...). Figure 1 and Figure 2As described above, system 100 (see...) Figure 1 This can be implemented by at least a portion of the reasoning and / or training logic 1115. (See reference...) Figure 1 Supercomputers (see Figures 14A-14D It can perform inference and / or training operations (e.g., at least a portion of inference and / or training logic 1115) for any of the noise reduction process 120, edge detection process 122, object detection process 124, optional thresholding process 126, decision logic 128, extraction process 130 and SGM process 132.

[0177] Computer System

[0178] Figure 15 This is a block diagram illustrating an exemplary computer system according to at least one embodiment. The exemplary computer system may be a system of interconnected devices and components, a system-on-a-chip (SoC), or some combination thereof formed with a processor, which may include an execution unit to execute instructions. In at least one embodiment, according to this disclosure, such as the embodiments described herein, computer system 1500 may include, but is not limited to, components such as processor 1502, whose execution unit includes logic to execute algorithms for process data. In at least one embodiment, computer system 1500 may include a processor, such as one available from Intel Corporation of Santa Clara, California. Processor family, Xeon TM , XScale TM and / or StrongARM TM , Core TM or Nervana TM A microprocessor may be used, although other systems (including PCs, engineering workstations, set-top boxes, etc.) with other microprocessors may also be used. In at least one embodiment, computer system 1500 may execute a version of the Windows operating system available from Microsoft Corporation of Redmond, Washington, although other operating systems (such as UNIX and Linux), embedded software, and / or graphical user interfaces may also be used.

[0179] The embodiments can be used in other devices, such as handheld devices and embedded applications. Some examples of handheld devices include cellular phones, Internet Protocol (IP) devices, digital cameras, personal digital assistants (PDAs), and handheld PCs. In at least one embodiment, the embedded application may include a microcontroller, a digital signal processor (DSP), a system-on-a-chip (SoC), a network computer (NetPC), a set-top box, a network hub, a wide area network (WAN) switch, or any other system that can execute one or more instructions according to at least one embodiment.

[0180] In at least one embodiment, the computer system 1500 may include, but is not limited to, a processor 1502, which may include, but is not limited to, one or more execution units 1508, to perform machine learning model training and / or inference according to the techniques described herein. In at least one embodiment, the computer system 1500 is a single-processor desktop or server system, but in another embodiment, the computer system 1500 may be a multiprocessor system. In at least one embodiment, the processor 1502 may include, but is not limited to, a Complex Instruction Set Computer (CISC) microprocessor, a Reduced Instruction Set Computing (RISC) microprocessor, a Very Long Instruction Word (VLIW) microprocessor, a processor implementing instruction set combinations, or any other processor device, such as a digital signal processor. In at least one embodiment, the processor 1502 may be coupled to a processor bus 1510, which can transmit data signals between the processor 1502 and other components in the computer system 1500.

[0181] In at least one embodiment, processor 1502 may include, but is not limited to, a Level 1 (L1) internal cache memory ("cache") 1504. In at least one embodiment, processor 1502 may have a single internal cache or multiple levels of internal cache. In at least one embodiment, the cache memory may reside outside of processor 1502. Depending on specific implementation and requirements, other embodiments may also include a combination of internal and external caches. In at least one embodiment, register file 1506 may store different types of data in various registers, including but not limited to integer registers, floating-point registers, status registers, and instruction pointer registers.

[0182] In at least one embodiment, an execution unit 1508, including but not limited to logic for performing integer and floating-point operations, is also located within the processor 1502. In at least one embodiment, the processor 1502 may further include a microcode ("ucode") read-only memory ("ROM") for storing microcode of certain macro instructions. In at least one embodiment, the execution unit 1508 may include logic for processing a packaged instruction set 1509. In at least one embodiment, by including the packaged instruction set 1509 in the instruction set of a general-purpose processor, along with the associated circuitry for executing the instructions, the packaged data in the processor 1502 can be used to perform operations used by numerous multimedia applications. In at least one embodiment, many multimedia applications can be executed more quickly and efficiently by using the full width of the processor's data bus to perform operations on the packaged data, which may eliminate the need to transfer smaller data units on the processor's data bus to perform one or more operations on one data element at a time.

[0183] In at least one embodiment, the execution unit 1508 may also be used in a microcontroller, embedded processor, graphics device, DSP, and other types of logic circuitry. In at least one embodiment, the computer system 1500 may include, but is not limited to, memory 1520. In at least one embodiment, memory 1520 may be a dynamic random access memory (DRAM) device, a static random access memory (SRAM) device, a flash memory device, or another storage device. In at least one embodiment, memory 1520 may store instructions 1519 and / or data 1521 represented by data signals that can be executed by processor 1502.

[0184] In at least one embodiment, the system logic chip may be coupled to processor bus 1510 and memory 1520. In at least one embodiment, the system logic chip may include, but is not limited to, a memory controller hub (MCH) 1516, and processor 1502 may communicate with MCH 1516 via processor bus 1510. In at least one embodiment, MCH 1516 may provide a high-bandwidth memory path 1518 to memory 1520 for instruction and data storage, as well as for storage of graphics commands, data, and textures. In at least one embodiment, MCH 1516 may initiate data signals between processor 1502, memory 1520, and other components in computer system 1500, and bridge data signals between processor bus 1510, memory 1520, and system I / O interface 1522. In at least one embodiment, the system logic chip may provide a graphics port for coupling to a graphics controller. In at least one embodiment, MCH 1516 can be coupled to memory 1520 via high-bandwidth memory path 1518, and graphics / video card 1512 can be coupled to MCH 1516 via Accelerated Graphics Port (AGP) interconnect 1514.

[0185] In at least one embodiment, the computer system 1500 may use the system I / O interface 1522 as a proprietary hub interface bus to couple the MCH 1516 to the I / O controller hub (ICH) 1530. In at least one embodiment, the ICH 1530 may provide direct connectivity to certain I / O devices via a local I / O bus. In at least one embodiment, the local I / O bus may include, but is not limited to, a high-speed I / O bus for connecting peripheral devices to the memory 1520, chipset, and processor 1502. Examples may include, but are not limited to, an audio controller 1529, a firmware hub (Flash BIOS) 1528, a wireless transceiver 1526, a data storage 1524, a conventional I / O controller 1523 including a user input and keyboard interface 1525, a serial expansion port 1527 (e.g., a Universal Serial Bus (USB) port), and a network controller 1534. In at least one embodiment, the data storage 1524 may include a hard disk drive, a floppy disk drive, a CD-ROM device, a flash memory device, or other mass storage device.

[0186] In at least one embodiment, Figure 15 A system including interconnected hardware devices or "chips" is shown, while in other embodiments, Figure 15 The SoC can be shown. In at least one embodiment, Figure 15The devices shown can be interconnected with proprietary interconnects, standardized interconnects (e.g., PCIe), or some combination thereof. In at least one embodiment, one or more components of the computer system 1500 are interconnected using a Computational Fast Link (CXL) interconnect.

[0187] The inference and / or training logic 1115 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 11A and / or Figure 11B Details are provided regarding the inference and / or training logic 1115. In at least one embodiment, the inference and / or training logic 1115 can... Figure 15 Used in systems for reasoning or predicting operations based at least in part on weight parameters calculated using neural network training operations, neural network functions and / or architectures or neural network use cases described herein.

[0188] In at least one embodiment, system 100 (see...) Figure 1 The processor 114 can be implemented at least partially by the computer system 1500. For example, the processor 114 can be implemented by the processor 1502 and / or the graphics / video card 1512, the interface 115 can be implemented at least partially by the network controller 1534, the memory 116 can be implemented by the memory 1520, and the bus 117 can be implemented at least partially by the processor bus 1510 and / or the AGP interconnect 1514. As another non-limiting example, instruction 118 (see...) Figure 1 It can be stored in memory 1520, executed by processor 1502 and / or graphics / video card 1512, and used to obtain optical flow map 110 (see...). Figure 1 and Figure 2 As described above, system 100 (see...) Figure 1 This can be implemented by at least a portion of the reasoning and / or training logic 1115. (See reference...) Figure 1 Computer System 1500 (see Figure 15 It can perform inference and / or training operations (e.g., at least a portion of inference and / or training logic 1115) for any of the noise reduction process 120, edge detection process 122, object detection process 124, optional thresholding process 126, decision logic 128, extraction process 130 and SGM process 132.

[0189] Figure 16 This is a block diagram illustrating an electronic device 1600 for utilizing a processor 1610 according to at least one embodiment. In at least one embodiment, the electronic device 1600 may be, for example, but not limited to, a laptop computer, tower server, rack server, blade server, laptop computer, desktop computer, tablet computer, mobile device, telephone, embedded computer, or any other suitable electronic device.

[0190] In at least one embodiment, the electronic device 1600 may, but is not limited to, a processor 1610 communicatively coupled to any suitable number or type of components, peripherals, modules, or devices. In at least one embodiment, the processor 1610 is coupled using a bus or interface, such as I... 2 C-bus, System Management Bus (SMBus), Low Pin Count (LPC) bus, Serial Peripheral Interface (SPI), High Definition Audio (HDA) bus, Serial Advanced Technology Accessory (SATA) bus, Universal Serial Bus (USB) (versions 1, 2, 3, etc.), or Universal Asynchronous Receiver / Transmitter (UART) bus. In at least one embodiment, Figure 16 The system shown includes interconnected hardware devices or "chips," while in other embodiments, Figure 16 An exemplary SoC can be shown. In at least one embodiment, Figure 16 The device shown can be interconnected with proprietary interconnects, standardized interconnects (e.g., PCIe), or some combination thereof. In at least one embodiment, Figure 16 One or more components are interconnected using Computational Fast Link (CXL) interconnects.

[0191] In at least one embodiment, Figure 16 It may include a display 1624, a touch screen 1625, a touchpad 1630, a near field communication unit (NFC) 1645, a sensor hub 1640, a thermal sensor 1646, a fast chipset (EC) 1635, a trusted platform module (TPM) 1638, a BIOS / firmware / flash memory (BIOS, FW Flash) 1622, a DSP 1660, a driver 1620 (e.g., a solid-state drive (SSD) or a hard disk drive (HDD)), a wireless local area network unit (WLAN) 1650, a Bluetooth unit 1652, a wireless wide area network unit (WWAN) 1656, a global positioning system (GPS) unit 1655, a camera (USB 3.0 camera) 1654 (e.g., a USB 3.0 camera), and / or a low-power double data rate (LPDDR) memory unit (LPDDR3) 1615 implemented in, for example, the LPDDR3 standard. These components can each be implemented in any suitable way.

[0192] In at least one embodiment, other components may be communicatively coupled to processor 1610 via the components described herein. In at least one embodiment, accelerometer 1641, ambient light sensor (ALS) 1642, compass 1643, and gyroscope 1644 may be communicatively coupled to sensor hub 1640. In at least one embodiment, thermal sensor 1639, fan 1637, keyboard 1636, and touchpad 1630 may be communicatively coupled to EC 1635. In at least one embodiment, speaker 1663, earphone 1664, and microphone (mic) 1665 may be communicatively coupled to audio unit (audio codec and Class D amplifier) ​​1662, which in turn may be communicatively coupled to DSP 1660. In at least one embodiment, audio unit 1662 may include, for example, but not limited to, audio encoder / decoder (codec) and Class D amplifier. In at least one embodiment, SIM card (SIM) 1657 may be communicatively coupled to WWAN unit 1656. In at least one embodiment, components such as WLAN unit 1650, Bluetooth unit 1652, and WWAN unit 1656 can be implemented as next-generation form factor (NGFF).

[0193] Inference and / or training logic 1115 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 11A and / or Figure 11B Details regarding the inference and / or training logic 1115 are provided. In at least one embodiment, the inference and / or training logic 1115 can be implemented in the system. Figure 16 It is used in the context of reasoning or predicting operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures or neural network use cases described herein.

[0194] In at least one embodiment, system 100 (see...) Figure 1 This can be implemented, at least in part, by electronic device 1600. For example, processor 114 can be implemented by processor 1610 and upstream hardware 104 can be implemented by camera 1654. As another non-limiting example, instruction 118 (see...) Figure 1 This can be executed by processor 1610 to obtain optical flow map 110 from the image captured by camera 1654 (see...). Figure 1 and Figure 2 As described above, system 100 (see...) Figure 1 This can be implemented by at least a portion of the reasoning and / or training logic 1115. (See reference...) Figure 1 Electronic equipment 1600 (see Figure 16It can perform inference and / or training operations (e.g., at least a portion of inference and / or training logic 1115) for any of the noise reduction process 120, edge detection process 122, object detection process 124, optional thresholding process 126, decision logic 128, extraction process 130 and SGM process 132.

[0195] Figure 17 A computer system 1700 according to at least one embodiment is shown. In at least one embodiment, the computer system 1700 is configured to implement various processes and methods described throughout this disclosure.

[0196] In at least one embodiment, the computer system 1700 includes, but is not limited to, at least one central processing unit (CPU) 1702 connected to a communication bus 1710 implemented using any suitable protocol, such as PCI (Peripheral Interconnect), Peripheral Component Interconnect Express (PCI-Express), AGP (Accelerated Graphics Port), HyperTransport, or any other bus or point-to-point communication protocol. In at least one embodiment, the computer system 1700 includes, but is not limited to, main memory 1704 and control logic (e.g., implemented in hardware, software, or a combination thereof), and data may be stored in main memory 1704 in the form of random access memory (RAM). In at least one embodiment, a network interface subsystem (network interface) 1722 provides an interface to other computing devices and networks for receiving data using the computer system 1700 and transferring data to other systems.

[0197] In at least one embodiment, the computer system 1700 includes, but is not limited to, an input device 1708, a parallel processing system 1712, and a display device 1706, which may be implemented using conventional cathode ray tube (CRT), liquid crystal display (LCD), light-emitting diode (LED) display, plasma display, or other suitable display technologies. In at least one embodiment, user input is received from the input device 1708 (such as a keyboard, mouse, touchpad, microphone, etc.). In at least one embodiment, each of the modules described herein may reside on a single semiconductor platform to form the processing system.

[0198] Inference and / or training logic 1115 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 11A and / or Figure 11B Details regarding the inference and / or training logic 1115 are provided. In at least one embodiment, the inference and / or training logic 1115 can be implemented in the system. Figure 17It is used to perform inference or prediction operations based at least in part on weight parameters calculated using neural network training operations, neural network functions and / or architecture or neural network use cases described herein.

[0199] In at least one embodiment, system 100 (see...) Figure 1 The processor 114 can be implemented at least partially by computer system 1700. For example, processor 114 can be implemented by CPU 1702 and / or one or more PPUs 1714, interface 115 can be implemented at least partially by network interface 1722, memory 116 can be implemented at least partially by main memory 1704, and bus 117 can be implemented at least partially by communication bus 1710, interconnect 1718, and / or switch 1720. As another non-limiting example, instruction 118 (see...) Figure 1 It can be stored in main memory 1704, executed by CPU 1702 and / or one or more PPUs 1714, and used to obtain optical flow map 110 (see...). Figure 1 and Figure 2 The upstream hardware 104 can be implemented by a camera or storage device.

[0200] As described above, system 100 (see Figure 1 This can be implemented by at least a portion of the reasoning and / or training logic 1115. (See reference...) Figure 1 Computer System 1700 (see Figure 17 It can perform inference and / or training operations (e.g., at least a portion of inference and / or training logic 1115) for any of the noise reduction process 120, edge detection process 122, object detection process 124, optional thresholding process 126, decision logic 128, extraction process 130 and SGM process 132.

[0201] Figure 18 A computer system 1800 according to at least one embodiment is illustrated. In at least one embodiment, the computer system 1800 includes, but is not limited to, a computer 1810 and a USB stick 1820. In at least one embodiment, the computer 1810 may include, but is not limited to, any number and type of processors (not shown) and memory (not shown). In at least one embodiment, the computer 1810 includes, but is not limited to, a server, a cloud instance, a laptop computer, and a desktop computer.

[0202] In at least one embodiment, the USB stick 1820 includes, but is not limited to, a processing unit 1830, a USB interface 1840, and USB interface logic 1850. In at least one embodiment, the processing unit 1830 can be any instruction execution system, apparatus, or device capable of executing instructions. In at least one embodiment, the processing unit 1830 can include, but is not limited to, any number and type of processing cores (not shown). In at least one embodiment, the processing unit 1830 includes an application-specific integrated circuit (“ASIC”) optimized to perform any amount and type of operations associated with machine learning. For example, in at least one embodiment, the processing unit 1830 is a tensor processing unit (“TPC”) optimized to perform machine learning inference operations. In at least one embodiment, the processing unit 1830 is a vision processing unit (“VPU”) optimized to perform machine vision and machine learning inference operations.

[0203] In at least one embodiment, the USB interface 1840 can be any type of USB connector or USB receptacle. For example, in at least one embodiment, the USB interface 1840 is a USB 3.0 Type-C receptacle for data and power. In at least one embodiment, the USB interface 1840 is a USB 3.0 Type-A connector. In at least one embodiment, the USB interface logic 1850 may include any amount and type of logic that enables the processing unit 1830 to interface with a device (e.g., computer 1810) via the USB connector 1840.

[0204] Inference and / or training logic 1115 is used to perform inference and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 1115 are combined herein. Figure 11A And / or 11B is provided. In at least one embodiment, the inference and / or training logic 1115 can be provided in the system. Figure 18 The operation is used to infer or predict based at least in part on weight parameters calculated using neural network training operations, neural network functions and / or architecture, or neural network use cases as described herein.

[0205] In at least one embodiment, system 100 (see...) Figure 1 The processor 114 may be at least partially implemented by the computer system 1800. For example, the processor 114 may be a component of the computer 1810 and / or may be implemented by the processing unit 1830, the memory 116 may be implemented by the computer 1810 and / or the USB stick 1820, and the bus 117 may be at least partially implemented by the USB interface 1840 and the USB interface logic 1850. As another non-limiting example, instruction 118 (see...) Figure 1The optical flow map 110 can be stored by computer 1810 and / or USB memory stick 1820, executed by computer 1810 and / or processing unit 1830, and used to obtain the optical flow map 110 (see...). Figure 1 and Figure 2 The upstream hardware 104 can be implemented by a camera or storage device.

[0206] As described above, system 100 (see Figure 1 This can be implemented by at least a portion of the reasoning and / or training logic 1115. (See reference...) Figure 1 Computer System 1800 (see Figure 18 It can perform inference and / or training operations (e.g., at least a portion of inference and / or training logic 1115) for any of the noise reduction process 120, edge detection process 122, object detection process 124, optional thresholding process 126, decision logic 128, extraction process 130 and SGM process 132.

[0207] Figure 19A An exemplary architecture is shown in which multiple GPUs 1910(1)-1910(N) are communicatively coupled to multiple multi-core processors 1905(1)-1905(M) via high-speed links 1940(1)-1940(N) (e.g., bus / point-to-point interconnect, etc.). In at least one embodiment, the high-speed links 1940(1)-1940(N) support communication throughput of 4GB / s, 30GB / s, 80GB / s, or higher. In at least one embodiment, various interconnect protocols may be used, including but not limited to PCIe 4.0 or 5.0 and NVLink 2.0. In the various figures, "N" and "M" represent positive integers, the values ​​of which may vary from figure to figure.

[0208] Furthermore, in at least one embodiment, two or more GPUs 1910 are interconnected via high-speed links 1929(1)-1929(2), which can be implemented using a protocol / link similar to or different from that used for high-speed links 1940(1)-1940(N). Similarly, two or more multi-core processors 1905 can be connected via high-speed link 1928, which can be a symmetric multiprocessor (SMP) bus operating at speeds of 20GB / s, 30GB / s, 120GB / s, or higher. Alternatively, similar protocols / links (e.g., via a common interconnect structure) can be used. Figure 19A This shows all communication between the various system components.

[0209] In at least one embodiment, each multi-core processor 1905 is communicatively coupled to processor memories 1901(1)-1901(M) via memory interconnects 1926(1)-1926(M), and each GPU 1910(1)-1910(N) is communicatively coupled to GPU memories 1920(1)-1920(N) via GPU memory interconnects 1950(1)-1950(N). In at least one embodiment, memory interconnects 1926 and 1950 may utilize similar or different memory access technologies. By way of example and not limitation, processor memories 1901(1)-1901(M) and GPU memories 1920 may be volatile memories, such as dynamic random access memory (DRAM) (including stacked DRAM), graphics DDR SDRAM (GDDR) (e.g., GDDR5, GDDR6), or high bandwidth memory (HBM), and / or may be non-volatile memories, such as 3D XPoint or Nano-RAM. In at least one embodiment, some portions of the processor memory 1901 may be volatile memory, while other portions may be non-volatile memory (e.g., using a two-level memory (2LM) hierarchy).

[0210] As described herein, although various multi-core processors 1905 and GPUs 1910 can be physically coupled to specific memories 1901 and 1920 respectively, and / or can implement a unified memory architecture, in which the virtual system address space (also known as the "effective address" space) is distributed among the various physical memories. For example, processor memories 1901(1)-1901(M) can each contain 64GB of system memory address space, and GPU memories 1920(1)-1920(N) can each contain 32GB of system memory address space, resulting in a total addressable memory size of 256GB when M=2 and N=4. N and M may also be other values.

[0211] Figure 19B Additional details are shown regarding the interconnection between a multi-core processor 1907 and a graphics acceleration module 1946 according to an exemplary embodiment. In at least one embodiment, the graphics acceleration module 1946 may include one or more GPU chips integrated on a line card coupled to the processor 1907 via a high-speed link 1940 (e.g., PCIe bus, NVLink, etc.). In at least one embodiment, the graphics acceleration module 1946 may optionally be integrated on a package or chip having the processor 1907.

[0212] In at least one embodiment, the processor 1907 includes multiple cores 1960A-1960D, each core having a translation back cover (TLB) 1961A-1961D and one or more caches 1962A-1962D. In at least one embodiment, the cores 1960A-1960D may include various other components (not shown) for executing instructions and processing data. In at least one embodiment, the caches 1962A-1962D may include level 1 (L1) and level 2 (L2) caches. Furthermore, one or more shared caches 1956 may be included in the caches 1962A-1962D and shared by the respective groups of cores 1960A-1960D. For example, one embodiment of the processor 1907 includes 24 cores, each core having its own L1 cache, twelve shared L2 caches, and twelve shared L3 caches. In this embodiment, two adjacent cores share one or more L2 and L3 caches. In at least one embodiment, the processor 1907 and the graphics acceleration module 1946 are connected to a system memory 1914, which may include... Figure 19A The processor memory in the memory is 1901(1)-1901(M).

[0213] In at least one embodiment, consistency of data and instructions stored in the various caches 1962A-1962D, 1956 and system memory 1914 is maintained via inter-core communication through the consistency bus 1964. In at least one embodiment, for example, each cache may have associated cache consistency logic / circuit to communicate via the consistency bus 1964 in response to the detection of a read or write to a particular cache line. In at least one embodiment, a cache snooping protocol is implemented via the consistency bus 1964 to snoop on cache accesses.

[0214] In at least one embodiment, proxy circuitry 1925 communicatively couples graphics acceleration module 1946 to coherence bus 1964, thereby allowing graphics acceleration module 1946 to participate in cache coherence protocols as a peer of cores 1960A-1960D. Specifically, in at least one embodiment, interface 1935 provides connectivity to proxy circuitry 1925 via high-speed link 1940, and interface 1937 connects graphics acceleration module 1946 to high-speed link 1940.

[0215] In at least one embodiment, the accelerator integrated circuit 1936 provides cache management, memory access, context management, and interrupt management services for a plurality of graphics processing engines 1931(1)-1931(N) of the graphics acceleration module 1946. In at least one embodiment, the graphics processing engines 1931(1)-1931(N) may each include a separate graphics processing unit (GPU). In at least one embodiment, the graphics processing engines 1931(1)-1931(N) may optionally include different types of graphics processing engines within the GPU, such as graphics execution units, media processing engines (e.g., video encoders / decoders), samplers, and blit engines. In at least one embodiment, the graphics acceleration module 1946 may be a GPU having a plurality of graphics processing engines 1931(1)-1931(N), or the graphics processing engines 1931(1)-1931(N) may be individual GPUs integrated on a general-purpose package, line card, or chip.

[0216] In at least one embodiment, the accelerator integrated circuit 1936 includes a memory management unit (MMU) 1939 for performing various memory management functions, such as virtual-to-physical memory translation (also known as effective-to-real memory translation), and a memory access protocol for accessing system memory 1914. In at least one embodiment, the MMU 1939 may also include a translation back buffer ("TLB") (not shown) for caching virtual / effective-to-physical / real address translations. In at least one embodiment, cache 1938 may store commands and data for efficient access by graphics processing engines 1931(1)-1931(N). In at least one embodiment, a fetch unit 1944 may be used to keep data stored in cache 1938 and graphics memory 1933(1)-1933(M) consistent with core caches 1962A-1962D, 1956 and system memory 1914. As previously mentioned, this task can be accomplished via proxy circuitry 1925 representing cache 1938 and graphics memory 1933(1)-1933(M) (e.g., sending updates related to modifications / accesses to cache lines on processor caches 1962A-1962D, 1956 to cache 1938 and receiving updates from cache 1938).

[0217] In at least one embodiment, a set of registers 1945 stores context data of threads executed by graphics processing engines 1931(1)-1931(N), and context management circuitry 1948 manages the thread context. For example, context management circuitry 1948 can perform save and restore operations to save and restore the context of individual threads during context switching (e.g., saving the first thread and storing the second thread so that the second thread can be executed by the graphics processing engine). For example, during context switching, context management circuitry 1948 can store the current register value into a designated area in memory (e.g., identified by a context pointer). The register value can then be restored when returning to the context. In at least one embodiment, interrupt management circuitry 1947 receives and processes interrupts received from system devices.

[0218] In at least one embodiment, virtual / effective addresses from graphics processing engine 1931 are translated into real / physical addresses in system memory 1914 via MMU 1939. In at least one embodiment, accelerator integrated circuit 1936 supports multiple (e.g., 4, 8, 16) graphics accelerator modules 1946 and / or other accelerator devices. In at least one embodiment, graphics accelerator module 1946 may be dedicated to a single application executing on processor 1907, or may be shared among multiple applications. In at least one embodiment, a virtualized graphics execution environment is presented, wherein resources of graphics processing engines 1931(1)-1931(N) are shared with multiple applications or virtual machines (VMs). In at least one embodiment, resources may be subdivided into "slices" based on processing requirements and priorities associated with VMs and / or applications, which are allocated to different VMs and / or applications.

[0219] In at least one embodiment, the accelerator integrated circuit 1936 acts as a bridge to the system of the graphics acceleration module 1946 and provides address translation and system memory caching services. Additionally, in at least one embodiment, the accelerator integrated circuit 1936 can provide virtualization facilities for the host processor to manage the virtualization, interrupt, and memory management of the graphics processing engines 1931(1)-1931(N).

[0220] In at least one embodiment, since the hardware resources of the graphics processing engines 1931(1)-1931(N) are explicitly mapped to the real address space seen by the host processor 1907, any host processor can directly address these resources using valid address values. In at least one embodiment, a function of the accelerator integrated circuit 1936 is to physically separate the graphics processing engines 1931(1)-1931(N) so that they appear as independent units to the system.

[0221] In at least one embodiment, one or more graphics memories 1933(1)-1933(M) are coupled to each graphics processing engine 1931(1)-1931(N), and N = M. In at least one embodiment, the graphics memories 1933(1)-1933(M) store instructions and data processed by each graphics processing engine 1931(1)-1931(N). In at least one embodiment, the graphics memories 1933(1)-1933(M) may be volatile memory, such as DRAM (including stacked DRAM), GDDR memory (e.g., GDDR5, GDDR6), or HBM, and / or may be non-volatile memory, such as 3DXPoint or Nano-RAM.

[0222] In at least one embodiment, to reduce data traffic on the high-speed link 1940, a biasing technique can be used to ensure that the data stored in the graphics memory 1933(1)-1933(M) is the data most frequently used by the graphics processing engine 1931(1)-1931(N) and preferably not used (at least infrequently used) by the cores 1960A-1960D. Similarly, in at least one embodiment, the biasing mechanism attempts to keep the data needed by the cores (and preferably not the graphics processing engine 1931(-1)-1931(N)) in the caches 1962A-1962D, 1956 and system memory 1914.

[0223] Figure 19C Another exemplary embodiment is shown, in which the accelerator integrated circuit 1936 is integrated within the processor 1907. In this embodiment, the graphics processing engines 1931(1)-1931(N) communicate directly with the accelerator integrated circuit 1936 via a high-speed link 1940 through interfaces 1937 and 1935 (which can also be any form of bus or interface protocol). In at least one embodiment, the accelerator integrated circuit 1936 can perform operations related to... Figure 19B The described operation is similar. However, due to its close proximity to the coherence bus 1964 and caches 1962A-1962D, 1956, it may have higher throughput. In at least one embodiment, the accelerator integrated circuit supports different programming models, including a dedicated process programming model (without graphics acceleration module virtualization) and a shared programming model (with virtualization), which may include a programming model controlled by the accelerator integrated circuit 1936 and a programming model controlled by the graphics acceleration module 1946.

[0224] In at least one embodiment, graphics processing engines 1931(1)-1931(N) are dedicated to a single application or process under a single operating system. In at least one embodiment, a single application can funnel requests from other applications to graphics processing engines 1931(1)-1931(N), thereby providing virtualization within a VM / partition.

[0225] In at least one embodiment, graphics processing engines 1931(1)-1931(N) can be shared by multiple VM / application partitions. In at least one embodiment, the shared model can use a hypervisor to virtualize graphics processing engines 1931(1)-1931(N) to allow each operating system to access them. In at least one embodiment, for a single-partition system without a hypervisor, the operating system owns graphics processing engines 1931(1)-1931(N). In at least one embodiment, the operating system can virtualize graphics processing engines 1931(1)-1931(N) to provide access to each process or application.

[0226] In at least one embodiment, the graphics acceleration module 1946 or the individual graphics processing engine 1931(1)-1931(N) uses a process handle to select a process element. In at least one embodiment, the process element is stored in system memory 1914 and can be addressed using the effective address to real address translation techniques described herein. In at least one embodiment, the process handle may be an implementation-specific value provided to the host process when registering its context with the graphics processing engine 1931(1)-1931(N) (i.e., invoking system software to add the process element to the process element linked list). In at least one embodiment, the lower 16 bits of the process handle may be the offset of the process element in the process element linked list.

[0227] Figure 19DAn exemplary accelerator integration slice 1990 is illustrated. In at least one embodiment, a "slice" includes a designated portion of the processing resources of an accelerator integrated circuit 1936. In at least one embodiment, the application is an effective address space 1982 in system memory 1914, which stores process element 1983. In at least one embodiment, process element 1983 is stored in response to a GPU call 1981 from an application 1980 executing on processor 1907. In at least one embodiment, process element 1983 contains the process state of the corresponding application 1980. In at least one embodiment, a job descriptor (WD) 1984 contained in process element 1983 may be a single job requested by the application, or it may contain a pointer to a job queue. In at least one embodiment, WD 1984 is a pointer to a job request queue in the effective address space 1982 of the application.

[0228] In at least one embodiment, the graphics acceleration module 1946 and / or the various graphics processing engines 1931(1)-1931(N) may be shared by all processes or a subset of processes in the system. In at least one embodiment, infrastructure may be included for setting process states and sending WD 1984 to the graphics acceleration module 1946 to begin operations in a virtualized environment.

[0229] In at least one embodiment, the dedicated process programming model is implementation-specific. In at least one embodiment, in this model, a single process owns either the graphics acceleration module 1946 or an individual graphics processing engine 1931. In at least one embodiment, when the graphics acceleration module 1946 is owned by a single process, the hypervisor initializes the accelerator integrated circuit 1936 for the owned partition; when the graphics acceleration module 1946 is assigned, the operating system initializes the accelerator integrated circuit 1936 for the owned process.

[0230] In at least one embodiment, during operation, the WD acquisition unit 1991 in the accelerator integration slice 1990 acquires the next WD 1984, which includes instructions for work to be performed by one or more graphics processing engines of the graphics acceleration module 1946. In at least one embodiment, data from the WD 1984 may be stored in register 1945 and used by the MMU 1939, interrupt management circuitry 1947, and / or context management circuitry 1948, as shown. For example, one embodiment of the MMU 1939 includes segment / page roaming circuitry for accessing segment / page tables 1986 within the OS virtual address space 1985. In at least one embodiment, the interrupt management circuitry 1947 may process an interrupt event 1992 received from the graphics acceleration module 1946. In at least one embodiment, when performing graphics operations, a valid address 1993 generated by graphics processing engines 1931(1)-1931(N) is translated into a real address by the MMU 1939.

[0231] In one embodiment, register 1945 is copied for each graphics processing engine 1931(1)-1931(N) and / or graphics acceleration module 1946, and register 1945 may be initialized by a hypervisor or operating system. In at least one embodiment, each of these copied registers may be included in the accelerator integration slice 1990. Exemplary registers that may be initialized by a hypervisor are shown in Table 1.

[0232]

[0233]

[0234] Table 2 shows exemplary registers that can be initialized by the operating system.

[0235]

[0236] In at least one embodiment, each WD 1984 is specific to a particular graphics acceleration module 1946 and / or graphics processing engine 1931(1)-1931(N). In at least one embodiment, it contains all the information required for the graphics processing engine 1931(1)-1931(N) to complete its work, or it may be a pointer to a memory location where the application has set up a command queue for the work to be completed.

[0237] Figure 19EAdditional details of an exemplary embodiment of the shared model are shown. This embodiment includes a hypervisor real address space 1998, in which a list of process elements 1999 is stored. In at least one embodiment, the hypervisor real address space 1998 can be accessed via a hypervisor 1996, which virtualizes the graphics acceleration module engine for operating system 1995.

[0238] In at least one embodiment, the shared programming model allows all processes or subsets of processes from all partitions or subsets of partitions in the system to use the graphics acceleration module 1946. In at least one embodiment, there are two programming models in which the graphics acceleration module 1946 is shared by multiple processes and partitions, namely, time-slice sharing and graphics-oriented sharing.

[0239] In at least one embodiment, in this model, the hypervisor 1996 owns the graphics acceleration module 1946 and makes its functionality available to all operating systems 1995. In at least one embodiment, for the graphics acceleration module 1946 to support virtualization through the hypervisor 1996, the graphics acceleration module 1946 may comply with certain requirements, such as (1) the job requests of the application must be autonomous (i.e., no state needs to be maintained between jobs), or the graphics acceleration module 1946 must provide a context saving and recovery mechanism, (2) the graphics acceleration module 1946 guarantees that the job requests of the application are completed within a specified amount of time, including any conversion errors, or the graphics acceleration module 1946 provides the ability to preempt job processing, and (3) when operating in a directed shared programming model, fairness between the processes of the graphics acceleration module 1946 must be ensured.

[0240] In at least one embodiment, application 1980 needs to make operating system 1995 system calls using the graphics acceleration module type, working descriptor (WD), permission mask register (AMR) value, and context save / restore region pointer (CSRP). In at least one embodiment, the graphics acceleration module type describes the target acceleration function for the system call. In at least one embodiment, the graphics acceleration module type can be a system-specific value. In at least one embodiment, the WD is specifically formatted for graphics acceleration module 1946 and can take the form of graphics acceleration module 1946 commands, valid address pointers to user-defined structures, valid address pointers to command queues, or any other data structure describing the work to be performed by graphics acceleration module 1946.

[0241] In at least one embodiment, the AMR value is the AMR state for the current process. In at least one embodiment, the value passed to the operating system is similar to that of the application that sets the AMR. In at least one embodiment, if the implementation of the accelerator integrated circuit 1936 (not shown) and the graphics acceleration module 1946 does not support the User Rights Mask Overwrite Register (UAMOR), the operating system may apply the current UAMOR value to the AMR value before passing the AMR in the hypervisor call. In at least one embodiment, the hypervisor 1996 may selectively apply the current Rights Mask Overwrite Register (AMOR) value before placing the AMR into the process element 1983. In at least one embodiment, CSRP is one of the registers 1945 that contains the effective address of a region in the effective address space 1982 of the application for the graphics acceleration module 1946 to save and restore the context state. In at least one embodiment, this pointer is optional if it is not necessary to save state between jobs or when a job is preempted. In at least one embodiment, the context save / restore region may be fixed system memory.

[0242] Upon receiving a system call, operating system 1995 can verify that application 1980 has been registered and granted permission to use graphics acceleration module 1946. Then, in at least one embodiment, operating system 1995 uses the information shown in Table 3 to invoke hypervisor 1996.

[0243]

[0244]

[0245] In at least one embodiment, upon receiving a hypervisor call, hypervisor 1996 verifies that operating system 1995 has been registered and granted permission to use graphics acceleration module 1946. Then, in at least one embodiment, hypervisor 1996 adds process element 1983 to a linked list of process elements of the corresponding graphics acceleration module 1946 type. In at least one embodiment, the process element may include the information shown in Table 4.

[0246]

[0247] In at least one embodiment, the management program initializes multiple accelerator integration slice 1990 registers 1945.

[0248] like Figure 19FAs shown, in at least one embodiment, a unified memory is used, which is addressable via a common virtual memory address space for accessing physical processor memories 1901(1)-1901(N) and GPU memories 1920(1)-1920(N). In this implementation, operations performed on GPUs 1910(1)-1910(N) utilize the same virtual / effective memory address space to access processor memories 1901(1)-1901(M) and vice versa, thereby simplifying programmability. In at least one embodiment, a first portion of the virtual / effective address space is allocated to processor memory 1901(1), a second portion to second processor memory 1901(N), a third portion to GPU memory 1920(1), and so on. In at least one embodiment, the entire virtual / effective memory space (sometimes referred to as the effective address space) is thus distributed across each of processor memory 1901 and GPU memory 1920, thereby allowing any processor or GPU to access that memory using a virtual address mapped to any physical memory.

[0249] In at least one embodiment, the bias / coherence management circuitry 1994A-1994E within one or more MMUs 1939A-1939E ensures cache coherence between the caches of one or more host processors (e.g., 1905) and the GPU 1910, and implements biasing techniques to indicate the physical memory in which certain types of data should be stored. In at least one embodiment, although in Figure 19F Several instances of bias / coherence management circuitry 1994A-1994E are shown, but bias / coherence circuitry can be implemented within the MMU of one or more host processors 1905 and / or within the accelerator integrated circuit 1936.

[0250] One embodiment allows GPU memory 1920 to be mapped as part of system memory and accessed using shared virtual memory (SVM) technology without suffering the performance drawbacks associated with full system cache coherence. In at least one embodiment, the ability to access GPU memory 1920 as system memory without the heavy overhead of cache coherence provides a favorable operating environment for GPU offloading. In at least one embodiment, this arrangement allows the host processor 1905 to software-set operands and access computation results without the overhead of conventional I / O DMA data copying. In at least one embodiment, such conventional copying includes driver calls, interrupts, and memory-mapped I / O (MMIO) accesses, all of which are less efficient than simple memory accesses. In at least one embodiment, the ability to access GPU memory 1920 without cache coherence overhead may be critical to the execution time of offloaded computations. In at least one embodiment, for example, in cases with high streaming write memory traffic, cache coherence overhead can significantly reduce the effective write bandwidth seen by GPU 1910. In at least one embodiment, the efficiency of operand setting, the efficiency of result access, and the efficiency of GPU computation may play a role in determining the effectiveness of GPU offloading.

[0251] In at least one embodiment, the selection of GPU bias and host processor bias is driven by a bias tracker data structure. In at least one embodiment, for example, a bias table can be used, which may be a page-granular structure (e.g., controlled at the memory page level) comprising 1 or 2 bits of memory pages attached to each GPU. In at least one embodiment, with or without a bias cache (e.g., for caching frequently / recently used entries in the bias table) in GPU 1910, the bias table can be implemented across one or more stolen memory ranges of GPU memory 1920. Alternatively, in at least one embodiment, the entire bias table can be maintained within the GPU.

[0252] In at least one embodiment, prior to actual access to GPU memory, an access to the bias table entry associated with each access to GPU-attached memory 1920 is performed, resulting in the following operations: In at least one embodiment, a local request from GPU 1910 to find its page in the GPU bias is directly forwarded to the corresponding GPU memory 1920. In at least one embodiment, a local request from the GPU to find its page in the host bias is forwarded to processor 1905 (e.g., via the high-speed link described herein). In at least one embodiment, a request from processor 1905 to find the requested page in the host processor bias completes a request similar to a normal memory read. Alternatively, a request to a GPU bias page can be forwarded to GPU 1910. In at least one embodiment, if the GPU is not currently using the page, the GPU may subsequently migrate the page to the host processor bias. In at least one embodiment, the page bias state can be changed through software-based mechanisms, hardware-assisted software mechanisms, or, in limited cases, purely hardware-based mechanisms.

[0253] In at least one embodiment, a mechanism for changing the bias state employs an API call (e.g., OpenCL), which subsequently invokes the GPU's device driver. The device driver then sends a message (or enqueues a command descriptor) to the GPU, instructing the GPU to change the bias state and, in some migration, performs a cache refresh operation on the host. In at least one embodiment, the cache refresh operation is used for migration from the host processor 1905 bias to the GPU bias, but not for the reverse migration.

[0254] In at least one embodiment, cache coherence is maintained by temporarily rendering GPU bias pages that the host processor 1905 cannot cache. In at least one embodiment, to access these pages, the processor 1905 may request access from the GPU 1910, which may or may not immediately grant access. Therefore, in at least one embodiment, to reduce communication between the processor 1905 and the GPU 1910, it is beneficial to ensure that the GPU bias pages are pages needed by the GPU rather than those needed by the host processor 1905, and vice versa.

[0255] In at least one embodiment, processor 114 may be implemented by one or more of multi-core processors 1905, GPU 1910, multi-core processor 1907, and / or graphics acceleration module 1946; memory 116 may be implemented by processor memory 1901, system memory 1914, and / or GPU memory 1920; and bus 117 may be implemented at least partially by memory interconnect 1926, high-speed link 1940, high-speed link 1929, and / or GPU memory interconnect 1950. As another non-limiting example, instruction 118 (see...) Figure 1 The optical flow map 110 may be stored in processor memory 1901, system memory 1914, and / or GPU memory 1920, executed by at least one of multi-core processor 1905, GPU 1910, multi-core processor 1907, and / or graphics acceleration module 1946, and used to obtain the optical flow map 110 (see...). Figure 1 and Figure 2 ).

[0256] As described above, system 100 (see Figure 1 This can be implemented by at least a portion of the reasoning and / or training logic 1115. (See reference...) Figure 1 One or more of the multi-core processor 1905, one or more of the GPU 1910, the multi-core processor 1907 and / or the graphics acceleration module 1946 (see also...) Figure 19A -D) can perform inference and / or training operations (e.g., at least a portion of inference and / or training logic 1115) for any of the noise reduction process 120, edge detection process 122, object detection process 124, optional thresholding process 126, decision logic 128, extraction process 130 and SGM process 132.

[0257] Figure 20 Exemplary integrated circuits and associated graphics processors according to various embodiments described herein are illustrated, which may be manufactured using one or more IP cores. In addition to the illustrations, at least one embodiment may include other logic and circuitry, including additional graphics processors / cores, peripheral interface controllers, or general-purpose processor cores.

[0258] Figure 20This is a block diagram illustrating an exemplary system on a chip integrated circuit 2000 that can be fabricated using one or more IP cores according to at least one embodiment. In at least one embodiment, the integrated circuit 2000 includes one or more application processors 2005 (e.g., CPUs), at least one graphics processor 2010, and may additionally include an image processor 2015 and / or a video processor 2020, any of which may be a modular IP core. In at least one embodiment, the integrated circuit 2000 includes peripheral or bus logic, which includes a USB controller 2025, a UART controller 2030, an SPI / SDIO controller 2035, and an I... 2 2S / I 2 2C controller 2040. In at least one embodiment, integrated circuit 2000 may include display device 2045 coupled to one or more of High Definition Multimedia Interface (HDMI) controller 2050 and Mobile Industrial Processor Interface (MIPI) display interface 2055. In at least one embodiment, storage may be provided by flash memory subsystem 2060, including flash memory and flash memory controller. In at least one embodiment, a memory interface may be provided via memory controller 2065 for accessing SDRAM or SRAM memory devices. In at least one embodiment, some integrated circuits also include embedded security engine 2070.

[0259] Inference and / or training logic 1115 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 11A and / or Figure 11B Details regarding the inference and / or training logic 1115 are provided. In at least one embodiment, the inference and / or training logic 1115 may be used in the integrated circuit 2000 to infer or predict operations based at least in part on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0260] In at least one embodiment, system 100 (see...) Figure 1 The processor 114 may be implemented at least partially by the system-on-chip integrated circuit 2000. In at least one embodiment, the processor 114 may be implemented by one or more processors 2005-2020, and the memory 116 may be implemented by an SDRAM or SRAM memory device. As another non-limiting example, instruction 118 (see...) Figure 1 It can be stored by an SDRAM or SRAM memory device, executed by at least one of processors 2005-2020, and used to obtain optical flow diagram 110 (see...). Figure 1 and Figure 2 As described above, system 100 (see...) Figure 1This can be implemented by at least a portion of the reasoning and / or training logic 1115. (See reference...) Figure 1 System-on-a-Chip Integrated Circuit 2000 (see...) Figure 20 It can perform inference and / or training operations (e.g., at least a portion of inference and / or training logic 1115) for any of the noise reduction process 120, edge detection process 122, object detection process 124, optional thresholding process 126, decision logic 128, extraction process 130 and SGM process 132.

[0261] Figures 21A-21B Exemplary integrated circuits and associated graphics processors according to various embodiments described herein are illustrated, which may be manufactured using one or more IP cores. In addition to the illustrations, at least one embodiment may include other logic and circuitry, including additional graphics processors / cores, peripheral interface controllers, or general-purpose processor cores.

[0262] Figures 21A-21B This is a block diagram illustrating an exemplary graphics processor used within a SoC according to embodiments described herein. Figure 21A An exemplary graphics processor 2110 of a system-on-a-chip according to at least one embodiment is shown, which can be manufactured using one or more IP cores. Figure 21B Further exemplary graphics processor 2140 of a system-on-a-chip according to at least one embodiment is shown, which can be manufactured using one or more IP cores. In at least one embodiment, Figure 21A The graphics processor 2110 is a low-power graphics processor core. In at least one embodiment, Figure 21B The graphics processor 2140 is a higher-performance graphics processor core. In at least one embodiment, each graphics processor 2110, 2140 may be... Figure 37 A variant of the 2010 graphics processor.

[0263] In at least one embodiment, the graphics processor 2110 includes a vertex processor 2105 and one or more fragment processors 2115A-2115N (e.g., 2115A, 2115B, 2115C, 2115D to 2115N-1 and 2115N). In at least one embodiment, the graphics processor 2110 may execute different shader programs via separate logic, such that the vertex processor 2105 is optimized to perform operations for the vertex shader program, while one or more fragment processors 2115A-2115N perform fragment (e.g., pixel) shading operations for fragments or pixels or shader programs. In at least one embodiment, the vertex processor 2105 performs the vertex processing stage of the 3D graphics pipeline and generates primitive and vertex data. In at least one embodiment, one or more fragment processors 2115A-2115N use the primitive and vertex data generated by the vertex processor 2105 to generate a framebuffer for display on a display device. In at least one embodiment, one or more fragment processors 2115A-2115N are optimized to execute fragment shader programs as provided in the OpenGL API, which can be used to perform operations similar to those of pixel shader programs provided in the Direct 3D API.

[0264] In at least one embodiment, the graphics processor 2110 additionally includes one or more memory management units (MMUs) 2120A-2120B, one or more caches 2125A-2125B, and one or more circuit interconnects 2130A-2130B. In at least one embodiment, one or more MMUs 2120A-2120B provide a virtual-to-physical address mapping for the graphics processor 2110, including for the vertex processor 2105 and / or fragment processors 2115A-2115N, which can reference vertex or image / texture data stored in memory, in addition to vertex or image / texture data stored in one or more caches 2125A-2125B. In at least one embodiment, one or more MMUs 2120A-2120B can be synchronized with other MMUs within the system, including with... Figure 20 One or more application processors 2005, graphics processors 2015, and / or video processors 2020 are associated with one or more MMUs, enabling each processor 2005-2020 to participate in a shared or unified virtual memory system. In at least one embodiment, one or more circuit interconnects 2130A-2130B enable the graphics processor 2110 to be connected to other IP cores within the SoC via the SoC's internal bus or via a direct connection.

[0265] In at least one embodiment, the graphics processor 2140 includes one or more shader cores 2155A-2155N (e.g., 2155A, 2155B, 2155C, 2155D, 2155E, 2155F to 2155N-1 and 2155N), such as Figure 21B As shown, it provides a unified shader core architecture, where a single core or type or core can execute all types of programmable shader code, including shader program code for implementing vertex shaders, fragment shaders, and / or compute shaders. In at least one embodiment, the number of shader cores can vary. In at least one embodiment, the graphics processor 2140 includes an inter-core task manager 2145, which acts as a thread dispatcher to assign execution threads to one or more shader cores 2155A-2155N and tile unit 2158 to accelerate tile-based rendering operations, where scene rendering operations are subdivided in image space, for example, to utilize local spatial consistency within the scene or optimize the use of internal caches.

[0266] Inference and / or training logic 1115 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 11A and / or Figure 11B Details regarding the inference and / or training logic 1115 are provided. In at least one embodiment, the inference and / or training logic 1115 may be integrated into an integrated circuit. Figure 21A and / or Figure 21B The above is used for inference or prediction operations based at least in part on weight parameters calculated using neural network training operations, neural network functions or architectures, or neural network use cases described herein.

[0267] In at least one embodiment, processor 114 may be implemented at least in part by graphics processor 2110 and / or graphics processor 2140. As another non-limiting example, instruction 118 (see...) Figure 1 This can be executed by graphics processor 2110 and / or graphics processor 2140 and used to obtain optical flow map 110 (see...). Figure 1 and Figure 2 As described above, system 100 (see...) Figure 1 This can be implemented by at least a portion of the reasoning and / or training logic 1115. (See reference...) Figure 1 Graphics processor 2110 and / or graphics processor 2140 (see Figure 21A and Figure 21BIt can perform inference and / or training operations (at least a portion of the inference and / or training logic 1115) for any of the noise reduction process 120, edge detection process 122, object detection process 124, optional thresholding process 126, decision logic 128, extraction process 130 and SGM process 132.

[0268] Figures 22A-22B Additional exemplary graphics processor logic according to embodiments described herein is illustrated. In at least one embodiment, Figure 22A It shows that it can be included in Figure 20 The graphics core 2200 within the graphics processor 2010, and in at least one embodiment, may be as follows: Figure 21B The Unified Shader Core 2155A-2155N shown is an example. Figure 22B A highly parallel general-purpose graphics processing unit (GPGPU) 2230 suitable for deployment on a multi-chip module is shown in at least one embodiment.

[0269] In at least one embodiment, the graphics core 2200 includes a shared instruction cache 2202, texture units 2218, and a cache / shared memory 2220, which are common to the execution resources within the graphics core 2200. In at least one embodiment, the graphics core 2200 may include multiple slices 2201A-2201N or partitions of each core, and the graphics processor may include multiple instances of the graphics core 2200. In at least one embodiment, slices 2201A-2201N may include supporting logic, including local instruction caches 2204A-2204N, thread schedulers 2206A-2206N, thread dispatchers 2208A-2208N, and a set of registers 2210A-2210N. In at least one embodiment, slices 2201A-2201N may include a set of additional functional units (AFU 2212A-2212N), floating-point units (FPU 2214A-2214N), integer arithmetic logic units (ALU 2216A-2216N), address calculation units (ACU 2213A-2213N), double-precision floating-point units (DPFPU 2215A-2215N), and matrix processing units (MPU 2217A-2217N).

[0270] In at least one embodiment, the FPU 2214A-2214N can perform single-precision (32-bit) and half-precision (16-bit) floating-point operations, while the DPFPU 2215A-2215N performs double-precision (64-bit) floating-point operations. In at least one embodiment, the ALU 2216A-2216N can perform variable-precision integer operations with 8-bit, 16-bit, and 32-bit precision, and can be configured for mixed-precision operations. In at least one embodiment, the MPU 2217A-2217N can also be configured for mixed-precision matrix operations, including half-precision floating-point operations and 8-bit integer operations. In at least one embodiment, the MPU 2217-2217N can perform various matrix operations to accelerate machine learning application frameworks, including enabling support for accelerated generalized matrix-to-matrix multiplication (GEMM). In at least one embodiment, the AFU 2212A-2212N can perform additional logical operations not supported by floating-point or integer units, including trigonometric operations (e.g., sine, cosine, etc.).

[0271] Inference and / or training logic 1115 is used to perform inference and / or training operations associated with one or more embodiments. This is combined with... Figure 11A and / or Figure 11B Details regarding inference and / or training logic 1115 are provided. In at least one embodiment, inference and / or training logic 1115 may be used in graphics core 2200 for inferring or predicting operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures or neural network use cases described herein.

[0272] In at least one embodiment, processor 114 may be implemented at least partially by graphics core 2200. As another non-limiting example, instruction 118 (see...) Figure 1 This can be executed by the graphics core 2200 and used to obtain the optical flow map 110 (see...). Figure 1 and Figure 2 As described above, system 100 (see...) Figure 1 This can be implemented by at least a portion of the reasoning and / or training logic 1115. (See reference...) Figure 1 Graphics core 2200 (see...) Figure 22A It can perform inference and / or training operations (e.g., at least a portion of inference and / or training logic 1115) for any of the noise reduction process 120, edge detection process 122, object detection process 124, optional thresholding process 126, decision logic 128, extraction process 130 and SGM process 132.

[0273] Figure 22BA general-purpose processing unit (GPGPU) 2230 is illustrated in at least one embodiment, which can be configured to enable highly parallel computational operations to be performed by a set of graphics processing units. In at least one embodiment, the GPGPU 2230 can be directly linked to other instances of the GPGPU 2230 to create a multi-GPU cluster to improve the training speed for deep neural networks. In at least one embodiment, the GPGPU 2230 includes a host interface 2232 for connection to a host processor. In at least one embodiment, the host interface 2232 is a PCI Express interface. In at least one embodiment, the host interface 2232 can be a vendor-specific communication interface or communication structure. In at least one embodiment, the GPGPU 2230 receives commands from the host processor and uses a global scheduler 2234 to allocate execution threads associated with those commands to a set of compute clusters 2236A-2236H. In at least one embodiment, the compute clusters 2236A-2236H share a cache memory 2238. In at least one embodiment, cache memory 2238 can be used as a higher-level cache within the cache memory of computing clusters 2236A-2236H.

[0274] In at least one embodiment, the GPGPU 2230 includes memories 2244A-2244B, which are coupled to computing clusters 2236A-2236H via a set of memory controllers 2242A-2242B. In at least one embodiment, memories 2244A-2244B may include various types of memory devices, including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM), which includes graphics double data rate (GDDR) memory.

[0275] In at least one embodiment, each of the computing clusters 2236A-2236H includes a set of graphics cores, for example... Figure 22A The graphics core 2200 may include various types of integer and floating-point logic units that can perform computational operations across a range of precisions, including precisions suitable for machine learning computations. For example, in at least one embodiment, at least a subset of the floating-point units in each computing cluster 2236A-2236H may be configured to perform 16-bit or 32-bit floating-point operations, while different subsets of the floating-point units may be configured to perform 64-bit floating-point operations.

[0276] In at least one embodiment, multiple instances of GPGPU 2230 can be configured as a computing cluster. In at least one embodiment, the communication used for synchronization and data exchange by computing clusters 2236A-2236H varies between embodiments. In at least one embodiment, multiple instances of GPGPU 2230 communicate via host interface 2232. In at least one embodiment, GPGPU 2230 includes an I / O hub 2239 that couples GPGPU 2230 to GPU link 2240, enabling direct connection to other instances of GPGPU 2230. In at least one embodiment, GPU link 2240 is coupled to a dedicated GPU-to-GPU bridge, which enables communication and synchronization between multiple instances of GPGPU 2230. In at least one embodiment, GPU link 2240 is coupled to a high-speed interconnect for sending and receiving data to and from other GPGPUs or parallel processors. In at least one embodiment, multiple instances of GPGPU 2230 reside in a separate data processing system and communicate via network devices accessible through host interface 2232. In at least one embodiment, GPU link 2240 may be configured to enable connection to a host processor other than or as a replacement for host interface 2232.

[0277] In at least one embodiment, the GPGPU 2230 can be configured to train a neural network. In at least one embodiment, the GPGPU 2230 can be used within an inference platform. In at least one embodiment, when the GPGPU 2230 is used for inference, the GPGPU 2230 may include fewer compute clusters 2236A-2236H compared to when the GPGPU 2230 is used to train a neural network. In at least one embodiment, the memory technology associated with the memories 2244A-2244B can differ between inference and training configurations, wherein a higher bandwidth memory technology is dedicated to the training configuration. In at least one embodiment, the inference configuration of the GPGPU 2230 can support inference-specific instructions. For example, in at least one embodiment, the inference configuration can provide support for one or more 8-bit integer dot product instructions, which can be used during the inference operation of the deployed neural network.

[0278] Inference and / or training logic 1115 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 11A and / or Figure 11BDetails regarding the inference and / or training logic 1115 are provided. In at least one embodiment, the inference and / or training logic 1115 may be used in the GPGPU 2230 for inferring or predicting operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architecture, or neural network use cases described herein.

[0279] In at least one embodiment, processor 114 may be implemented at least partially by GPGPU 2230 and memory 116 may be implemented at least partially by memory 2244A-2244B. As another non-limiting example, instruction 118 (see...) Figure 1 It can be stored in memory 2244A-2244B, executed by GPGPU 2230, and used to obtain optical flow map 110 (see...). Figure 1 and Figure 2 As described above, system 100 (see...) Figure 1 This can be implemented by at least a portion of the reasoning and / or training logic 1115. (See reference...) Figure 1 GPGPU 2230 (see Figure 22B It can perform inference and / or training operations (e.g., at least a portion of inference and / or training logic 1115) for any of the noise reduction process 120, edge detection process 122, object detection process 124, optional thresholding process 126, decision logic 128, extraction process 130 and SGM process 132.

[0280] Figure 23 A block diagram of a computer system 2300 according to at least one embodiment is shown. In at least one embodiment, the computer system 2300 includes a processing subsystem 2301 having one or more processors 2302 and a system memory 2304 communicating via an interconnect path that may include a memory hub 2305. In at least one embodiment, the memory hub 2305 may be a separate component within a chipset component or may be integrated within one or more processors 2302. In at least one embodiment, the memory hub 2305 is coupled to an I / O subsystem 2311 via a communication link 2306. In one embodiment, the I / O subsystem 2311 includes an I / O hub 2307 that enables the computer system 2300 to receive input from one or more input devices 2308. In at least one embodiment, the I / O hub 2307 enables a display controller to provide output to one or more display devices 2310A, the display controller being included in one or more processors 2302. In at least one embodiment, one or more display devices 2310A coupled to the I / O hub 2307 may include local, internal or embedded display devices.

[0281] In at least one embodiment, the processing subsystem 2301 includes one or more parallel processors 2312 coupled to the memory hub 2305 via a bus or other communication link 2313. In at least one embodiment, the communication link 2313 may use any of many standards-based communication link technologies or protocols, such as, but not limited to, PCI Express, or may be a vendor-specific communication interface or communication architecture. In at least one embodiment, one or more parallel processors 2312 form a compute-intensive parallel or vector processing system, which may include a large number of processing cores and / or processing clusters, such as a multi-core integrated (MIC) processor. In at least one embodiment, one or more parallel processors 2312 form a graphics processing subsystem that can output pixels to one or more display devices 2310A coupled via an I / O hub 2307. In at least one embodiment, the parallel processors 2312 may also include a display controller and a display interface (not shown) to enable direct connection to one or more display devices 2310B.

[0282] In at least one embodiment, system storage unit 2314 may be connected to I / O hub 2307 to provide a storage mechanism for computer system 2300. In at least one embodiment, I / O switch 2316 may be used to provide an interface mechanism to enable connectivity between I / O hub 2307 and other components, such as network adapter 2318 and / or wireless network adapter 2319 which may be integrated into the platform, and various other devices that can be added via one or more additional devices 2320. In at least one embodiment, network adapter 2318 may be an Ethernet adapter or another wired network adapter. In at least one embodiment, wireless network adapter 2319 may include one or more of Wi-Fi, Bluetooth, Near Field Communication (NFC), or other network devices including one or more wireless devices.

[0283] In at least one embodiment, the computer system 2300 may include other components not explicitly shown, such as USB or other port connections, optical storage drives, video capture devices, etc., which may also be connected to the I / O hub 2307. In at least one embodiment, the interconnection can be implemented using any suitable protocol (e.g., a PCI (Peripheral Component Interconnect) based protocol (e.g., PCI-Express) or other bus or point-to-point communication interfaces and / or protocols). Figure 23 The communication paths of the various components, such as NV-Link high-speed interconnect or interconnect protocols.

[0284] In at least one embodiment, one or more parallel processors 2312 include circuitry optimized for graphics and video processing, including, for example, video output circuitry, and constituting a graphics processing unit (GPU). In at least one embodiment, the parallel processors 2312 include circuitry optimized for general-purpose processing. In at least one embodiment, components of the computer system 2300 may be integrated with one or more other system elements on a single integrated circuit. For example, in at least one embodiment, the parallel processor 2312, memory hub 2305, processor 2302, and I / O hub 2307 may be integrated into a system-on-a-chip (SoC) integrated circuit. In at least one embodiment, components of the computer system 2300 may be integrated into a single package to form a system-in-package (SIP) configuration. In at least one embodiment, at least a portion of the components of the computer system 2300 may be integrated into a multi-chip module (MCM) that can interconnect with other MCMs to a modular computer system.

[0285] Inference and / or training logic 1115 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 11A and / or Figure 11B Details are provided regarding the inference and / or training logic 1115. In at least one embodiment, the inference and / or training logic 1115 can... Figure 23 The system 2300 is used for reasoning or predicting operations based at least in part on weight parameters calculated using neural network training operations, neural network functions and / or architectures or neural network use cases described herein.

[0286] In at least one embodiment, system 100 (see...) Figure 1 The processor 114 can be implemented at least partially by the computing system 2300. For example, the processor 114 can be implemented by one or more processors 2302 and / or one or more parallel processors 2312, the memory 116 can be a component of the system memory 2304, the interface 115 can be implemented at least partially by the I / O hub 2307, and the bus 117 can be implemented at least partially by the communication links 2306 and / or 2313. As another non-limiting example, instruction 118 (see...) Figure 1 It can be executed by at least one of processor 2302 and parallel processor 2312, and is used to obtain optical flow map 110 (see...). Figure 1 and Figure 2 The upstream hardware 104 can be implemented as one of the input devices 2308.

[0287] As described above, system 100 (see Figure 1 This can be implemented by at least a portion of the reasoning and / or training logic 1115. (See reference...) Figure 1 Computing System 2300 (see Figure 23 It can perform inference and / or training operations (e.g., at least a portion of inference and / or training logic 1115) for any of the noise reduction process 120, edge detection process 122, object detection process 124, optional thresholding process 126, decision logic 128, extraction process 130 and SGM process 132.

[0288] processor

[0289] Figure 24A A parallel processor 2400 according to at least one embodiment is illustrated. In at least one embodiment, various components of the parallel processor 2400 may be implemented using one or more integrated circuit devices, such as programmable processors, application-specific integrated circuits (ASICs), or field-programmable gate arrays (FPGAs). In at least one embodiment, the illustrated parallel processor 2400 is according to an exemplary embodiment. Figure 23 The variant of the 2312, which includes one or more parallel processors, is shown.

[0290] In at least one embodiment, the parallel processor 2400 includes a parallel processing unit 2402. In at least one embodiment, the parallel processing unit 2402 includes an I / O unit 2404 that enables communication with other devices, including other instances of the parallel processing unit 2402. In at least one embodiment, the I / O unit 2404 can be directly connected to other devices. In at least one embodiment, the I / O unit 2404 is connected to other devices using a hub or switch interface (e.g., a memory hub 2405). In at least one embodiment, the connection between the memory hub 2405 and the I / O unit 2404 forms a communication link 2413. In at least one embodiment, the I / O unit 2404 is connected to a host interface 2406 and a memory crossbar switch 2416, wherein the host interface 2406 receives commands for performing processing operations, and the memory crossbar switch 2416 receives commands for performing memory operations.

[0291] In at least one embodiment, when host interface 2406 receives a command buffer via I / O unit 2404, host interface 2406 can direct work operations to execute those commands to front end 2408. In at least one embodiment, front end 2408 is coupled to scheduler 2410, which is configured to assign commands or other work items to processing cluster array 2412. In at least one embodiment, scheduler 2410 ensures that processing cluster array 2412 is correctly configured and in an active state before assigning tasks to processing cluster array 2412. In at least one embodiment, scheduler 2410 is implemented via firmware logic executed on a microcontroller. In at least one embodiment, the microcontroller-implemented scheduler 2410 can be configured to perform complex scheduling and work assignment operations at both coarse and fine granular levels, thereby enabling fast preemption and context switching of threads executing on processing array 2412. In at least one embodiment, host software can demonstrate workloads for scheduling on processing array 2412 via one of multiple graphics processing paths. In at least one embodiment, the workload can then be automatically distributed on the processing array 2412 by the scheduler 2410 logic within the microcontroller, which includes the scheduler 2410.

[0292] In at least one embodiment, the processing cluster array 2412 may include up to "N" processing clusters (e.g., clusters 2414A, 2414B to 2414N), where "N" represents a positive integer (which may be an integer different from the integer "N" used in other diagrams). In at least one embodiment, each cluster 2414A-2414N of the processing cluster array 2412 can execute a large number of concurrent threads. In at least one embodiment, the scheduler 2410 may use various scheduling and / or work allocation algorithms to allocate work to the clusters 2414A-2414N of the processing cluster array 2412, which may vary depending on the workload generated by each type of program or computation. In at least one embodiment, scheduling may be handled dynamically by the scheduler 2410, or may be partially assisted by compiler logic during the compilation of program logic configured to be executed by the processing cluster array 2412. In at least one embodiment, different clusters 2414A-2414N of the processing cluster array 2412 may be assigned to process different types of programs or to perform different types of computations.

[0293] In at least one embodiment, the processing cluster array 2412 can be configured to perform various types of parallel processing operations. In at least one embodiment, the processing cluster array 2412 is configured to perform general-purpose parallel computing operations. For example, in at least one embodiment, the processing cluster array 2412 may include logic for performing processing tasks, including filtering video and / or audio data, performing modeling operations, including physical operations, and performing data transformations.

[0294] In at least one embodiment, the processing cluster array 2412 is configured to perform parallel graphics processing operations. In at least one embodiment, the processing cluster array 2412 may include additional logic to support the execution of such graphics processing operations, including but not limited to texture sampling logic for performing texture operations, as well as tessellation logic and other vertex processing logic. In at least one embodiment, the processing cluster array 2412 may be configured to execute shader programs related to graphics processing, such as, but not limited to, vertex shaders, tessellation shaders, geometry shaders, and pixel shaders. In at least one embodiment, the parallel processing unit 2402 may transfer data from system memory via I / O unit 2404 for processing. In at least one embodiment, during processing, the transferred data may be stored in on-chip memory (e.g., parallel processor memory 2422) and then written back to system memory.

[0295] In at least one embodiment, when the parallel processing unit 2402 is used to perform graphics processing, the scheduler 2410 may be configured to divide the processing workload into tasks of approximately equal size to better distribute graphics processing operations among the multiple clusters 2414A-2414N of the processing cluster array 2412. In at least one embodiment, portions of the processing cluster array 2412 may be configured to perform different types of processing. For example, in at least one embodiment, a first portion may be configured to perform vertex shading and topology generation, a second portion may be configured to perform tessellation and geometry shading, and a third portion may be configured to perform pixel shading or other screen-space operations to generate a rendered image for display. In at least one embodiment, intermediate data generated by one or more of the clusters 2414A-2414N may be stored in a buffer to allow intermediate data to be transferred between the clusters 2414A-2414N for further processing.

[0296] In at least one embodiment, the processing cluster array 2412 may receive processing tasks to be executed via a scheduler 2410, which receives commands defining the processing tasks from a front end 2408. In at least one embodiment, the processing task may include an index of data to be processed, such as surface (patch) data, raw data, vertex data, and / or pixel data, as well as state parameters and commands defining how the data is processed (e.g., what program to execute). In at least one embodiment, the scheduler 2410 may be configured to acquire an index corresponding to a task, or may receive an index from the front end 2408. In at least one embodiment, the front end 2408 may be configured to ensure that the processing cluster array 2412 is configured to be active before initiating the workload specified by an incoming command buffer (e.g., a batch buffer, push buffer, etc.).

[0297] In at least one embodiment, each of one or more instances of the parallel processing unit 2402 may be coupled to the parallel processor memory 2422. In at least one embodiment, the parallel processor memory 2422 may be accessed via a memory crossbar switch 2416, which may receive memory requests from the processing cluster array 2412 and the I / O unit 2404. In at least one embodiment, the memory crossbar switch 2416 may be accessed via a memory interface 2418. In at least one embodiment, the memory interface 2418 may include a plurality of partition units (e.g., partition units 2420A, 2420B to 2420N), each of which may be coupled to a portion (e.g., a memory cell) of the parallel processor memory 2422. In at least one embodiment, the plurality of partition units 2420A-2420N are configured to be equal to the number of memory units, such that the first partition unit 2420A has a corresponding first memory unit 2424A, the second partition unit 2420B has a corresponding memory unit 2424B, and the Nth partition unit 2420N has a corresponding Nth memory unit 2424N. In at least one embodiment, the number of partition units 2420A-2420N may not be equal to the number of memory units.

[0298] In at least one embodiment, memory cells 2424A-2424N may include various types of memory devices, including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM), including graphics double data rate (GDDR) memory. In at least one embodiment, memory cells 2424A-2424N may also include 3D stacked memory, including but not limited to high bandwidth memory (HBM). In at least one embodiment, rendering targets such as frame buffers or texture maps may be stored across memory cells 2424A-2424N, allowing partitioning cells 2420A-2420N to write portions of each rendering target in parallel, to efficiently utilize the available bandwidth of the parallel processor memory 2422. In at least one embodiment, local instances of the parallel processor memory 2422 may be excluded to facilitate a unified memory design that combines system memory with local cache memory.

[0299] In at least one embodiment, any of the clusters 2414A-2414N of the processing cluster array 2412 can process data to be written to any memory cell 2424A-2424N within the parallel processor memory 2422. In at least one embodiment, the memory crossbar switch 2416 can be configured to transfer the output of each cluster 2414A-2414N to any partition cell 2420A-2420N or another cluster 2414A-2414N, and the clusters 2414A-2414N can perform further processing operations on the output. In at least one embodiment, each cluster 2414A-2414N can communicate with the memory interface 2418 via the memory crossbar switch 2416 to read from or write to various external storage devices. In at least one embodiment, the memory crossbar switch 2416 has a connection to a memory interface 2418 for communication with I / O unit 2404, and a connection to a local instance of parallel processor memory 2422, thereby enabling processing units within different processing clusters 2414A-2414N to communicate with system memory or other memory not local to parallel processing unit 2402. In at least one embodiment, the memory crossbar switch 2416 may use virtual channels to separate traffic flows between clusters 2414A-2414N and partition units 2420A-2420N.

[0300] In at least one embodiment, multiple instances of the parallel processing unit 2402 may be provided on a single insert card, or multiple insert cards may be interconnected. In at least one embodiment, different instances of the parallel processing unit 2402 may be configured to interoperate, even if the different instances have different numbers of processing cores, different numbers of local parallel processor memories, and / or other configuration differences. For example, in at least one embodiment, some instances of the parallel processing unit 2402 may include higher-precision floating-point units relative to other instances. In at least one embodiment, a system combining one or more instances of the parallel processing unit 2402 or the parallel processor 2400 may be implemented in various configurations and form factors, including but not limited to desktop, laptop, or handheld personal computers, servers, workstations, game consoles, and / or embedded systems.

[0301] Figure 24B This is a block diagram of a partitioning unit 2420 according to at least one embodiment. In at least one embodiment, the partitioning unit 2420 is... Figure 24A This is an example of one of the partitioning units 2420A-2420N. In at least one embodiment, the partitioning unit 2420 includes an L2 cache 2421, a frame buffer interface 2425, and a ROP 2426 (raster operation unit). In at least one embodiment, the L2 cache 2421 is a read / write cache configured to perform load and store operations received from the memory crossbar switch 2416 and the ROP 2426. In at least one embodiment, the L2 cache 2421 outputs read misses and urgent write-back requests to the frame buffer interface 2425 for processing. In at least one embodiment, updates can also be sent to the frame buffer for processing via the frame buffer interface 2425. In at least one embodiment, the frame buffer interface 2425 communicates with memory cells in the parallel processor memory (such as...). Figure 24A The memory cells 2424A-2424N (e.g., within the parallel processor memory 2422) interact with one of them.

[0302] In at least one embodiment, ROP 2426 is a processing unit that performs raster operations such as stenciling, z-testing, blending, etc. In at least one embodiment, ROP 2426 then outputs processed graphics data stored in graphics memory. In at least one embodiment, ROP 2426 includes compression logic to compress depth or color data written to memory and decompress depth or color data read from memory. In at least one embodiment, the compression logic may be lossless compression logic utilizing one or more of a variety of compression algorithms. In at least one embodiment, the type of compression performed by ROP 2426 may vary based on the statistical characteristics of the data to be compressed. For example, in at least one embodiment, incremental color compression is performed based on depth and color data on a per-tile basis.

[0303] In at least one embodiment, ROP 2426 is included within each processing cluster (e.g., Figure 24A Clusters 2414A-2414N are used instead of partition units 2420. In at least one embodiment, read and write requests for pixel data are made via memory crossbar switch 2416 instead of pixel fragment data transfer. In at least one embodiment, the processed graphics data can be displayed on a display device (such as...). Figure 23 Displayed by one or more display devices 2310, routed by processor 2302 for further processing, or by... Figure 24A One of the processing entities within the parallel processor 2400 is routed for further processing.

[0304] Figure 24C This is a block diagram of a processing cluster 2414 within a parallel processing unit according to at least one embodiment. In at least one embodiment, the processing cluster is... Figure 24A An instance of one of the processing clusters 2414A-2414N. In at least one embodiment, the processing cluster 2414 can be configured to execute a number of threads in parallel, where a "thread" refers to an instance of a specific program executing on a particular set of input data. In at least one embodiment, Single Instruction Multiple Data (SIMD) instruction issuing technology is used to support the parallel execution of a large number of threads without providing multiple independent instruction units. In at least one embodiment, Single Instruction Multiple Threading (SIMT) technology is used to support the parallel execution of a large number of generally synchronous threads, which uses a common instruction unit configured to issue instructions to a set of processing engines within each processing cluster.

[0305] In at least one embodiment, the operation of the processing cluster 2414 can be controlled by a pipeline manager 2432 that assigns processing tasks to the SIMT parallel processors. In at least one embodiment, the pipeline manager 2432... Figure 24AThe scheduler 2410 receives instructions and manages the execution of these instructions via the graphics multiprocessor 2434 and / or texture unit 2436. In at least one embodiment, the graphics multiprocessor 2434 is an exemplary instance of a SIMT parallel processor. However, in at least one embodiment, the processing cluster 2414 may include various types of SIMT parallel processors with different architectures. In at least one embodiment, the processing cluster 2414 may include one or more instances of the graphics multiprocessor 2434. In at least one embodiment, the graphics multiprocessor 2434 can process data, and the data cross switch 2440 can be used to distribute the processed data to one of a number of possible destinations, including other shader units. In at least one embodiment, the pipeline manager 2432 can facilitate the distribution of processed data by specifying the destination of the processed data to be distributed via the data cross switch 2440.

[0306] In at least one embodiment, each graphics multiprocessor 2434 within the processing cluster 2414 may include the same set of functional execution logic (e.g., arithmetic logic units, load-memory units, etc.). In at least one embodiment, the functional execution logic may be configured in a pipelined manner, wherein new instructions may be issued before previous instructions complete. In at least one embodiment, the functional execution logic supports a variety of operations, including integer and floating-point arithmetic, comparison operations, Boolean operations, shift operations, and computation of various algebraic functions. In at least one embodiment, the same functional unit hardware may be used to perform different operations, and any combination of functional units may exist.

[0307] In at least one embodiment, instructions sent to the processing cluster 2414 constitute threads. In at least one embodiment, a group of threads executed across a set of parallel processing engines is a thread group. In at least one embodiment, the thread group executes a general program on different input data. In at least one embodiment, each thread within the thread group may be assigned to a different processing engine within the graphics multiprocessor 2434. In at least one embodiment, the thread group may include fewer threads than the number of processing engines within the graphics multiprocessor 2434. In at least one embodiment, when the number of threads included in the thread group is less than the number of processing engines, one or more processing engines may be idle during a loop that is processing the thread group. In at least one embodiment, the thread group may also include more threads than the number of processing engines within the graphics multiprocessor 2434. In at least one embodiment, when the thread group includes more threads than the number of processing engines within the graphics multiprocessor 2434, processing can be performed in consecutive clock cycles. In at least one embodiment, multiple thread groups can be executed simultaneously on the graphics multiprocessor 2434.

[0308] In at least one embodiment, the graphics multiprocessor 2434 includes an internal cache memory for performing load and store operations. In at least one embodiment, the graphics multiprocessor 2434 may forgo the internal cache and use a cache memory within the processing cluster 2414 (e.g., L1 cache 2448). In at least one embodiment, each graphics multiprocessor 2434 may also access partition units (e.g., Figure 24A The L2 cache is located within partition units 2420A-2420N, which are shared among all processing clusters 2414 and can be used to transfer data between threads. In at least one embodiment, the graphics multiprocessor 2434 can also access off-chip global memory, which may include one or more of local parallel processor memory and / or system memory. In at least one embodiment, any memory outside of the parallel processing unit 2402 can be used as global memory. In at least one embodiment, the processing cluster 2414 includes multiple instances of the graphics multiprocessor 2434, which can share common instructions and data that can be stored in the L1 cache 2448.

[0309] In at least one embodiment, each processing cluster 2414 may include a memory management unit (MMU) 2445 configured to map virtual addresses to physical addresses. In at least one embodiment, one or more instances of the MMU 2445 may reside in Figure 24A The memory interface 2418 is located within the MMU 2445. In at least one embodiment, the MMU 2445 includes a set of page table entries (PTEs) for mapping virtual addresses to physical addresses of tiles and optionally to cache line indices. In at least one embodiment, the MMU 2445 may include an address translation lookup buffer (TLB) or a cache that may reside within the graphics multiprocessor 2434, the L1 cache 2448, or the processing cluster 2414. In at least one embodiment, physical addresses are processed to allocate surface data access locality for efficient request interleaving between partition units. In at least one embodiment, cache line indices may be used to determine whether a request for a cache line is a hit or a miss.

[0310] In at least one embodiment, the processing cluster 2414 can be configured such that each graphics multiprocessor 2434 is coupled to a texture unit 2436 to perform texture mapping operations that determine texture sample locations, read texture data, and filter texture data. In at least one embodiment, texture data is read as needed from an internal texture L1 cache (not shown) or from an L1 cache within the graphics multiprocessor 2434, and texture data is also retrieved from an L2 cache, local parallel processor memory, or system memory. In at least one embodiment, each graphics multiprocessor 2434 outputs a processed task to a data crossbar switch 2440 to provide the processed task to another processing cluster 2414 for further processing or to store the processed task in an L2 cache, local parallel processor memory, or system memory via a memory crossbar switch 2416. In at least one embodiment, a preROP 2442 (pre-raster operation unit) is configured to receive data from the graphics multiprocessor 2434 and direct the data to a ROP unit, which can be associated with a partitioning unit (e.g., [missing information]). Figure 24A The PreROP 2442 unit is located together with the partitioning units 2420A-2420N. In at least one embodiment, the PreROP 2442 unit can perform optimizations for color blending, organize pixel color data, and perform address translation.

[0311] Inference and / or training logic 1115 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 11A and / or Figure 11B Details regarding inference and / or training logic 1115 are provided. In at least one embodiment, inference and / or training logic 1115 may be used in a graphics processing cluster 2414 to perform inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures or neural network use cases described herein.

[0312] Figure 24D A graphics multiprocessor 2434 according to at least one embodiment is illustrated. In at least one embodiment, the graphics multiprocessor 2434 is coupled to a pipeline manager 2432 of a processing cluster 2414. In at least one embodiment, the graphics multiprocessor 2434 has an execution pipeline including, but not limited to, an instruction cache 2452, an instruction unit 2454, an address mapping unit 2456, a register file 2458, one or more general-purpose graphics processing unit (GPGPU) cores 2462, and one or more load / store units 2466. In at least one embodiment, the GPGPU cores 2462 and the load / store units 2466 are coupled to a cache memory 2472 and a shared memory 2470 via a memory and cache interconnect 2468.

[0313] In at least one embodiment, instruction cache 2452 receives a stream of instructions to be executed from pipeline manager 2432. In at least one embodiment, instructions are cached in instruction cache 2452 and dispatched to instruction unit 2454 for execution. In one embodiment, instruction unit 2454 may dispatch instructions as thread groups (e.g., thread bundles), assigning each thread of the thread group to a different execution unit within GPGPU core 2462. In at least one embodiment, instructions can access any local, shared, or global address space by specifying an address within a unified address space. In at least one embodiment, address mapping unit 2456 may be used to translate addresses in the unified address space into different memory addresses that can be accessed by load / store unit 2466.

[0314] In at least one embodiment, register file 2458 provides a set of registers for functional units of graphics multiprocessor 2434. In at least one embodiment, register file 2458 provides temporary storage for operands of data paths connected to functional units of graphics multiprocessor 2434 (e.g., GPGPU core 2462, load / store unit 2466). In at least one embodiment, register file 2458 is partitioned among each functional unit, such that a dedicated portion of register file 2458 is allocated to each functional unit. In at least one embodiment, register file 2458 is partitioned among different thread bundles being executed by graphics multiprocessor 2434.

[0315] In at least one embodiment, each of the GPGPU cores 2462 may include a floating-point unit (FPU) and / or an integer arithmetic logic unit (ALU) for executing instructions of the graphics multiprocessor 2434. In at least one embodiment, the GPGPU cores 2462 may be architecturally similar or may differ in architecture. In at least one embodiment, a first portion of the GPGPU core 2462 includes a single-precision FPU and an integer ALU, while a second portion of the GPGPU core includes a double-precision FPU. In at least one embodiment, the FPU may implement the IEEE 754-2008 standard for floating-point algorithms or enable variable-precision floating-point algorithms. In at least one embodiment, the graphics multiprocessor 2434 may additionally include one or more fixed-function or special-function units to perform specific functions, such as copying rectangles or pixel blending operations. In at least one embodiment, one or more of the GPGPU cores 2462 may also include fixed-function or special-function logic.

[0316] In at least one embodiment, the GPGPU core 2462 includes SIMD logic capable of executing a single instruction on multiple sets of data. In one embodiment, the GPGPU core 2462 can physically execute SIMD4, SIMD8, and SIMD16 instructions and logically execute SIMD1, SIMD2, and SIMD32 instructions. In at least one embodiment, the SIMD instructions for the GPGPU core can be generated by a shader compiler at compile time or automatically generated when executing a program written and compiled for a Single Program Multiple Data (SPMD) or SIMT architecture. In at least one embodiment, multiple threads of a program configured for a SIMT execution model can be executed using a single SIMD instruction. For example, in at least one embodiment, eight SIMD threads performing the same or similar operations can be executed in parallel using a single SIMD8 logic unit.

[0317] In at least one embodiment, the memory and cache interconnect 2468 is an interconnect network connecting each functional unit of the graphics multiprocessor 2434 to the register file 2458 and the shared memory 2470. In at least one embodiment, the memory and cache interconnect 2468 is a cross-switch interconnect that allows the load / store unit 2466 to perform load and store operations between the shared memory 2470 and the register file 2458. In at least one embodiment, the register file 2458 can operate at the same frequency as the GPGPU core 2462, resulting in very low latency for data transfer between the GPGPU core 2462 and the register file 2458. In at least one embodiment, the shared memory 2470 can be used to enable communication between threads executing on functional units within the graphics multiprocessor 2434. In at least one embodiment, the cache memory 2472 can be used, for example, as a data cache to cache texture data communicated between functional units and texture units 2436. In at least one embodiment, the shared memory 2470 can also be used as a program-managed cache.

[0318] In at least one embodiment, in addition to the automatically cached data stored in cache memory 2472, the thread executing on GPGPU core 2462 can also programmatically store data in shared memory.

[0319] In at least one embodiment, a parallel processor or GPGPU, as described herein, is communicatively coupled to a host / processor core to accelerate graphics operations, machine learning operations, pattern analysis operations, and various general-purpose GPU (GPGPU) functions. In at least one embodiment, the GPU may be communicatively coupled to the host processor / core via a bus or other interconnect (e.g., high-speed interconnects such as PCIe or NVLink). In at least one embodiment, the GPU may be integrated with the core on a package or chip and communicatively coupled to the core via an internal processor bus / interconnect (i.e., within the package or chip). In at least one embodiment, regardless of how the GPU is connected, the processor core may assign work to the GPU in the form of a sequence of commands / instructions contained in a job descriptor. In at least one embodiment, the GPU then uses dedicated circuitry / logic to efficiently process these commands / instructions.

[0320] The inference and / or training logic 1115 is used to perform inference and / or training operations associated with one or more embodiments. The following is in conjunction with... Figure 11A and / or Figure 11B Details regarding the inference and / or training logic 1115 are provided. In at least one embodiment, the inference and / or training logic 1115 may be used in a graphics multiprocessor 2434 to perform inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures or neural network use cases described herein.

[0321] In at least one embodiment, processor 114 may be implemented at least partially by parallel processor 2400 and / or graphics multiprocessor 2434, and memory 116 may be implemented at least partially by shared memory 2470 and / or parallel processor memory 2422. As another non-limiting example, instruction 118 (see...) Figure 1 This can be executed by a parallel processor 2400 and / or a graphics multiprocessor 2434, and is used to obtain the optical flow map 110 (see...). Figure 1 and Figure 2 As described above, system 100 (see...) Figure 1 This can be implemented by at least a portion of the reasoning and / or training logic 1115. (See reference...) Figure 1 The parallel processor 2400 and / or the graphics multiprocessor 2434 (see Figure 24) can perform inference and / or training operations (e.g., at least a portion of the inference and / or training logic 1115) for any of the noise reduction process 120, edge detection process 122, object detection process 124, optional thresholding process 126, decision logic 128, extraction process 130 and SGM process 132.

[0322] Figure 25A multi-GPU computing system 2500 according to at least one embodiment is illustrated. In at least one embodiment, the multi-GPU computing system 2500 may include a processor 2502 coupled to a plurality of general-purpose graphics processing units (GPGPUs) 2506A-D via a host interface switch 2504. In at least one embodiment, the host interface switch 2504 is a PCI Express switch device that couples the processor 2502 to a PCI Express bus, through which the processor 2502 can communicate with the GPGPUs 2506A-D. In at least one embodiment, the GPGPUs 2506A-D may be interconnected via a set of high-speed P2P GPU-to-GPU links 2516. In at least one embodiment, the GPU-to-GPU links 2516 are connected to each of the GPGPUs 2506A-D via dedicated GPU links. In at least one embodiment, the P2P GPU links 2516 enable direct communication between each GPGPU 2506A-D without communication via the host interface switch 2504 to which the processor 2502 is connected. In at least one embodiment, when GPU-to-GPU traffic is directed to the P2P GPU link 2516, the host interface switch 2504 remains available for system memory access or, for example, communication with other instances of the multi-GPU computing system 2500 via one or more network devices. While in at least one embodiment, the GPGPUs 2506A-D are connected to the processor 2502 via the host interface switch 2504, in at least one embodiment, the processor 2502 includes direct support for the P2P GPU link 2516 and can be directly connected to the GPGPUs 2506A-D.

[0323] Inference and / or training logic 1115 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 11A and / or Figure 11B Details regarding inference and / or training logic 1115 are provided. In at least one embodiment, inference and / or training logic 1115 may be used in a multi-GPU computing system 2500 for performing inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures or neural network use cases described herein.

[0324] In at least one embodiment, system 100 (see...) Figure 1 The processor 114 can be implemented at least partially by the multi-GPU computing system 2500. In at least one embodiment, the processor 114 can be implemented at least partially by one or more of the processor 2502 and / or GPGPU 2506A-D. As another non-limiting example, instruction 118 (see...) Figure 1This can be executed by one or more of the processor 2502 and / or GPGPU 2506A-D, and used to obtain optical flow map 110 (see...). Figure 1 and Figure 2 As described above, system 100 (see...) Figure 1 This can be implemented by at least a portion of the reasoning and / or training logic 1115. (See reference...) Figure 1 Multi-GPU computing system 2500 (see Figure 25 It can perform inference and / or training operations (e.g., at least a portion of inference and / or training logic 1115) for any of the noise reduction process 120, edge detection process 122, object detection process 124, optional thresholding process 126, decision logic 128, extraction process 130 and SGM process 132.

[0325] Figure 26 This is a block diagram of a graphics processor 2600 according to at least one embodiment. In at least one embodiment, the graphics processor 2600 includes a ring interconnect 2602, a pipeline front end 2604, a media engine 2637, and graphics cores 2680A-2680N. In at least one embodiment, the ring interconnect 2602 couples the graphics processor 2600 to other processing units, said processing units including other graphics processors or one or more general-purpose processor cores. In at least one embodiment, the graphics processor 2600 is one of many processors integrated within a multi-core processing system.

[0326] In at least one embodiment, the graphics processor 2600 receives multiple batches of commands via a ring interconnect 2602. In at least one embodiment, the input commands are interpreted by a command streamer 2603 in a pipeline front-end 2604. In at least one embodiment, the graphics processor 2600 includes scalable execution logic for performing 3D geometry processing and media processing via graphics cores 2680A-2680N. In at least one embodiment, for 3D geometry processing commands, the command streamer 2603 provides the commands to the geometry pipeline 2636. In at least one embodiment, for at least some media processing commands, the command streamer 2603 provides the commands to a video front-end 2634, which is coupled to a media engine 2637. In at least one embodiment, the media engine 2637 includes a video quality engine (VQE) 2630 for video and image post-processing, and a multi-format encoding / decoding (MFX) engine 2633 for providing hardware-accelerated media data encoding and decoding. In at least one embodiment, the geometry pipeline 2636 and the media engine 2637 each generate an execution thread for thread execution resources provided by at least one graphics core 2680.

[0327] In at least one embodiment, the graphics processor 2600 includes scalable thread execution resources featuring graphics cores 2680A-2680N (which may be modular and sometimes referred to as core slices), each graphics core having multiple sub-cores 2650A-2650N, 2660A-2660N (sometimes referred to as core sub-slices). In at least one embodiment, the graphics processor 2600 may have any number of graphics cores 2680A. In at least one embodiment, the graphics processor 2600 includes graphics cores 2680A having at least a first sub-core 2650A and a second sub-core 2660A. In at least one embodiment, the graphics processor 2600 is a low-power processor with a single sub-core (e.g., 2650A). In at least one embodiment, the graphics processor 2600 includes multiple graphics cores 2680A-2680N, each graphics core including a set of first sub-cores 2650A-2650N and a set of second sub-cores 2660A-2660N. In at least one embodiment, each of the first sub-cores 2650A-2650N includes at least a first set of execution units 2652A-2652N and media / texture samplers 2654A-2654N. In at least one embodiment, each of the second sub-cores 2660A-2660N includes at least a second set of execution units 2662A-2662N and samplers 2664A-2664N. In at least one embodiment, each of the sub-cores 2650A-2650N and 2660A-2660N shares a set of shared resources 2670A-2670N. In at least one embodiment, the shared resources include shared cache memory and pixel operation logic.

[0328] Inference and / or training logic 1115 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 11A and / or Figure 11B Details are provided regarding the inference and / or training logic 1115. In at least one embodiment, the inference and / or training logic 1115 may be used in the graphics processor 2600 to perform inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures or neural network use cases described herein.

[0329] In at least one embodiment, processor 114 may be implemented at least in part by graphics processor 2600. For example, instruction 118 (see...) Figure 1 This can be executed by the graphics processor 2600 and used to obtain the optical flow map 110 (see...). Figure 1 and Figure 2 As described above, system 100 (see...) Figure 1This can be implemented by at least a portion of the reasoning and / or training logic 1115. (See reference...) Figure 1 Graphics processor 2600 (see...) Figure 26 It can perform inference and / or training operations (e.g., at least a portion of inference and / or training logic 1115) for any of the noise reduction process 120, edge detection process 122, object detection process 124, optional thresholding process 126, decision logic 128, extraction process 130 and SGM process 132.

[0330] Figure 27 This is a block diagram illustrating a microarchitecture for a processor 2700 according to at least one embodiment, the processor 2700 including logic circuitry for executing instructions. In at least one embodiment, the processor 2700 can execute instructions, including x86 instructions, ARM instructions, and special-purpose instructions for application-specific integrated circuits (ASICs). In at least one embodiment, the processor 2700 may include registers for storing packaged data, such as the 64-bit wide MMX registers used in Intel Corporation's Santa Clara, California-enabled microprocessors employing MMX technology. TM Registers. In at least one embodiment, MMX registers available in integer and floating-point forms can operate alongside packaged data elements accompanied by Single Instruction Multiple Data (SIMD) and Streaming SIMD Extensions (SSE) instructions. In at least one embodiment, a 128-bit wide XMM register associated with SSE2, SSE3, SSE4, AVX, or later (generally referred to as SSEx) technologies can hold such packaged data operands. In at least one embodiment, processor 2700 can execute instructions to accelerate machine learning or deep learning algorithms, training, or inference.

[0331] In at least one embodiment, processor 2700 includes an ordered front end ("front end") 2701 to fetch instructions to be executed and prepare instructions for later use in the processor pipeline. In at least one embodiment, front end 2701 may include several units. In at least one embodiment, instruction prefetcher 2726 fetches instructions from memory and provides the instructions to instruction decoder 2728, which in turn decodes or interprets the instructions. For example, in at least one embodiment, instruction decoder 2728 decodes the received instructions into one or more machine-executable so-called "micro-instructions" or "micro-operations" (also referred to as "micro-operations" or "micro-instructions"). In at least one embodiment, instruction decoder 2728 parses the instructions into opcodes and corresponding data and control fields, which can be used by the microarchitecture to perform operations according to at least one embodiment. In at least one embodiment, trace cache 2730 may assemble the decoded micro-instructions into a program-ordered sequence or trace in micro-instruction queue 2734 for execution. In at least one embodiment, when the trace cache 2730 encounters complex instructions, the microcode ROM 2732 provides the microinstructions required to complete the operation.

[0332] In at least one embodiment, some instructions may be converted into a single micro-operation, while others require several micro-operations to complete the entire operation. In at least one embodiment, if more than four micro-instructions are required to complete an instruction, the instruction decoder 2728 may access the microcode ROM 2732 to execute the instruction. In at least one embodiment, an instruction may be decoded into a small number of micro-instructions for processing at the instruction decoder 2728. In at least one embodiment, if multiple micro-instructions are required to complete the operation, the instructions may be stored in the microcode ROM 2732. In at least one embodiment, the trace cache 2730 references an entry point programmable logic array (PLA) to determine the correct micro-instruction pointer for reading a microcode sequence from the microcode ROM 2732 to complete one or more instructions, according to at least one embodiment. In at least one embodiment, after the microcode ROM 2732 has completed the micro-operation ordering of the instructions, the machine front end 2701 may resume fetching micro-operations from the trace cache 2730.

[0333] In at least one embodiment, the out-of-order execution engine ("out-of-order engine") 2703 can prepare instructions for execution. In at least one embodiment, the out-of-order execution logic has multiple buffers to smooth and reorder the instruction flow to optimize performance as instructions descend the pipeline and are scheduled for execution. In at least one embodiment, the out-of-order execution engine 2703 includes, but is not limited to, an allocator / register renamer 2740, a memory microinstruction queue 2742, an integer / floating-point microinstruction queue 2744, a memory scheduler 2746, a fast scheduler 2702, a slow / general-purpose floating-point scheduler ("slow / general-purpose FP scheduler") 2704, and a simple floating-point scheduler ("simple FP scheduler") 2706. In at least one embodiment, the fast scheduler 2702, the slow / general-purpose floating-point scheduler 2704, and the simple floating-point scheduler 2706 are also collectively referred to as "microinstruction schedulers 2702, 2704, and 2706". In at least one embodiment, the allocator / register renamer 2740 allocates the machine buffers and resources required for the sequential execution of each microinstruction. In at least one embodiment, the allocator / register renamer 2740 renames logical registers to entries in a register file. In at least one embodiment, the allocator / register renamer 2740 also allocates entries for each microinstruction in one of two microinstruction queues, a memory microinstruction queue 2742 for memory operations and an integer / floating-point microinstruction queue 2744 for non-memory operations, preceding the memory scheduler 2746 and microinstruction schedulers 2702, 2704, and 2706. In at least one embodiment, the microinstruction schedulers 2702, 2704, and 2706 determine when they are ready to execute a microinstruction based on the readiness of their dependent input register operand sources and the availability of the execution resource microinstructions that need to be completed. The fast scheduler 2702 of at least one embodiment can schedule on each half of the master clock cycle, while the slow / general-purpose floating-point scheduler 2704 and the simple floating-point scheduler 2706 can schedule once per master processor clock cycle. In at least one embodiment, microinstruction schedulers 2702, 2704, and 2706 arbitrate the scheduling port to schedule microinstructions for execution.

[0334] In at least one embodiment, execution block 2711 includes, but is not limited to, integer register file / tribute network 2708, floating-point register file / tribute network ("FP register file / tribute network") 2710, address generation units ("AGU") 2712 and 2714, fast arithmetic logic units ("fast ALU") 2716 and 2718, slow arithmetic logic unit ("slow ALU") 2720, floating-point ALU ("FP") 2722, and floating-point move unit ("FP move") 2724. In at least one embodiment, integer register file / tribute network 2708 and floating-point register file / bypass network 2710 are also referred to herein as "register files 2708, 2710". In at least one embodiment, AGUs 2712 and 2714, fast ALUs 2716 and 2718, slow ALU 2720, floating-point ALU 2722, and floating-point movement unit 2724 are also referred to herein as "execution units 2712, 2714, 2716, 2718, 2720, 2722, and 2724". In at least one embodiment, execution block 2711 may include, but is not limited to, any number (including zero) and type of register files, branch networks, address generation units, and execution units (in any combination).

[0335] In at least one embodiment, register networks 2708, 2710 may be arranged between microinstruction schedulers 2702, 2704, 2706 and execution units 2712, 2714, 2716, 2718, 2720, 2722, and 2724. In at least one embodiment, integer register file / tribute network 2708 performs integer operations. In at least one embodiment, floating-point register file / tribute network 2710 performs floating-point operations. In at least one embodiment, each of register networks 2708, 2710 may include, but is not limited to, a tribute network that can bypass or forward recently completed results not yet written to a register file to a new dependent object. In at least one embodiment, register networks 2708, 2710 may communicate data with each other. In at least one embodiment, integer register file / tribute network 2708 may include, but is not limited to, two separate register files, one register file for low-order 32-bit data and a second register file for high-order 32-bit data. In at least one embodiment, the floating-point register file / branch network 2710 may include, but is not limited to, entries with a width of 128 bits, since floating-point instructions typically have operands with a width of 64 to 128 bits.

[0336] In at least one embodiment, execution units 2712, 2714, 2716, 2718, 2720, 2722, and 2724 can execute instructions. In at least one embodiment, register networks 2708 and 2710 store integer and floating-point data operation values ​​that the microinstructions need to execute. In at least one embodiment, processor 2700 can be, but is not limited to, any number of execution units 2712, 2714, 2716, 2718, 2720, 2722, and 2724, and combinations thereof. In at least one embodiment, floating-point ALU 2722 and floating-point movement unit 2724 can perform floating-point, MMX, SIMD, AVX, and SSE or other operations, including specialized machine learning instructions. In at least one embodiment, floating-point ALU 2722 can be, but is not limited to, a 64-bit multiplication-64-bit floating-point divider to perform division, square root, and remainder micro-operations. In at least one embodiment, floating-point hardware can be used to process instructions involving floating-point values. In at least one embodiment, ALU operations can be passed to fast ALUs 2716 and 2718. In at least one embodiment, fast ALUs 2716 and 2718 can perform fast operations with an effective delay of half a clock cycle. In at least one embodiment, most complex integer operations are routed to slow ALU 2720, because slow ALU 2720 can include, but is not limited to, integer execution hardware for long-latency type operations, such as multipliers, shifters, flag logic, and branching. In at least one embodiment, memory load / store operations can be performed by AGUs 2712 and 2714. In at least one embodiment, fast ALU 2716, fast ALU 2718, and slow ALU 2720 can perform integer operations on 64-bit data operands. In at least one embodiment, fast ALU 2716, fast ALU 2718, and slow ALU 2720 can be implemented to support various data bit sizes, including sixteen, thirty-two, 128, 256, etc. In at least one embodiment, the floating-point ALU 2722 and the floating-point moving unit 2724 can be implemented to support a range of operands with various bit widths, for example, they can be combined with SIMD and multimedia instructions to operate on 128-bit wide packaged data operands.

[0337] In at least one embodiment, microinstruction schedulers 2702, 2704, and 2706 schedule dependent operations before the parent load completes execution. In at least one embodiment, since microinstructions can be speculatively scheduled and executed within processor 2700, processor 2700 may also include logic for handling memory misses. In at least one embodiment, if a data load miss occurs in the data cache, there may be a dependent operation running in the pipeline that temporarily deprives the scheduler of the correct data. In at least one embodiment, a replay mechanism tracks and re-executes instructions that use incorrect data. In at least one embodiment, it may be necessary to replay dependent operations and may allow independent operations to be completed. In at least one embodiment, the scheduler and replay mechanism of at least one embodiment of the processor may also be designed to capture instruction sequences for text string comparison operations.

[0338] In at least one embodiment, "register" can refer to an onboard processor storage location that can be used as part of an instruction that identifies operands. In at least one embodiment, registers can be those that can be used externally to the processor (from a programmer's perspective). In at least one embodiment, registers may not be limited to a particular type of circuitry. Rather, in at least one embodiment, registers can store data, provide data, and perform the functions described herein. In at least one embodiment, the registers described herein can be implemented using a variety of different techniques via circuitry within the processor, such as dedicated physical registers, dynamically allocated physical registers renamed using register renaming, combinations of dedicated and dynamically allocated physical registers, etc. In at least one embodiment, the integer register stores 32-bit integer data. The register file of at least one embodiment also includes eight multimedia SIMD registers for encapsulating data.

[0339] Inference and / or training logic 1115 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 11A and / or Figure 11B Details regarding the inference and / or training logic 1115 are provided. In at least one embodiment, some or all of the inference and / or training logic 1115 may be incorporated into execution block 2711 and other memories or registers shown or not shown. For example, in at least one embodiment, the training and / or inference techniques described herein may use one or more ALUs shown in execution block 2711. Furthermore, weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown) that configure the ALUs of execution block 2711 to execute one or more machine learning algorithms, neural network architectures, use cases, or training techniques described herein.

[0340] In at least one embodiment, processor 114 may be implemented at least in part by processor 2700. For example, instruction 118 (see...) Figure 1 This can be executed by processor 2700 and used to obtain optical flow map 110 (see...). Figure 1 and Figure 2 As described above, system 100 (see...) Figure 1 This can be implemented by at least a portion of the reasoning and / or training logic 1115. (See reference...) Figure 1 Processor 2700 (see Figure 27 It can perform inference and / or training operations (e.g., at least a portion of inference and / or training logic 1115) for any of the noise reduction process 120, edge detection process 122, object detection process 124, optional thresholding process 126, decision logic 128, extraction process 130 and SGM process 132.

[0341] Figure 28 A deep learning application processor 2800 according to at least one embodiment is illustrated. In at least one embodiment, the deep learning application processor 2800 uses instructions, which, if executed by the deep learning application processor 2800, cause the deep learning application processor 2800 to perform some or all of the processes and techniques described herein. In at least one embodiment, the deep learning application processor 2800 is an application-specific integrated circuit (ASIC). In at least one embodiment, the application processor 2800 performs matrix multiplication operations or is "hardwired" into hardware as a result of executing one or more instructions or both. In at least one embodiment, the deep learning application processor 2800 includes, but is not limited to, a processing cluster 2810(1)-2810(12), an inter-chip link ("ICL") 2820(1)-2820(12), an inter-chip controller ("ICC") 2830(1)-2830(2), a second-generation high-bandwidth memory ("HBM2") 2840(1)-2840(4), a memory controller ("Mem Ctrlr") 2842(1)-2842(4), and a high-bandwidth memory physical layer ("HBM"). PHY‖)2844(1)-2844(4), Management Controller Central Processing Unit (“Management Controller CPU”)2850, Serial Peripheral Interface, Internal Integrated Circuits and General Purpose Input / Output Blocks (“SPI, I2C, GPIO”)2860, Peripheral Component Interconnect Fast Controller and Direct Memory Access Block (“PCIe Controller and DMA”)2870, and Sixteen-Channel Peripheral Component Interconnect Fast Port (“PCIExpress x 16”)2880.

[0342] In at least one embodiment, the processing cluster 2810 can perform deep learning operations, including inference or prediction operations based on weight parameters computed using one or more training techniques, including those described herein. In at least one embodiment, each processing cluster 2810 can include, but is not limited to, any number and type of processors. In at least one embodiment, the deep learning application processor 2800 can include any number and type of processing cluster 2810. In at least one embodiment, the inter-chip link 2820 is bidirectional. In at least one embodiment, the inter-chip link 2820 and the inter-chip controller 2830 enable multiple deep learning application processors 2800 to exchange information, including activation information generated from executing one or more machine learning algorithms embodied in one or more neural networks. In at least one embodiment, the deep learning application processor 2800 can include any number (including zero) and type of ICL 2820 and ICC 2830.

[0343] In at least one embodiment, the HBM2 2840 provides a total of 32GB of memory. In at least one embodiment, the HBM2 2840(i) is associated with both the memory controller 2842(i) and the HBM PHY 2844(i), where "i" is any integer. In at least one embodiment, any number of HBM2 2840s can provide any type and total amount of high-bandwidth memory and can be associated with any number (including zero) and type of memory controller 2842 and HBM PHY 2844. In at least one embodiment, any number and type of blocks can replace SPI, I2C, GPIO2860, PCIe controller, and DMA 2870 and / or PCIe 2880 to implement any number and type of communication standards in any technically feasible manner.

[0344] Inference and / or training logic 1115 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 11A and / or Figure 11B Details regarding the inference and / or training logic 1115 are provided. In at least one embodiment, the deep learning application processor is used to train a machine learning model (e.g., a neural network) to predict or infer information provided to the deep learning application processor 2800. In at least one embodiment, the deep learning application processor 2800 is used to infer or predict information based on a trained machine learning model (e.g., a neural network) that has been trained by another processor or system or by the deep learning application processor 2800. In at least one embodiment, the processor 2800 may be used to perform one or more neural network use cases described herein.

[0345] In at least one embodiment, processor 114 may be implemented at least in part by deep learning application processor 2800. For example, instruction 118 (see...) Figure 1 This can be executed by the deep learning application processor 2800 and used to obtain the optical flow map 110 (see...). Figure 1 and Figure 2 As described above, system 100 (see...) Figure 1 This can be implemented by at least a portion of the reasoning and / or training logic 1115. (See reference...) Figure 1 Deep learning application processor 2800 (see...) Figure 28 It can perform inference and / or training operations (e.g., at least a portion of inference and / or training logic 1115) for any of the noise reduction process 120, edge detection process 122, object detection process 124, optional thresholding process 126, decision logic 128, extraction process 130 and SGM process 132.

[0346] Figure 29 This is a block diagram of a neuromorphic processor 2900 according to at least one embodiment. In at least one embodiment, the neuromorphic processor 2900 may receive one or more inputs from a source external to the neuromorphic processor 2900. In at least one embodiment, these inputs may be transmitted to one or more neurons 2902 within the neuromorphic processor 2900. In at least one embodiment, the neurons 2902 and their components may be implemented using circuitry or logic including one or more arithmetic logic units (ALUs). In at least one embodiment, the neuromorphic processor 2900 may include, but is not limited to, thousands upon thousands of instances of neurons 2902, but any suitable number of neurons 2902 may be used. In at least one embodiment, each instance of a neuron 2902 may include a neuron input 2904 and a neuron output 2906. In at least one embodiment, a neuron 2902 may generate an output that can be transmitted to the inputs of other instances of the neuron 2902. In at least one embodiment, the neuron input 2904 and the neuron output 2906 may be interconnected via synapses 2908.

[0347] In at least one embodiment, neuron 2902 and synapse 2908 may be interconnected, causing neuromorphic processor 2900 to operate to process or analyze information received by neuromorphic processor 2900. In at least one embodiment, neuron 2902 may send an output pulse (or "trigger" or "peak") when the input received through neuron input 2904 exceeds a threshold. In at least one embodiment, neuron 2902 may sum or integrate the signal received at neuron input 2904. For example, in at least one embodiment, neuron 2902 may be implemented as a leaky integral-triggered neuron, wherein if the summation (referred to as "membrane potential") exceeds a threshold, neuron 2902 may use a transfer function such as a sigmoid or threshold function to generate an output (or "trigger"). In at least one embodiment, the leaky integral-triggered neuron may sum the signal received at neuron input 2904 to a membrane potential and may apply an attenuation factor (or leakage) to reduce the membrane potential. In at least one embodiment, a leaking integral-triggered neuron may trigger if multiple input signals are received at neuron input 2904 quickly enough to exceed a threshold (i.e., before the membrane potential decays too low to trigger). In at least one embodiment, neuron 2902 may be implemented using circuitry or logic that receives input, integrates the input to the membrane potential, and decays the membrane potential. In at least one embodiment, the input may be averaged, or any other suitable transfer function may be used. Furthermore, in at least one embodiment, neuron 2902 may include, but is not limited to, comparator circuitry or logic that generates an output spike at neuron output 2906 when the result of applying the transfer function to neuron input 2904 exceeds a threshold. In at least one embodiment, once neuron 2902 is triggered, it can ignore previously received input information by, for example, resetting the membrane potential to 0 or another suitable default value. In at least one embodiment, once the membrane potential is reset to 0, neuron 2902 may resume normal operation after a suitable period of time (or recovery period).

[0348] In at least one embodiment, neurons 2902 can be interconnected via synapses 2908. In at least one embodiment, synapses 2908 can be operated to transmit signals from the output of a first neuron 2902 to the input of a second neuron 2902. In at least one embodiment, neurons 2902 can transmit information on more than one instance of synapses 2908. In at least one embodiment, one or more instances of neuron outputs 2906 can be connected via instances of synapses 2908 to instances of neuron inputs 2904 in the same neuron 2902. In at least one embodiment, an instance of neuron 2902 that produces an output to be transmitted on the instance of synapse 2908 can be referred to as a "presynaptic neuron". In at least one embodiment, an instance of neuron 2902 that receives input transmitted via an instance of synapse 2908 can be referred to as a "postsynaptic neuron". In at least one embodiment, regarding various instances of synapse 2908, since instances of neuron 2902 can receive input from one or more instances of synapse 2908 and can also transmit output through one or more instances of synapse 2908, a single instance of neuron 2902 can be both a "presynaptic neuron" and a "postsynaptic neuron".

[0349] In at least one embodiment, neurons 2902 may be organized into one or more layers. In at least one embodiment, each instance of neuron 2902 may have a neuron output 2906, which may fan out to one or more neuron inputs 2904 via one or more synapses 2908. In at least one embodiment, the neuron output 2906 of neuron 2902 in the first layer 2910 may be connected to the neuron input 2904 of neuron 2902 in the second layer 2912. In at least one embodiment, layer 2910 may be referred to as a "feedforward layer". In at least one embodiment, each instance of neuron 2902 in an instance of the first layer 2910 may fan out to each instance of neuron 2902 in the second layer 2912. In at least one embodiment, the first layer 2910 may be referred to as a "fully connected feedforward layer". In at least one embodiment, each instance of neuron 2902 in each instance of the second layer 2912 fan out to fewer than all instances of neuron 2902 in the third layer 2914. In at least one embodiment, the second layer 2912 may be referred to as a "sparsely connected feedforward layer". In at least one embodiment, neurons 2902 in the second layer 2912 may fan out to neurons 2902 in multiple other layers, including neurons 2902 fan out to the second layer 2912. In at least one embodiment, the second layer 2912 may be referred to as a "recurrent layer". In at least one embodiment, the neuromorphic processor 2900 may be any suitable combination of recurrent layers and feedforward layers, including but not limited to sparsely connected feedforward layers and fully connected feedforward layers.

[0350] In at least one embodiment, the neuromorphic processor 2900 may include, but is not limited to, a reconfigurable interconnect architecture or dedicated hardwired interconnects to connect synapses 2908 to neurons 2902. In at least one embodiment, the neuromorphic processor 2900 may include, but is not limited to, circuitry or logic that allows synapses to be assigned to different neurons 2902 as needed, depending on the neural network topology and neuron fan-in / fan-out. For example, in at least one embodiment, synapses 2908 may be connected to neurons 2902 using interconnect structures (such as on-chip networks) or via dedicated connections. In at least one embodiment, synaptic interconnects and their components may be implemented using circuitry or logic.

[0351] In at least one embodiment, processor 114 may be implemented at least partially by neuromorphic processor 2900. For example, instruction 118 (see...) Figure 1 This can be executed by the neuromorphic processor 2900 and used to obtain the optical flow map 110 (see [link]). Figure 1 and Figure 2 As described above, system 100 (see...) Figure 1This can be implemented by at least a portion of the reasoning and / or training logic 1115. (See reference...) Figure 1 Neuromorphic processor 2900 (see Figure 29 It can perform inference and / or training operations (e.g., at least a portion of inference and / or training logic 1115) for any of the noise reduction process 120, edge detection process 122, object detection process 124, optional thresholding process 126, decision logic 128, extraction process 130 and SGM process 132.

[0352] Figure 30 A processing system according to at least one embodiment is illustrated. In at least one embodiment, system 3000 includes one or more processors 3002 and one or more graphics processors 3008, and may be a single-processor desktop system, a multi-processor workstation system, or a server system having a large number of processors 3002 or processor cores 3007. In at least one embodiment, system 3000 is a processing platform incorporated within a system-on-a-chip (SoC) integrated circuit for use in mobile, handheld, or embedded devices.

[0353] In at least one embodiment, system 3000 may include or be integrated into a server-based gaming platform, including a game console, mobile game console, handheld game console, or online game console, which are game and media consoles. In at least one embodiment, system 3000 is a mobile phone, smartphone, tablet computing device, or mobile internet device. In at least one embodiment, processing system 3000 may also include components coupled to or integrated into a wearable device, such as a smartwatch, smart glasses, augmented reality, or virtual reality device. In at least one embodiment, processing system 3000 is a television or set-top box device having one or more processors 3002 and a graphical interface generated by one or more graphics processors 3008.

[0354] In at least one embodiment, each of the one or more processors 3002 includes one or more processor cores 3007 for processing instructions that, when executed, perform operations against the system and user software. In at least one embodiment, each of the one or more processor cores 3007 is configured to process a specific instruction sequence 3009. In at least one embodiment, the instruction sequence 3009 may facilitate Complex Instruction Set Computing (CISC), Reduced Instruction Set Computing (RISC), or computation via Very Long Instruction Word (VLIW). In at least one embodiment, each processor core 3007 may process a different instruction sequence 3009, which may include instructions that facilitate the emulation of other instruction sequences. In at least one embodiment, the processor core 3007 may also include other processing devices, such as a digital signal processor (DSP).

[0355] In at least one embodiment, processor 3002 includes cache memory 3004. In at least one embodiment, processor 3002 may have a single internal cache or more levels of internal caches. In at least one embodiment, the cache memory is shared among the various components of processor 3002. In at least one embodiment, processor 3002 also uses an external cache (e.g., a Level 3 (L3) cache or a last-level cache (LLC)) (not shown), which can be shared among processor cores 3007 using known cache coherence techniques. In at least one embodiment, processor 3002 further includes a register file 3006, which may include different types of registers (e.g., integer registers, floating-point registers, status registers, and instruction pointer registers) for storing different types of data. In at least one embodiment, register file 3006 may include general-purpose registers or other registers.

[0356] In at least one embodiment, one or more processors 3002 are coupled to one or more interface buses 3010 to transmit communication signals, such as address, data, or control signals, between the processors 3002 and other components in the system 3000. In at least one embodiment, the interface bus 3010 may be a processor bus, such as a version of the Direct Media Interface (DMI) bus. In at least one embodiment, the interface bus 3010 is not limited to the DMI bus and may include one or more peripheral component interconnect buses (e.g., PCI, PCI Express), memory buses, or other types of interface buses. In at least one embodiment, the processor 3002 includes an integrated memory controller 3016 and a platform controller hub 3030. In at least one embodiment, the memory controller 3016 facilitates communication between memory devices and other components of the processing system 3000, while the platform controller hub (PCH) 3030 provides connectivity to input / output (I / O) devices via a local I / O bus.

[0357] In at least one embodiment, memory device 3020 may be a dynamic random access memory (DRAM) device, a static random access memory (SRAM) device, a flash memory device, a phase-change memory device, or a device with suitable performance for use as processor memory. In at least one embodiment, memory device 3020 may be used as system memory of processing system 3000 to store data 3022 and instructions 3021 for use when one or more processors 3002 execute an application or process. In at least one embodiment, memory controller 3016 is also coupled to an optional external graphics processor 3012, which may communicate with one or more graphics processors 3008 of processor 3002 to perform graphics and media operations. In at least one embodiment, display device 3011 may be connected to processor 3002. In at least one embodiment, display device 3011 may include one or more internal display devices, such as in mobile electronic devices or laptop devices, or external display devices connected via a display interface (e.g., DisplayPort). In at least one embodiment, the display device 3011 may include a head-mounted display (HMD), such as a stereoscopic display device for virtual reality (VR) or augmented reality (AR) applications.

[0358] In at least one embodiment, the platform controller hub 3030 enables peripheral devices to connect to the storage device 3020 and the processor 3002 via a high-speed I / O bus. In at least one embodiment, the I / O peripheral devices include, but are not limited to, an audio controller 3046, a network controller 3034, a firmware interface 3028, a wireless transceiver 3026, a touch sensor 3025, and a data storage device 3024 (e.g., a hard disk drive, flash memory, etc.). In at least one embodiment, the data storage device 3024 may be connected via a storage interface (e.g., SATA) or via a peripheral bus, such as a peripheral component interconnect bus (e.g., PCI, PCIe). In at least one embodiment, the touch sensor 3025 may include a touchscreen sensor, a pressure sensor, or a fingerprint sensor. In at least one embodiment, the wireless transceiver 3026 may be a Wi-Fi transceiver, a Bluetooth transceiver, or a mobile network transceiver, such as a 3G, 4G, or LTE transceiver. In at least one embodiment, the firmware interface 3028 enables communication with the system firmware and may be, for example, a Unified Extensible Firmware Interface (UEFI). In at least one embodiment, the network controller 3034 can enable network connectivity to a wired network. In at least one embodiment, a high-performance network controller (not shown) is coupled to an interface bus 3010.

[0359] In at least one embodiment, the audio controller 3046 is a multi-channel high-definition audio controller. In at least one embodiment, the processing system 3000 includes an optional legacy I / O controller 3040 for coupling legacy (e.g., Personal System 2 (PS / 2)) devices to the system 3000.

[0360] In at least one embodiment, the platform controller hub 3030 may also be connected to one or more Universal Serial Bus (USB) controllers 3042, which connect input devices such as a keyboard and mouse combination 3043, a camera 3044, or other USB input devices.

[0361] In at least one embodiment, instances of the memory controller 3016 and platform controller hub 3030 may be integrated into a discrete external graphics processor, such as external graphics processor 3012. In at least one embodiment, the platform controller hub 3030 and / or the memory controller 3016 may be external to one or more processors 3002. For example, in at least one embodiment, system 3000 may include an external memory controller 3016 and a platform controller hub 3030, which may be configured as a memory controller hub and a peripheral controller hub in a system chipset communicating with processor 3002.

[0362] Inference and / or training logic 1115 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 11A and / or Figure 11B Details regarding the inference and / or training logic 1115 are provided. In at least one embodiment, some or all of the inference and / or training logic 1115 may be incorporated into the graphics processor 3008. For example, in at least one embodiment, the training and / or inference techniques described herein may use one or more ALUs embodied in a 3D pipeline. Furthermore, in at least one embodiment, the inference and / or training operations described herein may use, in addition to Figure 11A or Figure 11B The logic is performed using logic other than that shown. In at least one embodiment, the weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown), which configure the ALU of the graphics processor 3008 to execute one or more machine learning algorithms, neural network architectures, use cases, or training techniques described herein.

[0363] In at least one embodiment, system 100 (see...) Figure 1 The processor 114 may be implemented at least partially by system 3000. In at least one embodiment, the processor 114 may be implemented at least partially by processor 3002, processor core 3007, graphics processor 3008 and / or optional external graphics processor 3012, and the memory 116 may be implemented at least partially by memory device 3020. For example, instruction 118 (see...) Figure 1 This can be executed by processor 3002, processor core 3007, graphics processor 3008 and / or optional external graphics processor 3012, and is used to obtain optical flow map 110 (see...). Figure 1 and Figure 2 As described above, system 100 (see...) Figure 1 This can be implemented by at least a portion of the reasoning and / or training logic 1115. (See reference...) Figure 1 Processor 3002, processor core 3007, graphics processor 3008 and / or optional external graphics processor 3012 (see Processor 3002, processor core 3007, graphics processor 3008 and / or optional external graphics processor 3012) Figure 30 It can perform inference and / or training operations (e.g., at least a portion of inference and / or training logic 1115) for any of the noise reduction process 120, edge detection process 122, object detection process 124, optional thresholding process 126, decision logic 128, extraction process 130 and SGM process 132.

[0364] Figure 31This is a block diagram of a processor 3100 having one or more processor cores 3102A-3102N, an integrated memory controller 3114, and an integrated graphics processor 3108 according to at least one embodiment. In at least one embodiment, the processor 3100 may include additional cores, up to and including additional cores 3102N indicated by dashed boxes. In at least one embodiment, each processor core 3102A-3102N includes one or more internal cache units 3104A-3104N. In at least one embodiment, each processor core may also access one or more shared cache units 3106.

[0365] In at least one embodiment, internal cache units 3104A-3104N and shared cache unit 3106 represent a cache memory hierarchy within processor 3100. In at least one embodiment, cache memory units 3104A-3104N may include at least one level of instruction and data cache within each processor core and one or more levels of cache in a shared intermediate cache, such as Level 2 (L2), Level 3 (L3), Level 4 (L4), or other levels of cache, wherein the highest level of cache preceding external memory is classified as LLC. In at least one embodiment, cache coherence logic maintains coherence between the various cache units 3106 and 3104A-3104N.

[0366] In at least one embodiment, the processor 3100 may further include a set of one or more bus controller units 3116 and a system agent core 3110. In at least one embodiment, one or more bus controller units 3116 manage a set of peripheral buses, such as one or more PCI or PCIe buses. In at least one embodiment, the system agent core 3110 provides management functions for various processor components. In at least one embodiment, the system agent core 3110 includes one or more integrated memory controllers 3114 to manage access to various external memory devices (not shown).

[0367] In at least one embodiment, one or more processor cores 3102A-3102N include support for multi-threaded concurrent processing. In at least one embodiment, system agent core 3110 includes components for coordinating and operating cores 3102A-3102N during multi-threaded processing. In at least one embodiment, system agent core 3110 may additionally include a power control unit (PCU) including logic and components for regulating one or more power states of processor cores 3102A-3102N and graphics processor 3108.

[0368] In at least one embodiment, processor 3100 further includes a graphics processor 3108 for performing graph processing operations. In at least one embodiment, graphics processor 3108 is coupled to a shared cache unit 3106 and a system proxy core 3110 including one or more integrated memory controllers 3114. In at least one embodiment, system proxy core 3110 further includes a display controller 3111 for driving graphics processor outputs to one or more coupled displays. In at least one embodiment, display controller 3111 may also be a separate module coupled to graphics processor 3108 via at least one interconnect, or it may be integrated within graphics processor 3108.

[0369] In at least one embodiment, the ring-based interconnect unit 3112 is used to couple internal components of the processor 3100. In at least one embodiment, alternative interconnect units, such as point-to-point interconnects, switched interconnects, or other technologies, may be used. In at least one embodiment, the graphics processor 3108 is coupled to the ring interconnect 3112 via I / O link 3113.

[0370] In at least one embodiment, I / O link 3113 represents at least one of a variety of I / O interconnects, including packaged I / O interconnects that facilitate communication between various processor components and high-performance embedded memory modules 3118 (e.g., eDRAM modules). In at least one embodiment, each of the processor cores 3102A-3102N and the graphics processor 3108 uses the embedded memory module 3118 as a shared last-level cache.

[0371] In at least one embodiment, processor cores 3102A-3102N are homogeneous cores executing a common instruction set architecture. In at least one embodiment, processor cores 3102A-3102N are heterogeneous in terms of instruction set architecture (ISA), with one or more processor cores 3102A-3102N executing a common instruction set, while one or more other processor cores 3102A-3102N execute a subset of the common instruction set or a different instruction set. In at least one embodiment, processor cores 3102A-3102N are heterogeneous in terms of microarchitecture, with one or more cores having relatively high power consumption coupled to one or more power cores having lower power consumption. In at least one embodiment, processor 3100 may be implemented on one or more chips or implemented as a SoC integrated circuit.

[0372] Inference and / or training logic 1115 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 11A and / or Figure 11BDetails regarding the inference and / or training logic 1115 are provided. In at least one embodiment, some or all of the inference and / or training logic 1115 may be incorporated into the processor 3100. For example, in at least one embodiment, the training and / or inference techniques described herein may use one or more ALUs embodied in... Figure 31 The 3D pipeline, graphics core 3102, shared functional logic, or other logic are included. Furthermore, in at least one embodiment, the inference and / or training operations described herein can use, except... Figure 11A or Figure 11B The logic is performed using logic other than that shown. In at least one embodiment, the weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown), which configure the ALU of processor 3100 to execute one or more machine learning algorithms, neural network architectures, use cases, or training techniques described herein.

[0373] In at least one embodiment, system 100 (see...) Figure 1 The processor 114 can be implemented at least partially by the processor 3100. In at least one embodiment, the processor 114 can be implemented at least partially by the processor 3100. For example, instruction 118 (see...) Figure 1 This can be executed by processor 3100 and used to obtain optical flow map 110 (see...). Figure 1 and Figure 2 As described above, system 100 (see...) Figure 1 This can be implemented by at least a portion of the reasoning and / or training logic 1115. (See reference...) Figure 1 Processor 3100 (see Figure 31 It can perform inference and / or training operations (e.g., at least a portion of inference and / or training logic 1115) for any of the noise reduction process 120, edge detection process 122, object detection process 124, optional thresholding process 126, decision logic 128, extraction process 130 and SGM process 132.

[0374] Figure 32 This is a block diagram of a graphics processing unit 3200, which may be a discrete graphics processing unit or a graphics processing unit integrated with multiple processing cores. In at least one embodiment, the graphics processing unit 3200 communicates with registers on the graphics processing unit 3200 and commands placed in memory via a memory-mapped I / O interface. In at least one embodiment, the graphics processing unit 3200 includes a memory interface 3214 for accessing memory. In at least one embodiment, the memory interface 3214 is an interface to local memory, one or more internal caches, one or more shared external caches, and / or to system memory.

[0375] In at least one embodiment, the graphics processor 3200 further includes a display controller 3202 for driving display output data to the display device 3220. In at least one embodiment, the display controller 3202 includes a combination of hardware for one or more overlay planes of the display device 3220 and multi-layer video or user interface elements. In at least one embodiment, the display device 3220 may be an internal or external display device. In at least one embodiment, the display device 3220 is a head-mounted display device, such as a virtual reality (VR) display device or an augmented reality (AR) display device. In at least one embodiment, the graphics processor 3200 includes a video codec engine 3206 for encoding, decoding, or transcoding media into, from, or between one or more media encoding formats, including but not limited to Moving Picture Experts Group (MPEG) formats (e.g., MPEG-2), Advanced Video Coding (AVC) formats (e.g., H.264 / MPEG-4 AVC, and SMPTE 421M / VC-1), Joint Picture Experts Group (JPEG) formats (e.g., JPEG) and MotionJPEG (MJPEG) formats.

[0376] In at least one embodiment, the graphics processor 3200 includes a block image transfer (BLIT) engine 3204 to perform two-dimensional (2D) rasterizer operations, including, for example, bit boundary block transfer. However, in at least one embodiment, one or more components of a graphics processing engine (GPE) 3210 are used to perform 2D graphics operations. In at least one embodiment, the GPE 3210 is a computational engine for performing graphics operations, including three-dimensional (3D) graphics operations and media operations.

[0377] In at least one embodiment, GPE 3210 includes a 3D pipeline 3212 for performing 3D operations, such as rendering 3D images and scenes using processing functions that manipulate 3D primitive shapes (e.g., rectangles, triangles, etc.). In at least one embodiment, the 3D pipeline 3212 includes programmable and fixed function elements that perform various tasks and / or generate execution threads to the 3D / media subsystem 3215. While the 3D pipeline 3212 can be used to perform media operations, in at least one embodiment, GPE 3210 also includes a media pipeline 3216 for performing media operations such as video post-processing and image enhancement.

[0378] In at least one embodiment, the media pipeline 3216 includes fixed-function or programmable logic units for performing one or more specialized media operations, such as video decoding acceleration, video deinterlacing, and video encoding acceleration, replacing or representing the video codec engine 3206. In at least one embodiment, the media pipeline 3216 also includes a thread generation unit for generating threads to execute on the 3D / media subsystem 3215. In at least one embodiment, the generated threads perform computations of media operations on one or more graphics execution units included in the 3D / media subsystem 3215.

[0379] In at least one embodiment, the 3D / media subsystem 3215 includes logic for executing threads generated by the 3D pipeline 3212 and the media pipeline 3216. In at least one embodiment, the 3D pipeline 3212 and the media pipeline 3216 send thread execution requests to the 3D / media subsystem 3215, which includes thread dispatch logic for arbitrating various requests and dispatching them to available thread execution resources. In at least one embodiment, the execution resources include an array of graphics execution units for processing 3D and media threads. In at least one embodiment, the 3D / media subsystem 3215 includes one or more internal caches for thread instructions and data. In at least one embodiment, the subsystem 3215 also includes shared memory, which includes registers and addressable memory for sharing data between threads and storing output data.

[0380] Inference and / or training logic 1115 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 11A and / or Figure 11B Details regarding the inference and / or training logic 1115 are provided. In at least one embodiment, some or all of the inference and / or training logic 1115 may be incorporated into the processor 3200. For example, in at least one embodiment, the training and / or inference techniques described herein may use one or more ALUs included in the 3D pipeline 3212. Furthermore, in at least one embodiment, the inference and / or training operations described herein may use, except for... Figure 11A or Figure 11B The logic other than that shown is used to perform this task. In at least one embodiment, the weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown), which configure the ALU of the graphics processor 3200 to execute one or more machine learning algorithms, neural network architectures, use cases, or training techniques described herein.

[0381] In at least one embodiment, system 100 (see...) Figure 1The processor 114 can be implemented at least partially by the graphics processor 3200. In at least one embodiment, the processor 114 can be implemented at least partially by the graphics processor 3200. For example, instruction 118 (see...) Figure 1 This can be executed by the graphics processor 3200 and used to obtain the optical flow map 110 (see...). Figure 1 and Figure 2 As described above, system 100 (see...) Figure 1 This can be implemented by at least a portion of the reasoning and / or training logic 1115. (See reference...) Figure 1 Graphics processor 3200 (see...) Figure 32 It can perform inference and / or training operations (e.g., at least a portion of inference and / or training logic 1115) for any of the noise reduction process 120, edge detection process 122, object detection process 124, optional thresholding process 126, decision logic 128, extraction process 130 and SGM process 132.

[0382] Figure 33 This is a block diagram of a graphics processing engine 3310 of a graphics processor according to at least one embodiment. In at least one embodiment, the graphics processing engine (GPE) 3310 is... Figure 32 The version of GPE 3210 shown is illustrated. In at least one embodiment, the media pipeline 3316 is optional and may not be explicitly included in the GPE 3310. In at least one embodiment, a separate media and / or image processor is coupled to the GPE 3310.

[0383] In at least one embodiment, GPE 3310 is coupled to or includes command stream converter 3303, which provides command streams to 3D pipeline 3312 and / or media pipeline 3316. In at least one embodiment, command stream converter 3303 is coupled to memory, which may be system memory, or one or more of internal cache memory and shared cache memory. In at least one embodiment, command stream converter 3303 receives commands from memory and sends the commands to 3D pipeline 3312 and / or media pipeline 3316. In at least one embodiment, the commands are instructions, primitives, or micro-operations retrieved from a circular buffer that stores commands for 3D pipeline 3312 and media pipeline 3316. In at least one embodiment, the circular buffer may further include a batch command buffer storing multiple commands in batches. In at least one embodiment, commands for 3D pipeline 3312 may further include references to data stored in memory, such as, but not limited to, vertex and geometry data for 3D pipeline 3312 and / or image data and memory objects for media pipeline 3316. In at least one embodiment, the 3D pipeline 3312 and the media pipeline 3316 process commands and data by performing operations or by dispatching one or more execution threads to the graphics core array 3314. In at least one embodiment, the graphics core array 3314 includes one or more graphics core blocks (e.g., one or more graphics cores 3315A, one or more graphics cores 3315B), each block including one or more graphics cores. In at least one embodiment, each graphics core includes a set of graphics execution resources, which include general-purpose and graphics-specific execution logic for performing graphics and computational operations, and fixed-function texture processing and / or machine learning and artificial intelligence acceleration logic, including... Figure 11A and Figure 11B The reasoning and / or training logic in 1115.

[0384] In at least one embodiment, the 3D pipeline 3312 includes fixed functions and programmable logic for processing one or more shader programs, such as vertex shaders, geometry shaders, pixel shaders, fragment shaders, compute shaders, or other shader programs, by processing instructions and dispatching execution threads to the graphics core array 3314. In at least one embodiment, the graphics core array 3314 provides a unified execution resource block for processing shader programs. In at least one embodiment, the multipurpose execution logic (e.g., execution units) within the graphics cores 3315A-3315B of the graphics core array 3314 includes support for various 3D API shader languages ​​and can execute multiple concurrently running threads associated with multiple shaders.

[0385] In at least one embodiment, the graphics core array 3314 further includes execution logic for performing media functions, such as video and / or image processing. In at least one embodiment, in addition to graphics processing operations, the execution unit also includes general-purpose logic programmable to perform parallel general-purpose computing operations.

[0386] In at least one embodiment, output data can be output to memory in a unified return buffer (URB) 3318, the output data being generated by a thread executing on the graphics core array 3314. In at least one embodiment, the URB 3318 can store data from multiple threads. In at least one embodiment, the URB 3318 can be used to send data between different threads executing on the graphics core array 3314. In at least one embodiment, the URB 3318 can also be used for synchronization between threads on the graphics core array 3314 and fixed-function logic within shared-function logic 3320.

[0387] In at least one embodiment, the graphics core array 3314 is scalable, such that it includes a variable number of graphics cores, each having a variable number of execution units based on the target power and performance level of the GPE 3310. In at least one embodiment, the execution resources are dynamically scalable, such that they can be enabled or disabled as needed.

[0388] In at least one embodiment, the graphics core array 3314 is coupled to shared function logic 3320, which includes multiple resources shared among the graphics cores in the graphics core array 3314. In at least one embodiment, the shared functions performed by the shared function logic 3320 are embodied in hardware logic units that provide dedicated supplementary functions to the graphics core array 3314. In at least one embodiment, the shared function logic 3320 includes, but is not limited to, a sampler unit 3321, a math unit 3322, and inter-thread communication (ITC) logic 3323. In at least one embodiment, one or more caches 3325 are included in or coupled to the shared function logic 3320.

[0389] In at least one embodiment, shared functionality is used if the demand for dedicated functionality is insufficient to be contained within the graphics core array 3314. In at least one embodiment, a single instance of the dedicated functionality is used within shared functionality logic 3320 and shared among other execution resources within the graphics core array 3314. In at least one embodiment, a specific shared functionality may be included within shared functionality logic 3326 within the graphics core array 3314, said specific shared functionality being widely used within shared functionality logic 3320 of the graphics core array 3314. In at least one embodiment, shared functionality logic 3326 within the graphics core array 3314 may include some or all of the logic within shared functionality logic 3320. In at least one embodiment, all logic elements within shared functionality logic 3320 may be replicated within shared functionality logic 3326 of the graphics core array 3314. In at least one embodiment, shared functionality logic 3320 is excluded to support shared functionality logic 3326 within the graphics core array 3314.

[0390] Inference and / or training logic 1115 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 11A 11B and / or 11B provide details about the inference and / or training logic 1115. In at least one embodiment, some or all of the inference and / or training logic 1115 may be incorporated into the graphics processor 3310. For example, in at least one embodiment, the training and / or inference techniques described herein may use one or more ALUs embodied in the 3D pipeline 3312, graphics core 3315, shared function logic 3326, shared function logic 3320, or... Figure 33 In other logic within the [process]. Furthermore, in at least one embodiment, the inference and / or training operations described herein can use [other methods besides...]. Figure 11A or Figure 11B The logic is performed using logic other than that shown. In at least one embodiment, the weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown), which configure the ALU of the graphics processor 3310 to execute one or more of the machine learning algorithms, neural network architectures, use cases, or training techniques described herein.

[0391] In at least one embodiment, system 100 (see...) Figure 1 The processor 114 can be implemented at least partially by the GPE 3310 (e.g., incorporated into the graphics processor 3200). In at least one embodiment, the processor 114 can be implemented at least partially by the GPE 3310. For example, instruction 118 (see...) Figure 1 This can be performed by GPE 3310 and used to obtain optical flow map 110 (see...). Figure 1 and Figure 2 As described above, system 100 (see...) Figure 1 This can be implemented by at least a portion of the reasoning and / or training logic 1115. (See reference...) Figure 1 GPE3310 (see Figure 33 It can perform inference and / or training operations (e.g., at least a portion of inference and / or training logic 1115) for any of the noise reduction process 120, edge detection process 122, object detection process 124, optional thresholding process 126, decision logic 128, extraction process 130 and SGM process 132.

[0392] Figure 34 This is a block diagram of the hardware logic of a graphics processor core 3400 according to at least one embodiment described herein. In at least one embodiment, the graphics processor core 3400 is included within a graphics core array. In at least one embodiment, the graphics processor core 3400 (sometimes referred to as a core slice) may be one or more graphics cores within a modular graphics processor. In at least one embodiment, the graphics processor core 3400 is an example of a graphics core slice, and the graphics processor described herein may include multiple graphics core slices based on target power and performance envelopes. In at least one embodiment, each graphics core 3400 may include a fixed-function block 3430, also referred to as a sub-slice, coupled to a plurality of sub-cores 3401A-3401F, which includes modules of general-purpose and fixed-function logic.

[0393] In at least one embodiment, the fixed-function block 3430 includes a geometry and fixed-function pipeline 3436, which, for example, may be shared by all sub-cores of the graphics processor 3400 in a lower-performance and / or lower-power graphics processor implementation. In at least one embodiment, the geometry and fixed-function pipeline 3436 includes a 3D fixed-function pipeline, a video front-end unit, a thread generator and a thread dispatcher, and a unified return buffer manager that manages a unified return buffer.

[0394] In at least one fixed embodiment, the fixed functional block 3430 further includes a graphics SoC interface 3437, a graphics microcontroller 3438, and a media pipeline 3439. In at least one embodiment, the graphics SoC interface 3437 provides an interface between the graphics core 3400 and other processor cores in the on-chip integrated circuit system. In at least one embodiment, the graphics microcontroller 3438 is a programmable subprocessor configurable to manage various functions of the graphics processor 3400, including thread dispatch, scheduling, and preemption. In at least one embodiment, the media pipeline 3439 includes logic that facilitates decoding, encoding, preprocessing, and / or post-processing of multimedia data, including image and video data. In at least one embodiment, the media pipeline 3439 implements media operations via requests for computation or sampling logic within subcores 3401-3401F.

[0395] In at least one embodiment, the SoC interface 3437 enables the graphics core 3400 to communicate with a general-purpose application processor core (e.g., a CPU) and / or other components within the SoC, including memory hierarchy elements such as shared last-level cache, system RAM, and / or embedded on-chip or packaged DRAM. In at least one embodiment, the SoC interface 3437 also enables communication with fixed-function devices within the SoC (e.g., a camera imaging pipeline) and enables the use and / or implementation of global memory atoms that can be shared between the graphics core 3400 and the CPU within the SoC. In at least one embodiment, the graphics SoC interface 3437 also implements power management control for the graphics processor core 3400 and enables interfacing between the clock domain of the graphics processor core 3400 and other clock domains within the SoC. In at least one embodiment, the SoC interface 3437 enables the receipt of command buffers from a command stream converter and a global thread dispatcher, configured to provide commands and instructions to each of one or more graphics cores within the graphics processor. In at least one embodiment, when a media operation is to be performed, commands and instructions may be dispatched to the media pipeline 3439, or when a graphics processing operation is to be performed, they may be assigned to the geometry and fixed-function pipeline (e.g., geometry and fixed-function pipeline 3436, and / or geometry and fixed-function pipeline 3414).

[0396] In at least one embodiment, the graphics microcontroller 3438 can be configured to perform various scheduling and management tasks on the graphics core 3400. In at least one embodiment, the graphics microcontroller 3438 can perform graphics and / or compute workload scheduling on various graphics parallel engines within the execution unit (EU) arrays 3402A-3402F, 3404A-3404F in subcores 3401A-3401F. In at least one embodiment, host software executing on the CPU core of the SoC including the graphics core 3400 can submit a workload to one of multiple graphics processor paths, which invokes scheduling operations on the appropriate graphics engine. In at least one embodiment, the scheduling operation includes determining which workload should be run next, submitting the workload to a command stream converter, preempting existing workloads running on the engine, monitoring the progress of the workload, and notifying the host software when the workload is completed. In at least one embodiment, the graphics microcontroller 3438 may also facilitate a low-power or idle state of the graphics core 3400, thereby providing the graphics core 3400 with the ability to save and restore registers across low-power state transitions within the graphics core 3400, independent of the operating system and / or the graphics driver software on the system.

[0397] In at least one embodiment, the graphics core 3400 may have up to N more or fewer modular sub-cores than the illustrated sub-cores 3401A-3401F. For each group of N sub-cores, in at least one embodiment, the graphics core 3400 may further include shared functional logic 3410, shared and / or cache memory 3412, geometry / fixed-function pipeline 3414, and additional fixed-function logic 3416 to accelerate various graphics and computational processing operations. In at least one embodiment, the shared functional logic 3410 may include logic units (e.g., samplers, mathematical and / or inter-thread communication logic) that can be shared by each of the N sub-cores within the graphics core 3400. In at least one embodiment, the shared and / or cache memory 3412 may be the last-level cache of the N sub-cores 3401A-3401F within the graphics core 3400, and may also be used as shared memory accessible by multiple sub-cores. In at least one embodiment, a geometry / fixed function pipeline 3414 may be included to replace the geometry / fixed function pipeline 3436 within the fixed function block 3430, and similar logic units may be included.

[0398] In at least one embodiment, the graphics core 3400 includes additional fixed-function logic 3416, which may include various fixed-function acceleration logics for use by the graphics core 3400. In at least one embodiment, the additional fixed-function logic 3416 includes additional geometry pipelines for use in position-only shading. In position-only shading, there are at least two geometry pipelines, and in the full geometry pipeline and culling pipeline within the geometry and fixed-function pipelines 3414, 3436, it is an additional geometry pipeline that can be included in the additional fixed-function logic 3416. In at least one embodiment, the culling pipeline is a trimmed version of the full geometry pipeline. In at least one embodiment, the full pipeline and the culling pipeline can execute different instances of the application, each with a separate environment. In at least one embodiment, position-only shading can hide long culling runs of discarded triangles, thereby allowing shading to be completed earlier in some cases. For example, in at least one embodiment, the culling pipeline logic in the additional fixed-function logic 3416 can execute the position shader in parallel with the main application and typically generates critical results faster than the full pipeline because the culling pipeline acquires and occludes the positional attributes of vertices without performing rasterization and rendering pixels to the framebuffer. In at least one embodiment, the culling pipeline can use the generated critical results to compute visibility information for all triangles, regardless of whether those triangles were culled. In at least one embodiment, the full pipeline (which may be referred to as the replay pipeline in this case) can consume visibility information to skip culled triangles and only occlude the visible triangles that are ultimately passed to the rasterization stage.

[0399] In at least one embodiment, the additional fixed-function logic 3416 may also include machine learning acceleration logic, such as fixed-function matrix multiplication logic, for implementing optimizations for machine learning training or inference.

[0400] In at least one embodiment, each graphics subcore 3401A-3401F includes a set of execution resources that can be used to perform graphics, media, and computational operations in response to requests from the graphics pipeline, media pipeline, or shader program. In at least one embodiment, the graphics subcore 3401A-3401F includes multiple EU arrays 3402A-3402F, 3404A-3404F, thread dispatch and inter-thread communication (TD / IC) logic 3403A-3403F, 3D (e.g., texture) samplers 3405A-3405F, media samplers 3406A-3406F, shader processors 3407A-3407F, and shared local memory (SLM) 3408A-3408F. In at least one embodiment, each of the EU arrays 3402A-3402F and 3404A-3404F includes multiple execution units, which are general-purpose graphics processing units capable of servicing graphics, media, or computational operations, performing floating-point and integer / fixed-point logic operations, including graphics, media, or computational shader programs. In at least one embodiment, the TD / IC logic 3403A-3403F performs local thread dispatch and thread control operations for the execution units within the subcore and facilitates communication between threads executing on the execution units of the subcore. In at least one embodiment, the 3D samplers 3405A-3405F can read data associated with textures or other 3D graphics into memory. In at least one embodiment, the 3D samplers can read texture data differently based on the sampling state and texture format configured and associated with a given texture. In at least one embodiment, the media samplers 3406A-3406F can perform similar read operations based on the type and format associated with the media data. In at least one embodiment, each graphics subcore 3401A-3401F may alternatively include a unified 3D and media sampler. In at least one embodiment, threads executing on execution units within each subcore 3401A-3401F may utilize shared local memory 3408A-3408F within each subcore, enabling threads executing within a thread group to utilize a common pool of on-chip memory for execution.

[0401] Inference and / or training logic 1115 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 11A and / or Figure 11B Details regarding the inference and / or training logic 1115 are provided. In at least one embodiment, some or all of the inference and / or training logic 1115 may be incorporated into the graphics processor 3400. For example, in at least one embodiment, the training and / or inference techniques described herein may be used in the 3D pipeline, graphics microcontroller 3438, geometry and fixed-function pipelines 3414 and 3436, or... Figure 34One or more ALUs embodied in other logic within the [the document]. Furthermore, in at least one embodiment, the inference and / or training operations described herein can use [other methods besides...]. Figure 11A or Figure 11B The logic other than that shown is used to perform the task. In at least one embodiment, the weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown), which configure the ALU of the graphics processor 3400 to execute one or more of the machine learning algorithms, neural network architectures, use cases or training techniques described herein.

[0402] In at least one embodiment, system 100 (see...) Figure 1 The processor 114 may be implemented at least partially by the graphics processor core 3400. In at least one embodiment, the processor 114 may be implemented at least partially by the graphics processor core 3400. For example, instruction 118 (see...) Figure 1 This can be executed by the graphics processor core 3400 and used to obtain the optical flow map 110 (see...). Figure 1 and Figure 2 As described above, system 100 (see...) Figure 1 This can be implemented by at least a portion of the reasoning and / or training logic 1115. (See reference...) Figure 1 GPE3310 (see Figure 33 It can perform inference and / or training operations (e.g., at least a portion of inference and / or training logic 1115) for any of the noise reduction process 120, edge detection process 122, object detection process 124, optional thresholding process 126, decision logic 128, extraction process 130 and SGM process 132.

[0403] Figure 35A and Figure 35B The diagram illustrates thread execution logic 3500 of an array of processing elements including a graphics processor core, according to at least one embodiment. Figure 35A At least one embodiment is shown in which thread execution logic 3500 is used. Figure 35B Exemplary internal details of a graphics execution unit 3508 according to at least one embodiment are shown.

[0404] like Figure 35AAs shown, in at least one embodiment, thread execution logic 3500 includes a shader processor 3502, a thread dispatcher 3504, an instruction cache 3506, a scalable execution unit array including multiple execution units 3507A-3507N and 3508A-3508N, a sampler 3510, a data cache 3512, and a data port 3514. In at least one embodiment, the scalable execution unit array can be dynamically scaled, for example, based on the computational requirements of the workload, by enabling or disabling one or more execution units (e.g., any one of execution units 3508A-N or 3507A-N). In at least one embodiment, the scalable execution units are interconnected via an interconnect structure linking to each execution unit. In at least one embodiment, thread execution logic 3500 includes one or more connections to memory (such as system memory or cache memory) via one or more of the instruction cache 3506, data port 3514, sampler 3510, and execution units 3507 or 3508. In at least one embodiment, each execution unit (e.g., 3507A) is an independent programmable general-purpose computing unit capable of executing multiple concurrent hardware threads, processing multiple data elements in parallel for each thread. In at least one embodiment, the array of execution units 3507 and / or 3508 is scalable to include any number of individual execution units.

[0405] In at least one embodiment, execution units 3507 and / or 3508 are primarily used to execute shader programs. In at least one embodiment, shader processor 3502 can process various shader programs and dispatch execution threads associated with the shader programs via thread dispatcher 3504. In at least one embodiment, thread dispatcher 3504 includes logic for arbitrating thread initialization celebrations from the graphics and media pipeline and for instantiating requested threads on one or more execution units 3507 and / or 3508. For example, in at least one embodiment, the geometry pipeline can dispatch vertex, tessellation, or geometry shaders to thread execution logic for processing. In at least one embodiment, thread dispatcher 3504 can also handle runtime thread generation requests from executing shader programs.

[0406] In at least one embodiment, execution units 3507 and / or 3508 support an instruction set that includes native support for many standard 3D graphics shader instructions, enabling shader programs in graphics libraries (e.g., Direct3D and OpenGL) to execute with minimal conversion. In at least one embodiment, the execution unit supports vertex and geometry processing (e.g., vertex programs, geometry programs, and / or vertex shaders), pixel processing (e.g., pixel shaders, fragment shaders), and general processing (e.g., computation and media shaders). In at least one embodiment, each execution unit 3507 and / or 3508 includes one or more arithmetic logic units (ALUs) capable of performing multiple-issue single-instruction multiple-data (SIMD) operations, and multithreaded operation enables an efficient execution environment despite higher latency memory access. In at least one embodiment, each hardware thread within each execution unit has a dedicated high-bandwidth register file and associated independent thread states. In at least one embodiment, execution is multiple issues per clock cycle to a pipeline capable of integer, single-precision, and double-precision floating-point operations, SIMD branching functions, logical operations, a priori operations, and other operations. In at least one embodiment, while waiting for data from one of the memory or shared functions, dependency logic within execution units 3507 and / or 3508 causes the waiting thread to sleep until the requested data is returned. In at least one embodiment, while the waiting thread is sleeping, hardware resources can be dedicated to processing other threads. For example, in at least one embodiment, during the latency associated with vertex shader operations, the execution unit can perform operations on the pixel shader, fragment shader, or another type of shader program (including different vertex shaders).

[0407] In at least one embodiment, each of execution units 3507 and / or 3508 operates on an array of data elements. In at least one embodiment, the plurality of data elements is an "execution size" or the number of instruction channels. In at least one embodiment, an execution channel is a logical unit for execution of data element access, masking, and flow control within an instruction. In at least one embodiment, the plurality of channels may be independent of the plurality of physical arithmetic logic units (ALUs) or floating-point units (FPUs) for a particular graphics processor. In at least one embodiment, execution units 3507 and / or 3508 support integer and floating-point data types.

[0408] In at least one embodiment, the execution unit instruction set includes SIMD instructions. In at least one embodiment, various data elements can be stored in registers as encapsulated data types, and the execution unit will process various elements based on the data size of those elements. For example, in at least one embodiment, when operating on a 256-bit wide vector, 256 bits of the vector are stored in registers, and the execution unit operates on the vector as four separate 64-bit encapsulated data elements (quad-word (QW) size data elements), eight separate 32-bit encapsulated data elements (double-word (DW) size data elements), sixteen separate 16-bit encapsulated data elements (word (W) size data elements), or thirty-two separate 8-bit data elements (byte (B) size data elements). However, in at least one embodiment, different vector widths and register sizes are possible.

[0409] In at least one embodiment, one or more execution units may be combined into a fused execution unit 3509A-3509N having thread control logic (3511A-3511N) for executing fused EUs, for example, fused execution unit 3507A and execution unit 3508A into a fused execution unit 3509A. In at least one embodiment, multiple EUs may be merged into a single EU group.

[0410] In at least one embodiment, the number of EUs in a fused EU group can be configured to execute separate SIMD hardware threads, and the number of EUs in the fused EU group may vary depending on the embodiment. In at least one embodiment, each EU can execute various SIMD widths, including but not limited to SIMD8, SIMD16, and SIMD32. In at least one embodiment, each fused graphics execution unit 3509A-3509N includes at least two execution units. For example, in at least one embodiment, the fused execution unit 3509A includes a first EU 3507A, a second EU 3508A, and thread control logic 3511A shared by the first EU 3507A and the second EU 3508A. In at least one embodiment, the thread control logic 3511A controls the threads executing on the fused graphics execution unit 3509A, thereby allowing each EU within the fused execution units 3509A-3509N to execute using a common instruction pointer register.

[0411] In at least one embodiment, one or more internal instruction caches (e.g., 3506) are included in the thread execution logic 3500 to cache thread instructions for the execution unit. In at least one embodiment, one or more data caches (e.g., 3512) are included to cache thread data during thread execution. In at least one embodiment, a sampler 3510 is included to provide texture sampling for 3D operations and media sampling for media operations. In at least one embodiment, the sampler 3510 includes dedicated texture or media sampling functions to process texture or media data during the sampling process before providing sampled data to the execution unit.

[0412] During execution, in at least one embodiment, the graphics and media pipeline sends thread initiation requests to thread execution logic 3500 via thread creation and dispatch logic. In at least one embodiment, once a set of geometric objects has been processed and rasterized into pixel data, pixel processor logic (e.g., pixel shader logic, fragment shader logic, etc.) within shader processor 3502 is invoked to further compute output information and cause the results to be written to output surfaces (e.g., color buffer, depth buffer, stencil buffer, etc.). In at least one embodiment, the pixel shader or fragment shader computes values ​​of various vertex attributes to be interpolated on the rasterized objects. In at least one embodiment, the pixel processor logic within shader processor 3502 then executes a pixel or fragment shader program provided by an application programming interface (API). In at least one embodiment, to execute the shader program, shader processor 3502 dispatches threads to execution units (e.g., 3508A) via thread dispatcher 3504. In at least one embodiment, shader processor 3502 uses texture sampling logic in sampler 3510 to access texture data in a texture map stored in memory. In at least one embodiment, arithmetic operations on the texture data and the input geometry data are performed to calculate pixel color data for each geometric segment, or one or more pixels are discarded for further processing.

[0413] In at least one embodiment, data port 3514 provides a memory access mechanism for thread execution logic 3500 to output processed data to memory for further processing on the graphics processor output pipeline. In at least one embodiment, data port 3514 includes or is coupled to one or more cache memories (e.g., data cache 3512) to cache data for memory access via the data port.

[0414] like Figure 35BAs shown, in at least one embodiment, the graphics execution unit 3508 may include an instruction fetch unit 3537, a general-purpose register file array (GRF) 3524, an architecture register file array (ARF) 3526, a thread arbiter 3522, a send unit 3530, a branch unit 3532, a set of SIMD floating-point units (FPUs) 3534, and a set of dedicated integer SIMD ALUs 3535. In at least one embodiment, the GRF 3524 and ARF 3526 include a set of general-purpose register files and architecture register files associated with each concurrent hardware thread that may be active in the graphics execution unit 3508. In at least one embodiment, the architecture state of each thread is maintained in the ARF 3526, while data used during thread execution is stored in the GRF 3524. In at least one embodiment, the execution state of each thread, including the instruction pointer of each thread, may be stored in thread-specific registers in the ARF 3526.

[0415] In at least one embodiment, the graphics execution unit 3508 has an architecture that is a combination of simultaneous multithreading (SMT) and fine-grained interleaved multithreading (IMT). In at least one embodiment, the architecture has a modular configuration that can be fine-tuned at design time based on a target number of simultaneous threads and the number of registers per execution unit, wherein execution unit resources are logically allocated for executing multiple simultaneous threads.

[0416] In at least one embodiment, the graphics execution unit 3508 can jointly issue multiple instructions, each of which can be a different instruction. In at least one embodiment, the thread arbiter 3522 of the graphics execution unit thread 3508 can dispatch instructions to one of the sending unit 3530, the branching unit 3532, or the SIMD FPU 3534 for execution. In at least one embodiment, each execution thread can access 128 general-purpose registers in the GRF 3524, where each register can store 32 bytes and can be accessed as a SIMD 8-element vector of 32-bit data elements. In at least one embodiment, each execution unit thread can access 4KB of the GRF 3524, although the embodiments are not limited thereto, and more or fewer register resources may be provided in other embodiments. In at least one embodiment, although the number of threads per execution unit may also vary depending on the embodiment, a maximum of seven threads can be executed simultaneously. In at least one embodiment where seven threads can access 4KB, the GRF 3524 can store a total of 28KB. In at least one embodiment, the flexible addressing mode can allow registers to be addressed together to efficiently build wider registers or rectangular block data structures representing strides.

[0417] In at least one embodiment, memory operations, sampler operations, and other longer-latency system communications are scheduled via "send" instructions executed by message sending unit 3530. In at least one embodiment, branch instructions are dispatched to branch unit 3532 to facilitate SIMD divergence and eventual convergence.

[0418] In at least one embodiment, the graphics execution unit 3508 includes one or more SIMD floating-point units (FPUs) 3534 to perform floating-point operations. In at least one embodiment, the one or more FPUs 3534 also support integer computation. In at least one embodiment, the one or more FPUs 3534 can perform up to M 32-bit floating-point (or integer) operations in SIMD, or up to 2M 16-bit integer or 16-bit floating-point operations in SIMD. In at least one embodiment, at least one FPU provides extended mathematical capabilities to support high-throughput a priori mathematical functions and double-precision 64-bit floating-point operations. In at least one embodiment, a set of 8-bit integer SIMD ALUs 3535 is also present and can be specifically optimized to perform operations related to machine learning computations.

[0419] In at least one embodiment, an array of multiple instances of the graphics execution unit 3508 may be instantiated in a graphics sub-core group (e.g., a sub-slice). In at least one embodiment, the execution unit 3508 may execute instructions across multiple execution channels. In at least one embodiment, each thread executing on the graphics execution unit 3508 executes on a different channel.

[0420] The inference and / or training logic 1115 is used to perform inference and / or training operations associated with one or more embodiments. The following is in conjunction with... Figure 11A and / or Figure 11B Details are provided regarding the inference and / or training logic 1115. In at least one embodiment, some or all of the inference and / or training logic 1115 may be incorporated into the thread execution logic 3500. Furthermore, in at least one embodiment, additional... Figure 11A or Figure 11B The logic other than that shown is used to perform the inference and / or training operations described herein. In at least one embodiment, weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown), which configure the ALU of thread execution logic 3500 to execute one or more machine learning algorithms, neural network architectures, use cases, or training techniques described herein.

[0421] In at least one embodiment, system 100 (see...) Figure 1This can be implemented at least partially using thread execution logic 3500. In at least one embodiment, processor 114 can be implemented at least partially by graphics execution unit 3508. For example, instruction 118 (see...) Figure 1 This can be executed by the graphics execution unit 3508 and used to obtain the optical flow map 110 (see...). Figure 1 and Figure 2 As described above, system 100 (see...) Figure 1 This can be implemented by at least a portion of the reasoning and / or training logic 1115. (See reference...) Figure 1 Thread execution logic 3500 and / or graphics execution unit 3508 (see...) Figure 35A and Figure 35B It can perform inference and / or training operations (e.g., at least a portion of inference and / or training logic 1115) for any of the noise reduction process 120, edge detection process 122, object detection process 124, optional thresholding process 126, decision logic 128, extraction process 130 and SGM process 132.

[0422] Figure 36 A parallel processing unit ("PPU") 3600 according to at least one embodiment is illustrated. In at least one embodiment, the PPU 3600 is configured with machine-readable code that, ...

Claims

1. A method for generating optical flow maps, comprising: Obtain an input image and a reference image, wherein the input image includes multiple image regions; Obtaining at least one penalty map, wherein the at least one penalty map includes a set of penalty values ​​for each of the plurality of image regions, and for each of the plurality of image regions, the set of penalty values ​​includes penalty values ​​for each of a plurality of directions intersecting with the image regions, wherein obtaining the at least one penalty map includes at least: The input image is denoised. An edge map is generated by performing edge detection on the denoised input image, and An object map is generated by performing object detection on at least one of the denoised input image or the edge map; and An optical flow map is generated for the input image and the reference image, at least in part, based on the at least one penalty map.

2. The method of claim 1, wherein the at least one penalty graph includes a first penalty graph and a second penalty graph. The optical flow map is generated using a semi-global matching (SGM) process that includes a first penalty value and a second penalty value, to determine the cost of a specific image region, a specific disparity, and a specific direction among the plurality of image regions. The first penalty value is obtained using the specific image region and the first penalty map in the specific direction, and The second penalty value is obtained using the specific image region and the second penalty map in the specific direction.

3. The method of claim 2, wherein the optical flow map is generated in the following manner: Determine multiple costs, and determine at least one of the multiple costs for one of the multiple image regions, one of the multiple parallaxes, and one of the multiple directions; Multiple cumulative costs for each unique pair are obtained by summing the costs determined for a unique pair of one of the multiple image regions and one of the multiple disparities. For each of the plurality of image regions, select the minimum cumulative cost among the plurality of cumulative costs obtained for the image region, the minimum cumulative cost having been obtained for a selected disparity among the plurality of disparities; as well as For each of the plurality of image regions, a value is placed in the optical flow map at a position corresponding to the image region, at least in part based on the selected parallax.

4. The method of claim 1, wherein the at least one penalty graph includes a first penalty graph and a second penalty graph. The first penalty graph is determined at least in part based on the object graph, and The second penalty map is determined at least in part based on the edge map.

5. The method of claim 1, wherein obtaining the at least one penalty graph further comprises: A thresholded edge map is generated by performing a thresholding process on the edge map to remove any edges with a width less than a threshold. The at least one penalty graph includes a first penalty graph and a second penalty graph. The first penalty graph is determined at least in part based on the object graph, and The second penalty map is determined at least in part based on both the edge map and the thresholded edge map.

6. The method of claim 1, wherein the plurality of image regions comprises at least one of a set of feature points or a set of pixels.

7. The method of claim 1, wherein the input image comprises a plurality of pixels, and The plurality of image regions include a subset of the plurality of pixels, the subset comprising fewer than all of the plurality of pixels.

8. A system for generating optical flow maps, comprising: One or more circuits are configured to: obtain a plurality of image regions from an input image; determine a set of penalty values ​​for each of the plurality of image regions; and create an optical flow map for the input image and a reference image, at least in part based on the set of penalty values ​​determined for each of the plurality of image regions, wherein, for each of the plurality of image regions, the set of penalty values ​​includes a penalty value for each of a plurality of directions intersecting the image regions. The one or more circuits thereon are configured to determine the set of penalty values ​​for each of the plurality of image regions at least in the following manner: Denoise the input image; An edge map is generated by performing edge detection on the denoised input image; and An object map is generated by performing object detection on at least one of the denoised input image or the edge map.

9. The system of claim 8, further comprising: Upstream hardware connected to the one or more circuits, the upstream hardware providing the input image to the one or more circuits.

10. The system of claim 8, wherein the one or more circuits comprise: At least one processor; as well as The memory stores instructions that can be executed by the at least one processor, and when executed by the at least one processor, the instructions cause the at least one processor to: obtain the plurality of image regions, determine the set of penalty values ​​for at least one of the plurality of image regions, and create the optical flow map.

11. The system of claim 10, wherein when the instructions are executed by the at least one processor, the instructions cause the at least one processor to generate the optical flow map in such a way as: Multiple costs are determined for the multiple image regions, the multiple disparities between the input image and the reference image, and the unique combination of the multiple directions; Multiple cumulative costs are obtained by summing those costs determined for each unique pair of the multiple costs for one of the multiple image regions and at least one of the multiple disparities; Select the minimum cumulative cost from a plurality of cumulative costs obtained for one or more of the plurality of image regions, wherein the minimum cumulative cost has been obtained for a selected disparity from the plurality of disparities; as well as For each of the portions of the plurality of image regions, a value is placed in the optical flow map at a position corresponding to the image region, at least in part based on the selected parallax.

12. The system of claim 11, wherein the one or more circuits are further configured to determine a specific cost for a specific combination of the plurality of costs for a specific image region including the plurality of image regions, a specific parallax of the plurality of parallaxes, and a specific direction of the plurality of directions, by means of: When the first disparity metric value of the specific image region differs by at most a predetermined amount from the second disparity metric value of the adjacent image regions in the plurality of image regions along the specific direction, the first penalty value in the set of penalty values ​​is added to the matching item; and When the first disparity metric value differs from the second disparity metric value along the specific direction by more than the predetermined amount, a second penalty value from the set of penalty values ​​is added to the matching item, wherein the first penalty value is less than the second penalty value.

13. The system of claim 12, wherein one or more circuits are further configured to determine that the particular cost equals the match when the first disparity metric is equal to the second disparity metric.

14. The system of claim 8, wherein the set of penalty values ​​includes a first set of penalty values ​​and different second sets of penalty values. Each of the first set of penalty values ​​is determined at least in part based on the object graph, and Each of the second set of penalty values ​​is determined at least in part based on the edge map.

15. The system of claim 8, wherein the one or more circuits are configured to further determine a set of penalty values ​​for each of the plurality of image regions by: thresholding the edge map to produce a thresholded edge map, wherein edges thinner than a threshold have been removed, and wherein, The set of penalty values ​​includes a first set of penalty values ​​and different second sets of penalty values. Each of the first set of penalty values ​​is determined at least in part based on the object graph, and Each of the second set of penalty values ​​is determined at least in part based on both the edge map and the thresholded edge map.

16. The system of claim 10, wherein the at least one processor comprises one or more parallel processing units.

17. The system of claim 10, wherein when the instructions are executed by the at least one processor, the instructions cause the at least one processor to extract a set of feature points from the input image as the plurality of image regions.

18. The system of claim 10, wherein when the instructions are executed by the at least one processor, the instructions cause the at least one processor to extract a set of pixels from the input image as the plurality of image regions.

19. The system of claim 18, wherein the input image comprises a plurality of pixels, and the set of pixels comprises fewer than all of the plurality of pixels.

20. One or more non-transitory processor-readable media, comprising instructions executable by at least one processor, wherein the instructions, when executed by the at least one processor, cause the at least one processor to: A set of disparity values ​​is obtained for each of a plurality of image regions in the input image, wherein for each of the plurality of image regions, the set of disparity values ​​includes disparity values ​​for each of a plurality of directions intersecting with the image region; and An optical flow map is created using the set of disparity values ​​for the input image and different reference images; The at least one processor obtains the set of disparity values ​​for each of the plurality of image regions at least in the following manner: Denoise the input image; Perform edge detection on the denoised input image to generate an edge map; and An object map is generated by performing object detection on at least one of the denoised input image or the edge map.

21. One or more non-transitory processor-readable media as claimed in claim 20, wherein when the instructions are executed by the at least one processor, the instructions cause the at least one processor to generate the optical flow map in such a way as: Multiple costs are determined for the multiple image regions, the multiple disparities between the input image and the reference image, and the unique combination of the multiple directions; Multiple cumulative costs are obtained by summing those costs determined for each unique pair of the multiple costs for one of the multiple image regions and one of the multiple disparities; For each of at least a portion of the plurality of image regions, select the minimum cumulative cost among the plurality of cumulative costs obtained for the image region, the minimum cumulative cost having been obtained for a selected disparity among the plurality of disparities; as well as For each of the portions of the plurality of image regions, a value is stored in the optical flow map at a location corresponding to the image region, at least in part based on the selected parallax.

22. One or more non-transitory processor-readable media as claimed in claim 21, wherein when the instructions are executed by the at least one processor, the instructions cause the at least one processor to store the set of disparity values ​​as a first disparity map and a second disparity map, and to determine a specific cost among the plurality of costs for a specific combination of a specific image region including the plurality of image regions, a specific disparity among the plurality of disparities, and a specific direction among the plurality of directions, by: The first value and the second value are obtained from the first disparity map and the second disparity map, respectively; When the first disparity metric value of the specific image region differs by at most a predetermined amount from the second disparity metric value of the adjacent image regions among the plurality of image regions along the specific direction, the first value is added to the matching item; and When the first disparity metric value differs from the second disparity metric value by more than a predetermined amount along the specific direction, the second value is added to the matching item, wherein the first disparity value is less than the second disparity value.

23. One or more non-transitory processor-readable media as claimed in claim 22, wherein the at least one processor is further configured to determine that the particular cost equals the match when the first disparity metric is equal to the second disparity metric.

24. One or more non-transitory processor-readable media as claimed in claim 20, wherein when the instructions are executed by the at least one processor, the instructions cause the at least one processor to store the set of disparity values ​​as a first disparity map and a second disparity map, and the at least one processor further obtains the set of disparity values ​​for each of the plurality of image regions in such a manner as follows: Assigning values ​​to the first disparity map, at least in part, based on the object map; and Values ​​are assigned to the second disparity map, at least in part, based on the edge map.

25. One or more non-transitory processor-readable media as claimed in claim 20, wherein when the instructions are executed by the at least one processor, the instructions cause the at least one processor to store the set of disparity values ​​as a first disparity map and a second disparity map, and the at least one processor further obtains the set of disparity values ​​for each of a plurality of image regions by: The edge map is thresholded to remove any edges that are not thicker than a threshold to produce a thresholded edge map; Values ​​are assigned to the first disparity map, at least in part, based on the object map, and Values ​​are assigned to the second disparity map based at least in part on the edge map and the thresholded edge map.

26. One or more non-transitory processor-readable media as claimed in claim 20, wherein when the instructions are executed by the at least one processor, the instructions cause the at least one processor to extract a set of feature points or a set of pixels from the input image as the plurality of image regions.

Citation Information

Patent Citations

  • KR1018041570000B1