Convolutional acceleration using activated static vectorization

By utilizing static data locations and shifted multiplier-accumulator units, the inefficiencies and power consumption associated with convolution operations in machine learning networks are mitigated, enhancing computational efficiency.

JP2026525387APending Publication Date: 2026-07-30QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
QUALCOMM INC
Filing Date
2024-05-02
Publication Date
2026-07-30

AI Technical Summary

Technical Problem

Convolution operations in machine learning networks are computationally and power-intensive due to the high number of data reads and transfers required, which can be inefficient and increase power consumption.

Method used

Implementing static data locations for input data that persist across multiple convolution cycles, reducing data toggling and using shifted multiplier and accumulator units to perform MAC operations efficiently.

Benefits of technology

Reduces the number of data reads and transfers, thereby decreasing power consumption and computational workload during convolution operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026525387000001_ABST
    Figure 2026525387000001_ABST
Patent Text Reader

Abstract

A system and technology for processing image data are provided. At each of multiple positions of a convolution kernel along a row of image data, a first value enclosed by the convolution kernel can be obtained and stored using the respective memory location associated with each position in the multiple positions. Based on each of these first values, the cumulative value corresponding to the convolution output for each position in the multiple positions can be updated. At each of the multiple positions, multiple second values ​​enclosed by the convolution kernel can be obtained. The multiple second values ​​include a subset of each first value and an additional second value. The memory location used to store first values ​​not included in the multiple second values ​​can be updated to store the additional second value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to image processing using machine learning networks. For example, aspects of the present disclosure relate to systems and techniques for performing image processing using one or more machine learning systems that implement convolutional operations.

Background Art

[0002] Many devices and systems are capable of capturing a scene by generating an image (or frame) of the scene and / or video data (including multiple frames). For example, a camera, or a device including a camera, can capture a sequence of frames of a scene (e.g., a video of the scene). In some cases, the sequence of frames can be processed to perform one or more functions, output for display, or output for processing and / or consumption by other devices, among many applications.

[0003] Artificial neural networks attempt to reproduce the logical inferences performed by the biological neural networks that make up the brains of animals using computer technology. Deep neural networks, such as convolutional neural networks, are widely used in a number of applications, especially object detection, object classification, object tracking, big data analysis, and the like. For example, a convolutional neural network can extract high-level features, such as the shape of a face, from an input image and use these high-level features to output, for example, the probability that the input image contains a particular object.

Summary of the Invention

[0004] The following provides a simplified overview of one or more embodiments disclosed herein. Therefore, this overview should not be considered a broad overview of all intended embodiments, nor should it be considered to identify the main or important elements of all intended embodiments, or to define the scope associated with any particular embodiment. Accordingly, the sole purpose of this overview is to provide, in a simplified form, certain concepts relating to one or more embodiments of the mechanisms disclosed herein, prior to the detailed descriptions presented below.

[0005] Systems, methods, apparatus, and computer-readable media for image processing are disclosed. According to at least one exemplary embodiment, a method for processing image data is provided. The method includes: obtaining each first value enclosed by the convolution kernel at each of a plurality of positions of the convolution kernel along a row of image data; storing each of the first values ​​using the respective memory location associated with each of the plurality of positions of the convolution kernel; updating a cumulative value corresponding to the convolution output for each of the plurality of positions of the convolution kernel based on each of the first values ​​stored using the respective memory location; obtaining a plurality of second values ​​enclosed by the convolution kernel at each of the plurality of positions, which include a subset of each first value and additional second values ​​not included in each of the first values; and updating a memory location used to store the first values ​​not included in the plurality of second values ​​with the additional second values ​​so that the memory location stores the additional second values.

[0006] In another exemplary embodiment, a device for processing image data is provided. The device includes at least one memory and at least one processor coupled to the at least one memory, the at least one processor being configured to take each first value enclosed by the convolution kernel at each of a plurality of positions of the convolution kernel along a row of image data, store each of the first values ​​using the respective memory location associated with each of the plurality of positions of the convolution kernel, update the cumulative value corresponding to the convolution output for each of the plurality of positions of the convolution kernel based on each of the first values ​​stored using the respective memory location, take a plurality of second values ​​enclosed by the convolution kernel at each of the plurality of positions, including a subset of each first value and additional second values ​​not included in each of the first values, and update the memory location used to store the first values ​​not included in the plurality of second values ​​with the additional second values ​​so that the memory location stores the additional second values.

[0007] In another exemplary embodiment, a non-temporary computer-readable storage medium contains a stored instruction, which, when executed by at least one processor, causes at least one processor to retrieve each first value enclosed by the convolution kernel at each of a plurality of positions of the convolution kernel along a row of image data; to store each of the first values ​​using the respective memory location associated with each of the plurality of positions of the convolution kernel; to update a cumulative value corresponding to the convolution output for each of the plurality of positions of the convolution kernel based on each of the first values ​​stored using the respective memory location; to retrieve a plurality of second values ​​enclosed by the convolution kernel at each of the plurality of positions, which include a subset of each first value and additional second values ​​not included in each of the first values; and to update a memory location used to store the first values ​​not included in the plurality of second values ​​with the additional second values ​​so that the memory location stores the additional second values.

[0008] In another exemplary embodiment, a device for processing image data is provided. The device includes means for obtaining each first value enclosed by the convolution kernel at each of a plurality of positions of the convolution kernel along a row of image data; means for storing each of the first values ​​using each memory location associated with each of the plurality of positions of the convolution kernel; means for updating cumulative values ​​corresponding to the convolution output for each of the plurality of positions of the convolution kernel based on each of the first values ​​stored using each memory location; means for obtaining a plurality of second values ​​enclosed by the convolution kernel at each of the plurality of positions, including a subset of each first value and additional second values ​​not included in each first value; and means for updating a memory location used to store the first values ​​not included in the plurality of second values ​​with the additional second values ​​so that the memory location stores the additional second values.

[0009] The embodiments are generally described in detail with reference to the drawings and this specification and include methods, apparatus, systems, computer program products, non-temporary computer-readable media, user devices, user equipment, wireless communication devices, and / or processing systems as shown in the drawings and this specification.

[0010] Some embodiments include a device having a processor configured to perform one or more of the operations summarized above. Further embodiments include a processing device used in the device, comprising processor-executable instructions for performing any of the operations summarized above. Further embodiments include a non-temporary processor-readable storage medium storing processor-executable instructions configured to cause the device's processor to perform any of the operations summarized above. Further embodiments include a device having means for performing any of the functions summarized above.

[0011] The above provides a fairly broad overview of the features and technical advantages of the embodiments of this disclosure so that the following “Modes for Carrying Out the Invention” may be better understood. Additional features and advantages are described below. The concepts and specific embodiments disclosed may readily be used as a basis for modifying or designing other structures to accomplish the same objectives of this disclosure. Such equivalent structures shall not deviate from the scope of the appended claims. The characteristics of the concepts disclosed herein will be better understood from the following description by examining both their configuration and method of operation in relation to the accompanying drawings, along with their relevant advantages. Each of the drawings is provided for illustrative and explanatory purposes and is not provided to define any limitation of the claims. The above, along with other features and embodiments, will become clearer by referring to the following specification, claims, and accompanying drawings.

[0012] This summary is not intended to identify the main or essential features of the claimed subject matter, nor is it intended to be used independently to determine the scope of the claimed subject matter. The subject matter should be understood by referring to the entire specification of this patent, any or all of the drawings, and the appropriate portions of each claim.

[0013] The accompanying drawings are provided to aid in describing various aspects of the disclosure and are provided solely for illustrative purposes, not to limit those aspects. A more detailed description of the features of the disclosure listed above can be obtained by referring to the aspects partially shown in the accompanying drawings, so as to a more detailed understanding of those aspects. However, it should be noted that the accompanying drawings only illustrate certain typical aspects of the disclosure, and therefore the description may be incorporated into other equally effective aspects and should not be considered to limit the scope of the disclosure. The same reference numerals in different drawings may identify the same or similar elements. [Brief explanation of the drawing]

[0014] [Figure 1] Several examples illustrate exemplary implementations of a system-on-a-chip (SoC). [Figure 2A] This document presents one example of a fully connected neural network, based on several implementations. [Figure 2B] This document presents one example of a locally connected neural network, based on several implementations. [Figure 2C] This document presents one example of a convolutional neural network, using several implementations. [Figure 3] This document presents one example of a 3x1 convolution, based on several embodiments. [Figure 4] This document presents one example of a 3x1 convolution using static data locations and static output accumulator locations, based on several embodiments. [Figure 5]This example demonstrates one instance of 3x1 convolution using a shifted output accumulator location, based on several embodiments. [Figure 6] This example demonstrates one instance of 3x3 convolution using a shifted output accumulator location, based on several embodiments. [Figure 7] This flowchart illustrates one example of a process for processing image and / or video data, based on several embodiments. [Figure 8] This is a block diagram showing one example of a deep learning network, based on several implementations. [Figure 9] This is a block diagram showing one example of a convolutional neural network, based on several implementations. [Figure 10] This figure shows an exemplary system architecture that implements some of the embodiments described herein. [Modes for carrying out the invention]

[0015] Specific embodiments of this disclosure are provided below for illustrative purposes only. Alternative embodiments can be devised without departing from the scope of this disclosure. In addition, well-known elements of this disclosure are not described in detail or are omitted so as not to obscure the relevant details of this disclosure. As will be apparent to those skilled in the art, some of the embodiments described herein can be applied independently, and some of them can be applied in combination. For illustrative purposes, specific details are provided in the following specification to provide a complete understanding of the embodiments of this application. However, it will be apparent that various embodiments can be practiced without these specific details. The figures and specification are not intended to be restrictive.

[0016] The following description provides exemplary embodiments and is not intended to limit the scope, applicability, or configuration of the present disclosure. Rather, the following description of the exemplary embodiments provides those skilled in the art with a feasible description for implementing the exemplary embodiments. It should be understood that various changes can be made to the functions and configurations of the elements without departing from the scope of the present application as described in the appended claims.

[0017] As described above, a machine learning system (e.g., a neural network system or model) can be used to perform various tasks, such as, but not limited to, detection and / or recognition (e.g., scene or object detection and / or recognition, face detection and / or recognition, etc.), depth estimation, pose estimation, image reconstruction, classification, three-dimensional (3D) modeling, high-density regression tasks, data compression and / or restoration, and image processing, among other tasks. Further, the machine learning model can be versatile and can achieve high-quality results in various tasks.

[0018] In some cases, the machine learning system can be implemented based on performing a plurality of convolution operations (e.g., using one or more convolutional layers). Convolution is an operation on two functions that creates a third function representing how one shape is modified by the other. Convolution and / or convolutional layers can be used to process pixel data and can be implemented in various image processing and / or video processing machine learning tasks. As used herein, "image" or "image data" may refer to a frame of pixel data having a horizontal resolution (e.g., number of horizontal pixels) and a vertical resolution (e.g., number of vertical pixels). The frame of pixel data can be associated with a still image or photograph and / or can be associated with a frame of video data.

[0019] A machine learning network can perform a convolution operation using a convolutional kernel or filter that slides along input features to generate a plurality of feature maps. The convolution operation can be implemented based on the hierarchical nature of the data being processed. For example, instead of processing an entire image or input at once, a convolutional neural network (CNN) decomposes the image into smaller and simpler features, which are represented by filters applied to different regions of the image to extract relevant information (e.g., corresponding feature maps for different image regions). As the CNN progresses through its convolutional layers, the feature maps are combined and assembled into increasingly complex structures that are used to learn increasingly abstract representations of the input.

[0020] In an exemplary CNN implementation, the input to the CNN can be a tensor having a shape given by (number of inputs) × (input height) × (input width) × (input channels). For example, after passing through a convolutional layer, the input image can be abstracted into a feature map (also referred to as an activation map, for example) having a shape given by (number of inputs) × (height of feature map) × (width of feature map) × (channels of feature map). The dimensions of the feature map can be made smaller than those of the input image. The convolutional layer convolves the input and passes the output result to the next layer. Each convolutional neuron processes data only for its receptive field. In some aspects, convolution can be performed to reduce the number of free parameters within the machine learning network (e.g., by reducing the number of free parameters, the machine learning network can become deeper).

[0021] A convolutional kernel can be implemented as a matrix of weights that slides across the input data provided to the CNN. For example, a 3x1 convolutional kernel may be a 3x1 matrix of weights that slides across the input data. For instance, a 3x1 convolutional kernel can be used to process an image based on a slide across rows of pixels contained in the image. At each step, the 3x1 convolutional kernel can process pixels in three columns within one row (e.g., 3x1) and perform element-wise multiplication with each pixel currently contained in the 3x1 window of the convolutional kernel. The results of the element-wise multiplications are summed to a single output value for the current sliding window of the convolution. The convolutional kernel can further repeat the above process across any location where the convolutional kernel slides, transforming the first-size 2D matrix into a second, smaller-size 2D matrix.

[0022] As mentioned above, convolution operations can be used to implement a variety of machine learning operations, including 2D convolution, 3D convolution, depth convolution, group convolution, and trans-convolution (e.g., transposed convolution). Convolution operations can be used for image processing, image segmentation, object detection and / or classification, etc. Convolution operations can be power-intensive and / or computationally intensive. For example, each output data of a convolution operation may be generated based on performing multiple multiply-accumulate (MAC) operations. The amount of MAC operations per output can depend on the convolution type, kernel size, channel size, etc. For example, a 3x3 convolution with 64 input channels requires 3 MAC operations per output data of the convolution operation. * 3 * This utilizes 64 = 576 MAC operations (for example, the element-wise multiplication and subsequent addition performed for each step of sliding a 3x3 convolution kernel uses 576 MAC operations).

[0023] In some cases, approximately 30% of the total power consumption associated with performing a convolution operation may be used to transfer data for local storage (e.g., RAM) to compute units (CUs). For example, in the above example of a 3x3 convolution with 64 input channels, each output value of the convolution utilizes 576 MAC operations. Each MAC operation requires 1 byte of activation data and 1 byte of weight data, and a total of 1,152 bytes of data reads may be required to compute a single output value for a 3x3 convolution with 64 input channels.

[0024] Systems and techniques that can be used to reduce power consumption and / or computational workload associated with performing convolution operations may be beneficial. Systems and techniques that can perform data fetching and transfer related to convolution operations more efficiently may also be beneficial. Reducing the amount of data fetching (e.g., data reading) and / or data transfer related to convolution operations may also be beneficial.

[0025] Systems, apparatus, processes (also called methods), and computer-readable media (collectively referred to as “Systems and Techniques”) for processing images (e.g., image data or video data) using convolutional machine learning networks are described herein. For example, a convolutional machine learning network (e.g., a CNN) can perform convolution without performing a complete refresh of the input data provided in each clock cycle of the convolution. In some embodiments, Systems and Techniques may be used to perform convolution with reduced data toggling. For example, Systems and Techniques described herein may be used to perform image processing using convolutional operations by reducing the number of data reads and / or the number of data transfers. In some embodiments, Systems and Techniques may utilize static data locations for input data that persists for multiple cycles of the convolutional operation. For example, the input data contained in the second location of the first position of the convolution kernel (e.g., the second location from the left) and the first location of the second position of the convolution kernel (e.g., the leftmost location) may be stored in the same memory location(s) associated with performing multiplier-accumulator (MAC) operations in the first convolution cycle corresponding to the first location of each convolution kernel position for each row of input data (e.g., image data), and in the second convolution cycle corresponding to the second location of each convolution kernel position for each row of input data.

[0026] In some embodiments, at least a portion of the memory locations(s) used to store each input data in a particular convolutional cycle may be static memory locations that are reused in multiple convolutional cycles, based on the fact that each input data is enclosed (e.g., processed) by a convolutional kernel location corresponding to a particular convolutional cycle and then enclosed by a convolutional kernel location corresponding to the next convolutional cycle. In some cases, memory locations used to store input data enclosed by a convolutional kernel location associated with the current convolutional cycle but not by a convolutional kernel location associated with the next convolutional cycle may be replaced (e.g., overwritten) by input data enclosed by a convolutional kernel location associated with the next convolutional cycle but not by a convolutional kernel location associated with the current convolutional cycle. In some embodiments, the replaced input data may be located at the trailing edge of the convolutional kernel in the current convolutional cycle, and the replaced input data may be located at the leading edge of the convolutional kernel in the next convolutional cycle.

[0027] Static data locations can be associated with local storage (e.g., RAM, cache, other memory, etc.) of the compute unit (CU) used to implement the convolution operation. For example, a particular pixel of an input image can be used to compute multiple different output location values ​​of a convolution (e.g., a particular pixel may be included in multiple different sliding window positions of the convolution kernel). In some embodiments, the system and technique can use static locations to store pixel data (e.g., RAM, cache, other memory, etc.), and the static pixel data is used to determine multiple different output location values ​​of the convolution. For each clock cycle (e.g., each "step" of a sliding convolution kernel), a portion of the static pixel data can be replaced. For example, a portion of the static pixel data that is no longer used by any MAC operations for the current position of the sliding convolution kernel can be replaced with an updated value corresponding to the pixel newly covered by the current position of the sliding convolution kernel.

[0028] In some embodiments, input data associated with a particular memory location may be provided to the same multiplier calculation unit (CU) in each convolution cycle. The multiplier CU may be used to implement a multiplier-accumulator (MAC) operation associated with one or more convolution outputs for a particular row of the input image data. In some cases, each multiplier CU may perform MAC operations on different convolution output locations in each convolution cycle of multiple convolution cycles associated with a particular row of the input image data. For example, the multiplier CU may be implemented as a shifted or moving multiplier CU that performs MAC operations on different convolution output locations in each convolution cycle.

[0029] In some cases, each multiplier CU may be associated with a corresponding accumulator or accumulator buffer for storing the cumulative value associated with the MAC operation for a particular convolution output location in each convolution cycle. In some embodiments, each convolution output location may use the same accumulator or accumulator buffer in each convolution cycle. For example, a first accumulator or accumulator buffer may correspond to a first convolution output location and receive the output from a first multiplier CU in a first convolution cycle, the output from a second multiplier CU in a second convolution cycle, and so on. Data switching can be performed between multiple multiplier CUs and multiple accumulators or accumulator buffers so that each accumulator or accumulator buffer receives input from a different multiplier CU in each convolution cycle.

[0030] In another embodiment, each multiplier CU may provide an output to a different accumulator or accumulator buffer in each convolution cycle. For example, the multiplier CU may be implemented as a shifted or moving multiplier CU that performs MAC operations on different convolution output locations in each convolution cycle, and the accumulator may be implemented as a shifted or moving accumulator that updates the accumulated value for different convolution output locations in each convolution cycle. In some embodiments, the multiplier CU and accumulator may be shifted by the same amount between each convolution cycle. For example, in a first convolution cycle, the first multiplier CU may provide an output to a first accumulator buffer, the output corresponding to a first convolution output location. After the first convolution cycle and before the second convolution cycle, the accumulated value in the first accumulator may be written to and replaced with the accumulated value in the second accumulator. A second convolution cycle may be performed to update the second accumulator with the output of the second multiplier CU, where both the second multiplier CU and the second accumulator correspond to the second convolution output location in the first convolution cycle and the first convolution output location in the second convolution cycle.

[0031] Various aspects of this disclosure will be explained with reference to the figures.

[0032] Figure 1 shows an exemplary implementation of a system-on-a-chip (SOC) 100 which may include a central processing unit (CPU) 102 or a multi-core CPU configured to perform one or more of the functions described herein. Among some of the information, parameters or variables (e.g., neural signals and synaptic weights), system parameters associated with computing devices (e.g., weighted neural networks), delays, frequency bin information, and task information may be stored in memory blocks associated with the neural processing unit (NPU) 108, memory blocks associated with the CPU 102, memory blocks associated with the graphics processing unit (GPU) 104, memory blocks associated with the digital signal processor (DSP) 106, memory block 118, and / or distributed across multiple blocks. Instructions executed in the CPU 102 may be loaded from program memory associated with the CPU 102 or from memory block 118.

[0033] The SOC100 may also include a GPU104, a DSP106, a connectivity block 110 which may include fifth-generation (5G) connectivity, fourth-generation long-term evolution (4G LTE) connectivity, Wi-Fi connectivity, USB connectivity, Bluetooth connectivity, and additional processing blocks adapted to specific functions, such as a multimedia processor 112 capable of detecting and recognizing gestures. In some implementations, the NPU is implemented within the CPU 102, DSP106, and / or GPU 104. The SOC100 may also include one or more sensors 114, image signal processors (ISPs) 116, and / or storage 120.

[0034] The SOC100 may be based on the ARM instruction set. In one aspect of this disclosure, the instructions loaded into the CPU102 may include code for retrieving the stored multiplication result in a lookup table (LUT) corresponding to the multiplication product of the input values ​​and filter weights. The instructions loaded into the CPU102 may also include code for disabling the multiplier during the multiplication operation of the multiplication product when a lookup table hit is detected for the multiplication product. In addition, the instructions loaded into the CPU102 may include code for storing the calculated multiplication product of the input values ​​and filter weights when a lookup table miss is detected for the multiplication product.

[0035] The SOC100 and / or its components may be configured to perform image processing using machine learning techniques according to the embodiments of this disclosure described herein. For example, the SOC100 and / or its components may be configured to perform disparity estimation improvements for pairs of images (e.g., stereo image pairs, each containing a left image and a right image). The SOC100 may be part of one computing device or multiple computing devices. In some embodiments, the SOC100 may be part of an electronic device (or multiple devices), such as a camera system (e.g., digital camera, IP camera, video camera, security camera, etc.), a telephone system (e.g., smartphone, mobile phone, conferencing system, etc.), a desktop computer, an XR device (e.g., a head-mounted display, etc.), a smart wearable device (e.g., a smartwatch, smart glasses, etc.), a laptop or notebook computer, a tablet computer, a set-top box, a television, a display device, a system-on-a-chip (SoC), a digital media player, a game console, a video streaming device, a server, a drone, a computer in a car, an Internet of Things (IoT) device, or any other suitable electronic device(s)

[0036] In some implementations, the CPU 102, GPU 104, DSP 106, NPU 108, connectivity block 110, multimedia processor 112, one or more sensors 114, ISP 116, memory block 118, and / or storage 120 may be part of the same computing device. For example, in some cases, the CPU 102, GPU 104, DSP 106, NPU 108, connectivity block 110, multimedia processor 112, one or more sensors 114, ISP 116, memory block 118, and / or storage 120 may be integrated into a smartphone, laptop, tablet computer, smart wearable device, video gaming system, server, and / or any other computing device. In other implementations, the CPU 102, GPU 104, DSP 106, NPU 108, connectivity block 110, multimedia processor 112, one or more sensors 114, ISP 116, memory block 118, and / or storage 120 may be part of two or more separate computing devices.

[0037] Machine learning (ML) can be considered a subset of artificial intelligence (AI). ML systems can include algorithms and statistical models that computer systems can use to perform various tasks by relying on patterns and inference without using explicit instructions. An example of an ML system is a neural network (also called an artificial neural network), which may contain an interconnected group of artificial neurons (e.g., neuron models). Neural networks may be used in a variety of applications and / or devices, including, among others, image and / or video coding, image analysis and / or computer vision applications, Internet Protocol (IP) cameras, Internet of Things (IoT) devices, autonomous vehicles, and service robots.

[0038] Individual nodes within a neural network can emulate biological neurons by acquiring input data and performing simple operations on it. The results of these operations are selectively passed to other neurons. Each vector and node in the network is associated with a weight value, which constrains how the input data relates to the output data. For example, the input data of each node may be multiplied by the corresponding weight value, and the products may be summed. The sum of the products may be adjusted by an arbitrary bias, and an activation function may be applied to the result to obtain the node's output signal, or "output activation" (sometimes called a feature map or activation map). The weight values ​​may initially be determined by an iterative flow of training data through the network (for example, the weight values ​​are established during the training phase, when the network learns how to identify specific classes based on their typical input data characteristics).

[0039] In particular, there are different types of neural networks, such as convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), multilayer perceptron (MLP) neural networks, and transformer neural networks. For example, convolutional neural networks (CNNs) are a type of feedforward artificial neural network. A convolutional neural network may contain a collection of artificial neurons, each having a receptive field (e.g., a spatially localized region of the input space) and tiling the input space together. RNNs work on the principle of storing the output of a layer and feeding this output back into the input to help predict the outcome of the layer. GANs are a form of generative neural network that can learn patterns in input data so that the neural network model can reasonably generate a new synthetic output that is likely to be from the original dataset. A GAN can include two neural networks working together: a generative neural network that produces a synthesized output, and a discriminative neural network that evaluates the output for reliability. In an MLP neural network, data is fed into an input layer, and one or more hidden layers may provide a level of abstraction to the data. Predictions can then be made in the output layer based on the abstracted data.

[0040] Deep learning (DL) is an example of a machine learning technique and can be considered a subset of ML. Many DL methods are based on neural networks, such as RNNs or CNNs, and utilize multiple layers. The use of multiple layers in a deep neural network allows for the extraction of progressively higher levels of features from a given input of raw data. For example, the output of the first layer of an artificial neuron becomes the input to the second layer of the artificial neuron, the output of the second layer of the artificial neuron becomes the input to the third layer of the artificial neuron, and so on. Layers located between the inputs and outputs of the entire deep neural network are often called hidden layers. Hidden layers learn (e.g., are trained) to transform intermediate inputs from previous layers into somewhat more abstract and synthetic representations that can be provided to later layers until the final or desired representation is obtained as the final output of the deep neural network.

[0041] As mentioned above, a neural network is an example of a machine learning system that can include an input layer, one or more hidden layers, and an output layer. Data is provided from the input nodes of the input layer, processed by the hidden nodes of one or more hidden layers, and an output is produced through the output nodes of the output layer. Deep learning networks typically include multiple hidden layers. Each layer of a neural network can include a feature map or activation map, which may contain artificial neurons (or nodes). Feature maps can include filters, kernels, etc. Nodes can include one or more weights used to indicate the importance of one or more nodes in the layer. In some cases, a deep learning network can have a series of many hidden layers, with early layers used to determine simple, low-level characteristics of the input, and later layers building a hierarchy of more complex and abstract characteristics.

[0042] Deep learning architectures can learn feature hierarchies. When presented with visual data, for example, the first layer might learn to recognize relatively simple features such as edges in the input stream. In another example, when presented with auditory data, the first layer might learn to recognize spectral power at specific frequencies. A second layer, taking the output of the first layer as input, might learn to recognize combinations of features, such as simple shapes in the case of visual data, or combinations of sounds in the case of auditory data. For example, higher layers might learn to represent complex shapes in visual data or words in auditory data. Even higher layers might learn to recognize common visual objects or spoken phrases. Deep learning architectures can work particularly well when applied to problems with natural hierarchical structures. For example, classifying motorized vehicles might benefit from first learning to recognize wheels, windshields, and other features. These features can then be combined in different ways in higher layers to recognize cars, trucks, and airplanes.

[0043] Neural networks can be designed using various connectivity patterns. In a feedforward network, information is passed from lower layers to higher layers, with each neuron in a given layer transmitting information to neurons in higher layers. As mentioned above, a hierarchical representation can be constructed within the consecutive layers of a feedforward network. Neural networks can also have recursive or feedback (also called top-down) connectivity. In recursive connectivity, the output from a neuron in a given layer can be transmitted to another neuron in the same layer. Recursive architectures can be useful when recognizing patterns that span two or more chunks of input data delivered to the neural network in a sequence. Connectivity from a neuron in a given layer to a neuron in a lower layer is called feedback (or top-down) connectivity. Networks with a large number of feedback connections can be useful when the recognition of a higher-level concept can help identify specific lower-level features of the input.

[0044] The connections between layers in a neural network may be fully connected or locally connected. Figure 2A shows an example of a fully connected neural network 202. In the fully connected neural network 202, neurons in the first hidden layer may transmit their output to any neuron in the second hidden layer, so that each neuron in the second layer receives input from any neuron in the first layer. Figure 2B shows an example of a locally connected neural network 204. In the locally connected neural network 204, neurons in the first hidden layer may be connected to a limited number of neurons in the second hidden layer. More generally, the locally connected layers of the locally connected neural network 204 may be configured such that each neuron in the layer has the same or similar connectivity pattern but different connection strengths (e.g., 210, 212, 214, and 216). Since neurons in the upper layers of a given region can receive inputs that are modulated through training on properties of a limited portion of the total input to the network, the connectivity patterns of local connections can give rise to spatially distinct receptive fields within the upper layers.

[0045] An example of a locally connected neural network is a convolutional neural network. Figure 2C shows an embodiment of a convolutional neural network 206. The convolutional neural network 206 may be configured such that the connection strength (e.g., 208) associated with the input for each neuron in the second layer is shared. Convolutional neural networks may be suitable for problems where the spatial location of the input is meaningful. According to aspects of this disclosure, the convolutional neural network 206 may be used to perform one or more embodiments of video compression and / or restoration. Exemplary embodiments of deep learning networks are described in more detail with respect to the exemplary block diagram in Figure 8. Exemplary embodiments of convolutional neural networks are described in more detail with respect to the exemplary block diagram in Figure 9.

[0046] As described above, the systems and techniques described herein can be used to perform image processing using convolution operations by reducing the number of data reads and / or data transfers. In some embodiments, the systems and techniques can utilize static data locations for input data that persists for multiple cycles of the convolution operation. Static data locations may be associated with local storage (e.g., RAM, cache, other memory, etc.) of the compute unit (CU) used to implement the convolution operation. For example, a particular pixel of the input image can be used to compute multiple different output location values ​​of the convolution (e.g., a particular pixel may be in multiple different sliding window positions of the convolution kernel). In some embodiments, the systems and techniques can use static locations to store pixel data (e.g., RAM, cache, other memory, etc.), and the static pixel data is used to determine multiple different output location values ​​of the convolution. For each clock cycle (e.g., each "step" of the sliding convolution kernel), a portion of the static pixel data can be replaced. For example, portions of static pixel data that are no longer used by any MAC operations for the current position of the sliding convolution kernel can be replaced with updated values ​​corresponding to pixels newly covered by the current position of the sliding convolution kernel.

[0047] Figure 3 shows an embodiment of a 3x1 two-dimensional (2D) convolution 300 that uses static multiplier units and static accumulator locations to implement MAC operations based on receiving new activation data in each cycle (e.g., each clock cycle). For example, the exemplary convolution 300 is associated with 2D input data 310 containing multiple pixels. In some embodiments, the 2D input data 310 may be an image or video frame, etc.

[0048] By sliding a 3x1 convolution kernel (or filter) across each row of the input data 310 (for example, from left to right or right to left), the corresponding row of the convolution output 320 is generated. Each output location of the convolution output 320 is associated with a different sliding window position of the 3x1 convolution kernel across the rows of the input data 310.

[0049] For example, a total of six input locations in the input data 310 are used to generate four output locations in the convolution output 320. Location "0" in the convolution output 320 is based on a 3x1 convolution kernel location spanning locations "0", "1", and "2" in the input data 310. Location "1" in the convolution output 320 is based on a 3x1 convolution kernel location spanning locations "1", "2", and "3" in the input data 310. Location "2" in the convolution output 320 is based on a 3x1 convolution kernel location spanning locations "2", "3", and "4" in the input data 310, and location "3" in the convolution output 320 is based on a 3x1 convolution kernel location spanning locations "3", "4", and "5" in the input data 310.

[0050] In some embodiments, a CNN or other machine learning network can implement a convolution operation based on simultaneously calculating multiple output locations. For example, the four output locations of a convolutional output 320 can be calculated simultaneously using four corresponding computation units (CUs) to perform a MAC operation for each output location of the convolutional output 320.

[0051] The CU used to perform MAC operations on a specific convolution output location may include a multiplier (M) 340, an accumulator (A) 350, and / or a buffer (B) 360. In these embodiments, four multipliers 340, four accumulators 350, and four buffers 360 can be used to compute four output locations of the convolution output 320 simultaneously using a total of three cycles.

[0052] Each cycle is associated with corresponding input data. For example, the first cycle ("Cycle 1") is associated with the Cycle 1 input data 330(1), the second cycle ("Cycle 2") is associated with the Cycle 2 input data 330(2), and the third cycle ("Cycle 3") is associated with the Cycle 3 input data 330(3). Cycle input data 330(1)~(3) can be obtained as a subset of input data 310. For example, in Cycle 1, input data 0 / 1 / 2 / 3 are sent to the four multipliers 340, respectively. In Cycle 2, input data 1 / 2 / 3 / 4 are sent to the four multipliers 340. In Cycle 3, input data 2 / 3 / 4 / 5 are sent to the four multipliers 340.

[0053] The multiplier 340 may be part of a multiplier and adder tree associated with implementing MAC operations for convolution. In some embodiments, the multiplier 340 may receive or utilize one or more weight values ​​(for example, shown as "WGT K[7:0]" in Figure 3). The output of each multiplier (for example, of multiplier 340) may be provided to the corresponding accumulator (for example, of accumulator 350). The output of each accumulator 350 may be provided to the corresponding accumulator buffer 360.

[0054] After performing cycles 1, 2, and 3, the contents of the accumulator buffer 360 (e.g., accumulated data for each MAC) may be provided to the streaming mux 370. In some cases, the output of the streaming mux 370 may be provided to a post-processing engine 380 which can perform one or more post-processing operations (e.g., "BIAS / BN" in Figure 3) to generate a convolution output 390. In some embodiments, the convolution output 390 may be the same as the convolution output 320.

[0055] In some convolution techniques, the input data is completely cleared and refreshed between each cycle (for example, the first memory location stores input data "0" during the first cycle 330(1), then input data "1" during the second cycle 330(2), and then input data "2" during the third cycle 330(3)). The same input data value may also be read multiple times and transferred to memory over multiple cycles corresponding to generating the convolution output 390. The same input data value may be stored in different memory locations over multiple cycles 330(1)-(3). For example, input data "1" may be stored in two different memory locations within cycle 1 input data 330(1) and cycle 2 input data 330(2). Input data "2" and "3" are stored in three different memory locations within the cycle 1 input data 330(1), cycle 2 input data 330(2), and cycle 3 input data 330(3), respectively.

[0056] Refreshing the input data in each cycle may be associated with performing multiple data toggling operations in the data path of an exemplary 3x1 convolution 300. Increased data toggling may be associated with increased power consumption. Therefore, in some cases, it may be beneficial to reduce data toggling and thus reduce the power consumption associated with performing the convolution operation.

[0057] Figure 4 shows one embodiment of a 3×1 convolution 400 using one or more static input data locations and static output accumulator locations, according to several embodiments. For example, the 3×1 convolution 400 may be performed based on a first cycle ("Cycle 1") 430(1), a second cycle ("Cycle 2") 430(2), and a third cycle ("Cycle 3") 430(3). Each of the three cycles 430(1), 430(2), and 430(3) may be associated with corresponding cycle input data 410(1), 410(2), and 410(3), respectively. Each cycle input data 410(1) to 410(3) may be a subset of input data 410. In some embodiments, input data 410 may be the same as or similar to input data 310 in Figure 3.

[0058] In some embodiments, the convolution 400 may be performed using one or more static input data locations per cycle. For example, the respective memory locations associated with the local storage of input data "1", "2", and "3" (e.g., local to the CU containing the multiplier (M) and accumulator (A) for implementing MAC operations) may not change between cycle 1 and cycle 2 (e.g., they may be static). The respective memory locations associated with the local storage of input data "2" and "3" may not change between cycle 1, cycle 2, and cycle 3 (e.g., they may be static).

[0059] For example, between cycle 1 and cycle 2, the stored values ​​for input data "1", "2", and "3" are not updated based on the fact that these input data are used in both the MAC calculation in cycle 1 and the MAC calculation in cycle 2. Input data "0" is used in the MAC calculation in cycle 1 but is not needed in the MAC calculation in cycle 2. In some embodiments, only the memory locations associated with local storage of input data not needed for the MAC calculation in the current cycle are refreshed and updated with new input data values ​​(for example, corresponding to input data newly needed for the MAC calculation in the current cycle but not needed for the MAC calculation in previous cycles).

[0060] For example, the cycle 1 input data 410(1) includes the input data "0", which is necessary for the cycle 1 MAC calculation but not for the cycle 2 MAC calculation. The input data "4" is necessary for the cycle 2 MAC calculation but not for the cycle 1 MAC calculation. In some cases, the cycle 2 input data 410(2) does not include the input data "0" but includes the input data "4".

[0061] In some embodiments, each cycle input data 410(1)-(3) may be generated by replacing input data values ​​located outside the current convolution kernel position (e.g., the position of the convolution kernel in the current cycle) with input data values ​​located within the current convolution kernel position. In some cases, the replaced input data values ​​are located at the trailing edge of the sliding convolution kernel (e.g., the left edge of the convolution kernel in the case of sliding from left to right), and the updated or replaced input data values ​​are located at the leading edge of the sliding convolution kernel (e.g., the right edge of the convolution kernel in the case of sliding from right to left).

[0062] In some embodiments, the replaced input data values ​​(e.g., at the trailing edge of the convolution kernel position) and the replaced input data values ​​(e.g., at the leading edge of the convolution kernel position) in each cycle can be stored using the same memory location. For example, the replaced input data value "4" can overwrite the replaced input data value "0" in the same memory location to generate the cycle 2 input data 410(2) from the cycle 1 input data 410(1). In another embodiment, the replaced input data value "5" can overwrite the replaced input data value "1" in the same memory location to generate the cycle 3 input data 410(3) from the cycle 2 input data 410(2).

[0063] In the example 3x1 convolution 400 in Figure 4, instead of sending new data to each CU in each cycle (as described for example with respect to the example convolution 300 in Figure 3), the system and technique can use static memory locations to store 75% of the input data from the previous cycle, and update the memory locations corresponding to the remaining 25% of the input data from the previous cycle with the replacement input data values ​​of the current cycle.

[0064] In some cases, each cycle input data 410(1)-410(3) contains four columns and eight rows for a total of 32 input data values ​​per cycle. In some embodiments, the eight rows can represent different channels of the input data 410. Additional CUs (e.g., multipliers 440(1)-(4), accumulators 450(1)-(4), etc.) may be provided to process the 32 input data values ​​per cycle simultaneously. In some embodiments, four multipliers 440(1)-(4) and four accumulators 450(1)-(4) may be used to process four input values ​​in a single row simultaneously (e.g., corresponding to four different columns), and each of the eight row inputs of four values ​​is processed sequentially using the same four multipliers 440(1)-(4) and accumulators 450(1)-(4).

[0065] In some embodiments, at least a portion of the cycle input data memory locations remain constant (e.g., static or persistent) for at least two cycles of the exemplary convolution 400. In some embodiments, the system and technique can utilize shifting (e.g., moving) MACs (e.g., CUs, multipliers 440(1)-(4), accumulators 450(1)-(4), etc.) to generate the four convolution outputs X_0, X_1, X_2, and X_3 in Figure 4 (e.g., shown in cycle 3 as the output of accumulators 450(1)-(4) after three cycles have been completed).

[0066] For example, in cycle 1, multiplier 440(1) receives the input data "0" from cycle 1 input data 410(1). The output of multiplier 440(1) is stored in accumulator 450(1). In cycle 2, the mapping between multipliers and accumulators is shifted one column to the right. Multiplier 440(2) receives the input data "1" from cycle 2 input data 410(2), and the output of multiplier 440(2) is stored in the same accumulator 450(1). In cycle 3, the mapping between multipliers and accumulators is shifted one column to the right again. Multiplier 440(3) receives the input data "2" from cycle 3 input data 410(3), and the output of multiplier 440(3) is stored in the same accumulator 450(1).

[0067] At the end of three cycles of convolution 400, the output of accumulator 450(1) (e.g., the output of the corresponding accumulator buffer for accumulator 450(1)) is X_0, shown in Figure 4 as the first convolution output location 460(1). Based on shifting the mapping between multipliers 440(1)-(4) and accumulators 450(1)-(4) to correspond to the respective static input data memory locations used in each cycle input data 410(1)-(3) and the updated input data memory locations, the respective convolution outputs X_0, X_1, X_2, and X_3 in Figure 4 are the same as the convolution outputs 320 / 390 determined using the exemplary convolution 300 in Figure 3 (e.g., the convolution outputs determined using the new input data of each cycle without shifting the multipliers or CUs).

[0068] The convolution output X_3 460(4) can be generated using the same shift in the mapping between multipliers 440(1)~(4) and accumulators 450(1)~(4) over three cycles, and the convolution output X_3 460(4) is based on MAC operations using input data "3", "4", and "5". For example, in cycle 1, multiplier 440(4) receives input data "3" from cycle 1 input data 410(1). The output of multiplier 440(1) is stored in accumulator 450(4). In cycle 2, the mapping between the multiplier and accumulator is shifted one column to the right (for example, as described above). The convolution output X_3 460(4) utilizes the replaced input data "4" used to generate cycle 2 input data 410(2) by updating the memory location corresponding to the replaced input data "0" in cycle 1 input data 410(1). In cycle 2, the replacement input data "4" may be provided to multiplier 440(1), and the output of multiplier 440(1) may be stored in the same accumulator 450(4) based on the shifted mapping between the multiplier and accumulator used for cycle 2. At the end of cycle 2, the accumulator buffer associated with accumulator 450(4) contains the results of MAC operations based on input data "3" and input data "4".

[0069] In cycle 3, the mapping between the multiplier and the accumulator is again shifted one column to the right (for example, as described above). The convolution output X_3 460(4) utilizes the replaced input data "5" used to generate the cycle 3 input data 410(3) by updating the memory location corresponding to the replaced input data "1" in the cycle 2 input data 410(2). In cycle 3, the replaced input data "3" may be provided to the multiplier 440(2), and the output of the multiplier 440(2) may be stored in the same accumulator 450(4) based on the shifted mapping between the multiplier and the accumulator used for cycle 3. At the end of cycle 3, the accumulator buffer associated with accumulator 450(4) contains the result of the MAX operation based on the input data "3", "4", and "5" (for example, the fourth convolution output location value X_3 460(4)).

[0070] Figure 5 shows one embodiment of a 3×1 convolution 500 using one or more static input data locations with shifted CUs (e.g., shifted MAC calculation units, each containing its own multiplier and its own accumulator), according to several embodiments. For example, the 3×1 convolution 500 may be performed based on a first cycle ("cycle 1") 530(1), a second cycle ("cycle 2") 530(2), and a third cycle ("cycle 3") 530(3). Each of the three cycles 530(1), 530(2), and 530(3) may be associated with corresponding cycle input data 510(1), 510(2), and 510(3), respectively. Each cycle input data 510(1) to 510(3) may be a subset of input data 510. In some embodiments, input data 510 may be the same as or similar to input data 310 in Figure 3 and / or input data 410 in Figure 4.

[0071] In some embodiments, the exemplary convolution 400 in Figure 4 may be implemented based on shifting the MAC CU location each cycle using corresponding data switching to maintain the mapping between MAC outputs and accumulator inputs (for example, so that each accumulator output corresponds to one of four convolution outputs X_0, X_1, X_2, or X_3) and a static output accumulator location. In some embodiments, the exemplary convolution 500 in Figure 5 may be implemented based on shifting the MAC CU location and output accumulator location together. Based on shifting the MAC CU location and output accumulator location in each cycle of the convolution 500, the system and technique can reduce data toggling between MAC outputs and accumulator inputs. For example, using the shifted MAC CU and accumulator in Figure 5, each respective accumulator 550(1) to 550(4) receives data from a total of two locations. As will be explained in more detail below, the number of accumulator data inputs can be 2, independent of the number size of the convolution kernel (for example, in the case of n×m convolutions in addition to the exemplary 3×1 convolution 500). Using the shifted MAC CU and static accumulators in Figure 4, each respective accumulator 450(1) to 450(4) receives data from a total of 3 locations per row (for example, in the case of n×1 convolutions, each respective accumulator 450(1) to 450(4) receives data from n locations for each row in the cycle input data).

[0072] In some cases, the cycle input data 510(1), 510(2), and 510(3) in Figure 5 may be the same as the cycle input data 410(1), 410(2), and 410(3) in Figure 4, respectively. In cycle 1 (for example, 530(1) in Figure 5), input data "0" is provided to the first multiplier 540(1), input data "1" is provided to the second multiplier 540(2), input data "2" is provided to the third multiplier 540(3), and input data "3" is provided to the fourth multiplier 540(4). Multipliers 540(1) to (4) can correspond to their respective accumulators 550(1) to 550(4). For example, the output of the first multiplier 540(1) may be provided to the first accumulator 550(1), and so on.

[0073] At the end of each cycle, the accumulated values ​​in each accumulator 550(1)~(4) may be shifted to a different accumulator corresponding to a multiplier that will receive the next input data value used to calculate a particular convolution output value. For example, in cycle 1, the first multiplier 540(1) and the first accumulator 550(1) correspond to the first convolution output value X_0.

[0074] In cycle 2, the second multiplier 540(2) and the second accumulator 550(2) correspond to the first convolution output value X_0. In some embodiments, at the end of cycle 1 (for example, before cycle 2), the contents of accumulator 550(1) can be shifted to accumulator 550(2). At the start of cycle 2, accumulator 550(2) stores the MAC value based on the input data "0". In cycle 2, the input data "1" is provided to the multiplier 540(2), and the output of the multiplier 540(2) is stored in accumulator 550(2). After cycle 2 has been performed, accumulator 550(2) stores the MAC value based on the input data "0" and "1".

[0075] In cycle 3, the third multiplier 540(3) and the third accumulator 550(3) correspond to the first convolution output value X_0. In some embodiments, at the end of cycle 2 (e.g., before cycle 3), the contents of accumulator 550(2) may be shifted to accumulator 550(3). In cycle 3, the input data "2" is provided to the multiplier 540(3), and the output of the multiplier 540(3) is stored in accumulator 550(3). After cycle 3 has been performed, accumulator 550(3) stores a MAC value based on the input data "0", "1", and "2" (e.g., equal to the first convolution output X_0 560(1)).

[0076] In some embodiments, the accumulators 550(1) to (4) can implement the same shift at the end of each cycle of the exemplary convolution 500. For example, after cycles 1 and 2, the accumulated value stored in the first accumulator 550(1) may be shifted to and stored in the second accumulator 550(2). The accumulated value previously stored in the second accumulator 550(2) may be shifted to and stored in the third accumulator 550(3). The accumulated value previously stored in the third accumulator 550(3) may be shifted to and stored in the fourth accumulator 550(4). The accumulated value previously stored in the fourth accumulator 550(4) may be shifted to and stored in the first accumulator 550(1).

[0077] In some embodiments, each accumulator 550(1)-(4) may be used to implement the convolution 500 on the basis that each accumulator receives two respective inputs per cycle (e.g., every 1-3 cycles). For example, in each cycle, each accumulator 550(1)-(4) may receive a first input corresponding to the respective outputs of the multipliers 540(1)-(4). The second input to each accumulator 550(1)-(4) is the accumulated value stored in the adjacent accumulator. For example, in each cycle, the first accumulator 550(1) receives a first input corresponding to the output of the first multiplier 540(1), and at the end of the cycle, receives a second input corresponding to the accumulated value stored in the fourth accumulator 550(4). In each cycle, the second accumulator 550(2) receives a first input corresponding to the output of the second multiplier 540(2), and at the end of the cycle, receives a second input corresponding to the accumulated value stored in the first accumulator 550(1), and so on. Based on the implementation of the shifted MAC CU and accumulator pairs described above, the exemplary convolution 500 can be associated with a 2:1 mux.

[0078] Figure 6 shows one embodiment of a 3x3 convolution 600 using shifted MAC CU and output accumulator locations, according to several embodiments. The 3x3 convolution 600 can be performed using input data 610, which may be an image containing multiple pixels, each associated with multiple channels. For example, the input data 610 contains 6 columns and 3 rows (e.g., 18 values) for each of the multiple channels. For the 8 channels shown in Figure 6, the input data 610 totals 6 * 3 * 8 = 144 values.

[0079] The 3x3 convolution 600 can be implemented using a 3x3 convolution kernel, which performs MAC operations to compute element-wise multiplication and addition of a 3x3 subset of the input data 610 using the weight matrix of the 3x3 convolution kernel. For example, the first convolution output 660(1) can be determined using the 3x3 convolution kernel at position 612(1). Position 612(1) of the 3x3 convolution kernel corresponds to the input data 0 / 1 / 2 / 6 / 7 / 8 / 12 / 13 / 14. The first convolution output 660(1) can be determined based on the MAC of the input data 0 / 1 / 2 / 6 / 7 / 8 / 12 / 13 / 14 using the 3x3 weight matrix of the convolution kernel.

[0080] The fourth convolution output 660(4) can be determined using a 3x3 convolution kernel at position 612(4). Position 612(4) of the 3x3 convolution kernel corresponds to input data 3 / 4 / 5 / 9 / 10 / 11 / 15 / 16 / 17. The fourth convolution output 660(4) can be determined based on the MAC of input data 3 / 4 / 5 / 9 / 10 / 11 / 15 / 16 / 17 using the 3x3 weight matrix of the convolution kernel.

[0081] In some embodiments, the systems and techniques described herein can implement an exemplary 3x3 convolution 600 using multiple cycles, where each cycle corresponds to cycle input data which is a subset of the input data 610. For example, the first cycle ("cycle 1") may correspond to the first cycle input data 610(1), the second cycle ("cycle 2") may correspond to the second cycle input data 610(2), ..., and the ninth cycle ("cycle 9") may correspond to the ninth cycle input data 610(9).

[0082] In some embodiments, cycle input data 610(1), 610(2), and 610(3) may be the same as the respective cycle input data 510(1), 510(2), and 510(3) in Figure 5. Cycle input data associated with the same row within input data 610 can reuse input data values ​​across multiple cycles using static memory locations, as previously described with respect to Figures 4 and 5. For example, the memory location used to store input data "1" may be the same for cycle 1 data 610(1) and cycle 2 data 610(2), and the memory locations used to store input data "2" and "3" may be the same for cycle 1 data 610(1), cycle 2 data 610(2), and cycle 3 data 610(3), and so on.

[0083] In some embodiments, each convolution row of the input data 610 is associated with three cycles of MAC operations. For example, the first 6x1 row of input data 610 is used to perform four respective MAC operations (e.g., corresponding to each of the four convolution outputs 660(1) to 660(4)) using input data 0 / 1 / 2 / 3 / 4 / 5. The second 6x1 row is used to perform four respective MAC operations using input data 6 / 7 / 8 / 9 / 10 / 11. The third 6x1 row is used to perform four respective MAC operations using input data 12 / 13 / 14 / 15 / 16 / 17.

[0084] One or more memory locations used to store input data corresponding to a particular row may be the same for one or more cycles performed for that particular row. In some embodiments, when a new row begins, all memory locations used to store the input data of the previous row may be updated to store the input data of the current row. For example, cycles 1-3 correspond to the first row of input data 0 / 1 / 2 / 3 / 4 / 5 used to compute four convolutional outputs 660(1)-(4). Cycles 4-6 correspond to the second row of input data 6 / 7 / 8 / 9 / 10 / 11 used to compute four convolutional outputs 660(1)-(4). Cycles 7-9 correspond to the third row of input data 12 / 13 / 14 / 15 / 16 / 17 used to compute four convolutional outputs 660(1)-(4).

[0085] After cycle 3, each memory location used to store cycle 3 input data 610(3) may be updated (e.g., replaced) to store cycle 4 input data 610(4). After cycle 6, each memory location used to store cycle 6 input data 610(6) may be updated (e.g., replaced) to store cycle 7 input data 610(7).

[0086] Between cycles associated with the same row of input data 610 (for example, between cycles 1-2 and 2-3, cycles 4-5 and 5-6, and cycles 7-8 and 8-9), only one set of memory locations is updated to store the replaced input data values. For example, between cycles 1 and 2, input data "0" is replaced with input data "4", and between cycles 2 and 3, input data "1" is replaced with input data "5". Between cycles 4 and 5, input data "6" is replaced with input data "10", and between cycles 5 and 6, input data "7" is replaced with input data "11". Between cycles 7 and 8, input data "12" is replaced with input data "16", and between cycles 8 and 9, input data "13" is replaced with input data "17".

[0087] In some cases, a 3x3 convolution 600 can be implemented using shifted MAC CU and accumulator locations, as described above with respect to the exemplary 3x1 convolution 500 in Figure 5. For example, in cycle 1, the accumulated value corresponding to the convolution output X_0 660(1) may be stored in the first accumulator 650(1), ..., and the accumulated value corresponding to the convolution output X_3 660(4) may be stored in the fourth accumulator 650(4).

[0088] Before or during cycle 2, the accumulated value stored in the first accumulator 650(1) may be transferred to the second accumulator 650(2), ..., and the accumulated value stored in the fourth accumulator 650(4) may be transferred to the first accumulator 650(1). After cycle 2, the accumulated value stored in the second accumulator 650(2) corresponds to the first convolution output X_0 660(1), and the accumulated value stored in the first accumulator 650(1) corresponds to the fourth convolution output X_3 660(4). After cycle 3, the accumulated value stored in the third accumulator 650(3) corresponds to the first convolution output X_0 660(1), and the accumulated value stored in the second accumulator 650(2) corresponds to the fourth convolution output X_3 660(4).

[0089] After cycle 4, the accumulated value stored in the fourth accumulator 650(4) corresponds to the first convolution output X_0 660(1), and the accumulated value stored in the third accumulator 650(3) corresponds to the fourth convolution output X_3 660(4). After cycle 5, the accumulated value stored in the first accumulator 650(1) corresponds to the first convolution output X_0 660(1), and the accumulated value stored in the fourth accumulator 650(4) corresponds to the fourth convolution output X_3 660(4).

[0090] After cycle 9, the accumulated value stored in the first accumulator 650(1) corresponds to the complete convolution output X_0 660(1) and is based on a MAC operation using the input data 0 / 1 / 2 / 6 / 7 / 8 / 12 / 13 / 14 in the 3x3 convolution kernel at the first position 612(1). After cycle 9, the accumulated value stored in the fourth accumulator 650(1) corresponds to the complete convolution output X_3 660(4) and is based on a MAC operation using the input data 3 / 4 / 5 / 9 / 10 / 11 / 15 / 16 / 17 in the 3x3 convolution kernel at the fourth position 612(4).

[0091] Figure 7 is a flowchart of one embodiment of a process 700 for processing image data. The exemplary process 700 shows a specific sequence of operations, but the sequence can be modified without departing from the scope of the disclosure. For example, some of the operations shown may be performed in parallel or in different sequences that do not substantially affect the functionality of process 700. In other embodiments, different components of an exemplary device or system implementing process 700 may perform functions substantially simultaneously or in specific sequences.

[0092] In block 702, process 700 includes obtaining each first value enclosed by the convolution kernel at each of several positions of the convolution kernel along the rows of image data. For example, the rows of image data may be the same as or similar to the rows of input data 310 in Figure 3. Each first value enclosed by the exemplary 3×1 convolution kernel in Figure 3 at the first position may be the input data "0", each first value enclosed by the exemplary 3×1 convolution kernel in Figure 3 at the second position may be the input data "1", and so on. In some cases, the multiple positions of the convolution kernel along the rows of image data may correspond to sliding the convolution kernel across the rows of image data. For example, cycle 430(1) in Figure 4 may correspond to the first position of the convolution kernel along the row of image data 410, cycle 430(2) may correspond to the second position of the convolution kernel, and cycle 430(3) may correspond to the third position of the convolution kernel. In Figure 5, cycle 530(1) can correspond to the first position of the convolution kernel along the row of image data 510, cycle 530(2) can correspond to the second position of the convolution kernel, and cycle 530(3) can correspond to the third position of the convolution kernel. Cycles 1 to 9 in Figure 6 can correspond to nine different positions of the convolution kernel along the three rows of input image data 610 and / or 612(1) in Figure 6.

[0093] In some embodiments, the process 700 includes determining a convolution output for each of several locations in the convolution kernel based on multiple convolution cycles, where each convolution cycle corresponds to a different location within the convolution kernel. For example, the convolution output 320 in Figure 3 may be determined for each of several locations in the convolution kernel based on multiple convolution cycles 330(1) to 330(3). The convolution outputs 460(1) to 460(4) in Figure 4 may be determined for each of several locations in the convolution kernel along the image data 410 based on multiple convolution cycles 430(1) to 430(3). The convolution outputs 560(1) to 560(4) in Figure 5 may be determined for each of several locations in the convolution kernel along the image data 510 based on multiple convolution cycles 530(1) to 530(3).

[0094] In some embodiments, the first convolution cycle corresponds to each first value enclosed by the convolution kernel, and each of these first values ​​corresponds to a first location within the convolution kernel. For example, each first value may correspond to a first location that includes the leftmost location within the convolution kernel positioned on a row of image data. In some cases, the convolution output for each of the multiple locations of the convolution kernel along the rows of image data is associated with a different multiplier calculation unit (CU) in each convolution cycle of multiple convolution cycles. For example, the multiplier CU may be the same as or similar to one or more of the multiplier 340, accumulator 350, and / or buffer 360 in Figure 3 (e.g., collectively, the MAC CU). Multiplier CU may be the same as or similar to one or more of the multipliers CU440(1) to 440(4) in Figure 4, one or more of the multipliers CU540(1) to 540(4) in Figure 5, and / or one or more of the multipliers CU640(1) to 640(4) in Figure 6.

[0095] In block 704, process 700 includes storing each of the first values ​​using the respective memory locations associated with each of the multiple locations in the convolution kernel. For example, each of the first values ​​may be stored using one of the accumulators 350 and / or one of the buffers 360 in Figure 6, one of the accumulators 450(1) to 450(4) in Figure 4, one of the accumulators 550(1) to 550(4) in Figure 5, and / or one of the accumulators 650(1) to 650(4) in Figure 6.

[0096] In block 706, process 700 includes updating cumulative values ​​corresponding to the convolution output for each of the multiple locations in the convolution kernel, based on each respective first value stored using the respective memory location. In some cases, the cumulative value is updated by using each respective first value stored using the respective memory location, performing each multiplier-accumulator (MAC) operation using a compute unit (CU) for each respective memory location, and providing the output of each respective MAC operation for each respective memory location to the accumulator buffer associated with the CU. For example, each cumulative value stored in each of the accumulators 450(1) to 450(4) in Figure 4 may be updated in each of the convolution cycles 430(1) to 430(3) using the output of one of the multipliers 440(1) to 440(4) based on each respective first value. In another embodiment, each accumulated value stored in each of the accumulators 550(1) to 550(4) in Figure 5 may be updated in each of the convolution cycles 530(1) to 530(3) using the output of one of the multipliers 540(1) to 540(4) based on each of the respective first values.

[0097] In block 708, process 700 includes obtaining a plurality of second values ​​enclosed by a convolution kernel at each of the plurality of positions, the plurality of second values ​​including a subset of each first value and additional second values ​​not included in each first value. For example, in Figure 4, the plurality of positions of the 3×1 convolution kernel along the row of input data 410 containing input [0 1 2 3 4 5] may include a first position enclosing

[0012] , a second position enclosing

[0123] , a third position enclosing

[0234] , and a fourth position enclosing

[0345] .

[0098] Multiple first values ​​can be obtained for four positions in the 3×1 convolution kernel in Figure 4 during the first convolution cycle 430(1). The first first value is 0, the second first value is 1, the third first value is 2, and the fourth first value is 3.

[0099] Multiple second values ​​can be obtained for four positions in the 3×1 convolution kernel in Figure 4 during the second convolution cycle 430(2). The first second value is 1, the second second value is 2, the third second value is 3, and the fourth second value is 4.

[0100] Multiple second values ​​(e.g., 1, 2, 3, 4) include a subset of the first value (e.g., 0, 1, 2, 3) (1, 2, 3) and additional second values ​​(e.g., 4) that are not included in the multiple first values.

[0101] In some cases, the convolution output may be determined for each of multiple positions in the convolution kernel based on multiple convolution cycles, where each of the multiple convolution cycles corresponds to a different location within the convolution kernel. For example, the first convolution cycle corresponds to each first value enclosed by the convolution kernel at each position, and each first value corresponds to a first location within the convolution kernel. The first convolution cycle may be the same as or similar to the first convolution cycle 430(1) in Figure 4, which corresponds to each of the first values ​​0, 1, 2, and 3 enclosed by the 3×1 convolution kernel in Figure 4 at each of the four positions. Each of the first values ​​0, 1, 2, and 3 can correspond to the leftmost location within the convolution kernel at each of the four positions.

[0102] The second convolution cycle corresponds to a plurality of second values ​​enclosed by the convolution kernel at each position, each of which second value corresponds to a second location in the convolution kernel, and the second location is adjacent to the first location. For example, the second convolution cycle may be the same as or similar to the second convolution cycle 430(2) in Figure 4, which corresponds to each of the second values ​​1, 2, 3, and 4 enclosed by the 3×1 convolution kernel in Figure 4 at each of the four positions. Each of the second values ​​1, 2, 3, and 4 can correspond to the central location of the 3×1 convolution kernel in Figure 4 (for example, adjacent to the first location, which is the leftmost location of the exemplary 3×1 convolution kernel in Figure 4).

[0103] In some embodiments, the convolution output for each of multiple positions of the convolution kernel along a row of image data is associated with a different multiplier calculation unit (CU) in each convolution cycle of multiple convolution cycles. In some cases, the convolution output for a first of multiple positions is associated with a first multiplier CU in the first convolution cycle of multiple convolution cycles, which is associated with a first memory location used to store each first value for the first position of the convolution kernel, and a second multiplier CU in the second convolution cycle of multiple convolution cycles, which is associated with a second memory location used to store each second value that is part of a subset of the first value.

[0104] For example, the first position of the exemplary 3×1 convolution kernel in Figure 4 is associated with multiplier CU440(1) in cycle 430(1), with multiplier CU440(2) in cycle 430(2), and with multiplier CU440(3) in cycle 430(3). The second position of the exemplary 3×1 convolution kernel in Figure 4 is associated with multiplier CU440(2) in cycle 430(1), with multiplier CU440(3) in cycle 430(2), and with multiplier CU440(4) in cycle 430(3).

[0105] In some embodiments, the output of the first multiplier CU is stored in the first accumulator buffer for each convolution cycle of a plurality of convolution cycles, and the output of the second multiplier CU is stored in the second accumulator buffer for each convolution cycle of a plurality of convolution cycles. For example, the outputs of the multipliers CU540(1) to 540(4) in Figure 5 are stored in their respective accumulator buffers 550(1) to 550(4) for each of the convolution cycles 530(1) to 530(3). For example, the output of the first multiplier CU540(1) is stored in the first accumulator buffer 550(1) for each of the convolution cycles 530(1) to 530(3), and so on.

[0106] In some cases, the cumulative value stored in the first accumulator buffer at the end of each convolution cycle replaces the cumulative value stored in the second accumulator buffer at the end of each convolution cycle. For example, the cumulative value stored in the first accumulator buffer 550(1) replaces the cumulative value stored in the second accumulator buffer 550(2) at the end of the first convolution cycle 530(1), at the end of the second convolution cycle 530(2), and at the end of the third convolution cycle 530(3). In some embodiments, the cumulative value stored in the second accumulator buffer at the end of each convolution cycle replaces the cumulative value stored in the third accumulator buffer at the end of each convolution cycle. For example, the accumulated value stored in the second accumulator buffer 550(2) replaces the value stored in the third accumulator buffer 550(3) at the end of the first convolution cycle 530(1), at the end of the second convolution cycle 530(2), and at the end of the third convolution cycle 530(3).

[0107] In some embodiments, during each convolution cycle of multiple convolution cycles, each accumulator buffer corresponding to a plurality of multiplier CUs receives a first input from a corresponding memory location, which represents the pixel value in a row of image data, and a second input from an adjacent accumulator buffer among the plurality of accumulator buffers, which represents the accumulated value stored in the adjacent accumulator buffer. In some embodiments, each accumulator buffer receives the first input at the beginning of each convolution cycle and the second input at the end of each convolution cycle.

[0108] For example, during the first convolution cycle 530(1) in Figure 5, the first accumulator buffer 550(1) receives a first input from the multiplier CU540(1) corresponding to the pixel value 0. At the end of the first convolution cycle 530(1), the first accumulator buffer 550(1) receives a second input from the adjacent accumulator buffer 550(4). In another embodiment, during the first convolution cycle 530(1) in Figure 5, the second accumulator buffer 550(2) receives a first input from the multiplier CU550(2) corresponding to the pixel value 1. At the end of the first convolution cycle 530(1), the second accumulator buffer 550(2) receives a second input from the adjacent accumulator buffer 550(1).

[0109] In block 710, process 700 updates a memory location used to store a first value that is not included in a plurality of second values, with an additional second value, so that the memory location stores an additional second value. For example, in the first convolution cycle 430(1) in Figure 4, the memory location is used to store the input data value "0". The first convolution cycle 430(1) is associated with each of the first values ​​0, 1, 2, and 3 (for example, as described above). The second convolution cycle 430(2) is associated with each of the second values ​​1, 2, 3, and 4 (for example, as described above). The input data value "0" is a first value that is not included in a plurality of second values, and the input data value "4" is an additional second value that is not included in the plurality of first values.

[0110] At the end of the first convolution cycle 430(1), the memory location used to store the first value "0" which is not included in the second set of multiple values ​​may be replaced with an additional second value "4" which will be used for the second convolution cycle 430(2). The updated memory location then stores the value "4" at the start of the second convolution cycle 430(2). The memory locations used to store the respective values ​​1, 2, and 3 which are included in both the first and second sets of multiple values ​​remain unchanged between the first and second convolution cycles 430(1) and 430(2). After the second convolution cycle 430(2), the memory location used to store the value "1" is updated to store an additional value "5" which will be newly used by the third convolution cycle 430(3). The memory locations used to store values ​​2, 3, and 4 remain unchanged between the second convolution cycle 430(2) and the third convolution cycle 430(3).

[0111] In some embodiments, the process 700 includes updating a first accumulator buffer at the end of each convolution cycle of a plurality of convolution cycles, corresponding to the convolution output for a first location among a plurality of locations of the convolution kernel. For example, the first accumulator buffer may be updated based on the output of a first multiplier CU in the first convolution cycle of the plurality of convolution cycles. The first accumulator buffer may be further updated based on the output of a second multiplier CU in the second convolution cycle of the plurality of convolution cycles. In some cases, the accumulated value of the first accumulator buffer may be updated in each convolution cycle of the plurality of convolution cycles based on the respective outputs of different multiplier CUs. For example, the accumulated value of the first accumulator buffer may be updated for each of a plurality of locations within the first location of the convolution kernel along the rows of image data, based on the respective outputs of different multiplier CUs. In some cases, each convolution cycle of the plurality of convolution cycles corresponds to each of a plurality of locations within the first location of the convolution kernel along the rows of image data.

[0112] In some embodiments, the processes described herein (e.g., process 700, and / or any other processes described herein) may be carried out by a computing device, apparatus, or system. In one embodiment, process 700 may be carried out by a computing device or system having the computing device architecture 1000 of Figure 10. The computing device, apparatus, or system may include any suitable device, such as a mobile device (e.g., a mobile phone), a desktop computing device, a tablet computing device, a wearable device (e.g., a VR headset, an AR headset, AR glasses, a network-connected watch or smartwatch, or other wearable device), a server computer, an autonomous vehicle or computing device for an autonomous vehicle, a robotic device, a laptop computer, a smart TV, a camera, and / or any other computing device having the resource capacity to carry out the processes described herein, including process 700 and / or any other processes described herein. In some cases, a computing device or apparatus may include a variety of components, such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors, and / or other components (one or more) configured to perform steps of the process described herein. In some embodiments, a computing device may include a display, a network interface configured to communicate and / or receive data, any combination thereof, and / or other components (one or more). The network interface may be configured to communicate and / or receive Internet Protocol (IP) based data or other types of data.

[0113] Components of a computing device can be implemented in a circuit configuration. For example, a component may include and / or be implemented using one or more programmable electronic circuits (e.g., microprocessors, graphics processing units (GPUs), digital signal processors (DSPs), central processing units (CPUs), and / or other suitable electronic circuits), electronic circuits, or other electronic hardware, and / or may include and / or be implemented using computer software, firmware, or any combination thereof for performing the various operations described herein.

[0114] Process 700 is presented as a logical flow diagram, and its operation represents a set of actions that can be implemented in hardware, computer instructions, or a combination thereof. In the context of computer instructions, an operation represents a computer-executable instruction stored on one or more computer-readable storage media that, when executed by one or more processors, performs the described action. Generally, computer-executable instructions include routines, programs, objects, components, data structures, etc., that perform a particular function or realize a particular data type. The order in which the operations are described is not intended to be interpreted as limiting, and any number of the described operations can be combined in any order and / or in parallel to implement the process.

[0115] Furthermore, process 700 and / or any other processes described herein may be carried out under the control of one or more computer systems configured with executable instructions, and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that is executed collectively, by hardware, or in combination thereof on one or more processors. As described above, the code may be stored on a computer-readable or machine-readable storage medium in the form of a computer program, for example, containing multiple instructions that can be executed by one or more processors. The computer-readable or machine-readable storage medium may be non-temporary.

[0116] Figure 8 shows an exemplary embodiment of a deep learning neural network 800. The input layer 820 contains input data. In some cases, the input layer 820 may contain data representing pixels of an input video frame. The neural network 800 includes several hidden layers 822a, 822b through 822n. The hidden layers 822a, 822b through 822n contain a number of "n" hidden layers, where "n" is an integer greater than or equal to 1. The number of hidden layers can be the same as the number of layers required for a given application. The neural network 800 further includes an output layer 824 that provides an output obtained as a result of the processing performed by the hidden layers 822a, 822b through 822n. In an exemplary embodiment, the output layer 824 may provide a classification for objects in the input video frame. The classification may include a class that identifies the type of object (e.g., person, dog, cat, or other object).

[0117] Neural Network 800 is a multilayer neural network of interconnected nodes. Each node can represent one piece of information. The information associated with a node is shared between different layers, and each layer holds the information as it is processed. In some cases, Neural Network 800 may include a feedforward network, in which case there are no feedback connections where the output of the network is fed back to the network itself. In some cases, Neural Network 800 may include a recurrent neural network, which may have loops that allow information to be carried between nodes while being read at the input.

[0118] Information can be exchanged between nodes through node-to-node interconnections between various layers. Nodes in input layer 820 can activate a set of nodes in the first hidden layer 822a. For example, as shown in the figure, each input node in input layer 820 is connected to each node in the first hidden layer 822a. Nodes in hidden layers 822a, 822b through 822n can transform the information of each input node by applying an activation function to the information. The information derived from this transformation can then be passed to the nodes of the next hidden layer 822b, and those nodes can be activated, allowing them to perform their own specified functions. Illustrative functions include convolution, upsampling, data transformation, and / or any other suitable function. The output of hidden layer 822b can then activate the nodes of the next hidden layer, and so on. The output of the last hidden layer 822n can activate one or more nodes in output layer 824, providing outputs at those nodes. In some cases, a node in neural network 800 (e.g., node 826) is shown as having multiple output lines, but the node actually has a single output, and all lines shown as outputs from the node represent the same output value.

[0119] In some cases, each node, or the interconnections between nodes, may have weights, which are a set of parameters derived from training the neural network 800. Once the neural network 800 is trained, it can be called a trained neural network and can be used to classify one or more objects. For example, the interconnections between nodes may represent a piece of information learned about the nodes they interconnect. By having tunable numerical weights that can be adjusted (for example, based on the training dataset), the neural network 800 can be adaptive to inputs and can learn more as more data is processed.

[0120] The neural network 800 is pre-trained to process features from the data in the input layer 820 using various hidden layers 822a, 822b through 822n to provide an output via the output layer 824. In one embodiment where the neural network 800 is used to identify objects in an image, the neural network 800 can be trained using training data that includes both images and labels. For example, training images can be input into the network, in which case each training image has a label that indicates the class of one or more objects in each image (basically, it tells the network what the objects are and what features those objects have). In some embodiments, the training images may include a number 2 images, in which case the label for that image may be [0 0 1 0 0 0 0 0 0 0].

[0121] In some cases, the neural network 800 can adjust the node weights using a training process called backpropagation. Backpropagation may include a forward pass, loss function, backward pass, and weight updates. The forward pass, loss function, backward pass, and parameter updates are performed for one training iteration. This process can be repeated for each set of training images over a certain number of iterations until the neural network 800 is sufficiently trained so that the layer weights are precisely adjusted.

[0122] In embodiments for identifying objects within an image, the forward pass may include passing the training image through the neural network 800. Before the neural network 800 is trained, the weights are first randomized. The image may include, for example, an array of numbers representing the pixels of the image. Each number in the array may contain a value between 0 and 255, describing the pixel intensity at its position in the array. In some embodiments, the array may include a 28 × 28 × 3 array of numbers with 28 rows and 28 columns of pixels and three color components (such as a red component, a green component, and a blue component, or a luminance component and two saturation components).

[0123] For the first training iterations of Neural Network 800, the output is likely to contain values ​​that do not favor any particular class, due to the random selection of weights during initialization. For example, if the output is a vector of probabilities that an object belongs to various classes, the probability values ​​for each of the various classes may be equal or at least very similar (for example, for 10 possible classes, each class may have a probability value of 0.1). With these initial weights, Neural Network 800 is unable to determine low-level features and therefore cannot accurately determine what the object's classification may be. A loss function can be used to analyze the error in the output. Any suitable loss function definition can be used. One example of a loss function is the Mean Squared Error (MSE). MSE is,

[0124]

number

[0125] For the initial training images, the actual values ​​will differ significantly from the predicted output, resulting in a large loss (or error). The goal of training is to minimize the amount of loss so that the predicted output matches the training label. The neural network 800 can perform a reverse pass by determining which input (weight) contributed the most to the network's loss, and then adjust the weights so that the loss decreases and is eventually minimized.

[0126] To determine the weight that contributed most to the network loss, the derivative of the loss with respect to the weight (expressed as dL / dW, where W is the weight in a particular layer) can be calculated. After the derivative is calculated, weight updating can be performed by updating all the weights of the filter. For example, the weights can be updated so that they change in the opposite direction of the gradient. Weight updating is,

[0127]

number

[0128] The neural network 800 may include any suitable deep network. One example is a convolutional neural network (CNN) which includes an input layer and an output layer, with multiple hidden layers between the input and output layers. An example of a CNN is described below with reference to Figure 9. The hidden layers of the CNN include a series of convolutional layers, a nonlinear layer, a pooling layer (for downsampling), and a fully connected layer. The neural network 800 may include any other deep networks besides CNNs, such as autoencoders, deep belief networks (DBNs), and recurrent neural networks (RNNs).

[0129] Figure 9 shows an exemplary embodiment of a convolutional neural network 900 (CNN900). The input layer 920 of the CNN900 contains data representing an image. For example, this data may contain an array of numbers representing pixels in the image, where each number in the array contains a value from 0 to 255 that describes the pixel intensity at its position in the array. Using the previous embodiment from above, the array may contain a 28 × 28 × 3 array of numbers with 28 rows and 28 columns of pixels and three color components (e.g., red, green, and blue components, or luminance and two saturation components). The output can be obtained in the output layer 924 by passing the image through a convolutional hidden layer 922a, an optional nonlinear activation layer, a pooling hidden layer 922b, and a fully connected hidden layer 922c. Although only one of each hidden layer is shown in Figure 9, those skilled in the art will understand that multiple convolutional hidden layers, nonlinear layers, pooling hidden layers, and / or fully connected layers may be included in CNN900. As mentioned above, the output may represent objects of a single class or may include probabilities of the class that best describes the objects in the image.

[0130] The first layer of CNN900 is the convolutional hidden layer 922a. The convolutional hidden layer 922a analyzes the image data of the input layer 920. Each node in the convolutional hidden layer 922a is connected to a region of nodes (pixels) in the input image, called a receptive field. The convolutional hidden layer 922a can be thought of as one or more filters (each filter corresponding to a different activation map or feature map), and each convolutional iteration of a filter is a node or neuron in the convolutional hidden layer 922a. For example, the region of the input image covered by the filter in each convolutional iteration becomes the receptive field for that filter. In some embodiments, if the input image contains a 28x28 array and each filter (and its corresponding receptive field) is a 5x5 array, then there will be 24x24 nodes in the convolutional hidden layer 922a. Each connection between a node and its receptive field learns weights and, in some cases, an overall bias, so that each node learns to analyze its specific local receptive field within the input image. Each node in the hidden layer 922a will have the same weights and biases (called shared weights and shared biases). For example, a filter has an array of weights (numerical values) and the same depth as the input. The filter has a depth of 3 (according to the three color components of the input image) in the example of a video frame. The size of an exemplary embodiment of the filter array is 5 × 5 × 3, corresponding to the size of the node's receptive field.

[0131] The convolutional properties of the convolutional hidden layer 922a stem from the fact that each node in the convolutional layer is applied to its corresponding receptive field. For example, a filter in the convolutional hidden layer 922a can start in the upper-left corner of the input image array and convolve around that input image. As described above, each convolutional iteration of the filter can be considered a node or neuron in the convolutional hidden layer 922a. In each convolutional iteration, the filter value is multiplied by the corresponding numerical value of the original pixel value of the image (for example, a 5x5 filter array is multiplied by a 5x5 array of input pixel values ​​at the upper-left corner of the input image array). The sum of the multiplications from each convolutional iteration can be obtained for that iteration or node. This process then continues at the next location in the input image, according to the receptive field of the next node in the convolutional hidden layer 922a.

[0132] For example, the filter can be moved to the next receptive field by a certain step amount. The step amount can be set to 1 or another suitable amount. For example, if the step amount is set to 1, the filter will be moved one pixel to the right in each convolution iteration. By processing the filter at each unique location in the input volume, a numerical value representing the filtered result for that location is created, and as a result, a summation value is determined for each node of the convolutional hidden layer 922a.

[0133] The mapping from the input layer to the convolutional hidden layer 922a is called an activation map (or feature map). The activation map contains values ​​for each node that represent the filtering result at each location in the input volume. The activation map may contain an array containing various summations resulting from each iteration of the filter on the input volume. For example, if a 5x5 filter is applied to each pixel of a 28x28 input image (with a step amount of 1), the activation map will contain a 24x24 array. The convolutional hidden layer 922a may contain several activation maps to identify multiple features in the image. The embodiment shown in Figure 9 includes three activation maps. Using the three activation maps, the convolutional hidden layer 922a can detect three different types of features, each of which is detectable throughout the image.

[0134] In some embodiments, a nonlinear hidden layer can be applied after the convolutional hidden layer 922a. The nonlinear layer can be used to introduce nonlinearity into a system that was previously performing linear operations. An exemplary embodiment of a nonlinear layer is a rectified linear unit (ReLU) layer. The ReLU layer can apply the function f(x)=max(0,x) to all values ​​in the input volume, thereby changing all negative activations to 0. Thus, ReLU can enhance the nonlinear properties of CNN900 without affecting the receptive field of the convolutional hidden layer 922a.

[0135] A pooling hidden layer 922b can be applied after the convolutional hidden layer 922a (and, if used, after the nonlinear hidden layer). The pooling hidden layer 922b is used to simplify the information in the output from the convolutional hidden layer 922a. For example, the pooling hidden layer 922b can take each activation map output from the convolutional hidden layer 922a and use a pooling function to generate a condensed activation map (or feature map). Maximum pooling is one example of a function performed by the pooling hidden layer. Other forms of pooling functions, such as average pooling, L2 norm pooling, or other preferred pooling functions, are used by the pooling hidden layer 922a. A pooling function (e.g., a maximum pooling filter, an L2 norm filter, or other preferred pooling filter) is applied to each activation map contained within the convolutional hidden layer 922a. In the embodiment shown in Figure 9, three pooling filters are used for the three activation maps in the convolutional hidden layer 922a.

[0136] In some embodiments, max pooling can be used by applying a max pooling filter (e.g., having a size of 2x2) to the activation map output from the convolutional hidden layer 922a with a step amount equal to the filter's dimension (e.g., a step amount of 2). The output from the max pooling filter contains the largest number in any sub-region that the filter surrounds and convolves. Using a 2x2 filter as an embodiment, each unit in the pooling layer can summarize a region of 2x2 nodes (each node being a value in the activation map) in the previous layer. For example, four values ​​(nodes) in the activation map will be analyzed by the 2x2 max pooling filter in each iteration of the filter, and the largest value from those four values ​​will be output as the "maximum" value. If such a max pooling filter is applied to an activation filter from the convolutional hidden layer 922a having a dimension of 24x24 nodes, the output from the pooling hidden layer 922b will be a 12x12 node array.

[0137] In some embodiments, an L2 norm pooling filter can also be used. The L2 norm pooling filter involves calculating the square root of the sum of squares of values ​​within a 2x2 region (or other preferred region) of the activation map (rather than calculating the maximum value as is done in max pooling), and using the calculated value as the output.

[0138] Intuitively, a pooling function (e.g., max pooling function, L2 norm pooling function, or other pooling functions) determines whether a given feature is found at any location within a region of an image, discarding exact location information. This can be done without affecting the feature detection results because, once a feature is found, its exact location is less important than its approximate location relative to other features. Max pooling (and other pooling methods) have the advantage of pooling far fewer features, thus reducing the number of parameters required in later layers of CNN900.

[0139] The final layer of connection in the network is a fully connected layer that connects every node from the pooling hidden layer 922b to each of the output nodes in the output layer 924. Using the example above, the input layer contains 28×28 nodes encoding the pixel intensity of the input image, the convolutional hidden layer 922a contains 3×24×24 hidden feature nodes based on the application of 5×5 local receptive fields (for filtering) to three activation maps, and the pooling layer 922b contains a layer of 3×12×12 hidden feature nodes based on the application of a max pooling filter to 2×2 regions across each of the three feature maps. Extending this embodiment, the output layer 924 may contain 10 output nodes. In such an embodiment, every node of the 3×12×12 pooling hidden layer 922b is connected to every node of the output layer 924.

[0140] The fully connected layer 922c can take the output of the preceding pooling layer 922b (which should represent the activation map of high-level features) and determine the features that are most correlated to a particular class. For example, the layer of the fully connected layer 922c can determine the high-level features that are most strongly correlated to a particular class and may contain weights (nodes) for those high-level features. By calculating the product between the weights of the fully connected layer 922c and the weights of the pooling hidden layer 922b, the probabilities for various classes can be obtained. For example, if CNN900 is used to predict that an object in a video frame is a person, there will be high values ​​in the activation map representing high-level features of a person (e.g., having two legs, having a face on top of the object, having two eyes on the upper left and upper right of the face, having a nose in the center of the face, having a mouth at the bottom of the face, and / or other features common to people).

[0141] In some embodiments, the output from output layer 924 may include an M-dimensional vector (M=10 in the preceding embodiments), where M may represent the number of classes the program must select when classifying objects in the image. Other exemplary outputs can also be provided. Each number in the N-dimensional vector may represent the probability that an object belongs to a particular class. In some cases, if a 10-dimensional output vector representing 10 different object classes is [0 0 0.05 0.8 0 0.15 0 0 0 0], then the vector indicates that there is a 5% probability that the image is an object of the third class (e.g., a dog), an 80% probability that the image is an object of the fourth class (e.g., a human), and a 15% probability that the image is an object of the sixth class (e.g., a kangaroo). The probability for a class can be considered a level of confidence that an object belongs to that class.

[0142] Figure 10 shows an exemplary computing device architecture 1000 of an exemplary computing device capable of implementing various techniques described herein. In some embodiments, the computing device may include a mobile device, a wearable device, an extended reality device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a personal computer, a laptop computer, a video server, a vehicle (or a computing device in a vehicle), or other devices. The components of the computing device architecture 1000 are shown to communicate electrically with each other using a connection 1005 such as a bus. The exemplary computing device architecture 1000 includes a processing unit (CPU or processor) 1010 and a computing device connection 1005 that connects various computing device components to the processor 1010, including computing device memory 1015 such as read-only memory (ROM) 1020 and random-access memory (RAM) 1025.

[0143] The computing device architecture 1000 may include a cache of high-speed memory directly connected to, adjacent to, or integrated as part of the processor 1010. The computing device architecture 1000 may copy data from memory 1015 and / or storage device 1030 to cache 1012 for rapid access by the processor 1010. In this way, the cache can provide a performance improvement that avoids delays in the processor 1010 while waiting for data. These and other engines may control or be configured to control the processor 1010 to perform various actions. Other computing device memories 1015 may also be available for use. Memory 1015 may include multiple different types of memory with different performance characteristics. The processor 1010 may include an arbitrary general-purpose processor and a dedicated processor with software instructions incorporated into the processor design, along with hardware or software services such as service 1 1032, service 2 1034, and service 3 1036 stored in storage device 1030, configured to control the processor 1010. The processor 1010 may be a self-contained system including multiple cores or processors, a bus, a memory controller, a cache, etc. The multicore processor may be symmetrical or asymmetrical.

[0144] To enable user interaction with the computing device architecture 1000, the input device 1045 can represent any number of input mechanisms, such as a microphone for speech, a touch-sensitive screen for gesture or graphical input, a keyboard, a mouse, motion input, or speech. The output device 1035 can also be one or more of some output mechanisms known to those skilled in the art, such as a display, projector, television, or speaker device. In some cases, the multimodal computing device can allow the user to provide multiple types of inputs for communication with the computing device architecture 1000. The communication interface 1040 can generally control and manage user input and computing device output. There are no restrictions on operating on any particular hardware configuration, and therefore, as improved hardware or firmware configurations are developed, the basic functions here can be easily replaced with them.

[0145] The storage device 1030 is a non-volatile memory and may be a hard disk or other type of computer-readable medium capable of storing computer-accessible data, such as a magnetic cassette, flash memory card, solid-state memory device, digital versatile disk, cartridge, random access memory (RAMs) 1025, read-only memory (ROM) 1020, and hybrids thereof. The storage device 1030 may include services 1032, 1034, and 1036 for controlling the processor 1010. Other hardware or software modules or engines are also conceivable. The storage device 1030 may be connected to the computing device connection 1005. In some embodiments, a hardware module performing a particular function may include software components stored on a computer-readable medium in relation to necessary hardware components such as the processor 1010, the connection 1005, and the output device 1035 in order to perform that function.

[0146] Aspects of this disclosure are applicable to any suitable electronic device (such as a security system, smartphone, tablet, laptop computer, vehicle, drone, or other device) that includes or is coupled with one or more active depth sensing systems. While the following describes devices having or being coupled with one optical projector, aspects of this disclosure are applicable to devices having any number of optical projectors and are therefore not limited to any particular device.

[0147] The term “device” is not limited to one or a specific number of physical objects (such as one smartphone, one controller, or one processing system). As used herein, a device may be any electronic device having one or more parts that can implement at least some parts of this disclosure. The following descriptions and examples use the term “device” to describe various aspects of this disclosure, but the term “device” is not limited to a specific configuration, type, or number of objects. Additionally, the term “system” is not limited to multiple components or a specific configuration. For example, a system may be mounted on one or more printed circuit boards or other boards and may have movable or static components. The following descriptions and examples use the term “system” to describe various aspects of this disclosure, but the term “system” is not limited to a specific configuration, type, or number of objects.

[0148] Specific details are provided in the above specification to provide a complete understanding of the embodiments and examples provided herein. However, those skilled in the art will understand that the embodiments can be practiced without these specific details. For the sake of clarity, in some cases the technology may be presented as including individual functional blocks, which include devices, device components, steps or routines in the way they are embodied in software, or combinations of hardware and software. Additional components other than those shown in the figures and / or described herein may also be used. For example, circuits, systems, networks, processes, and other components may be shown as components in the form of block diagrams, so as not to obscure the embodiments with unnecessary details. In other cases, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary details, so as not to obscure the embodiments.

[0149] Individual aspects may be described above as processes or methods represented as flowcharts, flow diagrams, data flow diagrams, structural diagrams, or block diagrams. While flowcharts can describe operations as sequential processes, many of these operations can also be performed in parallel or simultaneously. Furthermore, the order of operations can be rearranged. A process terminates when its operations are complete, but it may have additional steps not shown in the diagram. A process can correspond to a method, function, procedure, subroutine, subprogram, etc. If a process corresponds to a function, its termination may correspond to the function returning to a calling function or the main function.

[0150] The processes and methods described in the above examples can be implemented using computer-executable instructions stored in or available from computer-readable media. Such instructions may include instructions and data that cause, for example, a general-purpose computer, a dedicated computer, or a processing device to perform a certain function or group of functions, or otherwise configure them to perform them. The portion of computer resources used may be accessible over a network. Computer-executable instructions may be, for example, binary, or intermediate format instructions such as assembly language, firmware, or source code.

[0151] The term “computer-readable medium” includes, but is not limited to, portable or nonportable storage devices, optical storage devices, and various other media capable of storing, containing, or transporting instructions (one or more) and / or data. Computer-readable medium may include non-temporary media capable of storing data and that do not contain carrier waves and / or transient electronic signals propagating wirelessly or via wired connections. Examples of non-temporary media include, but are not limited to, magnetic disks or magnetic tapes, optical storage media such as flash memory, memory or memory devices, USB devices provided using magnetic disks or optical disks, flash memory, non-volatile memory, networked storage devices, compact disks (CDs) or digital versatile disks (DVDs), and any suitable combination thereof. Computer-readable medium may have code and / or machine-executable instructions stored on the computer-readable medium that may represent procedures, functions, subprograms, programs, routines, subroutines, modules, engines, software packages, classes, or any combination of instructions, data structures, or program statements. Code segments can be coupled to other code segments or hardware circuits by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc., can be passed, forwarded, or transmitted via any preferred means, including memory sharing, message passing, token passing, or network transmission.

[0152] In some embodiments, computer-readable storage devices, media, and memory may include cable or wireless signals, such as bitstreams. However, non-transient computer-readable storage media, as referred to, explicitly exclude media such as energy, carrier signals, electromagnetic waves, and signals themselves.

[0153] Devices implementing processes and methods in accordance with these disclosures may include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and may take on any of a variety of form factors. When implemented in software, firmware, middleware, or microcode, the program code or code segment (e.g., a computer program product) that performs the required task may be stored in a computer-readable or machine-readable medium. A processor(s) may perform the required task. Typical examples of form factors include laptops, smartphones, mobile phones, tablet devices or other small form factor personal computers, digital assistants, rack-mount devices, and standalone devices. The functionalities described herein may also be embodied in peripherals or add-in cards. Such functionalities may also, as a further example, be implemented between different chips on a circuit board or between different processes running within a single device.

[0154] Instructions, a medium for transmitting such instructions, computing resources for executing those instructions, and other structures supporting such computing resources are exemplary means for providing the functions described in this disclosure.

[0155] While the embodiments of this application are described in the above description with reference to those specific embodiments, those skilled in the art will recognize that this application is not limited thereto. Therefore, while exemplary embodiments of this application are described in detail herein, it should be understood that the concepts of the present invention can be embodied and employed in various other ways, and that, unless limited by the prior art, the appended claims are intended to be interpreted as including such variations. The various features and embodiments of this application described above may be used individually or in combination. Furthermore, embodiments may be used in any number of environments and applications other than those described herein without departing from the broader spirit and scope of this specification. Therefore, this specification and the drawings should be considered illustrative and not limiting. For illustrative purposes, the methods are described in a particular order. It should be understood that in alternative embodiments, the methods may be carried out in an order different from the order described.

[0156] Those skilled in the art will understand that the symbols or terms less than ("<") and greater than (">") used herein may be replaced with the symbols less than or equal to ("≦") and greater than or equal to ("≧"), respectively, without departing from the scope of this specification.

[0157] Where a component is described as “configured to” perform a particular operation, such configuration can be achieved, for example, by designing an electronic circuit or other hardware to perform that operation, by programming a programmable electronic circuit (e.g., a microprocessor or other suitable electronic circuit) to perform that operation, or by any combination thereof.

[0158] The phrase "connected to" refers to any component that is physically connected to another component, either directly or indirectly, and / or communicates with another component, either directly or indirectly (for example, connected to another component via a wired or wireless connection and / or other preferred communication interface).

[0159] The various exemplary logic blocks, modules, engines, circuits, and algorithmic steps described in relation to the embodiments disclosed herein may be implemented as electronic hardware, computer software, firmware, or a combination thereof. To clearly illustrate this hardware and software compatibility, various exemplary components, blocks, modules, engines, circuits, and steps are described above in general terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. A person skilled in the art may implement the described functionality in various ways for each specific application, but such a decision on implementation should not be construed as a cause for departure from the scope of this application.

[0160] The techniques described herein can also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques can be implemented in any of a variety of devices, such as general-purpose computers, wireless communication device handsets, or integrated circuit devices with multiple applications, including applications in wireless communication device handsets and other devices. Any feature described as a module or component can be implemented together in an integrated logic device, or separately as individual but interoperable logic devices. When implemented in software, the technique can be at least partially realized by a computer-readable data storage medium having program code that, when executed, performs one or more of the methods described above. The computer-readable data storage medium may also form part of a computer program product, which may include packaging materials. The computer-readable medium may include memory or data storage media such as random access memory (RAM), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), FLASH memory, and magnetic or optical data storage media. These techniques can be implemented, at least in part, by computer-readable communication media, such as propagating signals or waves, which carry or communicate program code in the form of instructions or data structures that can be accessed, read, and / or executed by a computer.

[0161] The program code can be executed by a processor, which may include one or more processors such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Such processors may be configured to implement any of the techniques described herein. The general-purpose processor may be a microprocessor, but alternatively, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors working with a DSP core, or any other such configuration. Thus, as used herein, the term “processor” may refer to any of the structures described herein, any combination thereof, or any other structure or device suitable for implementing the techniques described herein.

[0162] The claim language or other language that states “at least one of” a set and / or “one or more” of a set indicates that one member of that set, or multiple members of that set (in any combination), satisfy the claim. For example, the claim language that states “at least one of A and B” or “at least one of A or B” means A, B, or A and B. In another example, the claim language that states “at least one of A, B, and C” or “at least one of A, B, or C” means A, B, C, or A and B, or A and C, or B and C, or A and B and C. The language that states “at least one of” a set and / or “one or more” of a set does not limit the set to the items listed within it. For example, the wording of a claim that states "at least one of A and B" or "at least one of A or B" could mean A, B, or A and B, and additionally, could include items not listed in the set of A and B.

[0163] The language of a claim that includes phrases such as "at least one processor configured to perform X, Y, and Z" or other similar phrases indicates that one or more processors (in any combination) can perform the associated operations (one or more). For example, the wording of a claim that includes "at least one processor configured to perform X, Y, and Z" means that a single processor can perform operations X, Y, and Z, or that each of several processors can task with a subset of operations X, Y, and Z, or that a group of several processors can work together to perform operations X, Y, and Z, so that several processors perform X, Y, and Z together. In another example, the wording of a claim that includes "at least one processor configured to perform X, Y, and Z" may mean that any single processor can perform only at least one subset of operations X, Y, and Z.

[0164] Exemplary embodiments of this disclosure include:

[0165] Embodiment 1. An apparatus for processing image data, comprising at least one memory and at least one processor coupled to the at least one memory, wherein the at least one processor is configured to: acquire each first value enclosed by the convolution kernel at each of a plurality of positions of the convolution kernel along a row of image data; store each first value using the respective memory location associated with each of the plurality of positions of the convolution kernel; update a cumulative value corresponding to the convolution output for each of the plurality of positions of the convolution kernel based on each first value stored using the respective memory location; acquire a plurality of second values ​​enclosed by the convolution kernel at each of the plurality of positions, which include a subset of each first value and additional second values ​​not included in each first value; and update the memory location used to store the first values ​​not included in the plurality of second values ​​with the additional second values ​​so that the memory location stores the additional second values.

[0166] Embodiment 2. The apparatus according to Embodiment 1, wherein, in order to update the cumulative value, at least one processor is configured to use each first value stored using each memory location, perform each multiplier-accumulator (MAC) operation using a compute unit (CU) for each respective memory location, and provide the output of each respective MAC operation for each respective memory location to an accumulator buffer associated with the CU.

[0167] Embodiment 3. The apparatus according to Embodiment 1 or 2, wherein at least one processor is configured to determine a convolution output for each of a plurality of locations in a convolution kernel based on a plurality of convolution cycles, and each of the plurality of convolution cycles corresponds to a different location in the convolution kernel.

[0168] Embodiment 4. The apparatus according to Embodiment 3, wherein a first convolution cycle corresponds to each first value enclosed by a convolution kernel, each first value corresponds to a first location within the convolution kernel, a second convolution cycle corresponds to a plurality of second values ​​enclosed by a convolution kernel, each second value of the plurality of second values ​​corresponds to a second location within the convolution kernel, and the second location is adjacent to the first location.

[0169] Embodiment 5. The apparatus according to Embodiment 3 or 4, wherein the convolution output for each of multiple positions of a convolution kernel along a row of image data is associated with a different multiplier calculation unit (CU) in each convolution cycle of multiple convolution cycles.

[0170] Embodiment 6. The apparatus according to Embodiment 5, wherein the convolution output for a first position among a plurality of positions is associated with a first multiplier CU in a first convolution cycle among a plurality of convolution cycles, which is associated with a first memory location used to store each first value for the first position of the convolution kernel, and a second multiplier CU in a second convolution cycle among a plurality of convolution cycles, which is associated with a second memory location used to store each second value included in a subset of the first values.

[0171] Embodiment 7. The apparatus according to Embodiment 6, wherein the output of the first multiplier CU is stored in the first accumulator buffer in each convolutional cycle of the multiple convolutional cycles, and the output of the second multiplier CU is stored in the second accumulator buffer in each convolutional cycle of the multiple convolutional cycles.

[0172] Embodiment 8. The apparatus according to Embodiment 7, wherein the cumulative value stored in the first accumulator buffer at the end of each convolution cycle replaces the cumulative value stored in the second accumulator buffer at the end of each convolution cycle, and the cumulative value stored in the second accumulator buffer at the end of each convolution cycle replaces the cumulative value stored in the third accumulator buffer at the end of each convolution cycle.

[0173] Apparatus 9. The apparatus according to Apparatus 7 or 8, wherein during each convolution cycle of a plurality of convolution cycles, each accumulator buffer corresponding to a plurality of multiplier CUs receives a first input from a corresponding memory location, which indicates a pixel value in a row of image data, and a second input from an adjacent accumulator buffer among the plurality of accumulator buffers, which indicates an accumulated value stored in the adjacent accumulator buffer.

[0174] Embodiment 10. The apparatus according to Embodiment 9, wherein each accumulator buffer receives a first input at the start of each convolution cycle and a second input at the end of each convolution cycle.

[0175] Embodiment 11. The apparatus according to any one of embodiments 6 to 10, wherein at least one processor is configured to update a cumulative value corresponding to the convolution output for a first position among a plurality of positions of the convolution kernel using a first accumulator buffer at the end of each convolution cycle of a plurality of convolution cycles.

[0176] Embodiment 12. The apparatus according to Embodiment 11, wherein at least one processor is configured to update a first accumulator buffer based on the output of a first multiplier CU in a first convolution cycle of a plurality of convolution cycles, and to update the first accumulator buffer based on the output of a second multiplier CU in a second convolution cycle of a plurality of convolution cycles.

[0177] Embodiment 13. The apparatus according to Embodiment 11 or 12, wherein at least one processor is configured to update the cumulative value of a first accumulator buffer based on the respective outputs of different multiplier CUs in each convolutional cycle of a plurality of convolutional cycles.

[0178] Embodiment 14. The apparatus according to Embodiment 13, wherein at least one processor is configured to update the cumulative value of a first accumulator buffer based on the respective outputs of different multiplier CUs for each of a plurality of locations within a first position of a convolution kernel along a row of image data.

[0179] Embodiment 15. The apparatus according to Embodiment 14, wherein each convolutional cycle of a plurality of convolutional cycles corresponds to each of a plurality of locations within a first position of the convolutional kernel along a row of image data.

[0180] Embodiment 16. A method for processing image data, comprising: obtaining a first value enclosed by a convolution kernel at each of a plurality of positions of a convolution kernel along a row of image data; storing each of the first values ​​using the respective memory location associated with each of the plurality of positions of the convolution kernel; updating a cumulative value corresponding to the convolution output for each of the plurality of positions of the convolution kernel based on each of the first values ​​stored using the respective memory location; obtaining a plurality of second values ​​enclosed by a convolution kernel at each of the plurality of positions, which include a subset of the first values ​​and additional second values ​​not included in the first values; and updating a memory location used to store the first values ​​not included in the plurality of second values ​​with the additional second values ​​so that the memory location stores the additional second values.

[0181] Embodiment 17. The method of Embodiment 16, wherein updating the cumulative value includes using each first value stored using each memory location, performing each multiplier-accumulator (MAC) operation using a compute unit (CU) for each respective memory location, and providing the output of each MAC operation for each respective memory location to an accumulator buffer associated with the CU.

[0182] Embodiment 18. The method according to Embodiment 16 or 17, further comprising determining a convolution output for each of a plurality of locations in a convolution kernel based on a plurality of convolution cycles, wherein each of the plurality of convolution cycles corresponds to a different location in the convolution kernel.

[0183] Embodiment 19. The method according to Embodiment 18, wherein a first convolution cycle corresponds to each first value enclosed by a convolution kernel, each first value corresponds to a first location within the convolution kernel, a second convolution cycle corresponds to a plurality of second values ​​enclosed by a convolution kernel, each second value of the plurality of second values ​​corresponds to a second location within the convolution kernel, and the second location is adjacent to the first location.

[0184] Embodiment 20. The method according to Embodiment 18 or 19, wherein the convolution output for each of multiple positions of a convolution kernel along a row of image data is associated with a different multiplier computation unit (CU) in each convolution cycle of multiple convolution cycles.

[0185] Embodiment 21. The method according to Embodiment 20, wherein the convolution output for a first position among a plurality of positions is associated with a first multiplier CU in a first convolution cycle among a plurality of convolution cycles, which is associated with a first memory location used to store each first value for the first position of the convolution kernel, and a second multiplier CU in a second convolution cycle among a plurality of convolution cycles, which is associated with a second memory location used to store each second value that is part of a subset of the first values.

[0186] Embodiment 22. The method according to Embodiment 21, wherein the output of the first multiplier CU is stored in the first accumulator buffer in each convolutional cycle of the multiple convolutional cycles, and the output of the second multiplier CU is stored in the second accumulator buffer in each convolutional cycle of the multiple convolutional cycles.

[0187] Embodiment 23. The method according to Embodiment 22, wherein the cumulative value stored in the first accumulator buffer at the end of each convolution cycle replaces the cumulative value stored in the second accumulator buffer at the end of each convolution cycle, and the cumulative value stored in the second accumulator buffer at the end of each convolution cycle replaces the cumulative value stored in the third accumulator buffer at the end of each convolution cycle.

[0188] Embodiment 24. The method according to Embodiment 22 or 23, wherein during each convolution cycle of a plurality of convolution cycles, each accumulator buffer corresponding to a plurality of multiplier CUs receives a first input from a corresponding memory location, which indicates a pixel value in a row of image data, and a second input from an adjacent accumulator buffer among the plurality of accumulator buffers, which indicates an accumulated value stored in the adjacent accumulator buffer.

[0189] Embodiment 25. The method according to Embodiment 24, wherein each accumulator buffer receives a first input at the start of each convolution cycle and a second input at the end of each convolution cycle.

[0190] Embodiment 26. The method according to any one of embodiments 21 to 25, further comprising updating a first accumulator buffer at the end of each convolution cycle of a plurality of convolution cycles to an accumulated value corresponding to the convolution output for a first position among a plurality of positions of the convolution kernel.

[0191] Embodiment 27. The method according to Embodiment 26, further comprising updating a first accumulator buffer based on the output of a first multiplier CU in a first convolution cycle among a plurality of convolution cycles, and updating the first accumulator buffer based on the output of a second multiplier CU in a second convolution cycle among a plurality of convolution cycles.

[0192] Embodiment 28. The method according to Embodiment 26 or 27, further comprising updating the cumulative value of a first accumulator buffer based on the respective outputs of different multiplier CUs in each convolutional cycle of a plurality of convolutional cycles.

[0193] Embodiment 29. The method according to Embodiment 28, further comprising updating the cumulative value of a first accumulator buffer based on the respective outputs of different multiplier CUs for each of a plurality of locations within a first position of a convolution kernel along a row of image data.

[0194] Embodiment 30. The method according to Embodiment 29, wherein each convolution cycle of a plurality of convolution cycles corresponds to each of a plurality of locations within a first position of the convolution kernel along a row of image data.

[0195] Embodiment 31. A non-temporary computer-readable storage medium containing stored instructions, wherein, when executed by at least one processor, the instructions cause at least one processor to perform the operation described in any of Embodiments 1 to 15.

[0196] Embodiment 32. A non-temporary computer-readable storage medium containing stored instructions, wherein, when executed by at least one processor, the instructions cause at least one processor to perform the operation described in any of Embodiments 16 to 30.

[0197] Embodiment 33. An apparatus comprising one or more means for performing the operation described in any of Embodiments 1 to 15.

[0198] Embodiment 34. An apparatus comprising one or more means for performing the operations described in any of Embodiments 16 to 30.

Claims

1. A device for processing image data, At least one memory, At least one processor coupled to the at least one memory, At each of the multiple positions of the convolution kernel along the row of the image data, obtain the respective first value enclosed by the convolution kernel, The first value is stored using the respective memory location associated with each of the multiple locations of the convolution kernel, Based on each of the first values ​​stored using each of the aforementioned memory locations, the cumulative value corresponding to the convolution output for each of the plurality of locations in the convolution kernel is updated. Obtaining a plurality of second values ​​enclosed by the convolution kernel at each of the plurality of positions, wherein the plurality of second values ​​include a subset of the respective first values ​​and additional second values ​​not included in the respective first values, Updating a memory location used to store a first value not included in the plurality of second values ​​with an additional second value, wherein the updated memory location stores the additional second value. A processor configured to perform the following: A device equipped with the following features.

2. In order to update the cumulative value, at least one processor, For each memory location, a calculation unit (CU) is used to perform a multiplier-accumulator (MAC) operation using the respective first value stored in the respective memory location, For each memory location, the output of each MAC operation is provided to the accumulator buffer associated with the CU. The apparatus according to claim 1, configured as follows.

3. The apparatus according to claim 1, wherein the at least one processor is configured to determine the convolution output for each of the plurality of locations in the convolution kernel based on a plurality of convolution cycles, and each of the plurality of convolution cycles corresponds to a different location in the convolution kernel.

4. The first convolution cycle corresponds to each of the first values ​​enclosed by the convolution kernel, and each of the first values ​​corresponds to a first location within the convolution kernel. A second convolution cycle corresponds to the plurality of second values ​​enclosed by the convolution kernel, each of the plurality of second values ​​corresponds to a second location in the convolution kernel, and the second location is adjacent to the first location. The apparatus according to claim 3.

5. The apparatus according to claim 3, wherein the convolution output for each of the plurality of positions of the convolution kernel along the row of the image data is associated with a different multiplier calculation unit (CU) in each of the plurality of convolution cycles.

6. The convolution output for the first position among the plurality of positions is A first multiplier CU in a first convolution cycle among the plurality of convolution cycles, the first multiplier CU associated with a first memory location used to store the respective first values ​​for the first position of the convolution kernel, A second multiplier CU in the second convolution cycle of the plurality of convolution cycles, associated with a second memory location used to store each of the second values ​​included in the subset of the first value, and The apparatus according to claim 5, associated with the

7. The apparatus according to claim 6, wherein the output of the first multiplier CU is stored in a first accumulator buffer in each of the plurality of convolution cycles, and the output of the second multiplier CU is stored in a second accumulator buffer in each of the plurality of convolution cycles.

8. At the end of each convolution cycle, the cumulative value stored in the first accumulator buffer replaces the cumulative value stored in the second accumulator buffer at the end of each convolution cycle. The cumulative value stored in the second accumulator buffer at the end of each convolution cycle replaces the cumulative value stored in the third accumulator buffer at the end of each convolution cycle. The apparatus according to claim 7.

9. During each of the aforementioned convolution cycles, each of the accumulator buffers corresponding to the multiple multiplier CUs is, A first input from a corresponding memory location, which indicates the pixel value within the row of the image data, A second input from an adjacent accumulator buffer among the plurality of accumulator buffers, which indicates the accumulated value stored in the adjacent accumulator buffer, and The apparatus according to claim 7, which receives

10. The apparatus according to claim 9, wherein each accumulator buffer receives the first input at the start of each convolution cycle and receives the second input at the end of each convolution cycle.

11. The aforementioned at least one processor, The apparatus according to claim 6, wherein at the end of each of the plurality of convolution cycles, a first accumulator buffer is used to update the accumulated value corresponding to the convolution output for the first of the plurality of positions of the convolution kernel.

12. The aforementioned at least one processor, Based on the output of the first multiplier CU in the first convolution cycle among the plurality of convolution cycles, the first accumulator buffer is updated. The first accumulator buffer is updated based on the output of the second multiplier CU in the second convolution cycle among the plurality of convolution cycles. The apparatus according to claim 11, configured as follows.

13. The aforementioned at least one processor, The apparatus according to claim 11, configured to update the cumulative value of the first accumulator buffer based on the respective outputs of different multiplier CUs in each of the plurality of convolution cycles.

14. The aforementioned at least one processor, The apparatus according to claim 13, configured to update the cumulative value of the first accumulator buffer based on the output of each of different multiplier CUs for each of the plurality of locations within the first position of the convolution kernel along the row of the image data.

15. The apparatus according to claim 14, wherein each of the plurality of convolution cycles corresponds to each of the plurality of locations in the first position of the convolution kernel along the row of the image data.

16. A method for processing image data, The steps include obtaining a first value enclosed by the convolution kernel at each of the multiple positions of the convolution kernel along the row of the image data, The steps include storing each of the first values ​​using the respective memory locations associated with each of the plurality of positions in the convolution kernel, The steps include updating the cumulative value corresponding to the convolution output for each of the plurality of locations of the convolution kernel based on each of the first values ​​stored using each of the memory locations, A step of obtaining a plurality of second values ​​enclosed by the convolution kernel at each of the plurality of locations, wherein the plurality of second values ​​include a subset of the respective first values ​​and additional second values ​​not included in the respective first values. A step of updating a memory location used to store a first value not included in the plurality of second values ​​with an additional second value, wherein the updated memory location stores the additional second value. Methods that include...

17. The step of updating the cumulative value is, For each memory location, the calculation unit (CU) is used to perform a multiplier-accumulator (MAC) operation using the respective first values ​​stored in the respective memory location, and the calculation is performed accordingly. For each memory location, the steps include providing the output of each MAC operation to the accumulator buffer associated with the CU. The method according to claim 16, including the method described in claim 16.

18. The method according to claim 16, further comprising the step of determining the convolution output for each of the multiple locations of the convolution kernel based on a plurality of convolution cycles, wherein each of the plurality of convolution cycles corresponds to a different location in the convolution kernel.

19. The first convolution cycle corresponds to each of the first values ​​enclosed by the convolution kernel, and each of the first values ​​corresponds to a first location within the convolution kernel. A second convolution cycle corresponds to the plurality of second values ​​enclosed by the convolution kernel, each of the plurality of second values ​​corresponds to a second location in the convolution kernel, and the second location is adjacent to the first location. The method according to claim 18.

20. The method according to claim 18, wherein the convolution output for each of the plurality of positions of the convolution kernel along the row of the image data is associated with a different multiplier calculation unit (CU) in each of the plurality of convolution cycles.

21. The convolution output for the first position among the plurality of positions is A first multiplier CU in a first convolution cycle among the plurality of convolution cycles, the first multiplier CU associated with a first memory location used to store the respective first values ​​for the first position of the convolution kernel, A second multiplier CU in the second convolution cycle of the plurality of convolution cycles, associated with a second memory location used to store each of the second values ​​included in the subset of the first value, and The method according to claim 20, relating to the present invention.

22. The method according to claim 21, wherein the output of the first multiplier CU is stored in a first accumulator buffer in each of the plurality of convolution cycles, and the output of the second multiplier CU is stored in a second accumulator buffer in each of the plurality of convolution cycles.

23. At the end of each convolution cycle, the cumulative value stored in the first accumulator buffer replaces the cumulative value stored in the second accumulator buffer at the end of each convolution cycle. The cumulative value stored in the second accumulator buffer at the end of each convolution cycle replaces the cumulative value stored in the third accumulator buffer at the end of each convolution cycle. The method according to claim 22.

24. During each of the aforementioned convolution cycles, each of the accumulator buffers corresponding to the multiple multiplier CUs is, A first input from a corresponding memory location, which indicates the pixel value within the row of the image data, A second input from an adjacent accumulator buffer among the plurality of accumulator buffers, which indicates the accumulated value stored in the adjacent accumulator buffer, and The method according to claim 22, which allows receiving

25. The method according to claim 24, wherein each accumulator buffer receives the first input at the start of each convolution cycle and receives the second input at the end of each convolution cycle.

26. Step 1: At the end of each of the plurality of convolution cycles, update the cumulative value corresponding to the convolution output for the first position among the plurality of positions of the convolution kernel using the first accumulator buffer. The method according to claim 21, further comprising:

27. The steps include updating the first accumulator buffer based on the output of the first multiplier CU in the first convolution cycle among the plurality of convolution cycles, The steps include updating the first accumulator buffer based on the output of the second multiplier CU in the second convolution cycle among the plurality of convolution cycles, and The method according to claim 26, further comprising:

28. The step of updating the cumulative value of the first accumulator buffer based on the respective outputs of different multiplier CUs in each of the plurality of convolution cycles. The method according to claim 26, further comprising:

29. The step of updating the cumulative value of the first accumulator buffer based on the output of each different multiplier CU for each of the multiple locations within the first position of the convolution kernel along the row of the image data. The method according to claim 28, further comprising:

30. The method according to claim 29, wherein each convolution cycle of the plurality of convolution cycles corresponds to each of the locations of the plurality of locations in the first position of the convolution kernel along the row of the image data.