Efficient test time adaptation for improved video processing time consistency
By receiving video input and generating labels in the first layer of an artificial neural network, and combining it with an auxiliary network for online adaptation, the inconsistency problem in deep learning video processing is solved, achieving efficient and stable video processing results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- QUALCOMM INC
- Filing Date
- 2022-03-09
- Publication Date
- 2026-05-08
AI Technical Summary
Modern deep learning-based video processing methods or models may generate inconsistent outputs over time, leading to a decline in user experience and system stability, especially on mobile devices where computation costs are high and computation time is long.
By receiving video input in the first layer of an artificial neural network, generating the first label and updating the network, and combining the auxiliary network with the segmentation network for online adaptation, computational costs are reduced and the time consistency of video processing is improved.
It enables efficient video processing on mobile devices, reduces computing costs, improves output consistency, and enhances user experience and system stability.
Smart Images

Figure CN117223035B_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application claims priority to U.S. Patent Application No. 17 / 198,147, filed March 10, 2021, entitled “EFFICIENT TEST-TIME ADAPTATION FOR IMPROVED TEMPORAL CONSISTENCY IN VIDEO PROCESSING”, the disclosure of which is expressly incorporated herein by reference in its entirety. Background Technology
[0003] field
[0004] The various aspects of this disclosure generally relate to neural networks, and more specifically to video processing.
[0005] background
[0006] Artificial neural networks may include groups of interconnected artificial neurons (e.g., neuron models). Artificial neural networks may be computing devices or represented as methods to be performed by computing devices.
[0007] Neural networks consist of operands that consume tensors and produce tensors. Neural networks can be used to solve complex problems; however, the time it takes for a network to complete a task can be very long, due to the potentially enormous size of the network and the computational demands required to produce a solution. Furthermore, the computational cost of deep neural networks can be problematic because these tasks can be performed on mobile devices, which may have limited computing power.
[0008] A convolutional neural network (CNN) is a type of feedforward artificial neural network. A CNN can comprise an array of neurons, each with a receptive field, that collectively construct an input space. CNNs (such as deep convolutional neural networks (DCNs)) have numerous applications. Specifically, these neural network architectures are used in a variety of technologies, such as image recognition, pattern recognition, speech recognition, autonomous driving, object segmentation in video streams, video processing, and other classification tasks.
[0009] Modern deep learning-based video processing methods or models may generate inconsistent output over time. In some cases, inconsistencies can be observed as display flickering or other misalignments. Temporally inconsistent output (e.g., flickering) can degrade the user experience and enjoyment, as well as system stability and performance.
[0010] Overview
[0011] In one aspect of this disclosure, a method for processing video is provided. The method includes receiving video as input at a first layer of an artificial neural network (ANN). The method also includes processing a first frame of the video to generate a first label. Additionally, the method includes updating the artificial neural network based on the first label. The updating of the artificial neural network is performed while concurrently processing a second frame of the video.
[0012] In another aspect of this disclosure, an apparatus for processing video is provided. The apparatus includes a memory and one or more processors coupled to the memory. The processors are configured to receive video as input at a first layer of an artificial neural network (ANN). The processors are also configured to process a first frame of the video to generate a first label. Furthermore, the processors are configured to update the artificial neural network based on the first label. The update of the artificial neural network is performed while concurrently processing a second frame of the video.
[0013] In one aspect of this disclosure, an apparatus for processing video is provided. The apparatus includes means for receiving video as input at a first layer of an artificial neural network (ANN). The apparatus also includes means for processing a first frame of the video to generate a first label. Additionally, the apparatus includes means for updating the artificial neural network based on the first label. The updating of the artificial neural network is performed concurrently while processing a second frame of the video.
[0014] In a further aspect of this disclosure, a non-transient computer-readable medium is provided. Program code for processing video is encoded on this computer-readable medium. The program code is executed by a processor and includes code for receiving video as input at a first layer of an artificial neural network. The program code also includes code for processing a first frame of the video to generate a first label. Furthermore, the program code includes code for updating the artificial neural network based on the first label. The updating of the artificial neural network is performed while concurrently processing a second frame of the video.
[0015] In one aspect of this disclosure, a method for processing video is provided. The method includes receiving video as input at a first layer of a first artificial neural network and a second artificial neural network. The first artificial neural network has fewer channels than the second artificial neural network. The method also includes processing a first frame of the video via the first artificial neural network to generate a first label. The first artificial neural network provides intermediate features extracted from the first frame of the video to the second artificial neural network. Additionally, the method includes processing the first frame of the video via the second artificial neural network to generate a second label based on the intermediate features and the first frame. Furthermore, the method includes updating the first artificial neural network based on the first label while the second artificial neural network concurrently processes a second frame of the video.
[0016] In one aspect of this disclosure, an apparatus for processing video is provided. The apparatus includes a memory and one or more processors coupled to the memory. The processors are configured to receive video as input at a first layer of a first artificial neural network and a second artificial neural network. The first artificial neural network has fewer channels than the second artificial neural network. The processors are also configured to process a first frame of the video via the first artificial neural network to generate a first label. The first artificial neural network provides intermediate features extracted from the first frame of the video to the second artificial neural network. Additionally, the processors are configured to process the first frame of the video via the second artificial neural network to generate a second label based on the intermediate features and the first frame. Furthermore, the processors are configured to update the first artificial neural network based on the first label while the second artificial neural network concurrently processes a second frame of the video.
[0017] In one aspect of this disclosure, an apparatus for processing video is provided. The apparatus includes means for receiving video as input at a first layer of a first artificial neural network and a second artificial neural network. The first artificial neural network has fewer channels than the second artificial neural network. The apparatus also includes means for processing a first frame of the video via the first artificial neural network to generate a first label. The first artificial neural network provides intermediate features extracted from the first frame of the video to the second artificial neural network. Additionally, the apparatus includes means for processing the first frame of the video via the second artificial neural network to generate a second label based on the intermediate features and the first frame. Furthermore, the apparatus includes means for updating the first artificial neural network based on the first label while the second artificial neural network concurrently processes a second frame of the video.
[0018] In one aspect of this disclosure, a non-transient computer-readable medium is provided. Program code for processing video is encoded on the computer-readable medium. The program code is executed by a processor and includes code for receiving video as input at a first artificial neural network and a first layer of a second artificial neural network. The first artificial neural network has fewer channels than the second artificial neural network. The program code also includes code for processing a first frame of the video via the first artificial neural network to generate a first label. The first artificial neural network provides intermediate features extracted from the first frame of the video to the second artificial neural network. Additionally, the program code includes code for processing the first frame of the video via the second artificial neural network to generate a second label based on the intermediate features and the first frame. Furthermore, the program code includes code for updating the first artificial neural network based on the first label while the second artificial neural network concurrently processes a second frame of the video.
[0019] Additional features and advantages of this disclosure will be described below. Those skilled in the art will appreciate that this disclosure can be readily used as the basis for modifying or designing other structures for implementing the same purposes as this disclosure. Those skilled in the art will also recognize that such equivalent constructions do not depart from the teachings of this disclosure set forth in the appended claims. Novel features considered characteristic of this disclosure, in both their organization and manner of operation, along with further objects and advantages, will be better understood when considered in conjunction with the accompanying drawings. However, it is to be clearly understood that each drawing is provided for illustrative and descriptive purposes only and is not intended to be a definition of limitation of this disclosure. Brief description of the attached diagram
[0021] The features, nature, and advantages of this disclosure will become more apparent when understood in conjunction with the accompanying drawings, in which the same reference numerals are always used to indicate the subject.
[0022] Figure 1 An example implementation of a neural network using a system-on-a-chip (SoC) (including a general-purpose processor) according to certain aspects of this disclosure is explained.
[0023] Figure 2A , 2B 2C are illustrations explaining various aspects of the neural network according to this disclosure.
[0024] Figure 2D This is a diagram illustrating an exemplary deep convolutional network (DCN) according to various aspects of this disclosure.
[0025] Figure 3 This is a block diagram illustrating an exemplary deep convolutional network (DCN) according to various aspects of this disclosure.
[0026] Figure 4 This is a block diagram illustrating an exemplary software architecture that allows for the modularization of artificial intelligence (AI) functions.
[0027] Figure 5 This is a block diagram illustrating an example architecture for processing video according to various aspects of this disclosure.
[0028] Figure 6 This is a block diagram illustrating an example architecture for processing video according to various aspects of this disclosure.
[0029] Figure 7 This is a block diagram illustrating an example architecture for processing video according to various aspects of this disclosure.
[0030] Figure 8 This is a more detailed illustration of an example architecture for processing video according to various aspects of this disclosure.
[0031] Figure 9 and10 This is a flowchart illustrating methods for processing video according to various aspects of this disclosure.
[0032] Detailed description
[0033] The detailed description that follows, taken in conjunction with the accompanying drawings, is intended as a description of various configurations and is not intended to represent only the configurations in which the described concepts can be practiced. This detailed description includes specific details to provide a thorough understanding of the various concepts. However, it will be apparent to those skilled in the art that these concepts can be practiced without these specific details. In some instances, well-known structures and components are shown in block diagram form to avoid obscuring such concepts.
[0034] Based on this teaching, those skilled in the art will appreciate that the scope of this disclosure is intended to cover any aspect of this disclosure, whether implemented independently of or in combination with any other aspect of this disclosure. For example, any number of the aspects described may be used to implement an apparatus or method of practice. Furthermore, the scope of this disclosure is intended to cover such apparatus or methods practiced using other structures, functionalities, or structures and functionalities that complement or differ from the aspects of the described disclosure. It should be understood that any aspect of this disclosure may be implemented by one or more elements of the claims.
[0035] The word “exemplary” is used to mean “serving as an example, instance, or explanation.” Any aspect described as “exemplary” need not be construed as superior to or better than other aspects.
[0036] While specific aspects are described, numerous variations and substitutions of these aspects fall within the scope of this disclosure. Although some benefits and advantages of preferred aspects are mentioned, the scope of this disclosure is not intended to be limited to a particular benefit, use, or objective. Rather, the aspects of this disclosure are intended to be broadly applicable to different technologies, system configurations, networks, and protocols, some of which are illustrated as examples in the accompanying drawings and the following description of preferred aspects. The detailed description and accompanying drawings are merely illustrative and not limiting of this disclosure, the scope of which is defined by the appended claims and their equivalents.
[0037] Neural networks can be used to solve complex problems; however, the network can take a long time to complete a task because the network size and the amount of computation required to produce a solution can be enormous. Furthermore, the computational cost of deep neural networks can be problematic because these tasks can be performed on mobile devices, which may have limited computing power.
[0038] Conventional deep learning-based video processing methods or models may produce inconsistent outputs over time. In some cases, inconsistencies can be observed as display flickering or other misalignments. Temporally inconsistent output (e.g., flickering) can degrade the user experience and enjoyment, as well as system stability.
[0039] One reason for inconsistent outputs (e.g., predictions) is that, for example, when the output is around 0.5, the neural network may provide uncertain predictions. In such cases, the neural network can produce predictions arbitrarily. Consequently, for some frames in the video stream (between more precisely segmented frames), the segmentation accuracy may drop significantly. Thus, visually similar image regions processed by the neural network may lead to different predictions. Furthermore, the reduction in segmentation accuracy can negatively impact performance.
[0040] To address this issue, aspects of this disclosure relate to online (e.g., during testing) adaptation of the segmentation model. The segmentation network used to process the video can be updated while processing the video. In some aspects, an auxiliary network can be combined with the segmentation network to update the network as the segmentation network continues processing the video.
[0041] Figure 1 An example implementation of a system-on-a-chip (SOC) 100 is described, which may include a central processing unit (CPU) 102 or a multi-core CPU configured for video processing using an artificial neural network (e.g., a neural end-to-end network). Variables (e.g., neural signals and synaptic weights), system parameters associated with computing devices (e.g., a weighted neural network), latency, frequency slot information, and task information may be stored in memory blocks associated with a neural processing unit (NPU) 108, a memory block associated with the CPU 102, a memory block associated with a graphics processing unit (GPU) 104, a memory block associated with a digital signal processor (DSP) 106, a memory block 118, or may be distributed across multiple blocks. Instructions executed at the CPU 102 may be loaded from the program memory associated with the CPU 102 or from memory block 118.
[0042] The SoC 100 may also include additional processing blocks tailored to specific functions, such as a GPU 104, a DSP 106, a connectivity block 110 (which may include fifth-generation (5G) connectivity, fourth-generation LTE (4G LTE) connectivity, Wi-Fi connectivity, USB connectivity, Bluetooth connectivity, etc.), and a multimedia processor 112, for example, capable of detecting and recognizing gestures. In one implementation, an NPU 108 is implemented within the CPU 102, DSP 106, and / or GPU 104. The SoC 100 may also include a sensor processor 114, an image signal processor (ISP) 116, and / or a navigation module 120 (which may include a global positioning system).
[0043] SoC 100 may be based on the ARM instruction set. In one aspect of this disclosure, instructions loaded into general-purpose processor 102 may include code for receiving video as input at the first layer of an artificial neural network (ANN). General-purpose processor 102 may also include code for processing the first frame of the video to generate a first label. General-purpose processor 102 may further include code for updating the artificial neural network based on the first label. The update is performed during concurrent processing of the second frame of the video.
[0044] In one aspect of this disclosure, instructions loaded into the general-purpose processor 102 may include code for receiving video as input at a first artificial neural network (ANN) and a first layer of the second artificial neural network. The first artificial neural network has fewer channels than the second artificial neural network. The general-purpose processor 102 may also include code for processing a first frame of the video via the first artificial neural network to generate a first label. The first artificial neural network provides intermediate features extracted from the first frame of the video to the second artificial neural network. The general-purpose processor 102 may further include code for processing the first frame of the video via the second artificial neural network to generate a second label based on the intermediate features and the first frame. The general-purpose processor 102 may additionally include code for updating the first artificial neural network based on the first label while the second artificial neural network concurrently processes a second frame of the video.
[0045] Deep learning architectures perform object recognition tasks by learning to represent inputs at progressively higher levels of abstraction in each layer, thereby constructing useful feature representations of the input data. In this way, deep learning addresses a major bottleneck in traditional machine learning. Before deep learning, machine learning methods for object recognition problems often relied heavily on human-engineered features, perhaps combined with shallow classifiers. Shallow classifiers could be two-class linear classifiers, where a weighted sum of feature vector components is compared to a threshold to predict which class the input belongs to. Human-engineered features could be templates or kernels customized for a specific problem domain by engineers with domain expertise. In contrast, deep learning architectures can learn to represent features similar to those that human engineers might design, but this learning is achieved through training. Furthermore, deep networks can learn to represent and recognize novel types of features that humans might not have considered before.
[0046] Deep learning architectures can learn hierarchical levels of features. For example, if visual data is presented to the first layer, it can learn to recognize relatively simple features (such as edges) in the input stream. In another example, if auditory data is presented to the first layer, it can learn to recognize spectral power at specific frequencies. A second layer, taking the output of the first layer as input, can learn to recognize combinations of features, such as recognizing simple shapes in visual data or sound combinations in auditory data. For example, higher layers can learn to represent complex shapes in visual data or words in auditory data. Even higher layers can learn to recognize common visual objects or spoken phrases.
[0047] Deep learning architectures can perform particularly well when applied to problems with a naturally hierarchical structure. For example, the classification of motor vehicles can benefit from first learning to identify wheels, windshields, and other features. These features can then be combined in different ways at higher levels to identify cars, trucks, and airplanes.
[0048] Neural networks can be designed with various connectivity patterns. In feedforward networks, information is passed from lower layers to higher layers, where each neuron in a given layer communicates to neurons in higher layers. As mentioned above, hierarchical representations can be constructed in successive layers of a feedforward network. Neural networks can also have backflow or feedback (also known as top-down) connections. In a backflow connection, the output from a neuron in a given layer can be communicated to another neuron in the same layer. Backflow architectures can help identify patterns across more than one block of input data sequentially delivered to the neural network. Connections from neurons in a given layer to neurons in lower layers are called feedback (or top-down) connections. Networks with many feedback connections can be beneficial when the recognition of higher-level concepts can help discern specific lower-level features of the input.
[0049] The connections between layers of a neural network can be fully connected or partially connected. Figure 2A An example of a fully connected neural network 202 is explained. In a fully connected neural network 202, a neuron in the first layer can pass its output to each neuron in the second layer, so that each neuron in the second layer receives input from each neuron in the first layer. Figure 2B An example of a locally connected neural network 204 has been explained. In the locally connected neural network 204, neurons in the first layer can connect to a finite number of neurons in the second layer. More generally, the locally connected layers of the locally connected neural network 204 can be configured such that each neuron in a layer will have the same or similar connectivity pattern, but its connection strength can have different values (e.g., 210, 212, 214, and 216). The connectivity pattern of locally connected networks may produce spatially dissimilar receptive fields in higher layers because higher-layer neurons in a given region can receive inputs that are tuned to a restricted portion of the total input to the network through training.
[0050] An example of a locally connected neural network is a convolutional neural network. Figure 2C An example of a convolutional neural network 206 has been explained. The convolutional neural network 206 can be configured such that the connection strength associated with the input for each neuron in the second layer is shared (e.g., 208). Convolutional neural networks may be well-suited for problems where the spatial location of the input is meaningful.
[0051] One type of convolutional neural network is the deep convolutional network (DCN). Figure 2D A detailed example of a DCN 200 designed to recognize visual features from an image 226 input from an image capture device 230 (such as an in-vehicle camera) is explained. The DCN 200 of this example can be trained to identify traffic signs and the numbers provided on them. Of course, the DCN 200 can be trained for other tasks, such as identifying lane markings or traffic lights.
[0052] The DCN 200 can be trained using supervised learning. During training, an image (such as image 226 of a speed limit sign) can be presented to the DCN 200, and a forward pass can then be computed to produce output 222. The DCN 200 may include feature extraction segments and classification segments. Upon receiving image 226, a convolutional layer 232 may apply a convolutional kernel (not shown) to image 226 to generate a first set of feature maps 218. As an example, the convolutional kernel of the convolutional layer 232 may be a 5x5 kernel that generates a 28x28 feature map. In this example, since four different feature maps are generated in the first set of feature maps 218, four different convolutional kernels are applied to image 226 at the convolutional layer 232. The convolutional kernel may also be referred to as a filter or convolutional filter.
[0053] The first set of feature maps 218 can be subsampled by a max-pooling layer (not shown) to generate a second set of feature maps 220. The max-pooling layer reduces the size of the first set of feature maps 218. That is, the size of the second set of feature maps 220 (e.g., 14x14) is smaller than the size of the first set of feature maps 218 (e.g., 28x28). The reduced size provides similar information to subsequent layers while reducing memory consumption. The second set of feature maps 220 can be further convolved via one or more subsequent convolutional layers (not shown) to generate one or more subsequent sets of feature maps (not shown).
[0054] exist Figure 2D In the example, the second set of feature maps 220 is convolved to generate a first feature vector 224. Furthermore, the first feature vector 224 is further convolved to generate a second feature vector 228. Each feature of the second feature vector 228 may include a number corresponding to a possible feature of the image 226 (such as "sign", "60", and "100"). A softmax function (not shown) converts the numbers in the second feature vector 228 into probabilities. Thus, the output 222 of the DCN 200 is the probability that the image 226 includes one or more features.
[0055] In this example, the probabilities of “sign” and “60” in output 222 are higher than the probabilities of other features in output 222 (such as “30”, “40”, “50”, “70”, “80”, “90”, and “100”). Before training, output 222 generated by DCN 200 is likely incorrect. Therefore, the error between output 222 and the target output can be calculated. The target output is the ground truth of image 226 (e.g., “sign” and “60”). The weights of DCN 200 can then be adjusted so that output 222 of DCN 200 is more closely aligned with the target output.
[0056] To adjust the weights, the learning algorithm computes gradient vectors for each weight. The gradient indicates how much the error will increase or decrease as the weights are adjusted. At the top layers, the gradient directly corresponds to the values of the weights connecting the activated neurons in the penultimate layer to the neurons in the output layer. In lower layers, the gradient depends on the values of the weights and the computed error gradients from the higher layers. The weights can then be adjusted to reduce the error. This method of adjusting weights is called "backpropagation" because it involves a "back pass" in neural networks.
[0057] In practice, the error gradient of the weights may be calculated on a small number of examples, thus approximating the true error gradient. This approximation method is called stochastic gradient descent. Stochastic gradient descent can be repeated until the error rate achievable by the entire system stops decreasing or until the error rate reaches the target level. After learning, new images can be presented to the DCN, and the forward pass in the network produces an output 222, which can be considered an inference or prediction of the DCN.
[0058] Deep Belief Networks (DBNs) are probabilistic models that include multiple layers of hidden nodes. DBNs can be used to extract hierarchical representations of training datasets. DBNs can be obtained by stacking multiple layers of Restricted Boltzmann Machines (RBMs). RBMs are a class of artificial neural networks that can learn probability distributions on an input set. Because RBMs can learn probability distributions without information about which class each input should be classified into, they are often used in unsupervised learning. Using a hybrid unsupervised and supervised paradigm, the bottom RBM of a DBN can be trained unsupervised and used as a feature extractor, while the top RBM can be trained supervised (on the joint distribution of inputs from previous layers and the target class) and used as a classifier.
[0059] Deep convolutional networks (DCNs) are networks of convolutional networks configured with additional pooling and normalization layers. DCNs have achieved state-of-the-art performance on many tasks. DCNs can be trained using supervised learning, where both the input and output targets are known for many paradigms and are used to modify the network's weights using gradient descent.
[0060] DCNs can be feedforward networks. Furthermore, as mentioned above, the connections from neurons in the first layer of a DCN to the neuron group in the next higher layer are shared across neurons in the first layer. The feedforward and shared connections of a DCN can be used for fast processing. The computational burden of a DCN can be much smaller than, for example, a similarly sized neural network that includes backflow or feedback connections.
[0061] The processing of each layer in a convolutional network can be considered as a spatially invariant template or a fundamental projection. If the input is first decomposed into multiple channels, such as the red, green, and blue channels of a color image, then a convolutional network trained on that input can be considered three-dimensional, having two spatial dimensions along the axis of the image and a third dimension capturing color information. The output of the convolutional connections can be considered as forming a feature map in subsequent layers, where each element receives input from a range of neurons in the previous layer (e.g., feature map 220) and from each of those multiple channels. The values in the feature map can be further processed non-linearly (such as correction, max(0,x)). Values from neighboring neurons can be further pooled (which corresponds to downsampling) and provide additional local invariance and dimensionality reduction. Normalization, corresponding to whitening, can also be applied through lateral inhibition between neurons in the feature map.
[0062] The performance of deep learning architectures can improve as more labeled data points become available or as computational power increases. Modern deep neural networks are routinely trained with thousands of times more computational resources than were available to a typical researcher just fifteen years ago. New architectures and training paradigms can further boost the performance of deep learning. Corrected linear units reduce the training problem known as vanishing gradients. New training techniques reduce overfitting and thus enable larger models to achieve better generalization. Encapsulation techniques can abstract the data within a given receptive field and further improve overall performance.
[0063] Figure 3 This is a block diagram illustrating a Deep Convolutional Network 350. A Deep Convolutional Network 350 can include multiple layers of different types based on connectivity and weight sharing. For example... Figure 3 As shown, the deep convolutional network 350 includes convolutional blocks 354A and 354B. Each of the convolutional blocks 354A and 354B may be configured with a convolutional layer (CONV) 356, a normalization layer (LNorm) 358, and a max pooling layer (MAX POOL) 360.
[0064] Convolutional layer 356 may include one or more convolutional filters that can be applied to the input data to generate feature maps. Although only two convolutional blocks 354A and 354B are shown, this disclosure is not limited thereto, and any number of convolutional blocks 354A and 354B may be included in the deep convolutional network 350 according to design preferences. Normalization layer 358 may normalize the output of the convolutional filters. For example, normalization layer 358 may provide whitening or lateral suppression. Max pooling layer 360 may provide spatial downsampling aggregation to achieve local invariance and dimensionality reduction.
[0065] For example, the parallel filter set of the deep convolutional network can be loaded onto the CPU 102 or GPU 104 of the SoC 100 to achieve high performance and low power consumption. In an alternative embodiment, the parallel filter set can be loaded onto the DSP 106 or ISP 116 of the SoC 100. Additionally, the deep convolutional network 350 can access other processing blocks that may exist on the SoC 100, such as the sensor processor 114 and navigation module 120, respectively dedicated to sensors and navigation.
[0066] The deep convolutional network 350 may also include one or more fully connected layers 362 (FC1 and FC2). The deep convolutional network 350 may further include logistic regression (LR) layers 364. Weights (not shown) to be updated are located between each layer 356, 358, 360, 362, and 364 of the deep convolutional network 350. The output of each layer (e.g., 356, 358, 360, 362, and 364) can be used as input to a subsequent layer in the deep convolutional network 350 (e.g., 356, 358, 360, 362, and 364) to learn a hierarchical feature representation from the input data 352 (e.g., image, audio, video, sensor data, and / or other input data) supplied from the first convolutional block 354A. The output of the deep convolutional network 350 is a classification score 366 for the input data 352. The classification score 366 may be a set of probabilities, where each probability is the probability that the input data includes features from a feature set.
[0067] Figure 4 This is a block diagram illustrating an exemplary software architecture 400 that allows for modularization of artificial intelligence (AI) functionality. According to various aspects of this disclosure, by using this architecture, various processing blocks of a system-on-a-chip (SoC) 420 (e.g., CPU 422, DSP 424, GPU 426, and / or NPU 428) can be designed to support applications such as the disclosed adaptive rounding for post-training quantization for AI application 402.
[0068] AI application 402 can be configured to invoke functions defined in user space 404, such as those functions that provide detection and recognition of a scene indicating the current operating location of the device. For example, AI application 402 can configure the microphone and camera differently depending on whether the recognized scene is an office, lecture hall, restaurant, or an outdoor environment such as a lake. AI application 402 can make requests to compiled code associated with libraries defined in AI Function Application Programming Interface (API) 406. This request may ultimately rely on the output of a deep neural network configured to provide inferred responses based on, for example, video and location data.
[0069] The runtime engine 408 (which may be compiled code of the runtime framework) may further be accessible to the AI application 402. For example, the AI application 402 may cause the runtime engine to request inferences at specific time intervals or triggered by events detected by the application's user interface. When the runtime engine is prompted to provide an inference response, it may then signal to the operating system (OS) space (such as kernel 412) running on the SoC 420. The OS may then cause sequential quantization relaxation to be performed on the CPU 422, DSP 424, GPU 426, NPU 428, or some combination thereof. The CPU 422 may be directly accessible by the operating system, while other processing blocks may be accessed via drivers (such as drivers 414, 416, or 418 for the DSP 424, GPU 426, or NPU 428, respectively). In an exemplary example, a deep neural network may be configured to run on a combination of processing blocks (such as CPU 422, DSP 424, and GPU 426) or may run on the NPU 428.
[0070] Application 402 (e.g., an AI application) may be configured to invoke functions defined in user space 404, such as those providing the detection and recognition of a scene indicating the current operating location of the device. For example, application 402 may configure the microphone and camera differently depending on whether the recognized scene is an office, lecture hall, restaurant, or an outdoor environment such as a lake. Application 402 may make a request to compiled code associated with a library defined in the Scene Detection Application Programming Interface (API) 406 to provide an estimate of the current scene. This request may ultimately rely on the output of a differential neural network configured to provide scene estimates based on, for example, video and location data.
[0071] The runtime engine 408 (which may be compiled code of the runtime framework) may further be accessible to the application 402. For example, the application 402 may cause the runtime engine to request scene estimates at specific time intervals or triggered by events detected by the application's user interface. When causing the runtime engine to estimate the scene, the runtime engine may further signal to the operating system 410 (such as kernel 412) running on the SoC 420. The operating system 410 may then cause computations to be performed on the CPU 422, DSP 424, GPU 426, NPU 428, or some combination thereof. The CPU 422 may be directly accessible by the operating system, while other processing blocks may be accessed via drivers (such as drivers 414-488 for the DSP 424, GPU 426, or NPU 428, respectively). In an exemplary example, the differential neural network may be configured to run on a combination of processing blocks (such as CPU 422 and GPU 426) or may run on the NPU 428.
[0072] Various aspects of this disclosure relate to the porting of deep neural network models using adversarial function approximations.
[0073] Figure 5 This is a block diagram illustrating an example architecture 500 for processing video according to various aspects of this disclosure. For example... Figure 5 As shown, the segmentation network 502 receives video 504 as input. The segmentation network 502 can be, for example, a convolutional neural network, such as... Figure 3 The deep convolutional network 350 is shown. Video 504 is divided into frames (e.g., original t1, original t2, original t3, original t4, original t5, and original t6). Segmentation network 502 processes each frame of video 504 sequentially and generates output 506 in the form of segments (e.g., segment t1, segment t2, segment t3, segment t4, segment t5, and segment t6). Segmentation network 502 can also be updated while processing video 504. Figure 5 As shown, after each output segment is generated, the segmentation network 502 is updated via a backward pass (e.g., 508a-508f).
[0074] Figure 6 This is a block diagram illustrating an example architecture 600 for processing video according to various aspects of this disclosure. For example... Figure 6 As shown, the segmentation network 604 receives video 602 as input. The segmentation network 604 can be configured to... Figure 5 The segmentation network 502 operates similarly. The segmentation network 604 processes each frame of video 602 and generates an output segment 606. An argmax operation is performed on each output segment 606 to generate pseudo-labels 608 and output labels 612. Applying cross-entropy (CE) as a loss function, a negative log-likelihood can be calculated based on the pseudo-labels 608 and output labels 612. Furthermore, the segmentation network 604 can be updated via backpropagation (BP) to reduce the CE loss. In some aspects, argmax can be used to determine the output label 612 and can be supplemented by a confidence value calculated, for example, using softmax likelihood or the logit function. Additionally, in some aspects, the confidence value can be the best estimate or the second best estimate of the label ratio.
[0075] Figure 7 This is a block diagram illustrating an example architecture 700 for processing video according to various aspects of this disclosure. (See references) Figure 7Example architecture 700 includes a segmentation network 704 and an auxiliary network 706. The auxiliary network 706 may have an overall architecture similar to that of the segmentation network 704. However, the auxiliary network 706 is smaller than the segmentation network 704 (e.g., 1 / 10 the size of the segmentation network 704). For example, in some aspects, the auxiliary network 706 may be configured with fewer channels (e.g., the auxiliary network 706 may have 18 channels, while the segmentation network 704 may have 48 channels), or it may operate at a lower resolution than the segmentation network 704. In one example, the auxiliary network 706 can operate on a 32-bit CPU (e.g., ...). Figure 1 CPU 102) or GPU (e.g., Figure 1 The segmentation network 704 can operate on an 8-bit DSP (e.g., GPU 104) and can operate on an 8-bit DSP (e.g., GPU 104). Figure 1 It operates on the DSP 106.
[0076] In operation, segmentation network 704 and auxiliary network 706 each receive video 702 as input. Segmentation network 704 processes each frame of video 702 and generates output segment 710. Auxiliary network 706 processes each frame of video 702 and generates segment 708.
[0077] Additionally, an auxiliary network 706 provides intermediate features to a segmentation network 704. The segmentation network 704 processes these intermediate features, which are then aggregated with intermediate features generated in each layer of the segmentation network (indicated by a "+" sign) to compute output segments 710. An argmax operation is performed on each output segment 710 to generate pseudo-labels 712. Similarly, an argmax operation is performed on each segment 708 to generate output labels 714.
[0078] The pseudo-label 712 and output label 714 can be used to calculate the cross-entropy loss and, for example, update the auxiliary network 706 via backpropagation, while the segmentation network 704 continues to process the video 702 in the forward pass. Because the update is performed on the smaller auxiliary network 706 rather than the segmentation network 704, the computational cost is relatively low compared to... Figure 6 The computational cost of the segmentation network 604 can be significantly reduced.
[0079] Figure 8 This is an explanation based on the various aspects of this disclosure. Figure 7 A more detailed illustration of the example architecture 700 is shown in Figure 800. (Reference) Figure 8 The diagram illustrates the segmentation network 804 and the auxiliary network 806. Figure 8 In the example, the segmentation network 804 and the auxiliary network 806 can have similar architectures. For example, the segmentation network 804 and the auxiliary network 806 can be configured as an autoencoder. (As mentioned above...) Figure 7The auxiliary network 806 discussed can be smaller than the segmentation network 804 (e.g., 1 / 10 the size of the segmentation network 804). For example, in some aspects, the auxiliary network 806 can be configured with fewer channels (e.g., the auxiliary network 806 can have 18 channels, while the segmentation network 804 can have 48 channels). Additionally, in some aspects, the auxiliary network 806 can operate at a lower resolution than the segmentation network 804. For example, as... Figure 8 As shown, both the auxiliary network 806 and the segmentation network 804 receive video 802 as input. However, while the segmentation network 804 downsamples each frame of video 802 three times, the auxiliary network 806 downsamples each frame of video 802 four times, resulting in a lower resolution for each frame of the video compared to the segmentation network 804.
[0080] After downsampling the frames of video 802, segmentation network 804 and auxiliary network 806 each pass their respective lower-resolution frames through successive layers of convolutional filters (e.g., convolutional coding blocks) to extract features. Further downsampling can be performed via transition blocks (e.g., transition 1, transition 2, and transition 3) of segmentation network 804 and auxiliary network 806, respectively, to extract lower-resolution features.
[0081] Additionally, the auxiliary network 806 may provide intermediate features (e.g., 808, 810) to the segmentation network 804. These intermediate features (e.g., 808, 810) may be mapped (to account for a larger number of channels in the segmentation network 804) and combined or aggregated with the intermediate features of the segmentation network 804. The segmentation network 804 and the auxiliary network 806 each produce output feature sets (e.g., 812, 814), which may be upsampled, concatenated, and provided to the decoder (816, 818).
[0082] Decoders 816 and 818 then process the output features (e.g., 812, 814) to generate output segments 820 and 822. An argmax operation is performed on each of the output segments 820 and 822 to generate label 824.
[0083] Furthermore, the auxiliary network 806 can be updated using backpropagation. Notably, while the segmentation network 804 continues processing video 802, the auxiliary network 806 is updated online (e.g., during test time). Doing so reduces temporal inconsistencies (e.g., it can increase the intersection on the union metric). Thus, the flickering effect observed in video 802 can be reduced. Additionally, computational costs and processing time can be reduced.
[0084] Figure 9 The present disclosure describes a method 900 for processing video according to various aspects. For example... Figure 9As shown in the diagram, in box 902, method 900 receives video input at the first layer of an artificial neural network (ANN). The first artificial neural network can be a convolutional neural network, such as... Figure 3 A 350-degree deep convolutional network. (See reference...) Figure 5 The segmentation network 502, as discussed, receives video 504 as input.
[0085] In box 904, method 900 processes the first frame of the video to generate the first tag. (See reference...) Figure 5 The segmentation network 502 discussed here processes each frame of video 504 sequentially and generates output 506 in the form of segments (e.g., segments t1).
[0086] In box 906, method 900 updates the artificial neural network based on the first label, and this update is performed while concurrently processing the second frame of the video. For example, as Figure 5 As shown, after each output segment is generated, the segmentation network 502 is updated via backpropagation (e.g., 508a-508f).
[0087] Figure 10 The present disclosure describes 1000 methods for processing video according to various aspects. For example... Figure 10 As shown in box 1002, method 1000 receives video as input at the first layer of both a first artificial neural network (ANN) and a second artificial neural network. The first artificial neural network has fewer channels than the second artificial neural network. For example, refer to... Figure 7 Example architecture 700 includes a segmentation network 704 and an auxiliary network 706. Both segmentation network 704 and auxiliary network 706 receive video 702 as input. Additionally, auxiliary network 706 may have an architecture similar to that of segmentation network 704. For example, in some aspects, auxiliary network 706 may be configured with fewer channels (e.g., auxiliary network 706 may have 18 channels, while segmentation network 704 may have 48 channels), or it may operate at a lower resolution than segmentation network 704.
[0088] In box 1004, method 1000 processes the first frame of the video via a first artificial neural network to generate a first label. The first artificial neural network provides intermediate features extracted from the first frame of the video to a second artificial neural network. For example, as referenced... Figure 7 The auxiliary network 706, as discussed, processes each frame of the video and generates segments 708. The auxiliary network 706 provides intermediate features to the segmentation network 704. The segmentation network 704 processes the intermediate features, which are aggregated to compute output segments 710.
[0089] In box 1006, method 1000 processes the first frame of the video via a second artificial neural network to generate a second label based on intermediate features and the first frame. For example, segmentation network 704 processes each frame of video 702 and generates output segments 710.
[0090] In box 1008, method 1000 updates the first artificial neural network based on the first label when the second artificial neural network concurrently processes the second frame of the video. (See reference...) Figure 7 The pseudo-label 712 discussed herein can be used to update the auxiliary network 706, for example, via backpropagation, while the segmentation network 704 continues to process the video 702 in the forward pass.
[0091] In one aspect, the receiving device, processing device, and / or updating device may be a CPU 102, a GPU 104, a DSP 106 program memory associated with the CPU 102, a dedicated memory block 118, a fully connected layer 362, an NPU 428, and / or a routing connection processing unit 216 configured to perform the described functions. In another configuration, the aforementioned device may be any module or any equipment configured to perform the functions described by the aforementioned device.
[0092] The various operations of the methods described above can be performed by any suitable means capable of performing the corresponding functions. These means may include various hardware and / or software components and / or modules, including but not limited to circuits, application-specific integrated circuits (ASICs), or processors. Generally, in the cases where operations are illustrated in the accompanying drawings, those operations may have corresponding paired means with similar numbers plus functional components.
[0093] Examples of implementations are provided in the following numbered clauses:
[0094] 1. A method for processing video, comprising:
[0095] The video is received as input at the first layer of the artificial neural network (ANN);
[0096] Process the first frame of the video to generate the first tag; and
[0097] The artificial neural network is updated based on the first label, and this update is performed while the second frame of the video is being processed concurrently.
[0098] 2. The method of Clause 1 further includes: applying a first label to update the artificial neural network during a backward pass of the artificial neural network.
[0099] 3. The method of any one of clauses 1-2 further includes: generating a second tag based on the second frame, wherein concurrent processing is performed to reduce the time inconsistency between the first tag and the second tag.
[0100] 4. A method for processing video, comprising:
[0101] The video is received as input at the first layer of the first artificial neural network (ANN) and the second artificial neural network, with the first artificial neural network having fewer channels than the second artificial neural network.
[0102] The first frame of the video is processed by a first artificial neural network to generate a first label, and the first artificial neural network provides intermediate features extracted from the first frame of the video to a second artificial neural network.
[0103] The first frame of the video is processed via a second artificial neural network to generate a second label based on intermediate features and the first frame; and
[0104] The first artificial neural network is updated based on the first label when the second artificial neural network processes the second frame of the video concurrently.
[0105] 5. The method of Clause 4 further includes: generating a third label based on the second frame via a second artificial neural network, wherein concurrent processing is performed to reduce the temporal inconsistency between the second label and the third label.
[0106] 6. The method of any one of Clauses 4-5, wherein the first artificial neural network operates at a resolution lower than that of the second artificial neural network.
[0107] 7. An apparatus for processing video, comprising:
[0108] Memory; and
[0109] At least one processor coupled to the memory, the at least one processor being configured to:
[0110] The video is received as input at the first layer of the artificial neural network (ANN);
[0111] Process the first frame of the video to generate the first tag; and
[0112] The artificial neural network is updated based on the first label, and this update is performed while the second frame of the video is being processed concurrently.
[0113] 8. The apparatus of Clause 7, wherein the at least one processor is further configured to apply a first label to update the artificial neural network during a backward pass.
[0114] 9. The apparatus of any one of clauses 7-8, wherein the at least one processor is further configured to generate a second tag based on a second frame, and wherein concurrent processing is performed to reduce time inconsistencies between the first tag and the second tag.
[0115] 10. An apparatus for processing video, comprising:
[0116] Memory; and
[0117] At least one processor coupled to the memory, the at least one processor being configured to:
[0118] The video is received as input at the first layer of the first artificial neural network (ANN) and the second artificial neural network, with the first artificial neural network having fewer channels than the second artificial neural network.
[0119] The first frame of the video is processed by a first artificial neural network to generate a first label, and the first artificial neural network provides intermediate features extracted from the first frame of the video to a second artificial neural network.
[0120] The first frame of the video is processed via a second artificial neural network to generate a second label based on intermediate features and the first frame; and
[0121] The first artificial neural network is updated based on the first label when the second artificial neural network processes the second frame of the video concurrently.
[0122] 11. The apparatus of Clause 10, wherein the at least one processor is further configured to generate a third label based on a second frame via a second artificial neural network, and wherein concurrent processing is performed to reduce the temporal inconsistency between the second label and the third label.
[0123] 12. The apparatus of Clause 10, wherein the first artificial neural network operates at a resolution lower than that of the second artificial neural network.
[0124] 13. An apparatus for processing video, comprising:
[0125] A means for receiving the video as input at the first layer of an artificial neural network (ANN); means for processing the first frame of the video to generate a first label; and
[0126] The apparatus for updating the artificial neural network based on a first label is performed during concurrent processing of the second frame of the video.
[0127] 14. The apparatus of Clause 13 further includes: means for applying a first label in a backward pass of the artificial neural network to update the artificial neural network.
[0128] 15. The apparatus of any one of clauses 13-14 further includes: means for generating a second tag based on a second frame, wherein concurrent processing is performed to reduce time inconsistencies between the first tag and the second tag.
[0129] 16. An apparatus for processing video, comprising:
[0130] A device for receiving the video as input at the first layer of a first artificial neural network (ANN) and a second artificial neural network, wherein the first artificial neural network has fewer channels than the second artificial neural network.
[0131] A device for processing a first frame of the video via a first artificial neural network to generate a first label, wherein the first artificial neural network provides intermediate features extracted from the first frame of the video to a second artificial neural network;
[0132] A means for processing the first frame of the video via a second artificial neural network to generate a second label based on intermediate features and the first frame; and
[0133] A device for updating a first artificial neural network based on a first label when the second artificial neural network concurrently processes the second frame of a video.
[0134] 17. The apparatus of Clause 16 further includes: means for generating a third label based on a second frame via a second artificial neural network, wherein concurrent processing is performed to reduce the temporal inconsistency between the second label and the third label.
[0135] 18. The device of any of Clauses 16-17, wherein the first artificial neural network operates at a resolution lower than that of the second artificial neural network.
[0136] 19. A non-transient computer-readable medium having program code encoded thereon for processing video, the program code being executed by a processor and comprising:
[0137] Program code used to receive the video as input at the first layer of an artificial neural network (ANN);
[0138] Program code used to process the first frame of the video to generate the first tag; and
[0139] The program code used to update the artificial neural network based on the first label is executed while the second frame of the video is being processed concurrently.
[0140] 20. The non-transient computer-readable medium of Clause 19 further includes: program code for applying a first label to update the artificial neural network during a backward pass of the artificial neural network.
[0141] 21. A non-transient computer-readable medium as described in any of Clauses 19-20, further comprising: program code for generating a second tag based on a second frame, wherein concurrent processing is performed to reduce time inconsistencies between the first tag and the second tag.
[0142] 22. A non-transient computer-readable medium having program code encoded thereon for processing video, the program code being executed by a processor and comprising:
[0143] Program code for receiving the video as input at the first layer of a first artificial neural network (ANN) and a second artificial neural network, wherein the first artificial neural network has fewer channels than the second artificial neural network;
[0144] Program code for processing the first frame of the video via a first artificial neural network to generate a first label, wherein the first artificial neural network provides intermediate features extracted from the first frame of the video to a second artificial neural network;
[0145] Program code for processing the first frame of the video via a second artificial neural network to generate a second label based on intermediate features and the first frame; and
[0146] Program code for updating the first artificial neural network based on the first label when the second artificial neural network concurrently processes the second frame of the video.
[0147] 23. The non-transient computer-readable medium of Clause 22 further includes: program code for generating a third label based on a second frame via a second artificial neural network, wherein concurrent processing is performed to reduce the temporal inconsistency between the second label and the third label.
[0148] 24. A non-transient computer-readable medium as described in any of Clauses 22-23, wherein the first artificial neural network operates at a resolution lower than that of the second artificial neural network.
[0149] As used herein, the term "determine" encompasses a wide variety of actions. For example, "determine" can include calculation, computation, processing, derivation, research, searching (e.g., looking in a table, database, or other data structure), ascertainment, and similar actions. Additionally, "determine" can include receiving (e.g., receiving information), accessing (e.g., accessing data in memory), and similar actions. Furthermore, "determine" can include parsing, selecting, choosing, establishing, and similar actions.
[0150] As used in this article, the phrase “at least one of” a list of items refers to any combination of those items, including a single member. As an example, “at least one of a, b, or c” is intended to cover: a, b, c, ab, ac, bc, and abc.
[0151] The various illustrative logic blocks, modules, and circuits described in this disclosure can be implemented or executed using a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA) or other programmable logic device (PLD), discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the described functions. The general-purpose processor may be a microprocessor, but in alternatives, the processor may be any commercially available processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors cooperating with a DSP core, or any other such configuration.
[0152] The steps of the methods or algorithms described in this disclosure can be implemented directly in hardware, in a software module executed by a processor, or in a combination of both. The software module can reside in any form of storage medium known in the art. Some examples of usable storage media include random access memory (RAM), read-only memory (ROM), flash memory, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, removable disks, CD-ROMs, and so on. The software module may include a single instruction or many instructions, and may be distributed across several different code segments, across different programs, and across multiple storage media. The storage medium may be coupled to the processor so that the processor can read and write information from / to the storage medium. In an alternative, the storage medium may be integrated into the processor.
[0153] The methods disclosed herein include one or more steps or actions for achieving the described methods. These method steps and / or actions may be interchanged with each other without departing from the scope of the claims. In other words, unless a specific order of steps or actions is specified, the order and / or use of specific steps and / or actions may be modified without departing from the scope of the claims.
[0154] The described functionality can be implemented in hardware, software, firmware, or any combination thereof. If implemented in hardware, an example hardware configuration may include a processing system within the device. The processing system can be implemented using a bus architecture. Depending on the specific application and overall design constraints of the processing system, the bus may include any number of interconnect buses and bridges. The bus can link together various circuits, including processors, machine-readable media, and bus interfaces. The bus interface can be used, in particular, to connect network adapters and the like to the processing system via the bus. The network adapter can be used to implement signal processing functions. In some respects, user interfaces (e.g., keypads, displays, mice, joysticks, etc.) may also be connected to the bus. The bus can also link various other circuits, such as timing sources, peripherals, regulators, power management circuits, and similar circuits, which are well known in the art and will not be described further.
[0155] A processor is responsible for managing the bus and general processing, including executing software stored on a machine-readable medium. A processor may be implemented using one or more general-purpose and / or special-purpose processors. Examples include microprocessors, microcontrollers, DSP processors, and other circuit systems capable of executing software. Software should be interpreted broadly as instructions, data, or any combination thereof, whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise. As examples, a machine-readable medium may include random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, disks, optical disks, hard drives, or any other suitable storage medium, or any combination thereof. The machine-readable medium may be implemented in a computer program product. This computer program product may include packaging materials.
[0156] In hardware implementations, machine-readable media can be a separate part of the processing system from the processor. However, as those skilled in the art will readily appreciate, machine-readable media or any part thereof can be external to the processing system. As examples, machine-readable media may include transmission lines, carrier waves modulated by data, and / or computer components separate from the device, all accessible to the processor via a bus interface. Alternatively or additionally, machine-readable media or any part thereof may be integrated into the processor, such as caches and / or general-purpose register files. While the various components discussed may be described as having a specific location, such as local components, they can also be configured in various ways, such as certain components being configured as part of a distributed computing system.
[0157] The processing system can be configured as a general-purpose processing system having one or more microprocessors providing processor functionality, and external memory providing at least a portion of machine-readable medium, all linked to other supporting circuitry via an external bus architecture. Alternatively, the processing system may include one or more neuromorphic processors for implementing the described neuron and nervous system models. As another alternative, the processing system can be implemented using an application-specific integrated circuit (ASIC) with a processor, bus interface, user interface, supporting circuitry, and at least a portion of machine-readable medium integrated on a single chip, or using one or more field-programmable gate arrays (FPGAs), programmable logic devices (PLDs), controllers, state machines, gated logic, discrete hardware components, or any other suitable circuitry, or any combination of circuitry capable of performing the various functionalities described throughout this disclosure. Depending on the specific application and the overall design constraints imposed on the system, those skilled in the art will recognize how best to implement the functionality described with respect to the processing system.
[0158] Machine-readable media may include several software modules. These software modules include instructions that, when executed by a processor, cause the processing system to perform various functions. These software modules may include transfer modules and receive modules. Each software module may reside in a single storage device or be distributed across multiple storage devices. As an example, when a trigger event occurs, a software module may be loaded from a hard drive into RAM. During the execution of a software module, the processor may load some instructions into a cache to improve access speed. One or more cache lines may subsequently be loaded into a general-purpose register file for processor execution. In the context of the functionality of the software modules described below, it will be understood that such functionality is implemented by the processor when the processor executes the instructions from the software module. Furthermore, it should be understood that aspects of this disclosure result in improvements to the functionality of a processor, computer, machine, or other system implementing such aspects.
[0159] If implemented in software, the functions can be stored or transmitted as one or more instructions or codes on or through a computer-readable medium. Computer-readable media includes both computer storage media and communication media, encompassing any medium that facilitates the transfer of a computer program from one location to another. Storage media can be any available medium accessible to a computer. By way of example and not limitation, such computer-readable media may include RAM, ROM, EEPROM, CD-ROM or other optical disc storage, disk storage or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and is accessible to a computer. Additionally, any connection is also legitimately referred to as computer-readable media. For example, if the software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technology (such as infrared (IR), radio, and microwave), then that coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technology (such as infrared, radio, and microwave) is included in the definition of medium. The disks and discs used in this article include CDs, laser discs, optical discs, DVDs, floppy disks, and... Disks, where disks often magnetically reproduce data, and discs optically reproduce data using lasers. Therefore, in some aspects, computer-readable media may include non-transient computer-readable media (e.g., tangible media). Additionally, in other aspects, computer-readable media may include transient computer-readable media (e.g., signals). Combinations of the above should also be included within the scope of computer-readable media.
[0160] Therefore, some aspects may include computer program products for performing the operations given herein. For example, such computer program products may include computer-readable media on which instructions are stored (and / or encoded) that can be executed by one or more processors to perform the described operations. In some aspects, computer program products may include packaging materials.
[0161] Furthermore, it should be understood that modules and / or other suitable means for performing the described methods and techniques may be downloaded and / or otherwise obtained by the user terminal and / or base station where applicable. For example, such devices can be coupled to a server to facilitate the transfer of means for performing the described methods. Alternatively, the various methods described can be provided via a storage device (e.g., RAM, ROM, physical storage media such as CDs or floppy disks) so that the device can acquire the various methods once the storage device is coupled to or provided to the user terminal and / or base station. In addition, any other suitable techniques suitable for providing the described methods and techniques to the device may be utilized.
[0162] It will be understood that the claims are not limited to the precise configurations and components described above. Various modifications, substitutions, and variations can be made to the layout, operation, and details of the methods and apparatus described above without departing from the scope of the claims.
Claims
1. A processor-executed method for processing video, the processor-executed method being performed by at least one processor and comprising: The video is received as input at the first layer of the artificial neural network (ANN); Process the first frame of the video to generate a first output portion; The argmax function is applied to the first output portion to generate pseudo-tags for the first frame; as well as The artificial neural network is updated based on cross-entropy loss, which is calculated from a function of the first output portion and the pseudo-label of the first frame of the video.
2. The method executed by the processor as described in claim 1, further comprising: The cross-entropy loss is applied in the backpropagation of the artificial neural network.
3. The method executed by the processor as claimed in claim 1, further comprising generating a second output portion based on a second frame.
4. The method executed by the processor as claimed in claim 1, wherein the artificial neural network is updated during testing.
5. A processor-executed method for processing video, the processor-executed method being performed by at least one processor and comprising: The video is received as input at the first layer of the first artificial neural network (ANN) and the second artificial neural network, wherein the first artificial neural network has fewer channels than the second artificial neural network. The first frame of the video is processed by the first artificial neural network to generate a first label, and the first artificial neural network provides intermediate features extracted from the first frame of the video to the second artificial neural network; The first frame of the video is processed via the second artificial neural network to generate a second label based on the intermediate features and the first frame; as well as The first artificial neural network is updated based on the first label when the second artificial neural network concurrently processes the second frame of the video.
6. The method executed by the processor as described in claim 5, further comprising: A third label is generated based on the second frame via the second artificial neural network, wherein the concurrent processing is performed to reduce the time inconsistency between the second label and the third label.
7. The method executed by the processor of claim 5, wherein the first artificial neural network operates at a resolution lower than that of the second artificial neural network.
8. An apparatus for processing video, comprising: At least one memory; as well as At least one processor coupled to the at least one memory, the at least one processor being configured to: The video is received as input at the first layer of the first artificial neural network (ANN) and the second artificial neural network, wherein the first artificial neural network has fewer channels than the second artificial neural network. The first frame of the video is processed by the first artificial neural network to generate a first label, and the first artificial neural network provides intermediate features extracted from the first frame of the video to the second artificial neural network; The first frame of the video is processed via the second artificial neural network to generate a second label based on the intermediate features and the first frame; as well as The first artificial neural network is updated based on the first label when the second artificial neural network concurrently processes the second frame of the video.
9. The apparatus of claim 8, wherein the at least one processor is further configured to generate a third label based on the second frame via the second artificial neural network, and wherein the concurrent processing is performed to reduce the time inconsistency between the second label and the third label.
10. The apparatus of claim 8, wherein the first artificial neural network operates at a resolution lower than that of the second artificial neural network.
11. An apparatus for processing video, comprising: At least one memory; as well as At least one processor coupled to the at least one memory, the at least one processor being configured to: The video is received as input at the first layer of the artificial neural network (ANN); Process the first frame of the video to generate a first output portion; The argmax function is applied to the first output portion to generate pseudo-tags for the first frame; as well as The artificial neural network is updated based on cross-entropy loss, which is calculated from a function of the first output portion and the pseudo-label of the first frame of the video.
12. The apparatus of claim 11, wherein the at least one processor is further configured to apply the cross-entropy loss in the backpropagation of the artificial neural network.
13. The apparatus of claim 11, wherein the at least one processor is further configured to: The second output portion is generated based on the second frame.
14. The apparatus of claim 11, wherein the artificial neural network is updated during testing.
15. A non-transient computer-readable medium having program code encoded thereon for processing video, the program code being executed by a processor and comprising: Program code for receiving the video as input at the first layer of a first artificial neural network (ANN) and a second artificial neural network, wherein the first artificial neural network has fewer channels than the second artificial neural network. Program code for processing a first frame of the video via the first artificial neural network to generate a first label, wherein the first artificial neural network provides intermediate features extracted from the first frame of the video to a second artificial neural network; Program code for processing the first frame of the video via the second artificial neural network to generate a second tag based on the intermediate features and the first frame; as well as Program code for updating the first artificial neural network based on the first label when the second artificial neural network concurrently processes the second frame of the video.
16. The non-transient computer-readable medium of claim 15, further comprising: Program code for generating a third label based on the second frame via the second artificial neural network, wherein the concurrent processing is performed to reduce the time inconsistency between the second label and the third label.
17. The non-transient computer-readable medium of claim 15, wherein the first artificial neural network operates at a resolution lower than that of the second artificial neural network.
Citation Information
Patent Citations
Methods for training a CRNN and for semantic segmentation of an inputted video using said crnn
EP3608844A1
Segmenting generic foreground objects in images and videos
WO2018128741A1