Performing zoom operations on computing devices through convolutional neural network layers
By selecting multiple CNN layers and non-integer scaling ratios on a computing device to generate intermediate image frame sequences, the problem of time-consuming and power-intensive CNN processing is solved, achieving efficient and smooth image zoom animation.
Patent Information
- Application Number
- CN202380099480.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-28
- Publication Date
- 2026-02-13
AI Technical Summary
When performing image zoom operations on computing devices, existing technologies use CNN processing, which is time-consuming and power-intensive, resulting in unsmooth zoom animations and excessive battery drain.
By selecting multiple CNN layers on the computing device, utilizing non-integer scaling ratios and dynamic dimensional CNN layers, intermediate image frame sequences are generated to render zoom animations. The number of CNN layers is selected based on application requirements and the amount of available memory, optimizing the image scaling operation.
It improves the efficiency and performance of image zoom operations, reduces power consumption, and provides a smooth zoom animation experience.
Smart Images

Figure CN121532804A_ABST
Abstract
Description
BACKGROUND
[0001] Computing devices, such as smartphones, can utilize artificial intelligence (AI) for image processing using specialized AI processors or AI accelerators designed to efficiently run machine learning algorithms. AI processors / accelerators can also execute convolutional neural networks (CNNs) as part of a CNN-based AI image processing system. CNNs are well suited for a variety of image processing tasks, such as image recognition, object detection, resizing images, and image segmentation. While CNNs can be used to upscale images to support image zoom operations, CNN processing for image upsizing takes time to complete and, thus, if used directly, would result in a non-smooth zoom animation. CNN processing also consumes a significant amount of power, which can consume a large amount of battery power to upscale images. As a result, in many computing devices, runtime zoom or runtime upsizing operations are only available within specific applications executing on the computing device. SUMMARY
[0002] Various aspects include a method of performing a zoom operation on a computing device. Various aspects can include identifying a region of interest (ROI) of an image presented on a display of the computing device based on a zoom user input received by the computing device; determining a zoom ratio for a final zoomed presentation of the ROI; selecting two or more convolutional neural network (CNN) layers from a plurality of CNN layers, wherein each CNN layer includes a smaller zoom ratio than the zoom ratio for the final zoomed presentation of the ROI; generating an intermediate image according to each of the selected CNN layers, wherein a size of the ROI in each image is smaller than a size of the final zoomed presentation of the ROI and larger than a size of the ROI; and rendering an animation of a change in display size of the ROI using the sequence of intermediate size image frames of the ROI.
[0003] In some aspects, at least one of the selected CNN layers can include a non-integer zoom ratio. In some aspects, selecting two or more CNN layers from the plurality of CNN layers can include selecting a number of the plurality of CNN layers based on an application requirement of an application executing in the computing device that provides the image. In some aspects, selecting two or more CNN layers from the plurality of CNN layers can include selecting a number of the plurality of CNN layers based on an available amount of memory of the computing device. Some aspects can include generating the plurality of CNN layers based on an animation requirement of an image providing application executing in the computing device. In some aspects, at least one of the generated plurality of CNN layers can include a non-integer zoom ratio.
[0004] Further aspects can include a computing device having a processing system configured with processor-executable instructions for performing various operations corresponding to the methods discussed above. Further aspects can include a non-transitory processor-readable storage medium having stored thereon processor-executable instructions configured to cause a processing system to perform various operations corresponding to the method operations discussed above. Further aspects can include a computing device having various means for performing functions corresponding to the method operations discussed above. BRIEF DESCRIPTION OF DRAWINGS
[0005] The accompanying drawings, which are incorporated herein and constitute part of the specification, illustrate exemplary embodiments of the claims and, together with the general description given above and the detailed description given below, serve to explain features of the present disclosure.
[0006] Figure 1A is a conceptual diagram illustrating aspects of a method 100 for performing zoom operations on a computing device, in accordance with various embodiments.
[0007] Figure 1B is a conceptual diagram illustrating aspects of a method 150 for performing zoom operations on a computing device, in accordance with various embodiments.
[0008] Figure 2A and Figure 2B illustrates an example neural network suitable for implementation in a computing device, in accordance with various embodiments.
[0009] Figure 3A and Figure 3B illustrates example functional components that can be included in a convolutional neural network that can be implemented in a computing device configured to implement a general-purpose framework to accomplish continual learning, in accordance with various embodiments.
[0010] Figure 4 is a component block diagram illustrating an example computing system suitable for implementation as a system-in-a-package (SIP) implementation of various embodiments.
[0011] Figure 5A illustrates a method for performing zoom operations on a computing device, in accordance with various embodiments.
[0012] Figure 5B illustrates operations that can be performed as part of a method for performing zoom operations on a computing device, in accordance with various embodiments.
[0013] Figure 6 is a component block diagram illustrating an example computing device suitable for use in various embodiments.
[0014] Figure 7is a component block diagram illustrating an example wireless communication device suitable for use in various embodiments.
[0015] Figure 8 An example wearable computing device in the form of a smart watch is illustrated suitable for use in various embodiments. DETAILED DESCRIPTION
[0016] Various embodiments will be described in detail with reference to the drawings. Wherever possible, the same reference numbers will be used throughout the drawings to refer to the same or like parts. References made to particular examples and implementations are for illustrative purposes, and are not intended to limit the scope of the claims.
[0017] In general, various embodiments include methods for performing a zoom operation on an image or text on a display of a computing device in response to a user input for zooming the image or text and computing devices configured to implement the methods. Various embodiments can include responding to a user input for a zoom operation (“zoom user input”) by identifying a region of interest (ROI) of an image presented on a display of a computing device and determining a scaling ratio for a final zoomed- in presentation of the ROI. The zoom operation can include selecting two or more convolutional neural network (CNN) layers from a plurality of CNN layers, where each CNN layer includes a smaller scaling ratio than the scaling ratio for the final zoomed-in presentation of the ROI. The zoom operation can also include generating intermediate images using the selected CNN layers to provide a sequence of intermediate size image frames of the ROI, each intermediate size image frame having a size smaller than a size of the final zoomed-in presentation of the ROI and larger than a size of the ROI. The sequence of intermediate size image frames of the ROI can then be used to render an animation of a change in display size of the ROI. In this way, the computing device can be configured to implement dynamic dimensional CNN layer inference. The magnified image of the ROI at the final size or zoom factor can exhibit fine resolution achieved through the use of CNN image processing.
[0018] The term “computing device” can be used herein to refer to any or all of a quantum computing device, an edge computing device, an Internet access gateway, a modem, a router, a network switch, a residential gateway, an access point, an integrated access device (IAD), a mobile convergence product, a networking adapter, a multiplexer, a personal computer, a laptop computer, a tablet computer, a user equipment (UE), a smart phone, a personal or mobile multimedia player, a personal data assistant (PDA), a palm-top computer, a wireless electronic mail receiver, a multimedia Internet enabled cellular telephone, a game system (e.g., PlayStation® ™ , Xbox ™ , Nintendo Switch ™wearable devices (e.g., smart watches, head-mounted displays, fitness trackers, etc.), media players (e.g., DVD players, ROKU ™ , AppleTV ™ , etc.), digital video recorders (DVRs), automotive displays, portable projectors, 3D holographic displays, and other similar devices that include a display and a programmable processing system that can be configured to provide the functionality of the various embodiments.
[0019] The term "system on a chip" (SoC) is used herein to refer to a single integrated circuit (IC) chip that contains multiple resources or standalone processing systems integrated on a single substrate. A single SoC can contain circuitry for digital, analog, mixed-signal, and radio-frequency functionality. A single SoC can also include any number of general-purpose or special-purpose processors (e.g., network processors, digital signal processors, modem processors, video processors, etc.), memory blocks (e.g., ROM, RAM, Flash, etc.), and resources (e.g., timers, voltage regulators, oscillators, etc.), any one or more of which can function as a processing system. For example, a SoC can include an application processor that operates as the SoC's main processor, central processing unit (CPU), microprocessor unit (MPU), arithmetic logic unit (ALU), etc. A SoC can also include software for controlling the integrated resources and processors, as well as for controlling peripheral devices.
[0020] The term "system in a package" (SIP) can be used herein to refer to a single module or package that contains multiple resources, computing units, cores, processors, or processing systems on two or more IC chips, substrates, or SoCs. For example, a SIP can include a single substrate on which multiple IC chips or semiconductor dies are stacked in a vertical configuration. Similarly, a SIP can include one or more multi-chip modules (MCMs) on which multiple ICs or semiconductor dies are packaged into a unified substrate. A SIP can also include multiple independent SOCs coupled together via high-speed communication circuitry and packaged in close proximity, such as on a single motherboard, in a single UE, or in a single CPU device. The proximity of the SoCs facilitates high-speed communication as well as sharing of memory and resources.
[0021] The term“neural network” is used herein to refer to an interconnected group of processing nodes (e.g., neuron models, etc.) that collectively operate as a software application or process that controls the functioning of a computing device or produces a neural network inference. Individual nodes in a neural network can attempt to mimic biological neurons by receiving input data, performing simple operations on the input data to generate output data, and passing the output data (also referred to as“activations”) to the next node in the network. Each node can be associated with weight values that define or govern the relationship between input data and activations. The weight values can be determined during a training phase and iteratively updated as data flows through the neural network.
[0022] A deep neural network implements a hierarchical architecture in which the activations of a first layer of nodes become input to a second layer of nodes, the activations of the second layer of nodes become input to a third layer of nodes, and so on. Thus, computations in a deep neural network can be distributed across a population of processing nodes that make up a chain of computations. A deep neural network can also include activation functions and sub-functions between layers (e.g., rectified linear units that clip activations below zero, etc.). A first layer of nodes of a deep neural network can be referred to as an input layer. A final layer of nodes can be referred to as an output layer. Layers between the input layer and the final layer can be referred to as intermediate layers, hidden layers, or black box layers. Each layer in a neural network can have multiple inputs, and thus multiple previous layers or preceding layers. In other words, multiple layers can feed into a single layer. For ease of reference, some embodiments are described with reference to a single input or a single preceding layer. However, it should be understood that the operations disclosed and described in this application can apply to each input of a layer as well as multiple previous layers.
[0023] The term“convolutional neural network” (CNN) is used herein to refer to a deep neural network in which computations in at least one layer are structured as a convolution. Convolutional neural networks can also include multiple convolution-based layers, which allows the neural network to employ a hierarchy of very deep layers. In a convolutional neural network, a weighted sum of each output activation is computed based on a batch of inputs, and the same weight matrix (referred to as a“filter”) is applied to each output. These networks can also implement a fixed feed-forward structure in which all processing nodes that make up a chain of computations are used to process each task regardless of the input. In such feed-forward neural networks, all computations are performed as a sequence of operations on the outputs of previous layers. A final set of operations generates an overall inference result of the neural network, such as a probability that an image contains a particular object (e.g., a person, a cat, a watch, an edge, etc.) or information indicating that a proposed action should be taken.
[0024] The term“inference” can be used herein to refer to processes performed during runtime or during execution of a software application corresponding to a neural network. Inference can include traversing processing nodes in a neural network along a forward path to produce one or more values as an overall activation or overall“inference result.”
[0025] Computing devices can include displays and graphics hardware or subsystems that utilize machine learning techniques or artificial intelligence (AI) to improve the performance and capabilities of displays and graphics processing. In computing devices that utilize AI for image processing, the hardware can include specialized AI processing systems or AI accelerators designed to run machine learning algorithms more efficiently than other processing systems. CNNs are one of the most commonly used AI methods for image processing.
[0026] In computing devices with CNN-based AI image processing systems, the CNNs can run on the AI processing systems or accelerators to analyze and process images to be displayed on the screen of the computing device. Using AI image processing systems can improve the performance and efficiency of displays and graphics of computing devices. AI image processing systems can also improve the experience of users by supporting features such as improved image quality, real-time image enhancement, and more realistic virtual reality and augmented reality experiences. For example, AI image processing systems can be used to improve the realism of virtual reality and augmented reality experiences by improving 3D graphics rendering and object tracking.
[0027] Bilinear and bicubic scalars are interpolation scaling techniques that can be used to resize or scale images. For example, bilinear or bicubic scalars can be used to change the size of an image by increasing or decreasing the number of pixels that the image contains. Bilinear scaling can include using linear interpolation to determine the value of a new pixel based on the values of nearby pixels in the original image. Bilinear scaling is particularly useful when the original image is smooth and does not contain many fine details. Bilinear scaling can result in loss of detail and blurring when used to enlarge images with high levels of detail or texture. Bicubic scaling is a more complex technique that uses cubic interpolation to determine the value of a new pixel based on the values of nearby pixels in the original image. Bicubic scaling can produce better results than bilinear scaling, particularly when scaling up images with a wide range of image types, including those with high levels of detail or texture. Bicubic scaling can be slower or more processor-intensive than bilinear scaling. Both bilinear and bicubic scaling techniques can be used for image zooming, upscaling, and downscaling.
[0028] AI super-resolution is a technique that can be used to increase the resolution of an image or video. AI super-resolution is often used in image and video processing tasks to improve the visual quality of an image or video. AI super-resolution can be implemented by a CNN that takes a low-resolution image as input and generates a high-resolution version of the image as output. AI super-resolution can be used to improve the visual quality of an image and / or video and can be used for image zooming and upscaling. AI super-resolution can be slower than bilinear and bicubic scaling, but generally produces much higher image quality.
[0029] Some applications executing on computing devices, particularly mobile computing devices (e.g., smartphones), such as map applications and some web browser applications, provide a run-time zoom or run-time magnification functionality for content generated by the application. For example, some map applications are able to magnify map content within the map application. Such applications support the ability to magnify content within the application in response to user input to a touchscreen device or trackpad device (“zoom user input”). For example, zoom user input can include increasing the distance between two fingers in contact with the touchscreen (sometimes referred to as a “reverse pinch” or “spread” gesture), a double tap and drag gesture, or another suitable user input.
[0030] However, generally outside of such applications, computing devices typically do not provide run-time zoom functionality, in part because bilinear / bicubic image scaling and AI (CNN) based super-resolution are not well suited for the task. Bilinear or bicubic image scaling typically does not provide enough content detail, while AI (CNN) based super-resolution does not provide smooth animation and also consumes battery power at an undesirable rate.
[0031] For example, a CNN model can be executed or generated (e.g., in a graphics processing unit (GPU) or digital signal processor (DSP)) per frame to provide a hybrid zoom animation based on CNN super-resolution. A typical hybrid CNN super-resolution function incurs significant power consumption and can be difficult to provide smooth animation (e.g., without noticeable frame drops), especially on advanced displays providing high display rates of up to 120Hz or 144Hz.
[0032] A typical "super zoom" pipeline uses a fixed upscaling ratio (typically in the range of 1 to 5) that provides a balance between visual quality of the upscaled image and power consumption. CNN models (such as can be implemented in various computing devices) can be quantized and pre-compiled for various hardware (e.g., digital signal processors) and / or runtime software / hardware contexts. During the process of quantization and runtime optimization, CNN model layer weights (such as size, dimension, height, width, etc.) must remain fixed and constant. However, CNN models with fixed upscaling ratios are limited in their efficiency and performance in performing image processing tasks such as image upscaling.
[0033] Various embodiments include methods and processing systems configured to perform methods of performing zoom operations on a computing device by implementing one or more CNN models having a plurality of different dynamic dimension CNN layers, where each layer is suitable for a particular different scaling factor or ratio. In some embodiments, a processing system of a computing device can identify a ROI of an image presented on a display of the computing device in response to receiving a zoom user input of the computing device, such as a "pinch-out" gesture in the form of two fingertips touching a touch-sensitive display (i.e., a "touchscreen"). In some embodiments, the processing system can determine a scaling ratio for a final post-zoom presentation of the ROI based on various factors, such as a size of the ROI, a size of the display, etc.
[0034] In various embodiments, the processing system can employ or apply a "multi-headed" or "multi-layer" CNN, where each layer is trained to a different scaling ratio. Using such a multi-layer CNN approach enables the processing system to dynamically change the scaling ratio used to upscale the ROI in response to user input. In some embodiments, the processing system can select two or more CNN layers from a plurality of CNN layers, where each CNN layer includes a smaller scaling ratio than a scaling ratio for a final post-zoom presentation of the ROI. The processing system can use each of the selected CNN layers to generate an intermediate image to produce or generate a sequence of intermediate size image frames of the ROI, where a size of each intermediate image is smaller than a size of the final post-zoom presentation of the ROI and larger than a size of the ROI. The processing system can use the sequence of intermediate size image frames of the ROI to render an animation of a change in display size of the ROI. In some embodiments, at least one of the selected CNN layers includes a non-integer scaling ratio.
[0035] In some embodiments, the processing system can select a number of the plurality of CNN layers based on application requirements of an application executing in the computing device that provides the image. For example, an application that generates an image including the ROI can provide information to the image processor indicating requirements such as image generation frequency or speed (e.g., a number of intermediate images to be generated for a zoom animation).
[0036] In some embodiments, the processing system can select the number of CNN layers based on an available number of memory of the computing device. For example, a CNN layer can require a particular number of memory in which to perform and conduct image scaling operations. Based on the available number of memory of the computing device, the processing system can allocate memory for use by one or more CNN layers, which can be less than the total amount of available computing device memory, and based on the number of allocated memory, the computing device can select the number of CNN layers.
[0037] In some embodiments, the processing system can generate the plurality of CNN layers during runtime, such as when an application that generates images that can be upscaled is launched, or when such an application is executing, or in response to the computing device receiving a zoom user input. In some embodiments, the processing system can generate the plurality of CNN layers based on animation requirements of an image providing application executing in the computing device. For example, an application that generates images that include a ROI can provide information to the image processor indicating requirements for a minimum image resolution, such as intermediate images (e.g., images having various sizes between the original image size and the final scaled image size). In some embodiments, at least one of the generated CNN layers can include a non-integer scaling ratio.
[0038] Various embodiments improve the efficiency and performance of a processing system that performs zoom operations on image elements by improving the efficiency of a CNN used to perform image scaling operations, such as image upsampling.
[0039] Figure 1A is a conceptual diagram illustrating aspects of a method 100 for performing zoom operations on a computing device, in accordance with various embodiments. Components for performing the operations of the method 100 can include one or more processing systems (e.g., processors 410, 412, 414, 416, 418, 421, 422, 452, 460) in a computing device (e.g., 118, 400). Figure 4 in the computing device 118).
[0040] The processing system can identify a region of interest (ROI) 122 of an image 120 presented on a display of the computing device 118 based on a zoom user input received by the computing device 118 (e.g., received by a touchscreen input device of the computing device 118). In some embodiments, the processing system can use region histogram analysis and / or image segmentation to identify the ROI 122.
[0041] The processing system can determine a scaling ratio for the final post-zoomed rendering of the ROI 126. In some embodiments, the processing system can determine the scaling ratio for the final post-zoomed image based on one or both of image rendering requirements or inference speed of a frame-generating convolutional neural network (CNN) (“CNN model”) 124. In some embodiments, the processing system can determine (calculate, identify) dimensions of the ROI 122, such as width and height. In some embodiments, the processing system can determine dimensions of a display of the computing device 118, such as width and height, number of vertical and horizontal pixels, or other suitable dimensions. In some embodiments, the processing system can determine a time and / or number of frames required to expand the ROI 122 to fill the available display dimensions of the computing device 118.
[0042] In some embodiments, the processing system can identify dimensions of an image element within the ROI 122, such as a face or any other suitable image element. For example, the processing system can identify a face within the ROI 122 and can determine a time and / or number of frames required to expand the face to fill the display dimensions of the computing device 118.
[0043] In some embodiments, the image rendering requirements can include frames per second (FPS) generated by reducing the final post-zoomed rendering to a generated sequence of intermediate-sized image frames. In some embodiments, determining the scaling ratio for the final post-zoomed rendering of the ROI 122 can also be based on dimensions of the ROI 122 and dimensions of the display of the computing device 118. In some embodiments, determining the scaling factor for the final post-zoomed rendering of the ROI 122 can also be based on dimensions of an image element within the ROI 122 and dimensions of the display of the computing device 118. In some embodiments, the processing system can determine the scaling factor for the final post-zoomed rendering in response to detection and issuance of a user zoom input.
[0044] In some embodiments, the processing system can select a CNN layer 128a, 128b, 128c from a plurality of CNN layers. Each of the CNN layers 128a, 128b, 128c can include a scaling ratio that is smaller than the scaling ratio for the final post-zoomed rendering of the ROI 126. In some embodiments, at least one of the CNN layers 128a, 128b, 128c can include a non-integer scaling ratio (e.g., a scaling ratio of 1.2, 3.6, etc.). In some embodiments, each of the CNN layers 128a, 128b, 128c can be generated to provide a scaling ratio (which can be a non-integer scaling ratio) during a process of compilation and quantization of the CNN model 124. In some embodiments, the processing system quantization can use a scaling ratio that is based on a range of ratios or discrete ratios (which can be represented as r). For example, during the training process, the range of scaling ratios can be, for example, lx-6x, and during the quantization process, the scaling ratio for each CNN layer (e.g., 128a, 128b, 128c) can be quantized to a value such as 1.2, 2.1, 3.6, or another suitable scaling ratio. In operation, based on the scaling ratio, each CNN layer can include a different size. In some embodiments, each CNN layer can include a tensor shape with a plurality of indices such as (2130, 5130) x (64, 3, 3, 3), (1510, 3625) x (64, 3, 3, 3), or (710, 1710) x (64, 3, 3, 3). Each CNN layer can include weights of different sizes that are computed for the trained CNN model 124 during offline inference. In some embodiments, during runtime, the processing system can scale the scaling ratio based on the pipeline requirements (e.g., r ) select two or more CNN layers 128a, 128b, 128c from the plurality of CNN layers. For example, based on the requested scaling ratio of 1.2, the processing system can select a CNN layer with a size of (710, 1710) x (64, 3, 3, 3).
[0045] In some embodiments, the processing system can select the number of the plurality of CNN layers 128a, 128b, 128c based on application requirements of an application executing in the computing device 118 that provides the image. For example, an application that generates an image including an ROI can provide information to the image processor indicating requirements such as image generation frequency or speed (e.g., the number of intermediate images to be generated for a zoom animation).
[0046] In some embodiments, the processing system can select the number of the plurality of CNN layers 128a, 128b, 128c based on the available number of memory of the computing device. For example, a CNN layer can require a particular number of memory in which to execute and conduct image scaling operations. Based on the available number of memory of the computing device 118, the processing system can allocate memory for use by one or more CNN layers, which can be less than the total amount of available computing device memory, and based on the number of allocated memory, the processing system can select the number of the plurality of CNN layers.
[0047] The processing system can use each of the selected CNN layers 128a, 128b, 128c to generate an intermediate image 130, 132, 134 to produce or generate a sequence of intermediate size image frames of the ROI 122, where each intermediate image 130, 132, 134 is of a size that is less than the final zoomed after presentation size of the ROI 126 and greater than the size of the ROI 122. The processing system can use the sequence of intermediate size image frames 130, 132, 134 of the ROI to render an animation of the change in display size of the ROI.
[0048] Figure 1B is a conceptual diagram illustrating aspects of a method 150 for performing zoom operations on a computing device, in accordance with various embodiments. The components for performing the operations of the method 150 can include one or more processing systems (e.g., processors 410, 412, 414, 416, 418, 421, 422, 452, 460) in a computing device (e.g., 118, 400). Figure 4
[0049] In some embodiments, the processing system can generate the plurality of CNN layers 128a, 128b, 128c during runtime, such as when an application that can generate images that can be upscaled is launched, or when such an application is executing, or in response to the computing device receiving a zoom user input. In some embodiments, the processing system can generate the plurality of CNN layers 128a, 128b, 128c based on animation requirements of an image-providing application executing in the computing device. For example, an application that generates images including ROIs can provide information to the image processor indicating requirements of a minimum image resolution, such as for intermediate images (e.g., images having various sizes between the original image size and the final scaled image size). In some embodiments, at least one of the generated CNN layers can include a non-integer scaling ratio.
[0050] In some embodiments, the processing system can implement a residual dense network (RDN) 150. The RDN 150 can include a residual dense block (RDB) 152 and a feature learning model 154, such as a densely connected block. The RDB 152 can be configured to extract rich local features via densely connected convolutional layers. The feature learning model 154 can be configured to generate shared feature maps for arbitrary scaling factors. The RDN 150 can also include a meta-upscaling module 156. The meta-upscaling module 156 can be configured to receive a sequence of coordinate-dependent vectors and scaling-dependent vectors as input to generate weights for one or more CNN layers 128a, 128b, 128c. The meta-upscaling module 156 can also be configured to apply the CNN layers 128a, 128b, 128c to generate intermediate images 130, 132, 134.
[0051] Figure 2A and Figure 2B illustrates an example neural network 200 suitable for implementation in a computing device, in accordance with various embodiments. Referring to FIGS. 1-3 Figure 2A The neural network 200 can include an input layer 202, an intermediate layer 204, and an output layer 206. Each of the layers 202, 204, 206 can include one or more processing nodes that receive input values, perform computations based on the input values, and propagate results (activations) to the next layer.
[0052] In a feedforward neural network, such as neural network 200, all computations are performed as a sequence of operations on the outputs of the previous layer. The final set of operations generates the output of the neural network, such as a probability that an image contains a particular item (e.g., a dog, a cat, etc.) or information indicating that a proposed action should be taken. The final output of the neural network can correspond to the task that neural network 200 can be performing, such as determining whether an image contains a particular item (e.g., a dog, a cat, etc.). Many neural networks 200 are stateless. The output for an input is always the same, regardless of the sequence of inputs previously processed by the neural network 200.
[0053] Neural network 200 can include a fully connected (FC) layer, which is sometimes also referred to as a multilayer perceptron (MLP). In a fully connected layer, all outputs are connected to all inputs. The activation of each processing node is computed as a weighted sum of all inputs received from the previous layer.
[0054] An example computation performed by a processing node and / or neural network 200 can be: where W ij is a weight, x i is an input to the layer, y j is an output activation of the layer, f(•) is a non-linear function, and b is a bias, which can vary for each node (e.g., b j ). As another example, neural network 200 can be configured to receive pixels of an image in a first layer (i.e., input values) and generate outputs indicating the presence of different low-level features (e.g., lines, edges, etc.) in the image. At subsequent layers, these features can be combined to indicate the possible presence of higher-level features. For example, in the training of a neural network for image upsampling, an output layer can generate probability values that various lines, edges, colors, etc. are present in an upscaled (i.e., upsized) version of an input image. In this way, a neural network can use information available in a low-resolution small image to predict elements of the image that should be added to fill in the details between pixels when the image is upscaled.
[0055] The neural network 200 can be trained to upscale images using a database of images of different sizes and resolutions. During this training process, a small image can be provided as input, and a larger, higher resolution image of the same subject can be provided as the expected / desired output. Learning is accomplished by comparing the output generated by the neural network 200 to the expected / desired output and adjusting the weights. The difference between the expected / desired output and the output generated by the neural network 200 is referred to as the loss (L). The weights and biases in the neural network are then adjusted to make the output image closer to the provided expected output image, thereby reducing the loss. During training, the weights can be updated using a hill-climbing optimization process known as “gradient descent.” The gradient indicates how the weights should be changed in order to reduce the loss (L). The multiple of the gradient of the loss with respect to each weight, which can be the partial derivative of the loss (L) with respect to the weight, can be used to update the weights and biases. This learning process can be repeated for a large database of images. ij ). The gradient indicates how the weights should be changed in order to reduce the loss (L). The multiple of the gradient of the loss with respect to each weight, which can be the partial derivative of the loss (L) with respect to the weight, can be used to update the weights and biases. This learning process can be repeated for a large database of images. , , The partial derivative of the gradient can be used to update the weights and biases. This learning process can be repeated for a large database of images.
[0056] Figure 2B The way in which the partial derivative of the gradient is calculated through a process known as backpropagation is illustrated. Referring to FIGS. 1 through Figure 2B Backpropagation can operate by passing values backward through the network to calculate how the loss is affected by each weight. The backpropagation calculations can be similar to the calculations used when traversing the neural network 200 in the forward direction (i.e., during inference). To improve performance, the loss (L) from multiple sets of input data (“a batch”) can be collected and used in a single pass to update the weights. Many passes can be needed to train the neural network 200 with weights suitable for use during inference (e.g., at runtime or during execution of a software application).
[0057] The overall structure of the neural network 200, as well as the operation of the processing nodes, does not change as the neural network learns the task. After such training is complete, the neural network 200 can use the determined weights and biases to process any image for upsizing.
[0058] Figure 3A and Figure 3B Example functional components that can be included in a convolutional neural network 300 in accordance with various embodiments are illustrated, which can be implemented in a processing system configured to implement a general-purpose framework to accomplish continual learning.
[0059] Referring to FIGS. 1 through Figure 3BIn the example illustrated in FIG. 3, the convolutional neural network 300 can include a first layer 301 and a second layer 311. Each layer 301, 311 can include one or more activation functions. In the example illustrated in FIG. 3, each layer 301, 311 includes a convolutional functional component 302, 312, a non-linear functional component 304, 314, a normalization functional component 306, 316, a pooling functional component 308, 318, and a quantization functional component 310, 320. It will be appreciated that the functional components 302-310 or 312-320 can be implemented as part of a neural network layer, or outside of a neural network layer, in various embodiments. It will also be appreciated that the order of operations illustrated in FIG. 3 is merely an example, and is not intended to limit the various embodiments to any given order of operations. In various embodiments, the order and / or inclusion of operations of the functional components 302-310 or 312-320 can vary in any given layer. For example, the normalization operation by the normalization functional component 306, 316 can occur after the convolution by the convolutional functional component 302, 312 and before the non-linear operation by the non-linear functional component 304, 314. Figure 3A In the example illustrated in FIG. 3, the convolutional functional component 302, 312 can be a convolution operation. The convolutional functional component 302, 312 can be configured to generate a matrix of output activations referred to as a feature map. The feature map generated in each successive layer 301, 311 generally includes values representing successively higher-level abstractions of the input data (e.g., lines, shapes, objects, etc.). Figure 3A In the example illustrated in FIG. 3, the non-linear functional component 304, 314 can be a non-linear activation function. The non-linear functional component 304, 314 can be configured to introduce non-linearity into the output activations of its layer 301, 311. In various embodiments, this can be implemented via a sigmoid function, a hyperbolic tangent function, a rectified linear unit (ReLU), a leaky ReLU, a parametric ReLU, an exponential LU function, a maxout function, a swish, etc.
[0060] In the example illustrated in FIG. 3, the normalization functional component 306, 316 can be a normalization function. The normalization functional component 306, 316 can be configured to control the input distribution across layers to speed up training and improve the accuracy of the output or activations. For example, the distribution of inputs can be normalized to have zero mean and unit standard deviation. The normalization function can also use batch normalization (BN) techniques to further scale and shift the values to improve performance.
[0061] In the example illustrated in FIG. 3, the pooling functional component 308, 318 can be a pooling operation. The pooling functional component 308, 318 can be configured to reduce the dimensionality of the feature map generated by the convolutional functional component 302, 312 and / or otherwise allow the convolutional neural network 300 to be resistant to small shifts and distortions in values.
[0062] In the example illustrated in FIG. 3, the quantization functional component 310, 320 can be a quantization operation. The quantization functional component 310, 320 can be configured to reduce the precision of the output activations of its layer 301, 311. In various embodiments, this can be implemented via a uniform quantization, a uniform quantization with clipping, a uniform quantization with clipping and scaling, a non-uniform quantization, a non-uniform quantization with clipping, a non-uniform quantization with clipping and scaling, etc.
[0063] In the example illustrated in FIG. 3, the quantization functional component 310, 320 can be a quantization operation. The quantization functional component 310, 320 can be configured to reduce the precision of the output activations of its layer 301, 311. In various embodiments, this can be implemented via a uniform quantization, a uniform quantization with clipping, a uniform quantization with clipping and scaling, a non-uniform quantization, a non-uniform quantization with clipping, a non-uniform quantization with clipping and scaling, etc.
[0064] In some embodiments, the input to the first layer 301 can be structured to form a set of three-dimensional input feature maps 352 that form channels of an input feature map. For example, the convolutional neural network 300 can include a batch size of N three-dimensional feature maps 352 of height H and width W, each having an input feature map of C channels (illustrated as two-dimensional maps in C channels), and M three-dimensional filters 354, including C filters for each channel (also illustrated as two-dimensional filters for C channels). Applying 1 to M filters 354 to 1 to N three-dimensional feature maps 352 results in N output feature maps 356, including M channels of width F and height E. As illustrated, each channel can be convolved with a three-dimensional filter 354. The results of these convolutions can be summed across all channels to generate an output activation of the first layer 301 in the form of channels of output feature maps 356. Additional three-dimensional filters can be applied to the input feature maps 352 to create additional output channels, and multiple input feature maps 352 can be processed together as a batch to improve reuse of filter weights. The results of the output channels (e.g., a set of output feature maps 356) can be fed to a second layer 311 in the convolutional neural network 300 for further processing.
[0065] Figure 4 is an example component block diagram of an example computing system implemented as a system-in-a-package (SIP) 400 suitable for implementing various embodiments. Referring to FIG. 1 through Figure 4 Various embodiments can be implemented in various architectures for processing systems, including single processors, multiple processors, and multi-processor processing systems, systems-on-a-chip (SOCs), SIPs, and any combination thereof.
[0066] The SIP 400 can include two SOCs 402, 404, a clock 406, and a voltage regulator 408. In some embodiments, the first SOC 402 can operate as a central processing unit (CPU) of a computing device that executes arithmetic, logical, control, and input / output (I / O) operations specified by instructions of software applications by executing those instructions. In some embodiments, the second SOC 404 can operate as a specialized processing unit. For example, the second SOC 404 can operate as a specialized 5G processing unit responsible for managing high capacity, high speed (e.g., 5 Gbps, etc.), and / or ultra-high frequency short wavelength (e.g., 28 GHz millimeter wave spectrum, etc.) communications.
[0067] The first SOC 402 can include a digital signal processor (DSP) 410, a modem processor 412, a graphics processor 414, an application processor 416, one or more coprocessors 418 (e.g., vector co-processor) connected to one or more of these processors, memory 420, a deep processing unit (DPU) 421, an AI processor 422, system components and resources 424, an interconnect / bus module 426, one or more temperature sensors 430, a thermal management unit 432, and a thermal power envelope (TPE) component 434. The second SOC 404 can include a 5G modem processor 452, a power management unit 454, an interconnect / bus module 464, a plurality of millimeter wave transceivers 456, memory 458, and various additional processors 460, such as an application processor, a packet processor, etc.
[0068] Each processor 410, 412, 414, 416, 418, 421, 422, 452, 460 can include one or more cores, and each processor / core can perform operations independently of the other processors / cores. For example, the first SOC 402 can include a processor that executes a first type of operating system (e.g., FreeBSD, LINUX, OS X, etc.) and a processor that executes a second type of operating system (e.g., MICROSOFT WINDOWS 10). Additionally, any or all of the processors 410, 412, 414, 416, 418, 421, 422, 452, 460 can be included as part of a processor cluster architecture (e.g., a synchronous processor cluster architecture, an asynchronous or heterogeneous processor cluster architecture, etc.).
[0069] The first SOC 402 and the second SOC 404 can include various system components, resources, and custom circuitry for managing sensor data, analog-to-digital conversion, wireless data transmission, and for performing other specialized operations such as decoding data packets and processing encoded audio and video signals for presentation in a web browser. For example, the system components and resources 424 of the first SOC 402 can include power amplifiers, voltage regulators, oscillators, phase-locked loops, peripheral bridges, data controllers, memory controllers, system controllers, access ports, timers, and other similar components for supporting the processors and software clients running on the UE device. The system components and resources 424 can also include circuitry for interfacing with peripheral devices such as cameras, electronic displays, wireless communication devices, external memory chips, etc.
[0070] The first SOC 402 and the second SOC 404 can communicate via an interconnect / bus module 450. The various processors 410, 412, 414, 416, 418 can be interconnected to one or more memory elements 420, system components and resources 424, and a thermal management unit 432 via an interconnect / bus module 426. Similarly, the processor 452 can be interconnected to a power management unit 454, a millimeter wave transceiver 456, a memory 458, and various additional processors 460 via an interconnect / bus module 464. The interconnect / bus modules 426, 450, 464 can include an array of reconfigurable logic gates and / or implement a bus architecture (e.g., CoreConnect, AMBA, etc.). Communication can be provided by an advanced interconnect, such as a high-performance network-on-chip (NoC).
[0071] The first SOC 402 and / or the second SOC 404 can further include an input / output module (not illustrated) for communicating with resources external to the SOC, such as a clock 406, a voltage regulator 408, a screen sensor unit 415, and a wireless transceiver 466 (e.g., a cellular wireless transceiver, a Bluetooth transceiver, etc.). The resources external to the SOC (e.g., the clock 406, the voltage regulator 408, the screen sensor unit 415, and the wireless transceiver 466) can be shared by two or more of the internal SOC processors / cores.
[0072] In some embodiments, any or all of the processors 410, 412, 414, 416, 418, 421, 422, 452, 460 can implement a CNN-based AI image processing system, a bilinear-bicubic image processing pipeline, and / or a convolutional neural network super-resolution (CNN SR) image processing pipeline. For example, in some embodiments, the GPU 414 and / or the DPU 421 can implement the bilinear-bicubic pipeline, and / or the AI processor 422, the GPU 414, and / or the DSP 410 can implement the CNN SR pipeline.
[0073] In some implementations, any or all of processors 410, 412, 414, 416, 418, 421, 422, 452, and 460 may be configured to work in conjunction with screen sensor unit 415 to perform zoom operations on the display of a computing device, depending on the implementation. For example, screen sensor unit 415 may detect touch gestures from an end user and conditionally trigger layer dynamic seamless zoom and corresponding sequential zoom animation frames with continuously varying boost step factors (1.0--1.01—1.02--1.03-----xxxxx---------1.99, 2.00, 2.01-----xxxxx-----3.0). Screen sensor unit 415 may detect zoom user input on an image layer within the display presented on the electronic display of the computing device and transmit the information to GPU 414 and AI processor 422. GPU 414 may receive the zoom user input and use interpolation scaling techniques in a first image processing pipeline to adjust the display size of the image layer. In parallel, the AI processor 422 can use super-resolution techniques in a second image processing pipeline to adjust the display size of the image layer. The AI processor 422, GPU 414, and / or another processor (e.g., application processor 416) can utilize the output of the super-resolution technique from the second image processing pipeline to upscale the output of the interpolation scaling technique in the first image processing pipeline to generate upscaled frames of the image layer, and enable the electronic display of the computing device to display the upscaled frames of the image layer.
[0074] Figure 5A An example of method 500a for performing zoom operation on a computing device according to various implementation schemes is illustrated. Refer to Figures 1 to 12. Figure 5A The components used to perform the operations of method 500a may include a processing system as described herein. The processing system may include one or more processors (e.g., processors 410, 412, 414, 416, 418, 421, 422, 452, 460) and / or hardware elements in a computing device (e.g., 118, 400), any one or a combination of which may be configured to perform any of the operations of method 500a. Furthermore, one or more processors within the processing system may be configured with software or firmware to perform various operations of the method. To encompass any of the processors, hardware elements, and software elements that may be involved in performing method 500a, the component performing the method operations is generally referred to as the “processing system”.
[0075] In box 502, the processing system may identify the ROI of the image presented on the display of the computing device based on zoom user input received by the computing device.
[0076] In block 504, the processing system can determine a scaling ratio for the final post-zoomed presentation of the ROI. The processing can determine the scaling ratio based on a number of factors known to the processor, such as, but not limited to, the initial size of the ROI, the size of the display of the computing device, the ratio of the ROI size to the display size, the computing resources available to perform the scaling operation, the estimated computation time required to perform the scaling operation, and the like.
[0077] In block 506, the processing system can select two or more CNN layers from the plurality of CNN layers. In some embodiments, each CNN layer can include a smaller scaling ratio than the scaling ratio for the final post-zoomed presentation of the ROI. In some embodiments, at least one of the selected CNN layers can include a non-integer scaling ratio. In some embodiments, the processing system can select the number of CNN layers based on application requirements of an application executing in the computing device that provided the image. In some embodiments, the processing system can select the number of CNN layers based on an available number of memory of the computing device. In some embodiments, the processing system can select the number of CNN layers based on the determined scaling ratio for the final post-zoomed presentation of the ROI.
[0078] In block 508, the processing system can generate an intermediate image according to each of the selected CNN layers. The generated intermediate images can provide a sequence of intermediate size image frames of the ROI, where each image is smaller in size than the final post-zoomed presentation of the ROI and larger in size than the ROI.
[0079] In block 510, the processing system can render an animation of the change in display size of the ROI using the sequence of intermediate size image frames of the ROI. Rendering the animation can involve displaying even larger images sequentially. In some embodiments, the processing system can use a linear scaling technique to enlarge the intermediate size image frames to provide additional image frames in the animation sequence, further smoothing the increase in ROI size from frame to frame.
[0080] Figure 5B Operations 500b are illustrated that can be performed as part of method 500a of performing a zoom operation on a computing device, in accordance with various embodiments. Reference is made to FIGS. 1-4 for context. Figure 5BThe components for performing the operations of 500b can include a processing system as described herein. The processing system can include one or more processors (e.g., processors 410, 412, 414, 416, 418, 421, 422, 452, 460) and / or hardware elements in a computing device (e.g., 118, 400), any or combination of which can be configured to perform any of the operations of the operations of 500b. Further, one or more processors within the processing system can be configured with software or firmware to perform various operations of the method. To cover for any of the processors, hardware elements, and software elements that can be involved in performing the operations of 500b, the element performing the operations of the method is generally referred to as a “processing system.”
[0081] In block 520, the processing system can generate a plurality of CNN layers based on animation requirements of an image providing application executing in the computing device.
[0082] The processing system can identify a ROI of an image presented on a display of the computing device based on a zoom user input received by the computing device in block 502 of method 500a as described.
[0083] Figure 6 is a component block diagram of a computing device suitable for use with various embodiments. Reference is made to FIGS. 1 through Figure 6 In some embodiments, the computing device can be implemented in the form of a laptop computer 600. In various embodiments, the laptop computer 600 can include a touchpad (or trackpad) touch surface 617 that serves as a pointing device for the computer, and thus can receive pinch, spread, drag, scroll, flick gestures, etc., similar to those that can be implemented on a computing device equipped with a touchscreen display. The laptop computer 600 can include a processing system 602 coupled to volatile memory 612 and a mass storage device (such as a disk drive of flash memory) 613. Additionally, the laptop computer 600 can include one or more antennas 608 for transmitting and receiving electromagnetic radiation, which can be connected to a wireless data link and / or to a cellular telephone transceiver 616 coupled to the processing system 602. The laptop computer 600 can also include a transceiver 614 that uses various short-range communication protocols, such as the Institute of Electrical and Electronics Engineers (IEEE) 802.11 Bluetooth® wireless IEEEThe laptop computer 600 can also include a compressed optical disc (CD) drive 615 coupled to the processing system 602. The laptop computer 600 can include a touchpad 617, a keyboard 618, and a display 619 all coupled to the processing system 602. Other configurations of the laptop computer 600 can include a computer mouse or trackball coupled to the processing system (e.g., via a universal serial bus (USB) input) as is well known, which can also be used in connection with various embodiments.
[0084] Figure 7 is a component block diagram of a computing device suitable for use with various embodiments. Reference is made to FIGS. 1 through Figure 7 In some embodiments, the computing device can be implemented in the form of a smart phone 700. The smart phone 700 can include a first circuit 402 coupled to a second circuit 404. The first SoC 402 and the second SoC 404 can be coupled to internal memory 716, a display 712, and a speaker 714. The display 712 can include or incorporate a touch-sensitive input device configured to receive input, such as pinch, zoom, drag, scroll, flick gestures, and the like. The first circuit 402 and the second circuit 404 can also be coupled to at least one subscriber identity module (SIM) 740 and / or a SIM interface, which can store information supporting a first 5G NR subscription and a second 5G NR subscription, the first 5G NR subscription and the second 5G NR subscription supporting services over a 5G non-standalone (NSA) network.
[0085] The smart phone 700 can include an antenna 704 for transmitting and receiving electromagnetic radiation, which can be connected to a wireless transceiver 466 coupled to one or more processing systems in the first circuit 402 and / or the second circuit 404. The smart phone 700 can also include a menu selection button or rocker switch 720 for receiving user input. The smart phone 700 can also include a sound encoding / decoding (CODEC) circuit 710 that digitizes sound received from a microphone into data packets suitable for wireless transmission and decodes received sound data packets to generate analog signals that are provided to a speaker to generate sound. Also, one or more of the processing systems in the first circuit 402 and the second circuit 404, the wireless transceiver 466, and the CODEC 710 can include digital signal processor (DSP) circuitry 740.
[0086] Figure 8 is a component block diagram of a computing device suitable for use with various embodiments. Reference is made to FIGS. 1 through Figure 8 In some embodiments, the computing device can be implemented in the form of a smart watch 800.
[0087] The smart watch 800 can include a SoC 802 that includes two or more processing systems (e.g., an application processor, a low-power processor) coupled to internal memories 804 and 806. The internal memories 804, 806 can be volatile or non-volatile memories, and can also be secure and / or encrypted memories, or non-secure and / or non-encrypted memories, or any combination thereof. The SoC 802 can also be coupled to a touch screen display 820, such as a resistive-sensing touch screen, a capacitive-sensing touch screen, an infrared-sensing touch screen, etc. Additionally, the smart watch 800 can have one or more antennas 808 for transmitting and receiving electromagnetic radiation, which can be connected to one or more wireless data links 812, such as one or more Bluetooth® wireless data links that can be coupled to the SoC 802. The SoC 802 can also be coupled to a power supply 814, such as a rechargeable battery or a solar cell. ® Transceiver.
[0088] The smart watch 800 can also include physical and / or virtual buttons 822 and 810 for receiving user input, as well as a slide sensor 816 for receiving user input. The touch screen display 820 can be coupled to a touch screen interface module configured to receive signals from the touch screen display 820 indicative of the location on the screen where a user's fingertip or stylus is touching the surface, and output information to the SoC 802 about the coordinates of the touch event. The physical and / or virtual buttons 822 and the touch screen display 820 can be configured to receive inputs such as pinch, zoom, drag, scroll, flick gestures, etc. Furthermore, the SoC 802 can be configured with processor-executable instructions to correlate images rendered on the touch screen display 820 with the locations of touch events received from the touch screen interface module in order to detect when a user has interacted with a graphical interface icon, such as a virtual button.
[0089] The SoC 802 can be any programmable microprocessor, microcomputer, multiple processor chip, or processing system that can be configured by software instructions (applications) to perform a variety of functions, including the functions of the various embodiments. In some devices, multiple processing systems can be provided, such as one for wireless communication functions and one for running other applications. Typically, software applications can be stored in the internal memory before they are accessed and loaded into the SoC 802. The SoC 802 can include internal memory sufficient to store the application software instructions. In many devices, the internal memory can be a volatile or non-volatile memory (e.g., flash) or a hybrid of the two. For the purposes of this description, a general reference to memory refers to memory accessible to the SoC 802, including internal memory or removable memory inserted into the wearable device, as well as memory within the SoC 802 itself.
[0090] The processing system of the laptop computer 600, the smartphone 700, and the smartwatch 800 can be any programmable microprocessor, microcomputer, multiple processor chip, or processing system capable of executing software instructions (applications) configured to perform the functions of the various embodiments described. In some computing devices, multiple processors can be provided, such as one processing system dedicated to wireless communication functions and one processing system dedicated to running other applications within a first circuitry and a second circuitry. The software applications can be stored in memory and then accessed and loaded into the processing system. The processing system can include internal memory sufficient to store the application software instructions.
[0091] Specific implementation examples are described in the following paragraphs. While some of the following implementation examples are described in the form of example methods, further example implementations can include: example methods discussed in the following paragraphs implemented by a computing device including a processing system configured with processor-executable instructions to perform operations of the methods of the following implementation examples; example methods discussed in the following paragraphs implemented by a computing device including components to perform functions of the methods of the following specific implementation examples; and example methods discussed in the following paragraphs can be implemented as non-transitory processor-readable storage media having stored thereon processor-executable instructions configured to cause a processing system of a computing device to perform operations of the methods of the following specific implementation examples.
[0092] Example 1. A method of performing a zoom operation on a computing device, the method comprising: identifying a region of interest (ROI) of an image presented on a display of the computing device based on a zoom user input received by the computing device; determining a scaling ratio for a final post-zoom presentation of the ROI; selecting two or more convolutional neural network (CNN) layers from a plurality of CNN layers, each CNN layer including a smaller scaling ratio than the scaling ratio for the final post-zoom presentation of the ROI; generating an intermediate image from each of the selected CNN layers, including a sequence of intermediate size image frames of the ROI, each intermediate size image frame having a size smaller than a size of the final post-zoom presentation of the ROI and larger than a size of the ROI; and rendering an animation of a change in display size of the ROI using the sequence of intermediate size image frames of the ROI.
[0093] Example 2. The method of example 1, wherein at least one of the selected CNN layers includes a non-integer scaling ratio.
[0094] Example 3. The method of any of examples 1 and 2, wherein selecting two or more CNN layers from the plurality of CNN layers comprises selecting a number of the plurality of CNN layers based on an application requirement of an application executing in the computing device that provides the image.
[0095] Example 4. The method of any of examples 1-3, wherein selecting two or more CNN layers from the plurality of CNN layers comprises selecting a number of the plurality of CNN layers based on an available number of memory of the computing device.
[0096] Example 5. The method of any of examples 1-4, further comprising generating the plurality of CNN layers based on an animation requirement of an image-providing application executing in the computing device.
[0097] Example 6. The method of example 5, wherein at least one of the generated plurality of CNN layers comprises a non-integer scaling ratio.
[0098] As used in this application, the terms “component,” “module,” “system,” and the like are intended to refer to a computer-related entity, either hardware, firmware, a combination of hardware and software, software, or software in execution, configured to perform particular operations or functions. For example, a component can be, but is not limited to being, a process running on a processor, a processor, an object, an executable, a thread of execution, a program, and / or a computer. By way of illustration, both an application running on a computing device and the computing device can be referred to as a component. One or more components can reside within a process and / or thread of execution and a component can be localized, co-resident, and / or distributed across one or more processors or cores. Also, these components can execute from various non-transitory computer-readable media having various instructions stored thereon. Components can communicate via local and / or remote processes, function- or procedure-calls, electronic signals, data packets, memory reads / writes, and other known computer, processor, and / or process related communication methods.
[0099] The various embodiments illustrated and described are provided merely as examples of various features to an illustrative claim. However, features shown and described with respect to any given embodiment are not necessarily limited to the associated embodiment and can be used in conjunction with other illustrated and described embodiments or combined with features shown and described with respect to other embodiments. Further, claims are not intended to be limited to any one example embodiment. For example, one or more operations of a method can replace or be combined with one or more operations of a method.
[0100] The foregoing method descriptions and the process flow diagrams are provided merely as illustrative examples and are not intended to require or imply that the operations of the various embodiments must be performed in the order presented. As will be appreciated by one of ordinary skill in the art, the order of operations in the foregoing embodiments can be performed in any order. Words such as "thereafter," "then," "next," etc. are not intended to limit the order of the operations; these words are simply used to guide the reader through the description of the methods. Furthermore, any reference to claim elements in the singular, for example, using the articles "one," "the," or "said," is not
[0101] The various illustrative logical blocks, modules, circuits, and algorithm operations described in connection with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and operations have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans can implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the claims.
[0102] The hardware used to implement the various illustrative logics, logical blocks, modules, and circuits described in connection with the embodiments disclosed herein can be implemented or performed with a general purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field
[0103] In one or more embodiments, the described functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, these functions may be stored as one or more instructions or code on a non-transitory computer-readable medium or a non-transitory processor-readable medium. The operation of the methods or algorithms disclosed herein may be embodied in a processor-executable software module that may reside on a non-transitory computer-readable or processor-readable storage medium. A non-transitory computer-readable or processor-readable storage medium may be any storage medium accessible by a computer or processor. By way of example and not limitation, such non-transitory computer-readable or processor-readable media may include RAM, ROM, EEPROM, flash memory, CD-ROM or other optical disc storage devices, magnetic disk storage devices or other magnetic storage devices, or any other medium that may be used to store desired program code in the form of instructions or data structures and is accessible by a computer. As used herein, disks and optical discs include compact optical discs (CDs), laser discs, optical discs, digital versatile optical discs (DVDs), floppy disks, and Blu-ray discs, wherein disks typically reproduce data magnetically, while optical discs reproduce data optically using lasers. Combinations of the above are also included within the scope of non-transitory computer-readable and processor-readable media. Additionally, the operation of a method or algorithm may reside as a single line of code and / or instruction, or any combination or set of code and / or instructions, on a non-transitory processor-readable medium and / or computer-readable medium that may be incorporated into a computer program product.
[0104] The above description of the disclosed embodiments is provided to enable any person skilled in the art to implement or use the claims. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be applied to other embodiments without departing from the scope of the claims. Therefore, this disclosure is not intended to be limited to the embodiments shown herein, but should be granted the broadest scope consistent with the following claims and the principles and novel features disclosed herein.
Claims
1. A method for performing a zoom operation on a computing device, the method comprising: Regions of interest (ROIs) of the image presented on the display of the computing device are identified based on zoom user input received by the computing device. Determine the zoom ratio for the final zoomed-out representation of the ROI; Select two or more CNN layers from a plurality of convolutional neural network (CNN) layers, each CNN layer including a scaling ratio smaller than the scaling ratio presented after the final zoom for the ROI; Intermediate images are generated for each of the selected CNN layers, including a sequence of intermediate-sized image frames of the ROI, each intermediate-sized image frame being smaller than the final zoomed size of the ROI and larger than the size of the ROI. as well as The sequence of intermediate-sized image frames of the ROI is used to render an animation of the change in the display size of the ROI.
2. The method of claim 1, wherein at least one of the selected CNN layers includes a non-integer scaling ratio.
3. The method of claim 1, wherein selecting two or more CNN layers from the plurality of CNN layers comprises selecting the number of the plurality of CNN layers based on application requirements of an application executed in the computing device providing the image.
4. The method of claim 1, wherein selecting two or more CNN layers from the plurality of CNN layers comprises selecting the number of the plurality of CNN layers based on the available number of memories of the computing device.
5. The method of claim 1, further comprising generating the plurality of CNN layers based on animation requirements of an image-providing application executed in the computing device.
6. The method of claim 5, wherein at least one of the generated plurality of CNN layers includes a non-integer scaling ratio.
7. A computing device, the computing device comprising: Processing system, the processing system being configured to: Regions of interest (ROIs) of the image presented on the display of the computing device are identified based on zoom user input received by the computing device. Determine the zoom ratio for the final zoomed-out representation of the ROI; Select two or more CNN layers from a plurality of convolutional neural network (CNN) layers, each CNN layer including a scaling ratio smaller than the scaling ratio presented after the final zoom for the ROI; Intermediate images are generated for each of the selected CNN layers, including a sequence of intermediate-sized image frames of the ROI, each intermediate-sized image frame being smaller than the final zoomed size of the ROI and larger than the size of the ROI. as well as The sequence of intermediate-sized image frames of the ROI is used to render an animation of the change in the display size of the ROI.
8. The computing device of claim 7, wherein at least one of the selected CNN layers includes a non-integer scaling ratio.
9. The computing device of claim 7, wherein the processing system is further configured to select the number of the plurality of CNN layers based on application requirements of an application executed in the computing device providing the image.
10. The computing device of claim 7, wherein the processing system is further configured to select the number of the plurality of CNN layers based on the number of available memories of the computing device.
11. The computing device of claim 7, wherein the processing system is further configured to generate the plurality of CNN layers based on the animation requirements of an image-providing application executed in the computing device.
12. The computing device of claim 11, wherein the processing system is further configured such that at least one of the generated plurality of CNN layers includes a non-integer scaling ratio.
13. A computing device, the computing device comprising: Components for identifying regions of interest (ROI) of an image presented on the display of the computing device based on zoom user input received by the computing device; A component used to determine the zoom ratio for the final zoomed-out rendering of the ROI; A component for selecting two or more CNN layers from a plurality of convolutional neural network (CNN) layers, each CNN layer including a scaling ratio smaller than the scaling ratio presented after the final zoom for the ROI; A component for generating intermediate images based on each of the selected CNN layers, the intermediate images comprising a sequence of intermediate-sized image frames of the ROI, each intermediate-sized image frame being smaller than the final zoomed size of the ROI and larger than the size of the ROI; and A component for rendering an animation of the change in the display size of the ROI using the sequence of intermediate-sized image frames of the ROI.
14. The computing device of claim 13, wherein at least one of the selected CNN layers includes a non-integer scaling ratio.
15. The computing device of claim 13, wherein the means for selecting two or more CNN layers from the plurality of CNN layers includes means for selecting the number of the plurality of CNN layers based on application requirements of an application executed in the computing device providing the image.
16. The computing device of claim 13, wherein the means for selecting two or more CNN layers from the plurality of CNN layers includes means for selecting the number of the plurality of CNN layers based on the number of available memories of the computing device.
17. The computing device of claim 13, further comprising components for generating the plurality of CNN layers based on animation requirements of an image-providing application performed in the computing device.
18. The computing device of claim 17, wherein at least one of the generated plurality of CNN layers includes a non-integer scaling ratio.
19. A non-transitory processor-readable medium having processor-executable instructions stored thereon, the processor-executable instructions being configured to cause a processing system in a computing device to perform operations, the operations including: Regions of interest (ROIs) of the image presented on the display of the computing device are identified based on zoom user input received by the computing device. Determine the zoom ratio for the final zoomed-out representation of the ROI; Select two or more CNN layers from a plurality of convolutional neural network (CNN) layers, each CNN layer including a scaling ratio smaller than the scaling ratio presented after the final zoom for the ROI; Intermediate images are generated for each of the selected CNN layers, including a sequence of intermediate-sized image frames of the ROI, each intermediate-sized image frame being smaller than the final zoomed size of the ROI and larger than the size of the ROI. as well as The sequence of intermediate-sized image frames of the ROI is used to render an animation of the change in the display size of the ROI.
20. The non-transitory processor-readable medium of claim 19, wherein the stored processor-executable instructions are further configured to cause the processing system in the computing device to perform operations such that at least one of the selected CNN layers includes a non-integer scaling ratio.
21. The non-transitory processor-readable medium of claim 19, wherein the stored processor-executable instructions are further configured to cause the processing system in the computing device to perform operations such that selecting two or more CNN layers from the plurality of CNN layers includes selecting the number of the plurality of CNN layers based on application requirements of an application executed in the computing device providing the image.
22. The non-transitory processor-readable medium of claim 19, wherein the stored processor-executable instructions are further configured to cause the processing system in the computing device to perform operations such that selecting two or more CNN layers from the plurality of CNN layers includes selecting the number of the plurality of CNN layers based on the available number of memories in the computing device.
23. The non-transitory processor-readable medium of claim 19, wherein the stored processor-executable instructions are further configured to cause the processing system in the computing device to perform operations, the operations further comprising generating the plurality of CNN layers based on animation requirements of an image provided for an application executed in the computing device.
24. The non-transitory processor-readable medium of claim 23, wherein the stored processor-executable instructions are further configured to cause the processing system in the computing device to perform operations such that at least one of the generated plurality of CNN layers includes a non-integer scaling ratio.