System and method for machine-learned image conversion

The described system employs a trained neural network to upscale images in real-time, addressing the trade-off between speed and quality by using a separable block transform, resulting in efficient and high-quality image conversion.

JP2025094020AActive Publication Date: 2025-06-24NINTENDO CO LTD

Patent Information

Application Number
JP2025040061
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-03-25
Filing Date
2025-03-13
Publication Date
2025-06-24
Estimated Expiration
2041-03-24

Smart Images

  • Figure 2025094020000001_ABST
    Figure 2025094020000001_ABST
Patent Text Reader

Abstract

To provide a system and method for converting a newly improved image.SOLUTION: A method is configured to: (a) acquire a first image of a first resolution; (b) select pixels of a first block from the first image; (c) generate a plurality of first input channels on the basis of the pixels of the first block; (d) insert values from each of the plurality of first input channels into a first activation matrix; (e) apply the first activation matrix to a trained neural network to generate a second activation matrix; (f) generate a second image of a second resolution on the basis of a second activation matrix; and (g) output the second image to a display. The steps (b) to (g) are performed for each of a plurality of pixel blocks selected from the first image.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims priority to U.S. Application Nos. 16 / 829,950 and 16 / 830,032, the entireties of both of which are incorporated by reference herein.

Background Art

[0002] Technical Overview The technology described herein relates to machine learning and converting a dataset or signal using machine learning into another dataset or signal. More specifically, the technology described herein relates to applying a block transform to such a dataset or signal. An application example of the technology includes converting an image of a certain resolution into another (e.g., higher) resolution, which may be used in real - time application examples, for example, from images generated by a video game engine.

[0003] Introduction Machine learning can give a computer the ability to “learn” a particular task without explicitly programming the computer for that task. One type of machine learning system is called a convolutional neural network (CNN), which is a class of deep - learning neural networks. Using such networks (and other forms of machine learning), for example, it is possible to help automatically recognize whether a cat is in a photo. Learning is done by “training” a model using thousands or millions of photos and recognizing when a cat is in the photo. This can be a powerful tool, but the processing that results from using (and training) a trained model can still be computationally expensive when deployed in a real - time environment.

[0004] Image upscaling is a technique that enables an image generated at a first resolution (e.g., 540p resolution or 960×540 which is 0.5 megapixels) to be converted to a higher resolution (e.g., 1080p resolution, 1920×1080 which is 2.1 megapixels). Using this process, an image at the first resolution can be shown on a display with a higher resolution. Thus, for example, a 540p image can be displayed on a 1080p TV, and (depending on the nature of the upscaling process), the fidelity of the illustration can be enhanced to display the 540p image compared to directly displaying the 540p image on a 540 TV with traditional (e.g., linear) upscaling. Different techniques for image upscaling can present a trade-off between speed (e.g., how much time the process takes to convert a given image) and the quality of the upscaled image. For example, if the upscaling process is performed in real time (such as during a video game), the quality of the resulting upscaled image may be degraded.

[0005] Accordingly, in these technical fields, it is recognized that there is a continuous need for new and improved techniques, systems, and processes. SUMMARY OF THE INVENTION MEANS FOR SOLVING THE PROBLEM

[0006] Overview In one exemplary embodiment, a computer system is provided for converting an image from a first resolution to a second resolution by using a trained neural network 。The source image is divided into blocks, and context data is added to each pixel block. The context blocks are divided into channels, and each channel from the same context block is inserted into the same activation matrix. Next, the activation matrix is run or applied against a trained neural network to generate a modified (e.g., output) activation matrix. Next, an output channel is generated using the modified activation matrix to construct an image at a second resolution. These techniques can be performed during runtime and in real-time along with the generation of the source image.

[0007] In one exemplary embodiment, a computer system is provided for training a neural network for converting signal data (e.g., an image). For example, an image at a first resolution is converted to a second resolution. Target signal data (e.g., a target image) is stored in a database or other non-transitory medium. For images, they can be at the resolution that is the target resolution. The computer system includes a processing system having at least one hardware processor. During training for image conversion, the computer system is configured to divide a first image into a first plurality of pixel blocks. Each one of the first plurality of pixel blocks is divided into a plurality of separate output channels to form target output data. Based on one of the plurality of separate output channels, a second image at a second resolution is generated. A plurality of context blocks are generated from the second image. The plurality of context blocks are then divided into a plurality of separate input channels and used to train the neural network by using the plurality of separate input channels until the neural network converges to the target output data.

[0008] In one exemplary embodiment, a method for converting signal data using a neural network is provided. The method includes placing a plurality of values based on data from a plurality of samples from a source signal in an initial activation matrix. Next, a separable block transform is applied over a plurality of layers of the neural network. The separable block transform is based on at least one learned matrix of coefficients and is applied to an input activation matrix to generate a corresponding output activation matrix. The initial activation matrix is used as the input activation matrix for a first layer of the plurality of layers, and the input activation matrix for each subsequent layer is the output activation matrix of the previous layer. The method outputs an output activation matrix of a last layer of the neural network and generates a transformed signal based on the output activation matrix of the last layer.

[0009] In one exemplary embodiment, the method operates such that at least two of the rows or columns of the initial activation matrix correspond to superimposable data from each of the plurality of samples.

[0010] In one exemplary embodiment, a distributed computer game system is provided. The system includes a display device configured to output an image at a target resolution (e.g., for a video game or another application). The system includes a cloud-based computer system including a plurality of processing nodes. The processing nodes of the cloud system are configured to execute a first video game thereon and generate an image of the first video game at a first resolution. The processing nodes of the cloud system are configured to transmit image data based on the generated image. The system also includes a client computing device configured to receive the image data. The client computing device includes at least one hardware processor and is configured to execute a neural network based on the received image data to generate a target image. Execution of the neural network on the client device applies a separable block transform to a plurality of activation matrices each corresponding to different blocks of pixel data within the image represented by the image data. The target image is generated at the target resolution and output to the display device at the target resolution for display on the display device during gameplay of the first video game.

[0011] This summary is provided to introduce various concepts that will be further described below in the detailed description. This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. Rather, this summary is intended to provide an overview of the subject matter described in this document. Accordingly, the above features are merely examples, and it will be appreciated that other features, aspects, and advantages of the subject matter described herein will become apparent from the following detailed description, figures, and claims.

[0012] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawings will be provided by the Office upon request and payment of the necessary fee.

[0013] These and other features and advantages will be more fully and completely understood by reference to the following detailed description of exemplary and non-limiting illustrative embodiments in conjunction with the drawings.

Brief Description of the Drawings

[0014]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8A

Figure 8B

Figure 8C

Figure 8D

Figure 8E

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Best Mode for Carrying Out the Invention

[0015] Detailed Description In the following description, for purposes of explanation and not limitation, specific details of specific nodes, functional elements, techniques, protocols, etc. are set forth to provide an understanding of the described technology. It will be apparent to those skilled in the art that other embodiments may be implemented apart from the specific details and examples described below. In some cases, detailed descriptions of well-known methods, systems, devices, techniques, etc. are omitted so as not to obscure the description with unnecessary details.

[0016] In this detailed description, sections are used only to orient the reader with respect to the general subject matter of each section. As will be seen below, the description of many features spans multiple sections, and headings should not be construed as affecting the meaning of the description contained in any section.

[0017] In many places in this document, including but not limited to the descriptions of FIGS. 1 and 10, software modules, software components, software engines, and / or actions performed by such elements are described. This is done to ease the description, and whenever it is described in this document that a software module or the like performs some action, it should be understood that the action is actually performed by underlying hardware elements (such as a processor, hardware circuits, and / or memory devices, etc.) in accordance with instructions including the software module or the like. Further details regarding this are provided below, especially in the description of FIG. 13.

[0018] Summary Certain exemplary techniques herein relate to converting an input signal (e.g., a digital signal) to an output signal by use of a neural network. Examples of different types of signals can be images, audio, or other data that, according to certain exemplary embodiments discussed herein, can be sampled or otherwise segmented and converted to a converted signal.

[0019] FIG. 1 shows a block diagram of an exemplary computer system (e.g., a video game system) that a user can use to play a video game. The system is configured to implement the process shown in FIG. 2 such that an image generated by a game engine at a first resolution (e.g., 540p) is up-converted to a different resolution (e.g., 1080p). FIGS. 3 through 7 show different aspects of the process shown in FIG. 2. FIGS. 8A and 8B show non-limiting examples according to the techniques discussed in FIG. 2. FIGS. 8C through 8E are block diagrams showing different SBT architectures according to an exemplary embodiment. FIG. 9 shows a block diagram of a computer system having a neural network used to train the process shown in FIG. 2. FIG. 10 is an exemplary process that can be executed on the computer system of FIG. 9 to generate a trained neural network. FIGS. 11 through 12 are further detailed aspects of the process shown in FIG. 10. FIG. 13 is a block diagram of an exemplary computer system that can be used to implement or execute the processes shown in FIGS. 1 and / or 9, and / or FIGS. 2 and / or 10.

[0020] Description of Figure 1 FIG. 1 is a block diagram including an exemplary computer system according to an exemplary embodiment.

[0021] The game device 100 is an example of the computer system 1300 shown in FIG. 13. This specification The term "gaming" device is used in connection with an exemplary embodiment in the specification for ease of use, and any kind of computing device may be used. In fact, a "gaming" device as used herein may be a computing device (e.g., a mobile phone, a tablet, a home computer, etc.) that is being used (or will be used) at that time to play a video game. A non-limiting exemplary list of computing devices may include, for example, smart devices or mobile devices (e.g., smartphones), tablet computers, laptop computers, desktop computers, home console systems, video game console systems, home media systems, and other types of computer devices. As will be described in connection with FIG. 13, computers may vary in size, shape, functionality, etc. In one exemplary embodiment, the techniques discussed herein can be used with non-gaming applications. For example, they may be used with real-time video surveillance, web browsing, voice recognition, or other applications where it may be useful to convert one dataset to another. Additional examples and applications of the techniques herein are discussed below.

[0022] The gaming device 100 may include a CPU 102, a GPU 106, and a DRAM (Dynamic Random Access Memory) 104. The CPU 102 and the GPU 106 are examples of the processor 1302 in FIG. 13. The DRAM 104 is an example of the memory device 1304 in FIG. 13. Different types of CPUs, GPUs, DSPs, dedicated hardware accelerators (e.g., ASICs), FPGAs, and memory technologies (both volatile and non-volatile) may be used on the gaming device 100.

[0023] Examples of different types of CPUs include Intel CPU architectures (e.g., x86) and ARM (Advanced RISC Machine) architectures. Examples of different GPUs include discrete GPUs such as the NVIDIA V100 (which may include hardware support for matrix multiplication or tensor cores / accelerators) and integrated GPUs that may be found on a system-on-chip (SoC). An SoC can combine two or more of a CPU 102, GPU 106, and local memory such as registers, shared memory, or cache memory (also called static RAM or SRAM) on a single chip. DRAM 104 (also called dynamic RAM) is typically fabricated as a separate semiconductor and connected to the SoC through wiring. For example, the NVIDIA Tegra X1 SoC includes multiple CPUs, GPUs, north bridge controllers, south bridge controllers, and memory controllers all on a single SoC. In one example, the processing capabilities provided by the CPU, memory components, GPU, and / or other hardware components that make up a given gaming device may be different on other gaming devices. A gaming device may be portable, may be a console gaming machine, or may operate as a personal computer (e.g., a desktop or laptop computer system used to play video games).

[0024] A GPU can include many processing cores that operate in parallel. Each processing core, which is part of the GPU, can operate with corresponding hardware registers that store data used by the various processing cores therein. For example, NVIDIA's GPU architecture includes a number of 32-bit, 16-bit, and / or 8-bit registers that output data to the processing cores of the GPU. In one GPU architecture, the highest bandwidth memory may be available in the registers, followed by shared memory, then cache memory, and then DRAM. As discussed in more detail below, data regarding the dataset to be converted (e.g., an image to be upconverted) can be efficiently loaded into these registers to improve the efficiency when converting the dataset to another form (e.g., another resolution). In fact, by utilizing the hardware registers on the GPU for this operation, an exemplary upconversion process can be performed in real-time (e.g., less than 1 second, less than 1 / 30 second, or less than 1 / 60 second) and / or during the execution time of an application or game (e.g., without significant latency) without having to change how the initial image is generated at a lower resolution. upconversion process can be performed in real-time (e.g., less than 1 second, less than 1 / 30 second, or less than 1 / 60 second) and / or during the execution time of an application or game (e.g., without significant latency).

[0025] In one exemplary embodiment, the techniques herein may advantageously utilize NVIDIA's Tensor Cores (or other similar hardware). A Tensor Core is a hardware unit that can multiply two 16×16 FP16 matrices (or matrices of other sizes depending on the nature of the hardware) and then add a third FP16 matrix to the result using a fused multiply-accumulate operation to obtain an FP16 result. In one exemplary embodiment, a Tensor Core (or other processing hardware) can be used to multiply two 16×16 INT8 matrices (or matrices of other sizes depending on the nature of the hardware) and then add a third INT32 matrix to the result using a fused multiply-accumulate operation to obtain an INT32 result. The INT32 result can then be converted to INT8 by dividing by an appropriate normalization amount (which can be calculated during a training process, for example, as described in connection with FIG. 9). Such a conversion can be achieved, for example, using a low-cost integer right shift. Such hardware acceleration for the processing discussed herein (for example, in the context of separable block transforms) can be advantageous.

[0026] Returning to FIG. 1, the game device 100 may also be coupled to an input device 114 and a display device 116. Examples of input devices 114 include video game controllers, keyboards, mice, touch panels, sensors, and other components that can provide input used by a computer system (such as a game device) to execute application programs and / or video games provided on the computer system.

[0027] Examples of the display device 116 include a television, a monitor, an integrated display device (such as part of a mobile phone or a tablet), etc. In one example, the game device 100 can be configured to couple to different types of display devices. For example, the game device 100 may be coupled to an integrated display (such as part of a structure housing the game device 100) where an image can be output. The game device 100 can also be configured to output an image to a larger television or other display. In one exemplary embodiment, different display devices can natively display different resolutions. For example, the integrated display of the game device may have 500,000 pixels (such as a 540p display), and a separate display may have 2.1 million pixels (such as a 1080p display). Using the techniques herein, the game device 100 can be configured to output different images for games depending on which display device is the output destination of the game device. Thus, for example, a 540p image can be output to the integrated display when the integrated display is used, and a 1080p image can be output to the 1080p display when it is used.

[0028] In one exemplary embodiment, the computer system can dynamically switch the type of image being output based on conditions associated with the computer system. Such switching may occur while the user is playing a game (possibly with a short pause while switching between two modes). For example, if the computer system is running on battery (such as not plugged into a socket), the computer system can be configured not to use the exemplary image conversion process using the techniques discussed herein. However, if the computer system is plugged into an AC power source, the techniques discussed herein for upconverting the image to a higher resolution can be used or turned on for video games or other applications. The reason is that the techniques discussed herein use a higher percentage (such as up to 80, 90, or 95% or more) of the available processing power of the GPU being used This is because it can increase the power consumption of the GPU. Thus, if a computer system operates only on the battery of a mobile device while using, for example, the process shown in FIG. 2, the battery can be depleted more quickly. In this way, with such a technique, a user may be able to play games on the mobile device when, for example, commuting home from work. In this mode, the user uses the device's local display (e.g., 540p) for video games. However, when the user arrives home, the user can plug the mobile device into an outlet so that the mobile device no longer relies on its own battery power. Similarly, the user can connect the mobile device to a larger display, such as a 1080p display (like a TV). Such a connection can be wired (e.g., a DisplayPort or HDMI (registered trademark) cable) or wireless (e.g., Bluetooth (registered trademark) or WiFi (registered trademark)). When detecting one (or both) of these scenarios (e.g., an output display that can display a higher resolution and / or a non-battery power source for the computing system), the system dynamically initiates the image conversion process discussed with respect to FIG. 2, enabling the user to play games on their 1080p TV and view the games at a higher resolution. In one exemplary embodiment, the user may also manually initiate the process of image upconversion.

[0029] The techniques herein can advantageously provide performance with fewer constraints due to limited memory bandwidth than prior approaches. In other words, the architecture for converting the images (or, more generally, datasets) discussed herein may not be limited by a memory bandwidth bottleneck. This can be particularly true for real-time inference, which may be limited to one batch (e.g., instead of a typical training scenario that benefits from larger batches such as 256 in general). In other words, the techniques herein can enable nearly 100% utilization of the matrix multiplication hardware accelerator during the execution time of an application (such as a video game), thus increasing (e.g., maximizing) the overall performance per dollar spent on the hardware used for the conversion.

[0030] Returning to FIG. 1, the game device 100 stores and executes a video game application program 108. The video game application program includes a game engine 110 and a neural network 112. The game device 100 may also store image data (e.g., textures) and other types of assets (such as sounds, text, pre-rendered video, etc.) used by the video game application program 108 and / or the game engine 110 to generate or produce content for a video game (or other application), such as images for the game. Such assets may be included with the video game application program on a CD, DVD, or other physical medium, or may be downloaded via a network (such as the Internet) as part of a download package for the video game application program 108, for example.

[0031] The game engine 110 includes a program structure for generating an image to be output to the display 116. For example, the game engine 110 may include a program structure for managing and updating the positions of objects in a virtual space based on inputs provided from the input device 114. The data provided is used, for example, to render an image of the virtual space using a virtual camera. This image can be a source image generated at a first resolution (e.g., 540p). The source image is applied to the neural network 112, and the neural network 112 converts the source image into an up-converted image at a higher resolution than the original source image (e.g., 1080p) (e.g., the up-converted image is generated based on the application of the source image to the neural network 112). The up-converted image is then output to the display device 116 and displayed there. Further explanation of how the neural network is generated is provided in connection with FIG. 9.

[0032] In an exemplary embodiment, the time taken to up-convert the source image (e.g., generated by the game engine 110) is less than 1 / 60 of a second. Thus, if the game engine is generating 60 images per second that are intended to be displayed on the display 116, there may be little or no noticeable visual delay when outputting the up-converted image to the display instead of the source image. In this way, such techniques may make it possible to generate and display an up-converted image in real time from the original source image. For example, if a video game application is developed to generate images at a first resolution (e.g., 540p), the techniques herein may enable a visual improvement of the video game application such that the images can be output from the video game application at a higher resolution (e.g., 1080p) than originally intended.

[0033] For the purpose of illustration, the video game application program 108 is used, but it is recognized that other applications that provide video output can be substituted. Also, a neural network 112 is shown as part of the video game application program 108, but this may be provided separately. For example, it may be part of an operating system service that modifies or upconverts an image when the video game application program outputs an image.

[0034] In one exemplary embodiment, the "game device" may be a device hosted within a cloud-based environment (e.g., Amazon's AWS or Microsoft's Azure system). In such a scenario, the game (or other application program) may be hosted on a virtual machine within the cloud computer system, and the input device and display device may be local to the user. The user may also have a "thin" client application or computer that communicates with the cloud-based service (e.g., communicates data from the device and receives and displays the received image from the cloud on the television). In this type of implementation, the user input is passed from the user's computer / input device to the cloud-based computer system running the video game application 108. The image is generated by the game engine, converted (e.g., upconverted) by the neural network, and then sent to the user's display (or the computer that outputs the image to the display).

[0035] In one exemplary embodiment, a cloud-based system may utilize the upscaling capabilities on a "thin" client by rendering, compressing, and streaming compressed low-resolution (e.g., 540p) video / images to the client at lower server costs (and bandwidth), and enabling upscaling (e.g., neural network processing 112) on the client hardware. In one example, this may also include having the neural network address or compensate for compression artifacts. Thus, the features herein may advantageously reduce the use of bandwidth in a cloud-based gaming environment.

[0036] In one exemplary embodiment, a cloud-based system may operate dynamically with respect to the output display being used by the user. Thus, for example, a video game may natively output a 540p image. A first user may receive a 1080p image (e.g., upscaled from 540p) using the cloud system, and a second user may receive an image of a different resolution (e.g., a 720p image, a 4 k image, or a 1440p image) using the cloud system. Each instance of the video game application (and / or neural network) may be hosted within its own virtual machine or virtual container, thus enabling multiple different users to "play" the same video game by flexibly providing different options (e.g., output of images of different resolutions).

[0037] Cloud-based implementations may be useful in contexts where the user has access to a GPU capable of performing the techniques discussed herein.

[0038] In one exemplary embodiment, the GPU may instead be (or may include) an ASIC or FPGA that operates in a manner similar to a GPU.

[0039] In one exemplary embodiment, the game device 100 may be two or more computer systems.

[0040] It should also be recognized that the type of "application" or program or data source providing the source image is not limited to video games. In fact, other types of applications, including real-time image recognition from wildlife observation cameras, audio, language / text translation, images input from home security cameras, movies, and other TV programs, etc., may utilize the techniques herein.

[0041] For more general application examples such as image classification, for example, a traditional CNN implementation on GPU processing hardware may involve: 1) loading layer weights in high-speed memory (e.g., GPU registers or shared memory); 2) loading layer inputs from DRAM into registers; 3) multiplying the inputs by the weights using matrix multiplication implemented on the GPU; 4) applying a non-linear function; 5) storing the layer output in DRAM; and 6) repeating this process for each layer. The drawback of this approach is the movement to and from DRAM. For example, layer data (e.g., activations) may not fully fit into the relatively limited amount of high-speed memory (such as registers) typically used in the processing of neural network layers. Thus, in some cases, it may be necessary to transfer that data between different memory locations. Because layer data (e.g., activations, which can be a matrix of 960×540×16 values, i.e., corresponding to the resolution of a 540p source image combined with 16 channels in one example) may not fit into the GPU registers (or other "high-speed" memory). Therefore, main memory (DRAM104) may be used to store such information.

[0042] In one exemplary embodiment, fusion of different layers (e.g., "layer fusion") may be used such that calculations from one layer and the next can be realized through a single processing code (e.g., a CUDA kernel). A potential drawback of this approach is that, since the CNN is translation invariant, as more layers are fused, more inputs may be required to calculate a single output value. Thus, this type of implementation can bring valuable benefits by increasing the receptive field (the "seeing" ability of the end value of a neural network / depending on a wide range of inputs), but may also be accompanied by performance drawbacks.

[0043] In one exemplary embodiment, the approach to how data can be prepared and processed may be based on the nature of the underlying hardware that will perform the operations (e.g., matrix operations). In one exemplary embodiment, an image may be divided into blocks, the size of which may be based on the underlying hardware. One exemplary embodiment may be implemented on NVidia GPU hardware (e.g., Volta and Turing architectures), where the CUDA API exhibits hardware acceleration for 16×16 matrix multiplication. Due to this (and as discussed below), a block size of 4×4 may be used within the image to be transformed (these 16 pixels are mapped to the rows of a 16×16 matrix). Such an implementation allows the input to be divided into 16 inputs having 16 channels (in one example, fewer than 16 channels may be used, as discussed below), thus conforming to an atomic 16×16 matrix. This can then be stored in the registers of the GPU (or other "fast memory" that will handle matrix calculations). Of course, the GPU 106 for an exemplary block-based neural network architecture for a particular size (or even the CPU 102 if so designed) may be designed with a different size of atomic matrix, depending on the nature of the dimensions of the fastest available atomic multiplication hardware.

[0044] When the matrices remain in the register, the layers of a given pixel (or other types of data from the signal) can be "fused" together. This is because they remain in the register during processing. This will be discussed in more detail below in relation to Figure 2. In one exemplary embodiment, the activation matrix can remain in the internal memory of the hardware (e.g., GPU, CPU, DSP, FPGA, ASIC, etc.) that is performing matrix operations on the activation matrix. In other words, the data of a given activation matrix remains in the same semiconductor hardware (which can be the same silicon for silicon-based memory, or other materials such as gallium, germanium, etc. for other memory types) while the various layers of the neural network are applied to that activation matrix - for example, continuously transforming the activation matrix across multiple layers of the neural network.

[0045] Based on such blocks, a general transformation of a layer using a block matrix (where each of block W is a generic p×p matrix) can exist as follows.

[0046]

Number

[0047] In such a block matrix design, it is recognized that by insulating each block, the propagation of the receptive field that would otherwise occur (e.g., in the case of a normal CNN) can be prevented. Thus, the techniques herein can make it possible to fuse many layers (e.g., as many as desired) while still maintaining the locality of the data in question. Since the width of the data remains somewhat constant between the input and output of each layer, such fused layers may sometimes be referred to as "block towers".

[0048] From an inference perspective, this type of approach can be advantageous. This is because it can be implemented as a series of atomic-sized matrix multiplications as follows.

[0049]

Number

[0050] One potential problem is that by maintaining data in such a localized manner, the system may be prevented from benefiting from a wider receptive field (which can be beneficial in certain classification applications). Such problems can be addressed, at least in part, by introducing, for example, "block convolution" and "block pooling" layers, as follows.

[0051]

Number

[0052] where W i is a p×p matrix (in a typical example, p = 16). With such a formulation, this can be similar to traditional CNN matrix formulations, except that individual CNN filter weights (e.g., single real floating-point numbers) are replaced by block matrices. Or, put another way, the block techniques discussed herein may be considered a generalization of CNN, since when using 1×1 dimensional block matrices, the technique can be reduced to more traditional CNN formulations.

[0053] In one exemplary embodiment, an input signal (which can be, for example, an image) can be processed by a separable block transform (or block convolutional SBT) that is separable in a "translation-invariant manner". Thus, in the context of an image, if a signal S (e.g., a first image) is translated horizontally by 4 pixels or vertically by 4 pixels to a signal S' (e.g., a second image), the resulting 4×4 blocks of signals S and S' (corresponding to the activation matrix used as the input to the SBT network) will generally match (except at the boundaries of each image). If the blocks of S and S' are identical (again, except at the boundaries of the signals), the output blocks generated by applying SBT to S and S' will be the same. In other words, the transformed signals will also be the same, although there is a translational difference between SBT(S) and SBT(S'). Another way to view this is to calculate SBT (and / or block convolutional SBT as well) on the first block, then calculate it again using the same weights (the same learned L and R matrices) on neighboring blocks, and then calculate it again on neighboring blocks, and so on. Thus, the signal is processed in a "convolutional manner" by applying the same calculation while moving (e.g., translating) the input position along the input signal.

[0054]

Number

[0055] This is sometimes called the tensor product, and its (block) matrix is called the Kronecker product of L and R.

[0056]

Number

[0057] The left matrix L of dimension p×p (e.g., the pointwise transformation in MobileNet) processes all channels of a given data point in the same way for each data point. This is a general form, meaning that all of its coefficients are fully independently learnable.

[0058] The right matrix R of dimension p×p (e.g., the depthwise convolution transformation in MobileNet) processes all the pixels of a given channel in the same way for each channel. This is a general form, meaning that all of its coefficients can be learned completely independently.

[0059] The above formulation is symmetric and balanced and can be generally applied in several different cases. The form may be further modified to handle rectangular matrices (e.g., of size p×q) on both the L - side and the R - side. In other words, the input dimension of a layer can match the output dimension of the previous layer. However, it is recognized that by making the values of p and q multiples of the size of atomic - acceleration - enabled hardware matrix multiplication, the use of hardware resources can be increased / optimized and in one example can be optimal with respect to speed.

[0060] Advantageously, using the block shape and invariance between data points, they can be processed together in a single matrix multiplication. This results in the efficient use of resources in one exemplary embodiment.

[0061] It is recognized that a 3×3 convolutional kernel can be realized by adding nine point - wise 1×1 kernels. Therefore, the separate transformations discussed above can also be combined as follows.

[0062]

Number

[0063]

Number

[0064] The potential further benefit of the summation approach can be the performance of inference realization. When the formats of the input matrix and the output matrix of the LXR product are the same (e.g., 16×16 FP16 values), the code for realizing inference can be strictly limited to matrix multiplications that are executed one after another (e.g., by fused multiply-accumulate). With this type of approach, advantageously, operations can be performed without the need to align data, re-organize such data into other forms, or convert the data into other formats. This type of approach can also advantageously avoid adding or using unnecessary instructions. This is because the data is already in the correct format for each part of the sum. In one example, the number of LXR sums can be set as a dynamic parameter. This is because the input and output formats of the sum do not change (e.g., it can be assumed to be a 16×16 matrix as discussed in the examples herein). Thus, this can be a way to freely increase the weights while ensuring that the time required to load the weights remains hidden behind the time taken to perform matrix multiplications (which depends on, e.g., each specific hardware memory bandwidth and matrix multiplication speed), and thus it can be possible to learn / memorize more.

[0065] For training, a greater degree of flexibility can be applied to train large networks. Next, the network can be compressed by pruning the least necessary elements of each sum while retaining only the most relevant aspects obtained from the "lottery ticket" / "draw" of matrix initialization. This dynamic process can help, on a content-by-training basis, to determine how many multiplications are allocated to each layer under a given budget of processing time. Such a decision can be based on knowing a simple model of the inference time that is linear in the number of matrix multiplications. Next, such aspects can be combined to determine the number of layers (the number of layers can be as few as a dozen or so and usually is not a particularly large potential search space).

[0066] In certain exemplary embodiments, a greater number of channels may be used. In such channels, some of the separable block towers discussed herein may be calculated in parallel from the same input values (e.g., activation matrices), but using different learned weights (L and R matrices). Such an approach may be similar in some respects to grouped channels in convolutional neural networks.

[0067] In certain exemplary embodiments, to avoid isolating and maintaining the channels of each tower from others all the way to the end of the network, the outputs of all block towers can be stored together (e.g., in a memory such as DRAM or cache) and used together as inputs to another group of separable block towers. Such an implementation can rely less on DRAM bandwidth (e.g., as data is accessed more quickly through the cache memory) compared to equivalent convolutional neural network architectures. Put another way, a p*p SBT can use more activations as input than p*p by multiplying each of several p*p input activations by different p*p weight matrices and adding (e.g., term to term) all the results to a single p*p matrix that becomes the input activation matrix of the SBT. This aspect will be described in more detail below in connection with FIGS. 8C through 8E. In certain exemplary embodiments discussed herein, GPUs are discussed, but it will be appreciated that in certain exemplary embodiments, ASICs and FPGAs may also be designed and used instead of such GPUs.

[0068]

[0069] Description of Figure 2 ​FIG. 2 is a flowchart showing a machine-learned up-conversion process that can be executed on the computer system of FIG. 1 to convert a 540p image into a 1080p image. FIGS. 3 through 7 are discussed below to provide additional details regarding certain aspects of the up-conversion process shown in FIG. 2. Although images and pixel data are described in connection with the examples herein, it is recognized that other types of signals may be used in connection with the techniques herein. For example, each “pixel” within an image discussed herein may be considered data sampled from an entire signal (e.g., an image). Accordingly, techniques for converting a source signal (e.g., an image) into a converted or otherwise transformed signal (e.g., a higher resolution image) are discussed herein.

[0070] In step 200, a 540p source image 205 is rendered by the game engine 110. In certain exemplary embodiments, as discussed herein, the source image may be derived from other sources such as a real camera, movie, television program, broadcast television, etc. For example, using the techniques herein, a source 540p signal received for a television program (e.g., a live broadcast of a sports event) may be converted to a 1080p signal and output for display to a user. Further, although 540p is discussed in connection with the example of FIG. 2 (and elsewhere herein), the techniques may be applied to images of other sizes. The details of the neural network 112 (e.g., coefficients or L and R) used as part of the upconversion process are seen to change as the details of the source and / or the converted image change (e.g., when the resolution of such an image is adjusted). For example, a neural network for upconversion from 540p to 1080p is different from one that upconverts from 1080p to 1440p (e.g., 2560×1440). The examples shown in FIGS. 3-7 relate to converting a 540p image to a 1080p image, but it is also recognized that the techniques herein may be applied to other image sizes (e.g., 720p to 1080p; 480p to 1080p, 1080p to 1440p, 1080p to 4k / 3840×2160, 720p to 4k, etc.).

[0071] In certain exemplary embodiments, the initial image may be generated by rendering using motion vector information and / or depth information (e.g., z-buffer data) or otherwise. This information may be used to improve the resulting converted image quality. In certain exemplary embodiments, such information may be added to an activation matrix created based on each pixel block.

[0072] In certain exemplary embodiments, not an integer (or not in the same ratio horizontally and vertically) It is also possible to increase the (none) ratio according to the techniques discussed herein. For example, in the case of 720p to 1080p, the output block can be 6×6 pixels (with 3 channels, thus having 108 output values), which can still be easily adapted to the 16×16 = 256 output values of the SBT output discussed herein. According to one exemplary embodiment, additional ratios such as 7 / 3 can also be considered (for example, this can correspond to the conversion from 1920×1080 to 4480×2520). In such an exemplary embodiment, the source image can be divided into 3×3 blocks (with context data added) and trained to output 7×7 blocks (which will still fit into the 16×16 output blocks discussed herein). In one exemplary embodiment, it is possible to modify the application example of outputting an image at a resolution that is not very common nowadays. The techniques herein can handle upscaling, for example, using alternative ratios. For example, in one exemplary embodiment, an analog TV anamorphic distortion can be compensated for or addressed using a horizontal upscaling ratio of 8 / 7 (which can then be multiplied by some integer ratio).

[0073] In any case, in step 200, a 540p image 205 is generated (e.g., rendered) by a game engine or the like. Next, in step 210, the image is prepared. This aspect of the process will be described in more detail in connection with FIG. 3, but this aspect relates to splitting the image into separate input channels or input data 215. Advantageously, the input data 215 can be stored in the registers (e.g., 16-bit registers) of the GPU 106 at this point. Once the input data 215 is generated, it is stored in the registers of the GPU. The input data 215 (or the matrix of activations 225) can remain in the registers (or other internal memory) over the course of being applied to the neural network. By this kind of implementation, advantageously, during the processing performed by the neural network (e.g., where multiple matrices of activations across an image are processed by the GPU), the (relatively) slow DRAM 104 in the system 100 is bypassed. This is facilitated by forming the data to fit within the registers, so that the massively parallel processing provided by the GPU 106 can be used more effectively.

[0074] In an exemplary embodiment, other types of hardware other than the GPU may be used to handle the conversion of the input data 215 to the 1080p output data 245. Generally, it is preferable to hold such data in on-chip memory (e.g., the registers on a GPU handling deep learning applications, or an SRAM FPGA). Thus, once the input data 215 is placed in the registers (or similar high-speed memory), it can remain there until the 1080p output data 245 is generated (or the final matrix of activations is generated) and used to construct the final converted image (which can occur in the DRAM).

[0075] Returning to FIG. 2, next, in step 220, the input data 215 is reorganized into a matrix to generate a 16×16 matrix 225 of activations. This step will be discussed in more detail in connection with FIG. 4.

[0076] In step 230, the initial activation matrix 225 is executed through the trained neural network 112 to generate a 16×16 activation matrix 235 transformed by the neural network 112. As discussed herein, this may involve applying a separable block transform to the activation matrix. This aspect of the process is discussed in more detail in FIG. 5.

[0077] Once the activation matrix is executed through the neural network in step 230, in step 240, it is reorganized into blocks to generate 1080p output data 245. This aspect of the process is discussed in more detail in FIG. 6.

[0078] In step 250, the 1080p output data 245 is then reorganized into a 1080p image 255 and output to the display 116 in step 260. This aspect of the process is explained in more detail in FIG. 7. As noted above, the processing shown between steps 220 and 250 (including both of these steps) can be performed entirely within the registers of the GPU (or other internal memory) without the need to transfer data to DRAM (or other relatively "slow" memory). Thus, for example, a given activation matrix 225 can remain stored within the same semiconductor hardware (e.g., the same register or location in memory) while being executed through the neural network. Such processing may be applied to each matrix generated for corresponding pixel blocks of an image (or other signal), which can then be executed simultaneously across multiple hardware processors of, for example, the GPU (or other hardware resources).

[0079] Description of Figure 3 FIG. 3 is a flowchart showing an enlarged view of the image preparation section of the machine-learned up-conversion process of FIG. 2.

[0080] The 540p image 205 output from the game engine 110 is cut or divided into 4×4 pixel blocks at step 300. Block 302 represents one of the pixel blocks from the image, and 304 is one pixel within that block. Each pixel can be represented by different color values of RGB (explained in more detail at step 330). Although color values (e.g., RGB values) are discussed in relation to an exemplary embodiment, it is recognized that other types of data may be divided into blocks and stored. For example, the technique may be used in relation to a grayscale image where each pixel stores the amount of light for that pixel. In an exemplary embodiment, color information may be processed / provided by using the YUV or YCoCg format. In an exemplary embodiment, the luminance (Y) channel may be used along with the techniques discussed herein to thus process (e.g., upscale) this using a neural network.

[0081] In an exemplary embodiment, block sizes other than 4×4 may be used. For example, in an exemplary embodiment, 8×2 pixel blocks may be used. In one example, it may be advantageous for the size of the pixel block to be determined based on the dimensions of the hardware used for matrix multiplication or a multiple thereof. Thus, if the hardware acceleration supports 16×16 matrix multiplication, 4×4 or 8×2 blocks may be selected first. Such sizes may advantageously enable processing pixels separately along one dimension of the matrix while processing channels along the other dimension.

[0082] The selection of the block size may also be based on the amount of high-speed memory (e.g., registers, etc.) available in the system. By holding the matrix blocks and corresponding data in high-speed memory during neural network processing, performance increase can advantageously be facilitated (e.g., real-time or runtime image conversion becomes possible). Thus, the 4×4 block size may be appropriate for certain types of hardware, but other block sizes may also be used in relation to the techniques discussed herein.

[0083] In any case, at 310, each block from the original 540p image 205 is selected. Thus, in one exemplary embodiment, there may be more than 30,000 pixel blocks that undergo the processing described in FIG. 3 for a single 540p image. For example, by using the hardware resources of a GPU or other processor, subsequent processing for all of the pixel blocks may be performed in parallel. (For example, depending on the number of individual processing units within the overall system) In some cases, multiple groups may be processed sequentially. For example, a first group of pixel blocks (e.g., 15,000) may be processed in parallel, and then another group (the remaining 15,000) may be processed. From the user's perspective, the processing for all of the blocks may still be performed in parallel.

[0084] At 320, context data is added to the 4×4 pixel block to create an 8×8 context block 322. The context data may be based on, derived from, or be a function of the pixel values of the pixels in the image surrounding a given pixel block. In one example, the pixel data used for the context block may remain unchanged from the pixels outside of the 4×4 pixel block. In one exemplary embodiment, other (absolute or relative) context block sizes may be used. For example, a 12×12 context block may be used for a 4×4 pixel block. In one exemplary embodiment, pixel data along the diagonal axis may not be considered, and pixel data may be selected along the horizontal axis and / or the vertical axis. Thus, if the pixel block is represented as X1 - X4 as shown in the following table, horizontal values (A1 - A4) and vertical values (B1 - B4) may be added to the content block, while diagonal values (C1 - C4) are not used within the context block.

[0085] [Table 1]

[0086] In one implementation example, while one pixel along the diagonal can be used, two (or more) along the horizontal or vertical direction can be used within the context block. In an exemplary embodiment, multi-resolution data may be included within the context block to increase the receptive field along the direction of a "slightly inclined line". Aliasing can extend away from the block. For example, one layer can contain 4×4 blocks calculated as the average of an 8×8 context block, then 4×4 blocks calculated as the average of a 16×16 context block, and so on. Such data can help increase the receptive field at a limited cost with respect to the number of inputs.

[0087] At 330, the context block 322b is divided into four separate input channels 333, 334, 335, and 336. The numbers represented by each of the input channels indicate the configuration of that particular channel. Thus, each 1 shown in 322b of FIG. 3 is used to form the input channel 333, each 2 is used to form the input channel 334, and so on. Each of the numbers represents one of the RGB values of the corresponding pixel. In this way, each context block 332 is repeated for each of the red (R), green (G), and blue (B) values (or otherwise executed), or the context block simply stores three values per pixel. Thus, there are 12 input channels per pixel block created as a result of the image preparation step 210. Additionally, in this exemplary embodiment, there are three input channels per pixel (one for each of the R, B, G values of the pixel). The 12 input channels created for each pixel block form the input data 215. This process is repeated for all of the pixel blocks of a given image (or otherwise executed) and is typically achieved in parallel. As discussed herein, in an exemplary embodiment, multiple pixel blocks ( and / or context blocks) may be processed in parallel.

[0088] In one exemplary embodiment, the signal data of the source signal can be cut or divided into at least two blocks. In one example, such blocks can then be processed independently by using the SBT discussed herein.

[0089] Description of Figure 4 FIG. 4 is a flowchart showing an enlarged view of the rearrangement unit into a matrix of the machine-learned up-conversion process of FIG. 2.

[0090] In this flowchart, at step 410, the input data 215 for each pixel block (e.g., 12 input channels) is rearranged into a single 16×16 matrix 225. For example, the value of input channel 333a (having a red value of the "1" pixel value in context block 322b, for example) is inserted (e.g., added) into row 412 of matrix 225. The value of input channel 333b (the blue value of the same "1" pixel from the context block) is inserted into row 414. Then, the value of input channel 333c (the green value of the same "1" pixel from the context block) is inserted into row 416. This process is repeated for all 12 rows or done in some other way, and thus, in the 16×16 matrix 225 of activations, values from the sampled pixels of the source image (e.g., the source signal) are placed. Thus, the resulting 16×16 matrix can include data for a single pixel within multiple rows. For example, the pattern of data for each pixel used to feed rows 412, 414, and 416 can be superimposed from one pixel to the next. It is recognized that data can be inserted into the matrix on a column basis rather than a row basis as shown in FIG. 4. Thus, in one exemplary embodiment, columns can be used instead of the rows referred to herein.

[0091] Examples of superimposable patterns can include, for example, two adjacent blocks located horizontally in a 4×4 pixel block (e.g., after horizontal translation by 4 pixels). As another example, any two rows (e.g., of 4×1 pixels) within a 4×4 block of pixels may be superimposable. Similarly, a row of 4×1 pixels can be superimposed on a column of 1×4 pixels (after a 90° rotation). The following patterns of blocks are superimposable. Specifically, the pattern of X in the following table (taking into account rotation and symmetry) is superimposable with the pattern of the sample represented by Y.

[0092]

Table 2

[0093] Other types of data (e.g., different types of signals) may also be superimposable such that the individual pieces that make up a sample piece of data are divided or separated into separate channels. In other words, depending on the nature of the source signal (e.g., an image or some other data), at least two of the rows (or columns) in the initial activation matrix may correspond to similarly organized or structured data from each sample taken from the underlying source. In the case of an image, the similarly organized or structured data may be individual pixels (e.g., when multiple channels are used per pixel) or groups of pixels that follow the same shape but are at different locations in the image. In one exemplary embodiment, at least two of the rows or columns of the activation matrix may be generated based on a common pattern of data from each sample in the underlying source signal.

[0094] In one exemplary embodiment, since there are 12 input channels, in step 420, the remaining 4 rows of the 16-row matrix are set to zero (or set to values that are ignored during matrix processing) to create an activation matrix 225, and the matrix then undergoes neural network processing.

[0095] In one exemplary embodiment, data can be placed in all 16 (or however many rows there are in the activation matrix that will be used). In one exemplary embodiment, additional information can be placed in four additional rows (or "extra" rows that do not have initial color information). For example, the game engine 110 may supply depth information regarding an object or other aspect of the image in question. This information may be incorporated into additional rows of the 16x16 matrix. In one exemplary embodiment, motion information regarding an object or other aspect of an image may be supplied from the game engine 110 and incorporated into the 16x16 matrix.

[0096] Description of Figure 5 FIG. 5 is a flowchart showing an enlarged view of the neural network execution part of the machine-learned up-conversion process of FIG. 2. Execution of the neural network on the activation matrix may include the application of a separable block transform that utilizes the LXR operations discussed herein.

[0097] The activation matrix 225 is executed through the neural network 112. An example of how such a neural network can be trained is discussed in connection with FIG. 9. The output of such training can be a matrix of coefficients (L and R) "trained" on an exemplary training dataset.

[0098] As part of the neural network processing at step 230, the activation matrix 225 generated from the input channels is executed at step 410 through a separable block transform. The equation representing this process is shown in FIG. 5, where L and R are 16x16 matrices (e.g., each having 256 coefficients in the 16x16 matrix) generated using the training system discussed in FIG. 9.

[0099] L is a 16×16 pixel unit matrix (or other sample unit dependent scenario) and is multiplied on the left side. This applies a linear transformation to all channel values of each activation pixel (e.g., each sample data) that can be in each column in the activation matrix, independent of the pixel position (e.g., the same transformation for each pixel).

[0100] R is a 16×16 channel unit matrix and is multiplied on the right side. This applies a linear transformation to all pixel values of each activation channel (e.g., each row in the activation matrix), independent of the channel position (e.g., the same transformation for each channel).

[0101] The transformation can also be represented as follows.

[0102]

Equation

[0103] Here, k varies between 1 and p for a p*p matrix (e.g., p = 16 in the example discussed above). 2 Thus, in one exemplary embodiment, for example, k can be 16. This can provide more degrees of freedom for training in a more meaningful layer (e.g., regarding weights, coefficients of L and R matrices, etc.). In one example, pruning by removing LXR transformations one by one during the training time is also possible, reducing complexity while maintaining the quality of the final image. Such scenarios will be discussed in more detail in relation to the training process.

[0104] As part of the execution of the neural network, the activation function 420 is applied. This can be a ReLU (Rectified Linear Unit), for example, if the value is negative, it is set to 0. If the value is positive, it remains as is. Depending on the specific application, other types of activation functions (e.g., linear function, tanh function, two-variable function, sigmoid function, leaky, parameter-based, and different versions of ReLU such as ELU, Swish, etc.) may also be used. For example, image processing may use one type of activation function, and natural language processing may use another. In an exemplary embodiment, the type of activation function used for a given layer can distinguish between the layers. For example, (in relation to the examples discussed in FIGS. 2 and 5 for upconverting an image, for example), the ReLU activation function may be used from layer 1 to n - 1 (where n is the number of layers), and the sigmoid activation function may be used in the nth (e.g., last) layer.

[0105] This process generates the activation transformation post-matrix 425. This is represented as X n+1 as. The process shown in FIG. 5 may be repeated a predetermined number of times or for a predetermined number of layers (e.g., 4) or otherwise performed. In this way, the activation matrix is changed from the initial matrix 225 of activation to the final version of the activation matrix 235 by the application of various trained L and R matrices. In one exemplary embodiment, the number of layers may vary between 2 and 12 or between 3 and 8. In one exemplary embodiment, more layers may be used with the understanding that additional layers may sometimes degrade performance. Thus, the number of layers may be selected based on the needs of a particular application and the balance between the resulting quality of the converted image and the performance of the up-conversion process. As the hardware becomes faster (or the performance is less of a controlling factor), additional layers may be added. In one exemplary embodiment, the number of layers may be dynamically controlled by the neural network 112, the video game application 108 (or other application such as an operating system handling the conversion process). For example, the system may measure the amount of time taken to process an image and add or remove layers based on such measurement. For example, if the conversion process is taking too long for real-time processing, a pre-trained network with one or more fewer layers may be used). Such techniques may be beneficial in accommodating different types of hardware resources used by a given computing device.

[0106] The following pseudo-code may represent the 16×16 to 16×16 matrix multiplication shown in FIG. 5 (in this example, multiplying matrix "Right" by matrix "Left").

[0107] [Table 3]

[0108] Here, Result[i][j] is the coefficient at the i-th column and j-th row (initialized to 0 before the loop).

[0109] Using separable block transform (SBT) at 410 in certain exemplary embodiments can be considered as an alternative to using fully connected layers / linear layers. A linear layer (e.g., a fully connected layer) is a matrix multiplication of an unstructured vector of input activations that gives an unstructured vector of output activations. For example, a 256×256 linear layer can be represented by a 256×256 matrix of independent weights and is applicable to 256 unstructured independent inputs. A potential drawback of this number of coefficients within a layer is that it may have too many coefficients (e.g., degrees of freedom) to train or calculate in runtime (e.g., to provide real-time image processing). Thus, certain exemplary embodiments may advantageously replace such a linear layer, for example, with a "low-rank approximation". An example of this is SBT. In certain exemplary embodiments, the SBT layer may be represented as a sum of LXR products where 256 inputs and outputs are structured into 16×16 matrices (as shown above). As noted above, this generalized version may be represented as follows.

[0110] [Number]

[0111] A special case SBT that is similar to or equivalent to a linear layer may be generated using the SBT layer. Specifically, it is as follows.

[0112] [Number]

[0113] L i,j n The matrix is set to a special form where the coefficient l i,j for each of the coordinates i, j is set to 1 and all other coefficients are zero is set. When l i,j = 1 and other coefficients are zero, the L n X n product is the matrix X nThe result is to extract the i-th line of and reposition it to the j-th line while setting the remainder to zero. Thus, X n+1 Each of the resulting j-th lines is a general linear combination of all the lines and thus the coefficients of X n is. Put another way, X n+1 All 256 output values in the matrix are a linear combination of the 256 input values of the X n matrix, which is the same as a linear layer of 256×256 coefficients. Thus, this configuration uses (in the R n matrix) 16×16×16×16 = 256×256 free coefficients. With this in mind, in situations where a linear layer is used (for example, it can be used as a substitute), a separable block transform technique may be applied.

[0114] Compared to a linear layer, SBT can offer one or more of the following advantages. In one exemplary embodiment, SBT may be gradually pruned by removing individual LXR terms (for example, those that contribute least to the quality of the result). Each removed LXR term can reduce the complexity of training and runtime calculations, the total number of weights stored and transmitted, and the remaining learning cost.

[0115] In one exemplary embodiment, a 16×16 SBT can be trained from the beginning with a number of LXR terms less than 256. This can also reduce the number of weights to be learned and the number of training and runtime operations.

[0116] In one exemplary embodiment, for a 16×16 SBT, the multiplications involving the sum of less than 8 LXR terms are fewer than those of a linear layer. For reference, a 256×256 linear layer (thus, the multiplication of a 256×256 matrix and a vector of size 256) involves 256×256 = 2 16 multiplications. In contrast, a single SBT involves two 16×16 matrix multiplications and thus 2×16×16×16 = 2 13Therefore, the sum of k L×R terms is multiplied by k*2 13 Therefore, k<2 3 (For example, 8) is more efficient than the linear layer. There will be fewer things to do.

[0117] Benefits of SBT compared to linear layers may include allowing a reduction in the number of weights (e.g., in some weight reuse scheme). It is recognized that a reduction in the number of weights may impact (e.g., perhaps significantly impact) performance because it may reduce memory traffic for handling the weights. This allows more space in memory to be devoted to activations. Memory pressure, e.g., in the form of external memory bandwidth or internal memory size, may also be alleviated (e.g., reduced).

[0118] In one exemplary embodiment, for a 16x16 SBT, the weights (and therefore memory and training time) for sums of LXR terms less than 128 are less than in a linear layer. 16 while a single 16×16SBT term has a weight of 2×256=2 9 The weight of is taken, therefore, 2 7 Adding =128 makes the weights equal.

[0119] In an example embodiment, SBT may be used to replace larger linear layers (e.g., 1024 to 1024, such as those used in natural language processing) with 32x32 SBT layers, which allows for a smaller number of weights while maintaining an acceptable level of quality. Thus, technical implementations of the SBT techniques discussed herein may be used in a variety of different applications and scenarios to achieve increased efficiency with little or no (e.g., perceived) loss of quality of the transformed data.

[0120] In certain exemplary embodiments, the size of the sums learned by trial and error and / or by global pruning can be different for each layer. In certain exemplary embodiments, a smaller version of the SBT network can be trained through distillation from a larger, trained version of the SBT network.

[0121] Description of Figure 6 FIG. 6 is a flowchart showing an enlarged view of a rearrangement part into blocks of the machine-learned up-conversion process of FIG. 2. Once an activation 16×16 matrix 235 is generated by executing through neural network 112, it is then reconverted back into the form of multiple channels. Specifically, each row of activation matrix 235 (or more specifically the first 12 rows, since the last 4 are all set to zero) is rearranged into a corresponding block of one output channel. Thus, as shown in FIG. 6, the first row of activation matrix 235 is converted back into the first block 602a of 1080p output data 245 (e.g., the red value of the top-left sub-pixel). And the second row of activation matrix 235 is converted back into the second block 602b of 1080 output data 245 (e.g., the green value of the top-left sub-pixel of the same channel), and so on. Thus, all 12 blocks (4 sub-pixel channels per block * 3 channels per color value) of the corresponding 12 rows of activation matrix 235 create 12 output channels of 1080p output data 245.

[0122] Description of Figure 7 FIG. 7 is a flowchart showing an enlarged view of a rearrangement part into a post-conversion image of the 1080p output data of the machine-learned up-conversion process of FIG. 2. In step 710, the 1080p output data 245 (e.g., 12 output channels of 4×4 blocks) is combined into a single 8×8 pixel block 712.

[0123] FIG. 7 shows an example of how values from a block (e.g., the highlighted value 713 from block 602) can be used to generate corresponding pixel values 714 (also highlighted) in pixel block 712. This includes combining color values to create each pixel. Thus, the red, green, and blue values of 713 from each of the red, green, and blue blocks 602 (e.g., from 602a, 602b) are used to generate the RGB values of pixel 714 in pixel block 712. The remaining 63 pixels in the 8×8 block are generated in a similar manner. The resulting 8×8 pixel block 712 is then positioned within the entire 1080p image 255.

[0124] This process of assembling the 8×8 pixel blocks is repeated (e.g., in parallel) for each of the 1080p output data 245 generated for a single (original) 540p image. At 720, a 1080p image 255 is assembled from multiple 8×8 pixel blocks 712. Each of the 8×8 pixel blocks is positioned within the overall image (e.g., based on the order in which the source image was processed). Thus, if the source image is processed left to right and top to bottom, the output image is constructed in a similar manner. Alternatively, in one exemplary embodiment, the position data for each pixel block may be stored as metadata for each of the input channels 215 created when it was originally created, for example, to determine where to position the 8×8 pixel block.

[0125] Once the 1080p image 255 is created, it is then output at 260 or stored otherwise (e.g., in a frame buffer) and can ultimately be displayed on display device 116.

[0126] Description from Figure 8A to 8B Figures 8A through 8B show an exemplary image 802 that is 128×128 pixels. Image 802 is applied to a neural network 803 that has been trained according to techniques discussed herein (e.g., in connection with FIG. 10). After applying image 802 to neural network 803, an enlarged image 804 is generated. Image 804 is a version of image 802 enlarged to 256×256 pixels.

[0127] FIG. 8B includes a version of the image from FIG. 8A that has been “zoomed” to create a tiled 512×512 pixel version. As shown in FIG. 8B, an image 822 that is a zoomed version of image 802 includes artifacts not seen in an image 824 that is a zoomed version of image 804. It should be recognized that the images shown in FIGS. 8A and 8B are shown as examples.

[0128] Description from Figure 8C to Figure 8E FIG. 8C shows an exemplary block diagram of a single “block tower” according to an exemplary embodiment. FIGS. 8D and 8E are exemplary block diagrams showing how several block towers may be used according to an exemplary embodiment.

[0129] FIG. 8C shows a block diagram that, in some respects, corresponds to the example discussed in connection with FIGS. 2 through 7. Specifically, a block of pixels 830 is selected from a source image 832. For block 830, at 834, a 16×16 activation matrix 836 is prepared (as described, e.g., in connection with FIGS. 3 and 4). Next, the activation matrix 836 is run through an SBT network 838 to create an output matrix 840 (as shown, e.g., in FIG. 5). Next, an output pixel block 844 is created at 842 and then placed into a converted image 846 (as shown, e.g., in FIGS. 6 and 7).

[0130] It is recognized that it may be beneficial to use a larger number of channels and / or L&R matrices (e.g., 32×32 or 64×64) as it can provide more expressiveness during processing. However, the drawback of this approach is that such matrices may not fit into local "fast" memory (e.g., registers), and thus slower DRAM usage may be required during processing. Although larger-sized fast memory may be possible in the future, the underlying problem of not having "sufficiently" fast memory may still remain.

[0131] In one exemplary embodiment, two 16×16 SBT towers (e.g., L&R matrices) with a 16-channel corresponding activation matrix may be used. Such an implementation can (at least partially) address the need to have more and more local fast memory, while still potentially benefiting from the increased expressiveness (higher degrees of freedom) that can be provided by using an increased number of channels (e.g., 32 or 64). In such cases, the SBTs may be processed sequentially or in parallel. In such an implementation, a given activation matrix may be executed through multiple different SBTs, and the outputs may be combined or used together in one of several different ways.

[0132] FIG. 8D shows a block diagram of an aggregation example for using several SBTs. As in the example of FIG. 8C, the activation matrix 836 is created from blocks using the source image. However, in this example, the activation matrix is applied to a plurality of different SBT networks. Specifically, the activation matrix 836 is applied to SBTs 852A, 852B, 852C, and 852D. In other words, the same activation matrix (derived from the same underlying pixel blocks) can be processed by separate SBTs (e.g., L&R matrices). Such processing may be done sequentially, in parallel, or some combination thereof (e.g., two at a time). Each SBT processes the activation matrix 836 in a different manner to create (presumably) four different output matrices 854A, 854B, 854C, and 854D. These four outputs are then aggregated term by term to create a final (e.g., 16×16) output matrix, which is then processed as discussed in connection with FIG. 8C.

[0133] FIG. 8E is a block diagram of an alternative example for using several SBTs. This example is the same as the example shown in FIG. 8D except that, instead of aggregating the results from several SBTs, at 860, the resulting outputs can be stacked or integrated together to form a larger matrix. This type of implementation may be useful, for example, for handling larger output blocks 862 of the output image 864, which may of course benefit from a larger number of activations in the output activation matrix. Such techniques may be similar or equivalent to, for example, channel grouping / grouped convolutions as used in various CNN architectures (e.g., AlexNet, MobileNet, etc.).

[0134]

[0135] Description of Figure 9 FIG. 9 is a block diagram including an exemplary training computer system 900 according to an exemplary embodiment. The training computer system 900 is an example of the computer system 1300 shown in FIG. 13. In an exemplary embodiment, the computer system 900 and the computer system 100 may be the same system (e.g., a system used to play a video game may also be configured to train a neural network for that video game).

[0136] System 900 includes a dataset preparation module 902 used to prepare images (e.g., 1080p images) input from a training set database 906. The images are prepared and then used to train a neural network via a neural network trainer module 904 (e.g., determine the L&R coefficients including each layer of the sum of the L&R transforms discussed herein). The neural network trainer module 904 generates one or more trained neural networks, which are stored in a database 908. Next, the trained neural network 908 can be communicated to various game devices 1, 2, 3, 4, 5, etc. (each of which can be an example of a game device 100) via a network 912 (e.g., the Internet) or via a physical medium (such as a game cartridge). In one exemplary embodiment, one or more trained neural networks can be distributed along with a game obtained by a user. For example, a user may download a game from an online store or the like, and one of the components of the game may be a neural network for processing images generated by the game. Similarly, a game provided on a cartridge or other physical medium may include one or more neural networks that can be used by a user to transform images generated by the game. In one example, the same instance of a game (e.g., an individual download or an instance of a particular physical medium) may be provided to multiple neural networks so that the game can be output to different types of displays (e.g., 1080p in one case, 1440p in another case, 4k in another case, etc.).

[0137] As discussed herein, different types of neural networks can be generated and distributed to various gaming devices. Thus, for example, gaming device 1 may receive and use a neural network different from the neural networks received and used by gaming devices 2, 3, 4, and 5. In one exemplary embodiment, each game (or more generally each application) may have a corresponding neural network (or neural networks) generated for that game (e.g., by system 900). Thus, for example, a gaming device may store multiple different neural networks and use such different networks based on the game (or type of game) being played on the corresponding gaming device. In one exemplary embodiment, multiple games may share or use the same neural network. For example, one neural network may be generated for a first-person shooting game, and another neural network may be generated for a strategy game, etc. Thus, games may be grouped based on their "type". Such classification of types may be based on the genre of the game or on another criterion such as the type of rendering engine the game uses to generate images therein.

[0138] In one exemplary embodiment, the game engine (or other service that provides the conversion functionality to the game engine) may dynamically determine to select one from among various neural networks depending on the remaining time available to "prepare the current video frame". If the frame is rendered quickly, there may be more time to upscale with a high-quality and low-speed neural network (e.g., one that includes additional layers), but if the frame uses more of the typically available 16 ms (for both rendering of the frame and subsequent upscaling of the image at 60 frames per second), the engine can select a faster neural network (e.g., one with fewer layers). However, the slower ones do not necessarily provide higher image quality. Such determination may be made through the "testing" phase of a video game application program (e.g., where the game engine generates a number of exemplary images), and / or during normal game play.

[0139] Returning to FIG. 9, the training data set 906 includes a plurality of data sets used as "targets". Thus, when generating a neural network to convert a 540p image to a 1080p image, this may include different 1080p images that would be used to generate the neural network. In one exemplary embodiment, the types of 1080p images can be selected according to specific use cases. In the case of a video game, the images may be 1080p images natively generated by the game engine. In one exemplary embodiment, the images may be from the same game engine or game that the neural network is being used with. Thus, for example, game A may include a game engine capable of generating 1080p images. This can be beneficial as a different version of game A that generates game images at 540p can be produced. This is for example because other versions of game A are created for lower power hardware such as mobile devices. Thus, 1080p images for the training data set can be placed using the game engine of game A and a neural network that can be used in connection with other versions of game A can be trained using the images (e.g., the neural network can be made to output 1080p images even if that version was not originally designed for such images).

[0140] In one exemplary embodiment, the target image (e.g., a 1080p image if the network is trained to upconvert from 540p to 1080) must have high visual quality. Such images may be prepared in advance and need not be rendered "in real time" (e.g., at 30 or 60 frames per second). Such images can be rendered as sharp and clean, and using higher anti-aliasing settings. Advantageously, the images may be generated from the same game or game engine as the target for which the trained network will be used. In such a scenario, the statistics of the training data may exactly match the statistics of the runtime data, so that the generated neural network can be better optimized for such games.

[0141] In one exemplary embodiment, a default or "common" selection of images may be used. Such an implementation can provide a good cross-section across multiple games. Such an implementation may select target images that are of relatively good or high quality, have a relatively good level of diversity, and sharpness (e.g., no relatively visible aliasing). This type of approach may make it possible to use the full spectrum of available spatial frequencies.

[0142] In one exemplary embodiment, artificially generated images can be used, in which case such images are rendered in pairs of low-resolution and high-resolution images. In one exemplary embodiment, different types of images (e.g., pixel art) may be selected and enlarged (e.g., in which case such images may be disadvantaged by the lack of available high-resolution images and may not look visually good even when enlarged through the use of general-purpose neural networks).

[0143] In one exemplary embodiment, the training computer system may be implemented in a cloud-based computing system.

[0144] Description of Figure 10 FIG. 10 is a flowchart showing a process that can be implemented on the system shown in FIG. 9 to train a neural network that can be used in connection with an exemplary embodiment that includes the process shown in FIG. 2.

[0145] From the training dataset 906, a plurality of target images or training images are selected. When training the neural network to upconvert to 1080p, the images can be a population of 1080p images 1000.

[0146] At 1002, each of the images within this population is passed to the dataset preparation module 902 to prepare a training dataset for use in training the neural network. This has two sub-processes. The first is to prepare the 1080p image into 1080p output data 1006. This aspect is discussed in FIG. 11. The second is to prepare the 540p image (or other image used as the source image) into 540p input data 1004. This aspect is discussed in FIG. 12. The processes discussed in FIGS. 11 and 12 may be repeated for each image used within the training dataset or otherwise performed. In an exemplary embodiment, the images may be streamed (e.g., the preparation process may proceed concurrently with the training process). In an exemplary embodiment, the preparation of the images may be batched (e.g., 256 or cropped sub-parts of such images may be prepared in a training batch before being used as data for one step within the neural network training process).

[0147] Next, at 1008, using the 540p input data 1004, train the neural network at 1010 until the training result converges to a coverage range that is close enough to the 1080p output data 1006. In other words, until the set of coefficients (e.g., L&R) converges to an acceptable approximation of the 1080p output data from the initial 540p input data. The training process is repeated until this convergence is reached (e.g., within an error threshold or because the error value has not decreased over a number of iterations greater than a threshold number of iterations).

[0148] Once converged, the trained neural network weights (e.g., the coefficients of the L&R matrix, which may sometimes be referred to herein as the trained neural network) 910 can be stored in a database within the system 900 and / or communicated to other computer systems (e.g., game devices 1, 2, 3, 4, 5, etc.).

[0149] In one exemplary embodiment, the techniques associated with the SBT network may enable a favorable environment for pruning. This is because each individual sum element (e.g., LXR) can be removed without interfering with the rest of the architecture, and because even if a connection remains, other connections do not directly depend on this particular term. In other words, each LXR term can be considered a single "branch" of the architecture that can be removed without disturbing the rest of the network. This type of approach can be advantageous. This is because each channel is generally used as an input to the next layer downstream, so removing channels in the remaining network can have negative consequences with respect to quality and / or performance.

[0150] In one exemplary embodiment, the determination of which LXR term (e.g., each SBT term) to remove (e.g., pruning) can be based on calculating the global loss with and without each LXR term (e.g., the result of the calculation of L*X*R as an individual term, or a part of the sum of LXR products), and then removing the term that has the least impact on the global loss. Thus, terms below a certain threshold may be removed, or the bottom x% of the terms may be removed (e.g., 1% or 5% may be removed), and then the process can be restarted until a given size or error target is reached.

[0151] In one exemplary embodiment, pruning may be performed on the SBT network by calculating or otherwise determining the loss gradient for each SBT term and removing the SBT term with the lowest gradient (or the terms with the lowest percentage).

[0152] Description of Figure 11 FIG. 11 is a flowchart showing an enlarged view of how 1080p image data is prepared as part of the neural network training process shown in FIG. 10.

[0153] Each 1080p image 1000 is cut into 8×8 pixel blocks at 1110. Next, at 1120, each pixel block (1122) is selected. The pixel block is then divided at 1130. FIG. 11 shows that the pixel block 1122 is divided into separate input channels for the step at 1130. As shown in FIG. 11, the corresponding numbered pixel values in the pixel block 1122 are assigned to the corresponding input channels. Each channel includes three separate input channels per RGB value of the source pixels. Thus, 12 input channels are created (e.g., 1132, 1134, 1136, 1138, etc., each having RGB) and used as 1080p output data. This process is repeated for each 1080p image to create multiple 1080p output data, and this output data is used during the training process of the neural network (e.g., to determine when the neural network converges).

[0154] Description of Figure 12 FIG. 12 is a flowchart showing an enlarged view of how 540p input data is prepared as part of the neural network training process shown in FIG. 10.

[0155] The 540p input data 1004 is prepared from the 1080p output data 1006 generated as shown in FIG. 11. Specifically, at 1210, a single 540p image 1212 is created using one of the output channels from the 1080p output data 1006.

[0156] From the created image, the process is similar to that shown in FIG. 3 where an image 540p to be upconverted is prepared at a certain point. Specifically, at 1220, each 4×4 pixel block (1214) within the created 540p image (which may correspond to the color channel of 1132 in FIG. 11, for example) is selected.

[0157] At 1230, next, context data is added around the 4×4 pixel block to create an 8×8 context block 1232a. The context data can be derived in a manner similar to that described above in connection with FIG. 2. At 1240, (which may be the same as context block 1232 a, but with a change in activation indexing) context block 1232b is split into four separate input channels. Each input channel includes three channels for the respective RGB values of the pixels included in the channel. As shown in FIG. 12, the input channels are created such that 1 in 1232b is mapped to channel 1242, 2 is mapped to channel 1244, and so on. The resulting twelve input channels constitute 540p input data 1004 (e.g., a 16×16 matrix), and using this input data, a neural network is trained during the training process discussed in connection with FIG. 10.

[0158] Using the techniques described above, low-resolution input can be generated by downsampling high-resolution input through point sampling (e.g., nearest neighbor). However, in other exemplary implementations, other downsampling methods may be used.

[0159] In one exemplary embodiment, an image rendered at high speed (e.g., at 60 fps etc.) by a real-time game engine can, of course, be similar to the image resulting from point-sampled downsampling. This is because each pixel value is calculated independently of its neighboring pixels. Thus, training a neural network using point-sampled data may be likely to better conform to the expansion of the game engine output. It can help the game engine in one exemplary embodiment to run faster. This is because the time-consuming anti-aliasing steps that are costly during traditional rendering stages can be skipped. Rather, such anti-aliasing can be handled more efficiently by the exemplary neural network techniques discussed herein.

[0160] Point sampling as part of the downsampling for the training process can provide additional benefits. The critically sampled signal is a discrete signal resulting from a continuous signal. In a continuous signal, the frequency reaches the maximum allowable frequency according to the Shannon-Nyquist sampling theorem (i.e., the signal frequency must not exceed half of the sampling frequency f), while the continuous signal can still be perfectly reconstructed from the discrete signal without any loss.

[0161] In the case of a high-resolution image, if such an image is critically sampled along the spatial frequency, the calculation of the spectrum of the entire signal (e.g., using the discrete Fourier transform) uses the entire allowable spectrum (e.g., from 0 to f / 2). When lower-resolution input data is prepared, the normal sampling theory first removes the high frequencies of the spectrum (e.g., any between f / 4 and f / 2) using a low-pass filter, which can then lead to a two-fold reduction using point sampling. Then, the resulting image will respect the sampling theorem by having frequencies below half of the (new) signal spatial sampling frequency f' (which is f / 2).

[0162] When the local spectrum is then calculated (e.g., for a 4×4 or 8×8 pixel block), the effective frequencies of the spectrum can mainly be located in the lower part (between 0 and f / 4) or the higher part (between f / 4 and f / 2) of the spectrum. If point sampling is used without first using a low-pass filter, the high frequencies (between f / 4 and f / 2) are not removed, but rather can be "folded" into the lower part of the spectrum (between 0 and f / 4, which in the newly downsampled signal will be between 0 and f' / 2).

[0163] Next, the neural network can reconstruct the signal in a non-linear (e.g., learned) way using the context information. For example, they learn whether the spectrum comes from the actual low frequencies and thus should be reconstructed as the low frequencies of the upsampled signal, or whether it comes from the high part of the spectrum and thus should be reconstructed as the high frequencies of the upsampled signal. In this way, in some cases, by using downsampling with point sampling in the training stage, up to twice as much information can be packed into the same memory space compared to conventional sampling techniques. In some cases, the high-resolution images used during training may be prepared according to techniques similar to those discussed above (e.g., using frequencies beyond the sampling limit). However, this is on the assumption that the image is not inappropriately resampled later through the display process.

[0164] In this way, in some cases, by using downsampling with point sampling in the training stage, up to twice as much information can be packed into the same memory space compared to conventional sampling techniques. In some cases, the high-resolution images used during training may be prepared according to techniques similar to those discussed above (e.g., using frequencies beyond the sampling limit). However, this is on the assumption that the image is not inappropriately resampled later through the display process.

[0165] Additional Exemplary Embodiments The processes discussed above generally relate to two-dimensional (e.g., image) data (e.g., signals). The techniques in this specification (e.g., the use of SBT) can also be applied to data or signals in other dimensions, such as 1D (e.g., speech recognition, anomaly detection in time series, etc.) signals and 3D (e.g., video, 3D texture) signals. The techniques can also be applied in other types of 2D areas, such as, for example, image classification, object detection and image segmentation, face recognition, style transfer, pose estimation, etc.

[0166] The processes discussed in relation to FIGS. 2 and 9 relate to upconverting an image from 540p to 1080. However, in other scenarios including 1) conversion to a different resolution (e.g., from 480p to 720p or 1080p and variations thereof), 2) downconverting an image to a different resolution, 3) converting an image without changing the resolution, 4) an image having other values regarding how the image is represented (e.g., grayscale), the techniques discussed herein may be used.

[0167] In one exemplary embodiment, the techniques herein may be applied to process an image to provide an anti-aliasing ability (e.g., in real time and / or during the execution time of an application / video game). In such an example, the size of the image remains the same before and after, but anti-aliasing is applied to the final image. Training for such a process may proceed by obtaining relatively low-quality images (e.g., rendered without anti-aliasing) and those rendered with high-quality anti-aliasing (or the level of anti-aliasing desired for a given application or use case) and training a neural network (e.g., L&R as discussed above).

[0168] Other examples of applications with a fixed resolution (e.g., converting an image from x resolution to x resolution) may include noise removal (e.g., related to the ray tracing process used by a rendering engine in a game engine). Another application example of the techniques herein may include inverse convolution in the context of, for example, removing image blur.

[0169] During the execution time, using the source image, prepare the input channels in a manner similar to that shown in FIG. 3. Specifically, each image is divided into blocks (e.g., 4×4), and context data is added to these blocks to create 8×8 context blocks. Next, the subsequent context blocks are divided into four input channels and made to have three channel colors per channel to create twelve input channels. Next, these twelve input channels are rearranged into a 16×16 matrix of activations in a manner similar to that shown in FIG. 4. Next, execute the matrix of activations through the neural network. Here, perform a separable block transformation using the L and R matrices developed through the training discussed above.

[0170] Once the matrix of activations is transformed, the first three (or any three obtainable based on training) output channels (e.g., the RGB values corresponding to the "1" pixel) are rearranged into their respective blocks and combined into a single 4×4 block. Repeat this process for each of the original 4×4 blocks obtained from the source image. The transformed blocks are all combined, thereby creating the resulting image, which can then be output.

[0171] In an exemplary embodiment, a classification process (e.g., finding / identifying an object in an image) may be used in combination with the SBT technique discussed herein. For example, a given image may be divided into 4×4 pixel blocks, and a sliding 3×3 block kernel transformation can be applied to all of the image blocks. In one example, the kernel may have other sizes (e.g., the kernel can have other sizes such as 2×2, or be separable into 3×1 and then 1×3).

[0172] In this example, eight blocks (e.g., 3×3 surrounding blocks) surrounding a given block and the block itself are processed by SBT, and the results are summed into a single target block (e.g., corresponding to the position of the selected block). Thus, the 16×16 block values are summed term by term.

[0173] For blocks on the edge of the image, blocks outside the image may be ignored. In certain exemplary embodiments, one or more block convolution layers can be changed to various types of reduction layers. For example, max or average pooling may be used, or downsampling using strides or other similar techniques may be used.

[0174] In certain exemplary embodiments, the neural network may include one or more normalization layers. Such layers can be generated by using batch normalization, weight normalization, layer normalization, group normalization, instance normalization, batch instance normalization, etc.

[0175] In certain exemplary embodiments, layer fusion can be achieved between consecutive block convolution layers to further reduce the pressure on the memory bandwidth (e.g., DRAM).

[0176] In certain exemplary embodiments, residual connections (e.g., skip connections) can be added between SBT layers to facilitate the training of deeper models.

[0177] For stride implementation examples, the blocks of the output image can be 2 times, or even less, in the horizontal and vertical dimensions. Thus, when the block convolution layer is alternated with the block stride layer (e.g., several times), the final image can end with only a 16×16 activation block. In certain exemplary embodiments, the output neuron count can be matched to a number of classes (e.g., for classification application examples), and the final block can be used as the input to a traditional fully connected layer.

[0178] For a 16×16 matrix, when the number of classes is 16 or less, the output classes may be put into the diagonal coefficients of the matrix. This enables learning the equivalent of a fully-connected layer in the L and / or R matrices even using a single LXR element without a sum (in the SBT training). More generally, for a number of classes greater than 16 and up to 256, an SBT with up to 256 sum elements may be used (which is equivalent to a fully-connected network of 256 neurons). For a number of classes less than 256, the sum of less than 256 LXR terms is likely to fit the problem well and the optimal number of terms can be found. In one exemplary embodiment, finding the optimal number of terms can be achieved by pruning the LXR sum. In one exemplary embodiment, finding the optimal number of terms is achieved by singular value decomposition (or matrix spectral decomposition) of a trained fully-connected layer to determine the number of "significant" singular values (e.g., those not close to zero) and train the corresponding number of LXR terms (e.g., two LXR terms for 32 significant singular values).

[0179] For pooling implementation examples, each group of 2×2 blocks is made into a single block by calculating the average (or maximum) of the corresponding terms. Thus, in one exemplary embodiment, the block convolutional layer may be alternated with the block pooling layer (e.g., several times) and the final image may end with a single block out of 16×16 activations. Similar to the stride implementation example, this final 16×16 activation may be used as the input to a traditional fully-connected layer by matching the output neuron count to the desired number of classes (e.g., for classification application examples).

[0180] It is recognized that the software implementation speed and / or the cost of dedicated acceleration hardware may be related to the activation accuracy. In other words, FP32 is more costly than FP16 which is more costly than INT8. In one exemplary embodiment, using INT8 may provide an attractive sweet spot regarding the trade-off between speed / quality and / or cost / quality. ​

[0181] In certain cases, the low - resolution and high - resolution outputs from the game engine may be used for training purposes (instead of, for example, downsampling). However, such an approach may lead to contradictions and / or may interfere with training. Images generated in such a manner may mitigate these problems when the rendering engine generating the images is "resolution - independent".

[0182] The specific exemplary embodiments discussed in relation to FIGS. 2 and 9 are given in the context of converting a 540p image to a 1080p image, but it is recognized that the techniques discussed herein may be applied to converting other resolutions to a new resolution. For example, whenever 540p is mentioned herein, similar techniques may always be applied to a 1080p source image. Also, whenever 1080p is mentioned in relation to the target image, the techniques discussed herein may always be applied to a 4k image (e.g., 3840×2160).

[0183] In certain exemplary embodiments, the conversion techniques discussed herein may operate in a two - stage process. In one example, a first image (e.g., a 1080p image) may be converted to, for example, an 8k image. Such a process may include first converting the 1080p image to a 4k image and then converting the resulting 4k image to an 8k image, in accordance with the techniques discussed herein.

[0184] Description of Figure 13 FIG. 13 is a block diagram of an exemplary computing device 1300 (which may also be referred to as, for example, an "arithmetic device", a "computer system", or an "arithmetic system") according to some embodiments. In some embodiments, the computing device 1300 includes one or more of one or more processors 1302, one or more memory devices 1304, one or more network interface devices 1306, one or more display interfaces 1308, and one or more user input adapters 1310. Additionally, in some embodiments, the computing device 1300 is connected to or includes one or more display devices 1312. Additionally, in some embodiments, the computing device 1300 is connected to or includes one or more input devices 1314. In some embodiments, the computing device 1300 may be connected to one or more external devices 1316. As described below, these elements (e.g., processor 1302, memory device 1304, network interface device 1306, display interface 1308, user input adapter 1310, display device 1312, input device 1314, external device 1316) are hardware devices (e.g., electronic circuits or combinations of circuits) configured to perform various different functions for and / or related to the computing device 1300.

[0185] In some embodiments, each or any of processors 1302 is, for example, a single-core processor or a multi-core processor, a microprocessor (which may sometimes be referred to as a central processing unit or CPU), a digital signal processor (DSP), a microprocessor associated with a DSP core, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) circuit, or a system-on-chip (SOC) (for example, an integrated circuit including other hardware components such as a CPU, a GPU, and memory and / or a memory controller (for example, a north bridge), an I / O controller (for example, a south bridge), a networking interface, etc.), or includes them. In some embodiments, each or any of processors 1302 uses an instruction set architecture such as x86 or advanced RISC machine (ARM). In some embodiments, each or any of processors 1302 is, for example, a graphics processing unit (GPU) which may be an electronic circuit designed to generate, for example, an image, or includes it. One or more processors 1302 may sometimes also be referred to as hardware processors, and in one example, one or more of processors 1302 may be used to form a processing system.

[0186] In some embodiments, each or any of the memory devices 1304 is or includes a random access memory (RAM) (such as dynamic RAM (DRAM) or static RAM (SRAM)), flash memory (such as based on NAND or NOR technology), hard disk, magneto-optical medium, optical medium, cache memory, register (such as holding instructions or data that can be executed by one or more of the processors 1302), or other types of devices that perform volatile or non-volatile storage of data and / or instructions (such as software executed on or by the processor 1302). The memory device 1304 is an example of a non-transitory computer-readable storage device. Memory devices as discussed herein may include memory provided on the same "die" as the processor (e.g., internal to the die on which the processor is located) and memory provided external to the die including the processor. Examples of "on-die" memory may include cache and registers while "off-die" or external memory may include DRAM. As discussed herein, on-die memory in the form of cache or registers may provide faster access at the expense of being more expensive to produce.

[0187] In some embodiments, each or any of the network interface devices 1306 includes one or more circuits (such as a baseband processor and / or a wired or wireless transceiver), and implements layer 1, layer 2, and / or higher layers for one or more wired communication technologies (such as Ethernet (registered trademark) (IEEE802.3)) and / or wireless communication technologies (such as Bluetooth (registered trademark), WiFi (e.g., IEEE802.11), GSM (registered trademark), CDMA2000, UMTS, LTE, LTE-Advanced (LTE-A), and / or other short-range (e.g., Bluetooth (registered trademark) Low Energy, RFID), medium-range, and / or long-range wireless communication technologies). The transceiver may comprise circuits for a transmitter and a receiver. The transmitter and the receiver may share a common housing and may share some or all of the circuits within the housing to perform transmission and reception. In some embodiments, the transmitter and the receiver of the transceiver may not share any common circuits at all and / or may be in the same or separate housings.

[0188] In some embodiments, each or any of the display interfaces 1308 receives data from the processor 1302 (e.g., via a discrete GPU, an integrated GPU, a CPU that performs graphical processing, etc.) for use in generating corresponding image data based on the received data, and / or outputs the generated image data to a display device 1312 that displays the image data (e.g., a high-definition multimedia interface (HDMI), a DisplayPort interface, a video graphics array (VGA) interface, a digital video interface (DVI), etc.) and is one or more circuits or includes such. Alternatively or in addition, in some embodiments, each or any of the display interfaces 1308 is or includes, for example, a video card, a video adapter, or a graphics processing unit (GPU). In other words, each or any of the display interfaces 1308 may include a processor used to generate image data therein. The generation of such images may be performed in conjunction with the processing performed by one or more of the processors 1302.

[0189] In some embodiments, each or any of the user input adapters 1310 is one or more circuits or includes such that receive and process user input data from one or more user input devices (1314) included in, attached to, or otherwise communicating with the computing device 1300, and output data to the processor 1302 based on the received input data. Alternatively or in addition, in some embodiments, each or any of the user input adapters 1310 is or includes, for example, a PS / 2 interface, a USB interface, a touch screen controller, etc. And / or the user input adapter 1310 facilitates input from the user input device 1314.

[0190] In some embodiments, display device 1312 may be a liquid crystal display (LCD) display, a light emitting diode (LED) display, or other types of display devices. In embodiments where display device 1312 is a component of computing device 1300 (e.g., the computing device and the display device are included in an integrated housing), display device 1312 may be a touch screen display or a non-touch screen display. In embodiments where display device 1312 is connected to computing device 1300 (e.g., external to computing device 1300 and communicating with computing device 1300 via wired and / or wireless communication technologies), display device 1312 may be, for example, an external monitor, a projector, a television, a display screen, etc.

[0191] In some embodiments, each or any of input devices 1314 is or includes a machine and / or electronic device that generates a signal provided to user input adapter 1310 in response to a physical phenomenon. Examples of input devices 1314 include, for example, a keyboard, a mouse, a trackpad, a touch screen, buttons, a joystick, sensors (e.g., an acceleration sensor, a gyro sensor, a temperature sensor, etc.). In some examples, one or more input devices 1314 generate a signal provided in response to a user input, such as by pressing a button or actuating a joystick. In other examples, one or more input devices generate a signal based on a detected physical quantity (e.g., force, temperature, etc.). In some embodiments, each or any of input devices 1314 is a component of the computing device (e.g., a button provided on a housing including processor 1302, memory device 1304, network interface device 1306, display interface 1308, user input adapter 1310, etc.).

[0192] In some embodiments, each or any of external devices 1316 includes a further computing device that communicates with computing device 1300 (e.g., another instance of computing device 1300). Examples include a server computer, a client computer system, a mobile computing device, a cloud It may include a base computer system, computing nodes, Internet of Things (IoT) devices, etc., all of which may communicate with the computing device 1300. Generally, the external device 1316 may include a device that communicates (e.g., electronically) with the computing device 1300. By way of example, the computing device 1300 may be a gaming device that communicates with a server computer system, which is an example of an external device 1316, over the Internet. Conversely, the computing device 1300 may be a server computer system that communicates with a gaming device, which is an exemplary external device 1316.

[0193] In various embodiments, the computing device 1300 includes one or each or any of the above-described elements (e.g., the processor 1302, the memory device 1304, the network interface device 1306, the display interface 1308, the user input adapter 1310, the display device 1312, the input device 1314), or two, or three, four, or more. Alternatively or in addition, in some embodiments, the computing device 1300 includes one or more of a processing system including the processor 1302, a memory or storage system including the memory device 1304, and a network interface system including the network interface device 1306.

[0194] The computing device 1300 can be arranged in many different ways in various embodiments. As a mere example, the computing device 1300 can be arranged such that the processor 1302 includes a multi (or single) - core processor, a first network interface device (for implementing, e.g., WiFi (registered trademark), Bluetooth (registered trademark), NFC, etc.), a second network interface device for implementing one or more cellular communication technologies (e.g., 3G, 4G LTE, CDMA, etc.), and a memory or storage device (e.g., RAM, flash memory, or hard disk). The processor, the first network interface device, the second network interface device, and the memory device can be integrated as part of the same SOC (e.g., one integrated circuit chip). As another example, the computing device 1300 can be arranged such that the processor 1302 includes two, three, four, five, or more multi - core processors, the network interface device 1306 includes a first network interface device for implementing Ethernet (registered trademark) and a second network interface device for implementing WiFi (registered trademark) and / or Bluetooth (registered trademark), and the memory device 1304 includes RAM and flash memory or a hard disk. As another example, the computing device 1300 can include a SoC having one or more processors 1302, a plurality of network interface devices 1306, a memory device 1304 including system memory and memory for application programs and other software, a display interface 13068 configured to output video signals, a display device 1312 integrated with the housing along with those mentioned and overlaid with a touch - screen input device 1314, and a plurality of input devices 1314 such as one or more joysticks, one or more buttons, and one or more sensors.

[0195] As previously noted, whenever a software module or software process is described in this document as performing any action, that action is in fact performed by underlying hardware elements in accordance with instructions that comprise the software module. Consistent with the foregoing, in various embodiments, each of the game device 100, game engine 110, neural network 112, input device 114, video game application 108, neural network trainer 904, dataset preparation module 902, or any combination thereof is implemented using the example of the arithmetic unit 1300 of FIG. 13. For the sake of clarity, however, for the remainder of this paragraph, each of these will be referred to individually as a "component." In such embodiments, the following applies to each component: (a) the elements of the arithmetic unit 1300 shown in FIG. 13 (i.e., one or more processors 1302 , one or more memory devices 1304, one or more network interface devices 1306, one or more display interfaces 1308, and one or more user input adapters 1310), or one or more display devices 1312, one or more input devices 1314, and / or external devices 1316, whether having or not having the above, any suitable combination or subset of the above is configured, adapted, and / or programmed to implement each or any combination of the features described herein as being performed by and / or included by an act, action, or component and / or within a component; (b) alternatively or in addition, to the extent that one or more software modules are described herein as being present within a component, in some embodiments, such software modules (as well as any data described herein as being handled by and / or used by the software modules) are stored in the memory device 1304 (e.g., in a volatile memory device such as RAM or an instruction register and / or in a non-volatile memory device such as a flash memory or a hard disk in various embodiments), and all acts described herein as being performed by the software modules are, as appropriate, performed by the processor 1302 together with other elements in and / or connected to the arithmetic unit 1300 (e.g., network interface device 1306, display interface 1308, user input adapter 1310, display device 1312, input device 1314, and / or external device 1316);(c) Alternatively or in addition, to the extent that components process and / or otherwise handle data as described herein, in some embodiments, such data is stored in a memory device 1304 (e.g., in some embodiments, in a volatile memory device such as RAM and / or in a non-volatile memory device such as flash memory or a hard disk), and / or as appropriate, processed / handled by the processor 1302 together with other elements in and / or connected to the computing device 1300 (e.g., network interface device 1306, display interface 1308, user input adapter 1310, display device 512, input device 1314, and / or external device 1316); (d) Alternatively or in addition, in some embodiments, the memory device 1302 stores instructions that, when executed by the processor 1302, cause the processor 1302 to perform, as appropriate, each or any combination of the acts described herein as being performed by a component and / or as included within a component and / or by any software module described herein together with other elements in and / or connected to the computing device 1300 (e.g., memory device 1304, network interface device 1306, display interface 1308, user input adapter 1310, display device 1312, input device 1314, and / or external device 1316).;

[0196] The above-described hardware configuration shown in FIG. 13 is given as an example, and the subject matter described herein may be utilized with a variety of different hardware architectures and elements. For example, in many of the figures of this document, individual functional / behavioral blocks are shown. In various embodiments, the functions of these blocks can be implemented using (a) individual hardware circuits, (b) application-specific integrated circuits (ASICs) specifically configured to perform the described functions / behavior, (c) one or more digital signal processors (DSPs) specifically configured to perform the described functions / behavior, (d) the hardware configuration described above with reference to FIG. 13, (e) through other hardware arrangements, architectures, and configurations, and / or through combinations of the techniques described in (a)-(e).

[0197] Technical Advantages of the Described Subject Matter In one exemplary embodiment, new techniques are provided for converting, transforming, or otherwise processing data from a source signal. Such techniques may include processing the data of the source signal in blocks and applying two separate learned matrices (e.g., pairs per layer of a trained neural network) to an activation matrix based on the blocked signal data, thereby generating an output matrix. One of the learned matrices is applied to the left side of the activation matrix and the other is applied to the right side. The sizes of the matrices (both the learned matrices and the activation matrix) may be selected to utilize hardware acceleration. The techniques may advantageously also process superimposable patterns of data (e.g., which may be pixels) from the source signal. The above-described hardware configuration shown in FIG. 13 is given as an example, and the subject matter described herein may be utilized with a variety of different hardware architectures and elements. For example, in many of the figures of this document, individual functional / behavioral blocks are shown. In various embodiments, the functions of these blocks can be implemented using (a) individual hardware circuits, (b) application-specific integrated circuits (ASICs) specifically configured to perform the described functions / behavior, (c) one or more digital signal processors (DSPs) specifically configured to perform the described functions / behavior, (d) the hardware configuration described above with reference to FIG. 13, (e) through other hardware arrangements, architectures, and configurations, and / or through combinations of the techniques described in (a)-(e).

[0198] In one exemplary embodiment, the arrangement of blocks of signal data (e.g., pixel data) can more effectively utilize the processing capacity of a processor (e.g., a GPU). For example, instead of leaving extra processing capacity (which may be considered wasted in terms of time and / or resources) unused, the GPU can operate at nearly 100% (e.g., at least 90 or 95 percent). Thus, according to some exemplary embodiments discussed herein (e.g., in connection with using separable block transforms versus conventional neural network approaches), something closer to the theoretical maximum processing throughput can be achieved.

[0199] In one exemplary embodiment, an image can be divided into blocks to improve how transforms are applied during the execution of a neural network. In one exemplary embodiment, the block size can be determined based on the smallest size matrix that can be used in the hardware (e.g., a GPU or ASIC, etc.) that is handling matrix operations. In one example, the atomic operations performed on input data from a 1080p source image are in a relatively fast time frame and can enable real-time image processing (e.g., exemplary atomic operations can be performed in less than about 0.04 ms).

[0200] The techniques herein enable a flexible approach in training models (e.g., neural networks) that can be tailored to different use cases. As an example, different neural networks can be trained to handle different types of games. One model may handle platformer games, and another model may handle first-person games. By using different models for different use cases (including specific models for specific games), the accuracy of the resulting images can be increased.

[0201] The techniques discussed in this specification can provide advantages for processing. For example, the processing can operate on relatively small grains, for example, by using 16×16×16 = 4096 multiplications per matrix multiplication. Thus, it is 2×4096 / 16 = 512 multiplications / pixel for each "atomic operation". And there are weights of 2×16×16 = 512, and thus it is 1KB per atomic operation in FP16. By increasing the width and depth of the network by multiples of the atomic operation, such processing can be scaled up as needed.

[0202] Advantageously, the techniques in this specification can operate with lower overhead on the DRAM of a computer system. This is because the data being computed during the application of the neural network to the activation matrix remains in the registers (e.g., internal memory) of the GPU (or other suitable hardware performing matrix operations).

[0203] In one exemplary embodiment, the techniques in this specification can provide for reducing the overall amount of storage space (e.g., file size) required to generate an image at a higher resolution size. For example, an application that generates an image at a higher resolution may also require assets (e.g., texture data) that are sized accordingly for the generation of such a high-resolution image. Thus, an exemplary application renders By reducing the image size, the size of the data used for such rendering can be similarly reduced, thus potentially using less memory or storage space. For example, the size of the textures used by the rendering engine can be reduced. Accordingly, the overall size required to distribute an application (such as a video game) can be reduced, making it suitable for relatively small physical media (e.g., regarding how much storage space is provided), and / or the bandwidth or amount of data required for downloading can be decreased. As an illustrative example, a video game designed to output images natively at 4K can have a total size of 60GB. However, if the size of the images generated by the video game engine is 1080p, the total size required for the video game can be reduced to, for example, 20GB. Even if the images are output at 1080p by the video game engine, such images can be converted to 4K images during runtime using the techniques described herein.

[0204] In one exemplary embodiment, the nature of how the data is prepared and the use of SBT can further leverage tensor hardware acceleration present in certain GPU hardware. For example, in NVIDIA's CUDA architecture, atomic neural network operations are 16×16 matrix multiplications. Some exemplary embodiments discussed herein are designed to work with this atomic operation. It is recognized that other types of hardware may have other sizes of atomic operations, and the techniques herein may also be adapted to such processing hardware.

[0205] Selected Terms In this document, whenever a given item is described as being present in "some embodiments", "various embodiments", "an embodiment", "an exemplary embodiment", "some exemplary embodiments", "the exemplary embodiment", or whenever any other similar language is used, it should be understood that the given item is present in at least one embodiment but not necessarily in all embodiments. Consistent with the above, whenever an action is described in this document as being "may", "can", or "could", a certain feature, element, or component is included "may", "can", or "could" within a given context, or is applicable to a given context, a given item has a given attribute "may", "can", or "could", or whenever any other similar phrase related to the terms "may", "can", or "could" is used, it should be understood that the given action, feature, element, component, attribute, etc. is present in at least one embodiment but not necessarily in all embodiments. The terms and phrases used in this document and their variations should be construed as non-limiting and non-restrictive unless otherwise specified. In the above example, "and / or" includes any and all combinations of one or more of the associated listed items (e.g., a and / or b means a, b, or a and b). The singular forms "a", "an", and "the" should be construed to mean "at least one", "one or more", etc. The term "example" is used to provide examples of the subject under discussion but is not an exhaustive or limiting list. The terms "comprise" and "include" (and other conjugations and other variations) identify the presence of the associated listed items but do not exclude the presence or addition of one or more other items. When an item is described as a "selected item", such a description should not be understood to indicate that the other items are not also selected items.

[0206] As used herein, the term "non-transitory computer-readable storage medium" includes registers, cache memories, ROMs, semiconductor memory devices (such as D-RAM, S-RAM, cache, or other RAMs), magnetic media such as flash memories, hard disks, magneto-optical media, optical media such as CD-ROMs, DVDs, or Blu-ray disks, or other types of devices for non-transitory electronic data storage. The term "non-transitory computer-readable storage medium" does not include transient propagating electromagnetic signals.

[0207] Additional Applications of the Described Subject Matter Process steps, algorithms, etc., including but not limited to referring to FIGS. 2 through 7 and FIGS. 10 through 12, may be described or claimed in a particular order, but such processes may be configured to operate in a different order. In other words, any sequence or order of steps that may be explicitly described or claimed in this document does not necessarily imply a requirement that the steps be performed in that order. Rather, the steps of the processes described herein may be performed in any order possible. Further, even if some steps are described or implied to occur non-simultaneously (e.g., because one step is described after another), they may be performed simultaneously (or in parallel). Additionally, the illustration of a process by its drawing in the figures does not imply that the illustrated process excludes other variations and modifications, does not imply that any of the illustrated processes or its steps are necessary, and does not imply that the illustrated process is preferred.

[0208] While various embodiments have been shown and described in detail, the claims are not limited to any particular embodiment or example. None of the above descriptions should be construed as implying that any particular element, step, range, or function is essential. All structural and functional equivalents to the elements of the above-described embodiments known to those of ordinary skill in the art are expressly incorporated herein by reference and are intended to be included. Further, the apparatus or method need not address every problem that is sought to be solved by the present invention, as it is included in the invention. No embodiment, feature, element, component, or step in this document is intended to be provided to the public.

Claims

1. 1. A computer system for converting an image to another resolution, the computer system comprising: a display device configured to display an image; and a processing system including at least one hardware processor, the processing system comprising: (a) acquiring a first image at a first resolution; (b) selecting a first block of pixels from the first image; (c) generating a first plurality of input channels based on the first block of pixels; (d) inserting values ​​from each of the first plurality of input channels into a first activation matrix; (e) applying the first activation matrix to the trained neural network to generate a second activation matrix; (f) generating a second image at a second resolution based on the second activation matrix; and (g) outputting the second image to a display; It is configured as follows: (b) through (g) are performed for each of a plurality of pixel blocks selected from the first image.

2. The computer system of claim 1 , wherein the second image is output to the display in real time relative to the generation of the first image.

3. 3. The computer system of claim 1, wherein applying the first activation matrix to the trained neural network comprises applying a separable block transform using the first activation matrix together with first and second matrices of learned coefficients of the trained neural network.

4. 4. The computer system of claim 3, wherein as part of the separable block transform, the first matrix is ​​multiplied on the left side of the activation matrix and the second matrix is ​​multiplied on the right side.

5. 5. The computer system of claim 1, wherein the trained neural network includes a plurality of distinct layers, the plurality of distinct layers being applied in succession to transform the first activation matrix into the second activation matrix.

6. 6. The computer system of claim 5, wherein each result of applying a different one of the layers to successive activation matrices is stored in an internal memory of a GPU included in the processing system.

7. 7. The computer system of claim 5 or 6, wherein a first layer of the plurality of layers is applied to the first activation matrix to generate a third activation matrix, and a second layer of the plurality of layers is applied to the third activation matrix to generate the second activation matrix.

8. The computer system of any one of claims 5 to 7, wherein the plurality of tiers is between three distinct tiers and eight distinct tiers.

9. the processing system includes a graphics processing unit (GPU); A computer system as claimed in any preceding claim, wherein results of matrix operations which are part of neural network processing are maintained in registers of the GPU during the neural network processing.

10. 10. The computer system of claim 1, wherein the steps (b) through (g) for the plurality of pixel blocks are performed in parallel using hardware acceleration to apply each activation matrix to the trained neural network.

11. 11. The computer system of claim 1, wherein the processing system is further configured to generate a plurality of images in association with execution of a video game or application program, each image generated being converted to another version of the second resolution and then displayed on the display.

12. 12. The computer system of claim 11, wherein the plurality of other versions of the plurality of images are displayed at least 30 times per second.

13. The computer system of any one of claims 1 to 12, wherein a size of the first block of pixels is based on a post hardware accelerated matrix multiplication size of a graphics processing unit that is part of the processing system.

14. 14. The computer system of claim 13, wherein the size is 4 pixels by 4 pixels and the post-hardware accelerated matrix multiplication size is 16x16.

15. The computer system of any one of claims 1 to 14, wherein each pixel in the first image is represented in the activation matrix by a separate RGB value of a corresponding pixel color.

16. The computer system of any one of claims 1 to 15, wherein at least some of the rows in the first activation matrix are set to zero.

17. 17. The computer system of claim 1, wherein at least some of the rows or columns in the first activation matrix are populated with motion or depth information generated in connection with generating the first image.

18. The processing system includes: generating a second plurality of output channels from the second activation matrix; combining the second plurality of output channels to form an output pixel block forming a portion of the second image. The computer system of any one of claims 1 to 17, further configured to:

19. 19. The computer system of claim 1, wherein the processing system is further configured to add context data around the first pixel block to create a first context block, and the first plurality of input channels are based on the first context block.

20. 20. The computer system of claim 19, wherein each pixel data from the first context block is split into a plurality of separate input channels of the first plurality of input channels, each of the plurality of separate input channels being a different color value of a corresponding pixel.

21. The computer system of any preceding claim, wherein the first resolution is lower than the second resolution.

22. 21. The computer system of claim 20, wherein the first resolution is 1080p and the second resolution is 4k.

23. 23. The computer system of claim 1, further comprising a non-transitory computer-readable storage medium configured to store a plurality of different trained neural networks, at least a first trained neural network trained to convert images in the first resolution to the second resolution and a second trained neural network trained to convert images in the first resolution to a third resolution different from the second resolution.

24. 24. The computer system of claim 1, further comprising a transceiver configured to communicate with another computer system to receive the trained neural network from the other computer system.

25. The computer system of any preceding claim, wherein each one of the first plurality of input channels is inserted into a different row of the first activation matrix.

26. The processing system includes a first processing system and a second processing system, the first processing system and the second processing system communicating with each other via a computer network, the first processing system: (a) is carried out, communicating data based on the first image to the second processing system; It is configured as follows: the second processing system is configured to perform (e), (f), and (g); The computer system of any preceding claim, wherein the display is located proximate to the second processing system.

27. the processing system is further configured to generate, in association with playing a video game, the first image at the first resolution; The computer system of any one of claims 1 to 26, wherein the second image is output for the video game.

28. A computer program product stored on a non-transitory storage medium, the computer program product for execution by a processing system including at least one hardware processor, the computer program product, when executed, causing the processing system to: (a) acquiring image data for a first image at a first resolution; (b) selecting a first block of pixels from the image data; and (c) generating a first plurality of input channels based on the first block of pixels; and (d) inserting values ​​from each of the first plurality of input channels into a first activation matrix; (e) applying the first activation matrix to the trained neural network to generate a second activation matrix; (f) generating a second image at a second resolution based on the second activation matrix; and (g) outputting the second image to a display device and displaying the second image on the display device; including an instruction to (b) through (g) are performed for each of a plurality of pixel blocks derived from the image data.

29. 1. A method executed on a computer system, the method comprising: (a) generating first image data at a first resolution; (b) selecting a first block of pixels from the first image data; and (c) generating a first plurality of input channels based on the first block of pixels; and (d) inserting values ​​from each of the first plurality of input channels into a first activation matrix; (e) applying the first activation matrix to the trained neural network to generate a second activation matrix; (f) generating a second image at a second resolution based on the second activation matrix; and (g) outputting the second image to a display device and displaying the second image on the display device; Equipped with A method wherein (b) through (g) are performed for each of a plurality of pixel blocks derived from the first image.

30. 1. A computer system for training a neural network to convert an image from a first resolution to a second resolution, the computer system comprising: a non-transitory computer readable storage device configured to store a plurality of target images, the plurality of images including a first image at the first resolution, the computer system further comprising: A processing system including at least one hardware processor, the processing system comprising: (a) dividing the first image into a first plurality of pixel blocks; (b) splitting each one of the first plurality of pixel blocks into a plurality of separate output channels to form target output data; (c) generating a second image at the second resolution based on one of the plurality of separate output channels; (d) generating a plurality of context blocks from the second image; (e) splitting the plurality of context blocks into a plurality of separate input channels; (f) training the neural network using the plurality of distinct input channels until the neural network converges to the target output data; The computer system is configured to:

31. 31. The computer system of claim 30, wherein each of the plurality of separate channels is comprised of a plurality of sub-channels.

32. 32. The computer system of claim 31, wherein each one of the plurality of sub-channels that make up a channel corresponds to a different color value that makes up an individual pixel in the first image.

33. 33. The computer system of claim 32, wherein each of the plurality of sub-channels corresponds to a different one of red, green, and blue values ​​of an RGB image.

34. 34. The computer system of claim 30, wherein the plurality of separate channels is four separate channels, one of the four plurality of channels being used to generate the second image.

35. 35. The computer system of claim 34, wherein each of the four separate channels is comprised of a number of sub-channels, each sub-channel corresponding to a different individual color value.

36. The computer system of any one of claims 30 to 35, wherein the size of each of the plurality of pixel blocks is the same as the size of each one of the plurality of context blocks.

37. 37. The computer system of claim 30, wherein the processing system is further configured to select a second plurality of pixel blocks from the generated second image, each one of the plurality of pixel blocks including data for a plurality of pixels from the generated second image, and each one of the plurality of context blocks being based on a corresponding one of the second plurality of pixel blocks.

38. 38. The computer system of claim 37, wherein context data is added to each pixel block to create a corresponding context block.

39. A computer system according to any one of claims 30 to 38, wherein each content block is divided into four separate input channels, each one of said plurality of separate input channels including a plurality of sub-channels.

40. The plurality of images includes a plurality of images with different target resolutions, and the processing system: (a2) dividing one of the plurality of images into a first plurality of pixel blocks; (b2) splitting each one of the first plurality of pixel blocks into a plurality of separate output channels to form target output data; (c2) generating a second image having a resolution lower than the resolution of the one of the plurality of images based on one of the plurality of separate output channels; (d2) generating a plurality of context blocks from the second image; (e2) splitting the plurality of context blocks into a plurality of separate input channels; (f2) training another neural network using the plurality of distinct input channels until the neural network converges to the target output data; 40. The computer system of any one of claims 30 to 39, further configured to:

41. the plurality of images includes a plurality of images generated by different game engines of different video games, each of the plurality of images having a resolution, and the processing system: (a2) dividing one of the plurality of images into a first plurality of pixel blocks; (b2) splitting each one of the first plurality of pixel blocks into a plurality of separate output channels to form target output data; (c2) generating a second image having a resolution lower than the resolution of the one of the plurality of images based on one of the plurality of separate output channels; (d2) generating a plurality of context blocks from the second image; (e2) splitting the plurality of context blocks into a plurality of separate input channels; (f2) training another neural network using the plurality of distinct input channels until the neural network converges to the target output data; 40. The computer system of any one of claims 30 to 39, further configured to:

42. a transceiver configured to receive a plurality of different requests for the video game enabled neural network from a plurality of different computing devices; The processing system includes: selecting, for each corresponding request, from among a plurality of different trained neural networks, at least one of the trained neural networks based on data included in the corresponding request; communicating, via the transceiver, the selected one of the trained neural networks to a requesting computing device. A computer system according to any one of claims 30 to 41, further configured to:

43. 43. The computer system of claim 42, wherein the data included in the corresponding request is an identifier for a particular video game.

44. 43. The computer system of claim 42, wherein the data included in the corresponding request indicates a target resolution.

45. The computer system of any one of claims 30 to 44, wherein the processing system is further configured to generate the first image by using a rendering engine for a video game.

46. 46. ​​The computer system of claim 45, wherein the trained neural network is stored in the non-transitory computer readable storage medium in association with the video game.

47. 47. The computer system of claim 46, wherein the trained neural network is communicated over a computer network to a gaming device being used to play the video game.

48. 48. The computer system of claim 30, wherein the trained neural network includes multiple separable block transform (SBT) terms across multiple layers of the trained neural network.

49. 49. The computer system of claim 48, wherein the processing system is further configured to prune the trained neural network by removing at least one SBT term from the plurality of SBT terms, and the pruned trained neural network is communicated to a plurality of computing devices for use on the plurality of computing devices.

50. The processing system includes: Calculating a first loss value from the trained neural network including a first SBT term; Calculating a second loss value from the trained neural network that does not include the first SBT term; calculating a difference between the first loss value and the second loss value; The method further comprises:

50. The computer system of claim 49, wherein the first SBT term is pruned based on the calculated difference.

51. A computer program product stored on a non-transitory storage medium, the computer program product for execution by a processing system including at least one hardware processor, the computer program product, when executed, causing the processing system to: acquiring a plurality of target images, the plurality of images including a first image at a first resolution, the computer program product, when executed, further comprising instructions to cause the processing system to: Dividing the first image into a first plurality of pixel blocks; splitting each one of the first plurality of pixel blocks into a plurality of separate output channels to form target output data; generating a second image at a second resolution based on one of the plurality of separate output channels; generating a plurality of context blocks from the second image; splitting the plurality of context blocks into a plurality of separate input channels; training the neural network using the plurality of separate input channels until the neural network converges to the target output data; A computer program product comprising instructions to cause a

52. 1. A method executed on a computer system, the method comprising: processing a plurality of target images, the plurality of images including a first image at a first resolution; Separating the first image into a first plurality of pixel blocks; splitting each one of the first plurality of pixel blocks into a plurality of separate output channels to form target output data; generating a second image at a second resolution based on one of the plurality of separate output channels; generating a plurality of context blocks from the second image; splitting the plurality of context blocks into a plurality of separate input channels; training the neural network using the plurality of separate input channels until the neural network converges to the target output data.

53. 1. A distributed computer game system, comprising: a display device configured to output an image at a target resolution; and a cloud-based computer system including a plurality of processing nodes, at least one of the processing nodes comprising: executing a first video game on at least one of the processing nodes to generate images for the first video game at a first resolution; Transmitting image data based on the generated image. It is further configured as follows: a client computing device configured to receive the image data, the client computing device including at least one hardware processor, the at least one hardware processor comprising: and configured to execute a neural network based on the received image data to generate a target image, the execution of the neural network applying a separable block transform to a plurality of activation matrices, each activation matrices corresponding to a different block of pixel data in the image represented by the image data, and the target image is generated at the target resolution, the at least one hardware processor further comprising: a target image output to said display device at said target resolution for display on said display device during gameplay of said first video game;

54. 54. The distributed computer game system of claim 53, wherein the first resolution is less than the target resolution.

55. 54. The distributed computer game system of claim 53, wherein the first resolution is the same as the target resolution.

56. 1. A method for transforming signal data using a neural network, the method comprising: (a) populating an initial activation matrix with values ​​based on data from a plurality of samples from a source signal; (b) applying a separable block transformation based on at least a first trained matrix and a second trained matrix to input activation matrices across multiple layers of the neural network to generate corresponding output activation matrices, wherein the initial activation matrix is ​​used as the input activation matrix for a first layer of the multiple layers, and the input activation matrix for each subsequent layer is the output activation matrix of a previous layer; and (c) outputting the output activation matrix of a final layer of the neural network to generate a transformed signal based on the output activation matrix of the final layer.

57. 57. The method of claim 56, wherein at least two of the rows or columns of the initial activation matrix correspond to superimposable data from each of the plurality of samples.

58. 58. A method according to claim 57, wherein the sample data in each row or column is for one colour of the colour values ​​which make up each pixel of the image.

59. 59. The method of any one of claims 56 to 58, wherein the initial activation matrix is ​​a p×p matrix, and the first trained matrix and the second trained matrix are p×p matrices of coefficients.

60. 60. The method of claim 59, wherein the input activation matrix, the output activation matrix, and the first trained matrix and the second trained matrix are 16x16 matrices.

61. The method of any one of claims 56 to 60, wherein the first trained matrix and the second trained matrix are a sample identity matrix and a channel identity matrix.

62. 62. The method of claim 61, wherein the sample identity matrix applies a transformation to all channel values ​​of each activation sample in the input activation matrix, independent of sample position.

63. 63. The method of claim 62, wherein each activation sample is a respective column of the initial activation matrix.

64. 62. The method of claim 61, wherein the channel identity matrix applies a transformation to all sample values ​​of each activated channel, independent of the channel position.

65. 65. The method of claim 64, wherein each activation channel is a row in the input activation matrix.

66. 66. The method of any one of claims 56 to 65, wherein the first matrix is ​​multiplied on the left side of the input activation matrix and the second matrix is ​​multiplied on the right side of the input activation matrix.

67. 67. The method of any one of claims 56 to 66, wherein the neural network includes at least one residual connection connecting one of the layers of the plurality of layers to another of the layers of the neural network.

68. A method according to any one of claims 56 to 67, wherein each input activation matrix and each output activation matrix are held within an internal memory of a graphics processing unit (GPU).

69. 69. The method of claim 68, wherein the internal memory is a register in the GPU.

70. A method according to any one of claims 56 to 69, wherein at least one input activation matrix and at least one output activation matrix are held within a single semiconductor hardware.

71. 71. The method of any one of claims 56 to 70, comprising performing (a) to (c) simultaneously on parallel processing hardware for each initial activation matrix for a different block of sampled data in the source signal, thereby generating the transformed signal.

72. 72. The method of any one of claims 56-71, wherein (a) through (c) are accomplished at least 30 times per second for 30.

73. 73. A method according to any one of claims 56 to 72, wherein at least two different activation functions are used across the layers of the neural network.

74. 74. The method of claim 73, wherein rectified linear units are used between at least two layers of the neural network.

75. 75. The method of any one of claims 56 to 74, wherein the neural network comprises 3 to 8 layers.

76. A method according to any one of claims 56 to 75, wherein each activation matrix is ​​populated with values ​​based on an nxm block of samples from the source signal.

77. A method according to any one of claims 56 to 76, wherein each activation matrix is ​​populated with values ​​based on a block of n samples from the source signal.

78. A method according to any one of claims 56 to 77, wherein each of the plurality of samples is the source signal and is a plurality of pixels from a source image at a first resolution, and the transformed signal is an image at a second resolution.

79. 80. The method of claim 78, wherein the first resolution and the second resolution are different resolutions.

80. pooling the plurality of output activation matrices to generate a second plurality of matrices; applying the second plurality of matrices as input matrices to a second neural network; The method of any one of claims 56 to 79, wherein the number of said plurality of output activation matrices is smaller than the number of said second plurality of matrices.

81. A method according to any one of claims 56 to 80, wherein at least two input activation matrices applied to the neural network are calculated at stride positions of the source signal.

82. A method according to any one of claims 56 to 81, wherein block convolution patterns of a neural network using separable block transforms are applied to the source signals in a translation-invariant manner.

83. A method according to any one of claims 56 to 82, wherein a neural network using a separable block transform is used in the context of transforming a signal in a convolutional manner instead of multiplying individual scalar activation values ​​by individual scalar weight values.

84. A method according to any one of claims 56 to 83, wherein the neural network includes at least one normalisation layer.

85. A computer program product stored on a non-transitory storage medium, the computer program product for execution by a processing system including at least one hardware processor, the computer program product, when executed, causing the processing system to: populating an initial activation matrix with a plurality of values ​​based on data from a plurality of samples from the source signal; and applying a separable block transformation based on at least a first trained matrix and a second trained matrix to an input activation matrix across a plurality of layers of the neural network to generate a corresponding output activation matrix, wherein the initial activation matrix is ​​used as the input activation matrix for a first layer of the plurality of layers, and the input activation matrix for each subsequent layer is the output activation matrix of a previous layer. The computer program product, when executed, causes the processing system to:

22. A computer program product comprising: instructions for outputting the output activation matrix of a final layer of the neural network; and generating a transformed signal based on the output activation matrix of the final layer.

86. 1. A computer system comprising: A processing system including at least one hardware processor, the processing system comprising: placing a plurality of values ​​based on data from a plurality of samples from the source signal into an initial activation matrix; configured to apply a separable block transformation based on at least a first trained matrix and a second trained matrix to an input activation matrix across a plurality of layers of the neural network to generate a corresponding output activation matrix, the initial activation matrix being used as the input activation matrix for a first layer of the plurality of layers, and the input activation matrix for each subsequent layer being the output activation matrix of a previous layer; and further configured for the processing system to: a computer system configured to output the output activation matrix of a final layer of the neural network to generate a transformed signal based on the output activation matrix of the final layer.

87. 87. The computer system of claim 86, wherein the source signal is a source image and the transformed signal is a transformed image.

88. 88. The computer system of claim 87, wherein the source image is in a first resolution and the transformed image is in a second resolution different from the first resolution.

89. A computer system according to any one of claims 86 to 88, wherein the source signal is generated in response to execution of a computer application program, and the transformed signal is output in association with execution of the computer application program.

90. 90. The computer system of claim 89, wherein the computer application program is a video game application program, and the source signal is generated by a rendering engine of the video game application program.

91. At least one of the hardware processors is a graphics processing unit (GPU), and each of the input activation matrix and the output activation matrix is ​​stored in an internal memory of the GPU. The computer system according to any one of claims 86 to 90,

92. The computer system of any one of claims 86 to 91, wherein at least one input activation matrix and at least one output activation matrix are held within a single semiconductor hardware of the processing system.

93. A distributed computer game system according to any one of claims 53 to 55, wherein the target image is output to a display in real time to receipt of the image data.

94. A distributed computer game system as described in any one of claims 53 to 55 and 93, wherein applying a separable block transform to the plurality of activation matrices comprises applying a separable block transform using first and second matrices of trained coefficients of the neural network together with corresponding activation matrices of the plurality of activation matrices.

95. 95. The distributed computer game system of claim 94, wherein as part of the separable block transform, the first matrix is ​​multiplied on the left side of the corresponding activation matrix and the second matrix is ​​multiplied on the right side.

96. A distributed computer game system as described in any one of claims 53 to 55 and 93 to 95, wherein the neural network includes a plurality of distinct layers, the plurality of distinct layers being applied in succession to transform an input activation matrix into an output activation matrix.

97. 97. The distributed computer game system of claim 96, wherein each result of applying a different one of the plurality of layers to successive activation matrices is stored in an internal memory of a GPU included in the client computing device.

98. 98. A distributed computer game system as described in claim 96 or 97, wherein a first layer of the plurality of layers is applied to the input activation matrix to generate a third activation matrix, and a second layer of the plurality of layers is applied to the third activation matrix to generate the output activation matrix.

99. A distributed computer game system according to any one of claims 96 to 98, wherein the plurality of tiers is between three distinct tiers and eight distinct tiers.

100. the client computing device is a graphics processing unit (GPU); A distributed computer game system as claimed in any one of claims 53 to 55 and 93 to 99, wherein results of matrix operations that are part of neural network processing are maintained in registers of the GPU during the neural network processing.

101. A distributed computer game system according to any one of claims 53 to 55 and 93 to 100, wherein the target image is output at least 30 times per second.

102. A distributed computer game system as claimed in any one of claims 53 to 55 and 93 to 101, wherein each pixel in the received image data is represented in a corresponding input activation matrix by separate RGB values ​​of the corresponding pixel's colour.

103. A distributed computer game system according to any one of claims 53 to 55 and 93 to 102, wherein at least some of the rows of the plurality of activation matrices are set to zero.

104. At least some of the rows or columns in the plurality of activation matrices include the cloud A distributed computer game system according to any one of claims 53 to 55 and 93 to 103, in which movement or depth information transmitted from a base computer system is arranged.

105. The at least one hardware processor: generating a second plurality of output channels from the second activation matrix; combining the second plurality of output channels to form an output pixel block forming a portion of the target image.

105. The distributed computer game system of any one of claims 53 to 55 and 93 to 104, further configured to:

106. A distributed computer game system as described in any one of claims 53 to 55 and 93 to 105, wherein the at least one hardware processor is configured to add context data around each block of pixel data to create a corresponding context block, and a first plurality of input channels for a first activation matrix is ​​based on a first of the context blocks.

107. 107. The distributed computer game system of claim 106, wherein each pixel data from the first context block is split into a plurality of separate input channels of the first plurality of input channels, each of the plurality of separate input channels being a different color value of a corresponding pixel.

108. A distributed computer game system according to any one of claims 53 to 55 and 93 to 107, wherein the first resolution is lower than the target resolution.

109. 109. The distributed computer game system of claim 108, wherein the target resolution is 4k.

110. 30. The computer program product of claim 28, wherein applying the first activation matrix to the trained neural network comprises applying a separable block transformation that uses first and second matrices of learned coefficients of the trained neural network together with the first activation matrix.

111. 111. The computer program product of claim 110, wherein as part of the separable block transform, the first matrix is ​​multiplied on the left side of the activation matrix and the second matrix is ​​multiplied on the right side.

112. 1. A method for transforming signal data using a neural network, comprising: (a) populating an initial activation matrix with values ​​based on data from a plurality of samples from a source signal; (b) applying a separable block transformation based on at least a first trained matrix and a second trained matrix to input activation matrices across multiple layers of the neural network to generate corresponding output activation matrices, wherein the initial activation matrix is ​​used as the input activation matrix for a first layer of the multiple layers, and the input activation matrix for each subsequent layer is the output activation matrix of a previous layer; and (c) outputting the output activation matrix of a final layer of the neural network to generate a transformed signal based on the output activation matrix of the final layer.

113. 113. The method of claim 112, wherein at least two of the rows or columns of the initial activation matrix correspond to superimposable data from each of the plurality of samples.

114. 114. A method according to claim 113, wherein the sample data in each row or column is for one colour of the colour values ​​which make up each pixel of the image.

115. 115. A method according to any one of claims 112 to 114, wherein the initial activation matrix is ​​a p×p matrix, and the first trained matrix and the second trained matrix are p×p matrices of coefficients.

116. 116. The method of claim 115, wherein the input activation matrix, the output activation matrix, and the first trained matrix and the second trained matrix are 16x16 matrices.

117. A method according to any one of claims 112 to 116, wherein the first trained matrix and the second trained matrix are a sample identity matrix and a channel identity matrix.

118. 118. The method of claim 117, wherein the sample identity matrix applies a transformation to all channel values ​​of each activation sample in the input activation matrix, independent of sample position.

119. 119. The method of claim 118, wherein each activation sample is a respective column of the initial activation matrix.

120. 118. The method of claim 117, wherein the channel identity matrix applies a transformation to all sample values ​​of each activated channel independent of the channel position.

121. 121. The method of claim 120, wherein each activation channel is a row in the input activation matrix.

122. The method of any one of claims 112 to 121, wherein the first matrix is ​​multiplied on the left side of the input activation matrix and the second matrix is ​​multiplied on the right side of the input activation matrix.

123. 123. The method of any one of claims 112 to 122, wherein the neural network includes at least one residual connection connecting one of the layers of the plurality of layers to another layer of the plurality of layers of the neural network.

124. A method according to any one of claims 112 to 123, wherein each input activation matrix and each output activation matrix are held within an internal memory of a graphics processing unit (GPU).

125. 125. The method of claim 124, wherein the internal memory is a register in the GPU.

126. The method of any one of claims 112 to 125, wherein at least one input activation matrix and at least one output activation matrix are held in a single semiconductor hardware.

127. 127. A method according to any one of claims 112 to 126, comprising performing (a) through (c) simultaneously on parallel processing hardware for each initial activation matrix for a different block of sampled data in the source signal, thereby generating the transformed signal.

128. 128. The method of any one of claims 112-127, wherein (a) through (c) are accomplished at least 30 times per second for 30.

129. A method according to any one of claims 112 to 128, wherein at least two different activation functions are used across the layers of the neural network.

130. 130. The method of claim 129, wherein rectified linear units are used between at least two layers of the neural network.

131. The method of any one of claims 112 to 130, wherein the neural network comprises 3 to 8 layers.

132. A method according to any one of claims 112 to 131, wherein each activation matrix is ​​populated with values ​​based on an nxm block of samples from the source signal.

133. A method according to any one of claims 112 to 132, wherein each activation matrix is ​​populated with values ​​based on a block of n samples from the source signal.

134. A method according to any one of claims 112 to 133, wherein each of the plurality of samples is the source signal and is a plurality of pixels from a source image at a first resolution, and the transformed signal is an image at a second resolution.

135. 135. The method of claim 134, wherein the first resolution and the second resolution are different resolutions.

136. pooling the plurality of output activation matrices to generate a second plurality of matrices; applying the second plurality of matrices as input matrices to a second neural network; The method of any one of claims 112 to 135, wherein the number of output activation matrices of the plurality is smaller than the number of matrices of the second plurality.

137. A method according to any one of claims 112 to 136, wherein at least two input activation matrices applied to the neural network are calculated at stride positions of the source signal.

138. A method according to any one of claims 112 to 137, wherein block convolution patterns of a neural network using separable block transforms are applied to the source signals in a translation-invariant manner.

139. A method according to any one of claims 112 to 138, wherein a neural network using a separable block transform is used in the context of transforming a signal in a convolutional manner, instead of multiplying individual scalar activation values ​​by individual scalar weight values.

140. A method according to any one of claims 112 to 139, wherein the neural network includes at least one normalisation layer.

141. A computer program product stored on a non-transitory storage medium, the computer program product for execution by a processing system including at least one hardware processor, the computer program product, when executed, causing the processing system to: constructing an initial activation matrix with a plurality of values ​​based on data from a plurality of samples from the source signal; and applying a separable block transformation based on at least a first trained matrix and a second trained matrix to an input activation matrix across a plurality of layers of the neural network to generate a corresponding output activation matrix, the initial activation matrix being , is used as the input activation matrix for a first layer of the plurality of layers, the input activation matrix for each subsequent layer being the output activation matrix of a previous layer, and the computer program product, when executed, causes the processing system to 22. A computer program product comprising: instructions for outputting the output activation matrix of a final layer of the neural network; and generating a transformed signal based on the output activation matrix of the final layer.

142. 1. A computer system comprising: A processing system including at least one hardware processor, the processing system comprising: placing a plurality of values ​​based on data from a plurality of samples from the source signal into an initial activation matrix; configured to apply a separable block transformation based on at least a first trained matrix and a second trained matrix to an input activation matrix across a plurality of layers of the neural network to generate a corresponding output activation matrix, the initial activation matrix being used as the input activation matrix for a first layer of the plurality of layers, and the input activation matrix for each subsequent layer being the output activation matrix of a previous layer; and further configured for the processing system to: a computer system configured to output the output activation matrix of a final layer of the neural network to generate a transformed signal based on the output activation matrix of the final layer.

143. 143. The computer system of claim 142, wherein the source signal is a source image and the transformed signal is a transformed image.

144. 144. The computer system of claim 143, wherein the source image is in a first resolution and the transformed image is in a second resolution different from the first resolution.

145. A computer system according to any one of claims 142 to 144, wherein the source signal is generated based on execution of a computer application program, and the transformed signal is output in association with execution of the computer application program.

146. 146. The computer system of claim 145, wherein the computer application program is a video game application program, and the source signal is generated by a rendering engine of the video game application program.

147. The computer system of any one of claims 142 to 146, wherein at least one of the hardware processors is a graphics processing unit (GPU), and each input activation matrix and output activation matrix is ​​held within an internal memory of the GPU.

148. The computer system of any one of claims 142 to 147, wherein the at least one input activation matrix and the at least one output activation matrix are held within a single semiconductor hardware of the processing system.

Citation Information

Patent Citations

  • Iterative multi-scale image generation using neural networks

    JP2020508504A

  • Image processing method and image reception device

    WO2018193333A1

Cited By

  • Deep learning model generation apparatus, deep learning model generation method, and program

    JP2025185108A