Systems and methods for lossy image and video compression and / or transmission utilizing a meta-network or neural network

By using the hyperparameters generated by the metanet to reconstruct images and videos, the problem of difficulty in effectively compressing data in the prior art is solved, and efficient and low-bandwidth image and video transmission is achieved.

CN114127788BActive Publication Date: 2025-05-09INTERDIGITAL VC HOLDINGS INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080031710.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-05-16
Filing Date
2020-04-29
Publication Date
2025-05-09
Estimated Expiration
2040-04-29

AI Technical Summary

Technical Problem

Existing video compression technologies are difficult to keep up with the exponential growth of data transmission volume, and high-efficiency video encoding (HVEC) and its improved versions have reached a diminishing point of return, making it difficult to produce significant further improvements.

Method used

The meta-network generates hyperparameters for the image-encoding network to reconstruct the desired image from the noisy image, thereby achieving lossy image and video compression.

Benefits of technology

The amount of data required for transmission is significantly reduced, the efficiency of image and video compression is improved, and a higher compression rate can be achieved without significantly increasing computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114127788B_ABST
    Figure CN114127788B_ABST
Patent Text Reader

Abstract

A system and method for lossy image and video compression that utilizes a meta-network to generate a set of hyperparameters necessary for an image coding network to reconstruct a desired image from a given noisy image. A system and method for lossy image and video compression and transmission that utilizes a neural network as a function for mapping a known noisy image to a desired or target image, thereby allowing only the hyperparameters of the function to be transmitted rather than a compressed version of the image itself. This allows a high quality approximation of the desired image to be reconstructed by any system receiving the hyperparameters, provided that the receiving system has the same noisy image and a similar neural network. The amount of data required to transmit an image of a given quality is significantly reduced compared to existing image compression techniques. Because video is simply a series of images, the application of such an image compression system and method allows video content to be transmitted at a rate greater than that of existing techniques for the same image quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data compression, and more particularly to the field of lossy image compression and / or transmission utilizing neural networks (e.g., non-generic neural networks). Background Art

[0002] Since the Internet became widely used more than two decades ago, the amount of data transmitted around the world has grown exponentially. Over the past decade, even as the amount of data transmitted continues to grow exponentially, video content has continued to account for a growing portion of that data. According to at least one study, video content accounted for 64% of global Internet traffic in 2014 and is expected to account for as much as 85% of global Internet traffic in 2019.

[0003] Given that the amount of data being transmitted continues to grow exponentially and that a large portion of global internet traffic is now video content, video compression technology has become critical. However, existing video compression technologies have not kept pace with the changes. The current dominant video compression technology, High Efficiency Video Coding (HVEC, also known as H.265), and its open source competitor AVI are incremental improvements on video compression technologies developed at the beginning of this century (i.e., H.264 / MPEG-4 AVC and Google's open source VP coding format). Although newer video compression standards have increased throughput over their older versions at the same quality level, this video compression approach may have reached a point of diminishing returns and is unlikely to produce significant further improvements.

[0004] A method for utilizing a non-generic neural network provides a significantly reduced amount of data to be transmitted for image compression or video frame compression, however, such methods described herein take a significant amount of time to optimize in order to find the best network configuration to compress a single image. Therefore, a meta-network that can operate on large data sets and train to further improve efficiency on a powerful platform such as a connected computing device in a data center can be trained to quickly find the structure of a neural network that can reconstruct an approximation of an image, thereby allowing many companies or institutions that need to transmit large amounts of video or images over a network to complete this task while utilizing a fraction of the bandwidth currently required even using current compression methods.

[0005] What is needed is a system and method for image and video compression that significantly exceeds the performance of current state-of-the-art image and video compression, such as without consuming a significant amount of time to compress an image or video. Summary of the invention

[0006] Thus, the inventors have conceived and reduced to practice a system and method for lossy image and video compression that utilizes a meta-network to generate a set of hyperparameters necessary for an image coding network to reconstruct a desired image from a given noisy image.

[0007] According to a preferred embodiment, a system for lossy image and video compression using a meta-network comprises: a meta-network engine, the meta-network engine comprising a processor, a memory and a first plurality of programming instructions stored in the memory, wherein the first plurality of programming instructions, when executed on the processor, causes the processor to: receive a desired image; receive a noisy image; receive a set of training images; use the set of training images to train a plurality of neural networks to reconstruct each of the set of training images by mapping the noisy image to each of the set of training images; store the parameters for each of the plurality of neural networks as a set of meta-network hyperparameters; use the set of meta-network hyperparameters as operating parameters for each of the plurality of neural networks; use the plurality of neural networks to map the noisy image to the desired image, thereby generating a second set of hyperparameters corresponding to a specific filter generated from the operation of each of the plurality of neural networks, such that the second set of hyperparameters, when applied to the noisy image using the neural network, produces an approximation of the desired image within an error less than a predetermined threshold; and store the second set of hyperparameters for use in future image mapping operations.

[0008] According to another preferred embodiment, a method for lossy image compression using a meta-network comprises the following steps: receiving a desired image; receiving a noisy image; receiving a set of training images; using the set of training images to train multiple neural networks to reconstruct each of the set of training images by mapping the noisy image to each of the set of training images; storing the parameters for each of the multiple neural networks as a set of meta-network hyperparameters; using the set of meta-network hyperparameters as operating parameters for each of the multiple neural networks; using the multiple neural networks to map the noisy image to the desired image, thereby generating a second set of hyperparameters corresponding to a specific filter generated from the operation of each of the multiple neural networks, so that the second set of hyperparameters, when applied to the noisy image using the neural network, will produce an approximation of the desired image within an error less than a predetermined threshold; and storing the second set of hyperparameters for use in future image mapping operations.

[0009] And therefore, the inventors have conceived and put into practice a system and method for lossy image and video compression that utilizes a neural network as a function that maps a known noisy image to a desired image, thereby allowing only the hyperparameters of the function to be transmitted rather than a compressed version of the image itself. This allows a high-quality approximation of the desired image to be reconstructed by any system that receives the hyperparameters, provided that the receiving system has the same noisy image and the same or similar neural network. The amount of data required to transmit an image of a given quality is significantly reduced compared to existing image compression techniques. Because video is just a series of images, applying such an image compression system and method allows image content to be transmitted at a rate greater than the prior art for the same image quality, and is expected to soon exceed the state of the art in video compression as well. For clarity, the following non-limiting overview of the invention is provided and should be interpreted consistently with the embodiments described in the following detailed description.

[0010] According to a preferred embodiment, a system for performing lossy image and video compression and lossy image and video transmission using a neural network is disclosed, the system comprising: an image compression engine, the image compression engine comprising a first processor, a first memory and a first plurality of programming instructions stored in the first memory, wherein the first plurality of programming instructions, when running on the first processor, causes the first processor to: receive a desired image; obtain a noisy image; use a first neural network to map the known noisy image to the desired image to find hyperparameters, so that when the hyperparameters are applied to the noisy image using the first neural network, an approximation of the desired image is produced within an error less than a predetermined threshold; and transmit the hyperparameters; and an image decompression engine, the image decompression engine comprising a second processor, a second memory and a second plurality of programming instructions stored in the memory, wherein the second plurality of programming instructions, when running on the second processor, causes the second processor to: receive the hyperparameters; obtain the noisy image; and use a second neural network to apply the hyperparameters to the noisy image to produce an approximation of the desired image within an error less than the predetermined threshold.

[0011] According to another preferred embodiment, a method for lossy image compression and lossy image transmission using a neural network is disclosed, the method comprising the following steps: receiving a desired image at a first computing device; acquiring a noisy image using the first computing device; mapping the noisy image to the desired image using a first neural network using the first computing device to find hyperparameters such that when the hyperparameters are applied to the noisy image using the first neural network, they produce an approximation of the desired image within an error less than a predetermined threshold; and transmitting the hyperparameters to a second computing device; and receiving the hyperparameters at the second computing device; acquiring the noisy image at the second computing device; and applying the hyperparameters to the noisy image using a second neural network using the second computing device to produce an approximation of the desired image within an error less than the predetermined threshold.

[0012] According to an aspect of the embodiment, the image compression engine also includes a dedicated 2D convolution processor to accelerate the operation of the first neural network.

[0013] According to an aspect of the embodiment, the image decompression engine also includes a dedicated 2D convolution processor to accelerate the operation of the second neural network. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The accompanying drawings illustrate several aspects and together with the description serve to explain the principles of the invention according to these aspects. It will be appreciated by those skilled in the art that the particular arrangements shown in the accompanying drawings are exemplary only and should not be construed as limiting the scope of the invention or the claims herein in any way.

[0015] FIG. 1 (Prior Art) is a method diagram illustrating or summarizing HVEC video compression and the similar BPG image compression method.

[0016] Figure 2 is a diagram showing the flow of objects from the compression phase to the decompression phase according to an embodiment, showing the functions used for the disclosed invention. Alternatively, Figure 2 An exemplary algorithm for implementing one aspect of the preferred embodiment may be presented.

[0017] Figure 3 is a system diagram of high-level components used in the operation of a system for lossy image compression utilizing a non-generic neural network in accordance with a preferred embodiment. Alternatively, Figure 3 An exemplary overall system diagram according to a preferred embodiment may be shown.

[0018] Figure 4 is a (eg, exemplary) system diagram of an image compression engine according to one aspect.

[0019] Figure 5is a system diagram of an image compression engine (eg, when used for image compression) utilizing a 2D convolution application specific integrated circuit ("ASIC") according to a preferred aspect or according to one embodiment (eg, another exemplary embodiment thereof).

[0020] Figure 6 is a flow chart illustrating the processing and compression of data through the system according to one aspect.

[0021] Figure 7 is a flow chart illustrating the processing and decompression of data through a system according to one aspect.

[0022] Figure 8 is a method diagram of high-level components used in the operation of a system for lossy image compression utilizing a non-generic neural network in accordance with a preferred embodiment.

[0023] Fig. 9 is a state diagram of two users, one encoding an image and one decoding an image, using the disclosed compression system according to a preferred aspect.

[0024] Fig.10 is a diagram illustrating a method for training a neural network with a single image (as opposed to learning from multiple images) in accordance with a preferred embodiment, thereby resulting in a non-generic method for image compression.

[0025] Fig.11 is a diagram showing a line drawn image and a static image (or noisy image) entering an image compression engine, identifying parameters that allow the static image (or noisy image) to be approximately converted to an initial non-static image (or initial non-noise image), and relayed through a network to a second system to regenerate an approximation of the original image from the static image (or noisy image).

[0026] Fig.12 is a system diagram of an image compression engine for decompressing an image according to a preferred aspect. Alternatively, Fig.12 A system diagram of an image decompression engine for decompressing an image according to one embodiment may be shown.

[0027] Fig.13 is a system diagram of an image compression engine utilizing a dedicated 2D convolution application specific integrated circuit ("ASIC"). Alternatively, Fig.13 A system diagram of an image decompression engine utilizing a dedicated 2D convolution application specific integrated circuit ("ASIC") may be shown.

[0028] Fig.14 is a system diagram of high-level components used in the operation of a system using a meta-network to implement image or video compression and decompression in accordance with a preferred embodiment.

[0029] Fig.15is a system diagram of a meta-network engine according to a preferred embodiment, which is used to train a meta-network with a set of training images, and to compress and find filters that can be used to convert a noisy image to an approximation of an input image.

[0030] Fig.16 is a data flow diagram of a system for lossy compression that utilizes a meta-network to train, compress, and send data for decompression to a system configured using a specific neural network, according to one embodiment.

[0031] Fig.17 is a system diagram of multiple individual networks communicating with each other within a meta-network to produce a sequence of convolutional filters to be applied to a noisy image to gradually transform the image into an approximation of the input image, according to one embodiment.

[0032] Fig.18 is a method diagram showing the steps required to losslessly compress images and videos using a meta-network.

[0033] Fig.19 is a flowchart of the steps taken by a single network within a meta-network to be trained on a set of images and produce a network that operates as part of a function g for neural network hyperparameter prediction of an image encoding network f according to one embodiment.

[0034] Fig. 20 is a flow chart of a process for multiple networks communicating within a meta-network for the purpose of cross-training and developing progressive filters to transform static images to help alleviate the vanishing gradient problem, according to one embodiment.

[0035] Fig.21 is a block diagram illustrating an exemplary hardware architecture of a computing device.

[0036] Fig. 22 is a block diagram illustrating an exemplary logical architecture of a client device.

[0037] Fig.23 is a block diagram illustrating an exemplary architectural arrangement of clients, servers, and external services.

[0038] Fig.24 is another block diagram illustrating an exemplary hardware architecture of a computing device. DETAILED DESCRIPTION

[0039] The present inventors have conceived and put into practice a system and method for lossy image and video compression using a meta-network.

[0040] And the inventors have conceived and put into practice a system and method for lossy image and video compression that utilizes a neural network as a function for mapping a desired image to a known noisy image, thereby allowing only the hyperparameters of the function to be transmitted rather than a compressed version of the image itself. This allows a high-quality approximation of the desired image to be reconstructed by any system that receives the hyperparameters, provided that the receiving system has the same noisy image and a similar neural network. The amount of data required to transmit an image of a given quality is significantly reduced compared to existing image compression techniques. Because video is just a series of images, the application of such image compression systems and methods is equally effective for video as for individual images, and can in time allow video content to be transmitted at a rate greater than the prior art for the same image quality.

[0041] One or more different aspects may be described in this application. In addition, for one or more aspects described herein, many alternative arrangements may be described; it should be understood that these are presented for illustrative purposes only and do not limit the aspects contained herein or the claims presented herein in any way. One or more arrangements may be widely applicable to many aspects, as can be easily seen from this disclosure. Generally speaking, these arrangements are described in sufficient detail to enable one or more aspects to be practiced by a person skilled in the art, and it should be understood that other arrangements may be utilized, and structural, logical, software, electrical and other changes may be made without departing from the scope of a particular aspect. The specific features of one or more aspects described herein may be described with reference to one or more specific aspects or figures forming part of this disclosure, and the specific arrangements of one or more aspects are shown by way of illustration. However, it should be understood that such features are not limited to the use in one or more specific aspects or drawings in which they are described with reference thereto. This disclosure is neither a literal description of all arrangements of one or more aspects, nor a list of features of one or more aspects that must exist in all arrangements.

[0042] The section headings in this patent application and the title of this patent application are provided for convenience only and are not to be construed as limiting the disclosure in any way.

[0043] Unless expressly specified otherwise, devices that are in communication with each other need not be in continuous communication with each other. In addition, devices that are in communication with each other may communicate directly or indirectly through one or more communication means or intermediaries (logical or physical).

[0044] The description of aspects with several components that communicate with each other does not imply that all such components are needed. On the contrary, a variety of optional components can be described to illustrate various possible aspects and to more fully illustrate one or more aspects. Similarly, although process steps, method steps, algorithms, etc. can be described in sequential order, unless otherwise specifically stated, such processes, methods, and algorithms can generally be configured to work in alternating order. In other words, any order or order of steps that can be described in this patent application does not indicate a requirement to perform these steps in this order in itself. The steps of the described process can be performed in any order practice. In addition, some steps can be performed simultaneously, although described or implied as not occurring at the same time (for example, because a step is described after another step). In addition, showing this process by depicting the process in the drawings does not mean that the process shown does not include other changes and modifications of this process, does not mean that any of the steps in the process shown or its steps are necessary for one or more aspects, and does not mean that the process shown is preferred. In addition, steps are generally described once in each aspect, but this does not mean that they must occur once, or that they can only occur once each time a process, method, or algorithm is executed or performed. Certain steps may be omitted in certain aspects or events, or certain steps may be performed more than once in a given aspect or event.

[0045] When a single device or article is described herein, it will be apparent that more than one device or article may be used in place of a single device or article. Similarly, where more than one device or article is described herein, it will be apparent that a single device or article may be used in place of more than one device or article.

[0046] The functionality or features of a device may instead be embodied by one or more other devices that are not explicitly described as having such functionality or features. Therefore, other aspects need not include the device itself.

[0047] For clarity, the techniques and mechanisms described or referenced herein will sometimes be described in singular form. However, it should be understood that, unless otherwise stated, a particular aspect may include multiple iterations of a technique or multiple instantiations of a mechanism. The process descriptions or boxes in the accompanying drawings should be understood to represent modules, sections, or code portions that include one or more executable instructions for implementing a specific logical function or step in the process. Alternative implementations are included within the scope of various aspects, wherein, for example, depending on the functions involved, the functions may not be performed in the order shown or discussed, including substantially simultaneously or in reverse order, as will be understood by those of ordinary skill in the art.

[0048] definition

[0049] As used herein, "artificial intelligence" or "AI" refers to a computer system or component that has been programmed in such a way that it can mimic some aspect or aspects of human cognitive functions associated with human intelligence, such as learning, problem solving, and decision making. Examples of current AI technology include understanding human speech, successfully competing in strategic games such as chess and Go, autonomous operation of vehicles, complex simulations, and interpretation of complex data such as images and video.

[0050] As used herein, "function", "image transformation function", "image transformation network", and "image transformation neural network" refer to the use of a neural network as a function for transforming an image or reconstructing an approximation of an image. The transformation is based on a mapping of a noise image (e.g., an input noise image) to a target image, and the adjustment of weights of different variables within the function, which may also be referred to as hyperparameters. Hyperparameters may also be used as input to a function along with a noise image to reconstruct an approximation of an image.

[0051] As used herein, "hyperparameters" refer to parameters of a function that maps a noise image to a target image (or desired image) within a specified error range. When the image is input into the function for mapping at the source location, the hyperparameters will be the output. The hyperparameters can then be transferred to the destination location, where the hyperparameters will be the input to the same (or similar) function with the same noise image, and the output of the function will be an approximation of the desired image at the destination location. (It should be noted that in the examples, the term "hyperparameter" used in function encoding terminology (as used in the examples herein) is equivalent to the term "weight" or "weight parameter" used in machine learning terminology).

[0052] "Target image" as used herein, which may also be interchangeably referred to as "target image" or "desired image" or "image", refers to a digital representation of an image. This digital representation may be in any image file format, of which a large number of image file formats already exist. Image files are generally classified into raster-based formats (i.e., formats in which pixel content or image regions are specified) or vector-based formats (i.e., formats in which shapes and their relationships are specified without regard to specific pixel, region, or image size). For exemplary purposes, a non-exhaustive and non-limiting list of raster-based formats includes Microsoft Windows bitmaps (bmp), CompuServe graphics interchange format (gif), Joint Photographic Experts Group JFIF format (jpg or jpeg), portable graphics network (png), and tagged image file formats (tif or tiff). For exemplary purposes, a non-exhaustive and non-limiting list of vector-based formats includes Adobe Illustrator files (ai), CorelDRAW image files (cdr), scalable vector graphics files (svg), and Microsoft Visio drawing files (vsd).

[0053] As used herein, "image conversion" refers to the act of converting any image, or even a blank or "empty" image file, into any other image. This can be accomplished by a number of different techniques in any combination, including specific pixel changes, the use of vector graphics based on mathematical equations (e.g., using equations to draw a curve will produce perfect resolution regardless of the scaling of the image itself because it is specified by equations, not represented by specific pixels of a given resolution), or examining and altering groups of pixels, such as for border detection and alteration. Changing the location of data pixels in an image but leaving the data unchanged in other respects (e.g., rotating an image) is another example of an image transformation.

[0054] "Machine learning" as used herein is an aspect of artificial intelligence in which a computer system or component can modify its behavior or understanding without being explicitly programmed to do so. Machine learning algorithms develop models of behavior or understanding based on information fed to them as training sets, and can modify these models based on new incoming information. An example of a machine learning algorithm is AlphaGo, the first computer program to defeat a human world champion in the game of Go. AlphaGo was not explicitly programmed to play Go. It was fed thousands of Go games and developed its own game models and game strategies.

[0055] As used herein, "neural network" refers to a computational model, architecture, or system composed of many simple, highly interconnected processing elements that process information through their dynamic state responses to external inputs and are therefore able to "learn" information by recognizing patterns or trends. Neural networks, sometimes also called "artificial neural networks," are based on our understanding of the structure and function of biological neural networks, such as the mammalian brain. Neural networks are a framework for applying machine learning algorithms.

[0056] As used herein, a "static noise image", also referred to as a "static image" or "noise image", refers to an image with random or pseudo-random content. In some embodiments, the random or pseudo-random content will be at the pixel level, or at the level of a group, zone or region of pixels. In some embodiments, the random or pseudo-random content may include a single bit for each pixel, representing the black or white color of the pixel. In some embodiments, the randomness attribute of the image as part of its attributes may be a 2D maximum information entropy or Shannon entropy. Shannon entropy refers to the amount of randomness or possible information that a given statistical model can have, expressed in bits. For example, a coin toss has an information entropy or Shannon entropy of 1, while m coin tosses have an entropy value of m, because a single coin toss has a value of 0 (back) or 1 (front), represented by 1 bit. The higher the entropy value of a given object, the more possible states it may be in, which makes it more difficult to predict its state, but easier to adapt to converting into many possible states and obtaining an estimate of another state, such as converting into an approximation of another image. In some embodiments, the random or pseudo-random content may also include grayscale or color information. In some embodiments, the random or pseudo-random content may include vector graphics or other forms of image representation instead of bitmaps or pixels.

[0057] As used herein, "video" refers to a digital representation of a sequence of images representing movement. This digital representation can be in any video file format, of which a large number of video file formats already exist. Video files typically contain encoded video (visual) data and audio (audio) data in a "container". This application focuses on the video data portion of a video file. For exemplary purposes, a non-exhaustive and non-restrictive list of video formats includes audio video interlacing (avi), moving picture experts group (mpg, mpeg, mp2, mp4, etc.), apple video format (m4v), and windows media video (wmv).

[0058] Conceptual Architecture

[0059] FIG. 1 (PRIOR ART) is a method diagram illustrating or summarizing HEVC video compression and similar BPG image compression methods. First, an image or video file is input into a compression codec or compression engine 110 before the data file is encoded with intra-frame blocks 120, which in the case of HEVC video file compression remain unchanged through the frames of the video, meaning that between frames of the video where groups of pixels do not change or change very little, the groups are not re-encoded in multiple frames of the video 130, thereby reducing the size of the video with only a small loss of motion in the frames. In the case of the BPG format for still images rather than the video for the HEVC video format, clusters of similar data, including color depth, are maintained for groups of pixels in the still image 130, which represents a reduction in the data required for images of equivalent size, resulting in a compression format that uses similar techniques for both video and still images 140.

[0060] Figure 2 201. It is a diagram showing an exemplary algorithm for implementing one aspect of the preferred embodiment. In the algorithm, a desired image 201 and a noise image 203 are input into a function f 202. Function f 202 is a neural network configured to map the noise image 203 to the desired image 201. The output of function f 202 is a set of hyperparameters P204, or "weights" of function f 202, which map the noise image 202 to the desired or target image 201 within a specified error range. The mathematical formula of the algorithm is:

[0061] Given: I, N, ε, find θ such that:

[0062] f: where f(θ|N):=I' and h(I',I)≤ε

[0063] The algorithm can be generally described as follows. Given a I , height H I and depth C I For a color image I of size W (which corresponds to the target or desired image 201), there is an algorithm 220 that can find the value of W by fine-tuning the hyperparameters θ of the known function f 202. N ×H N ×C N The algorithm 220 measures the error between two images by a function h, which can be one of several possible image comparison functions, such as a pixel-wise mean square error function. If for all images I there exists a vector with θ f f 202 makes the above statement true, then we can use {f,θ f,N} approximately represent all images. If function f 202 is fixed and its input is N, then the only thing missing to create I' is θ. If the file size of θ is smaller than the file size of I, then the algorithm has successfully losslessly compressed I because one only needs to send θ with a bitstream of size file size (θ) to form I' at the destination device from the triplet {f, θ, N} by inputting the noise image 206 and parameters θ 204 into f 205 on the destination / receiving side to obtain an approximate target image I' 207.

[0064] However, it is not obvious how to choose a function f that will efficiently find θ. The solution is to use a neural network, whose typical structure is convolution → bias → nonlinearity → repeated as a function f. By using a neural network as function f, the solution to finding its hyperparameters θ given I becomes a machine learning training problem that can be solved using, for example, stochastic gradient descent (SGD). In short, SGD is used to learn the hyperparameters (aka weights) that minimize the error function h, thereby obtaining θ for I.

[0065] After the hyperparameters θ 204 are output from the neural network 202, they can be input into a separate instance of the same or similar neural network 205 along with a separate instance of the same noise image 206 to produce an approximation of the desired image I 207. In this way, the system operates as a compressor and decompressor by generating one of the two objects based on what is received as input (hyperparameters θ 204 or an approximation of the image I 207). Instances of these objects, such as the decompression neural network 205 and associated data, can be run on different computers, for example, a first network endpoint with the desired image I 201, function f 202, and noise image 203 can be run on one computer such as a smartphone, laptop or other computing device, while another function f 205, noise image 206 can exist on another computing device, which may but not necessarily be over a network. For example, the file can be transferred via a portable memory device rather than over a network between two computers. It is important to note that the different instances of the neural network at both ends of the transfer do not have to be identical. Functionally similar neural networks will produce acceptable solutions even if their architecture, programming or other characteristics are different.

[0066] Figure 3is an exemplary overall system diagram according to a preferred embodiment. A network endpoint 301, which may be a laptop or desktop computer, a mobile phone, a workstation, or some other network-enabled computing device, may be connected via a network 306, such as the Internet, and possesses a desired image I 304 and a noise image N 305. On the network endpoint 301, both images may be used as input to an image compression engine 303, which is directly accessible by the network endpoint 301. The desired image I 304 may be one of many image formats, including .JPG / .JPEG, .PNG, .BMP, or other formats, the specific format of which is a non-critical element of the present invention, as new formats may be identified and interpreted in the future where applicable. The static noise image N 305 may similarly be in one of a variety of formats, and may be a single, unchanging image used in all embodiments of the system, or it may be one of several possible images and sent to the system as input from another source, whether it is a network endpoint 307 or some other source containing images. In some embodiments, the static noise image N 305 represents the image with the highest 2D Shannon entropy, which may be used to transform into other images as needed. The image compression engine 303 is an engine present on the network endpoint 301 that may obtain as inputs the desired image I 304 and the static noise image N 305, and execute the image transformation function 202 to generate hyperparameters of the function 202 that, when applied to the noise image N 350, allow reconstruction of a close approximation I' (read as "I-prime" and meaning an approximation of I) 310 of the desired image I 304. Specifically, the destination endpoint 307 may receive the hyperparameters 308 for input into another instance of the image decompression engine 309, also using the static noise image 311 as input, thereby allowing it to produce an approximation 310, I' of the desired image I 304.

[0067] Figure 4 FIG. 3 is an exemplary system diagram of the image compression engine 303 according to the preferred embodiment. Figure 3As shown, a noise image N 305 and a target image I 304, which represent the maximum entropy noise image and the initial image to be compressed, respectively, are input into a processor 430. The processor 430 is a common computing device for processing arithmetic and logical operations and is well known in the art. The processor 430 has a two-way communication with a volatile memory 440 (commonly referred to as a random access memory ("RAM")) of at least one bit for short-term storage of data to be accessed by the processor 430, which changes the data in the RAM and uses the data in the RAM to change its own process. This is generally understood to be one of the key interactions in modern computing. The processor 430 also communicates with a neural network 410, which runs as a function f in the algorithm 220, which runs as a framework for the machine learning algorithm 420 to process incoming data from the noise image 305 and the target image 304. After approximately reconstructing the input image 304, by using Figure 2 In algorithm 220, hyperparameters 302 of neural network 410 are output to network endpoint 301 for further use, including transmission to a destination endpoint, storage on a hard-drive or external storage device, or other common uses of compressed images.

[0068] Figure 5 is a system diagram of another exemplary embodiment of an image compression engine 500 utilizing a 2D convolution application specific integrated circuit ("ASIC") 560, according to one embodiment. In the depicted embodiment, the 2D convolution ASIC 560 is utilized to reduce the amount of processing time required by the system to fully compress an image. Figure 3 As shown, a noise image N 305 and a target image I 304, which represent the maximum entropy noise image and the initial image to be compressed, respectively, are input into a processor 530. The processor 530 is a common computing device for processing arithmetic and logical operations and is well known in the art. The processor 530 has a two-way communication with a volatile memory 540 (commonly referred to as a random access memory ("RAM")) for short-term storage of at least one bit of data to be accessed by the processor 530, and the processor changes the data in the RAM and uses the data in the RAM to change its own process. This is generally understood to be one of the key interactions in modern computing. The 2D convolution ASIC 511 also communicates bidirectionally with both the system processor 530 and the memory 540 to process the specialized processing of the convolutional neural network 510 data much faster than a general central processing unit can achieve. The processor 530 also communicates with the neural network 510, which runs as a function f in the algorithm 220, running as a framework for the machine learning algorithm 520 to process incoming data from the noise image 305 and the target image 304. After approximately reconstructing the input image 304, by using Figure 2In algorithm 220, hyperparameters 302 of neural network 510 are output to network endpoint 301 for further use, including transmission to a destination endpoint, storage on a hard-drive or external storage device, or other common uses of compressed images.

[0069] Figure 6 is a flow chart illustrating the processing and compression of data through the system according to one aspect. A desired image I may be input into the system (601), which is where the system will attempt to convert a noisy image into an image with certain hyperparameters. The noisy image is checked (602), which may produce positive or negative results as to whether the system already has a static noisy image 305. If a static noisy image does not currently exist in the system, the noisy image may be input into the system (603) via a network endpoint 307 that may have a static noisy image, or it may be automatically generated, or from some other source. If a static noisy image already exists, or after it is received or input into the system (603), the system is in a state where both the static noisy image N 305 and the desired image I 304 exist in the system. It is at this point that the image transformation function f (604) may be utilized or executed to be applied to the static noisy image N 305, the result of the transformation is checked for closeness to the desired image I (605), as defined by the error value ε, the hyperparameters of the neural network are output (607), and the system ends execution, ready to receive a new input image and begin execution again 601. If the image transformation function f 410 operating on the noisy image N 305 does not produce a result that is close enough to the desired image I 304, the neural network 410 may transform the variable hyperparameters θ 302 of the neural network f 410 (606) to try to produce a closer result until the output is close enough to the image I 304 (605).

[0070] Figure 7 is a flow chart illustrating the processing and decompression of data through the system according to one aspect. First, hyperparameters are received and input into the image compression engine 701, which are applied to the neural network 410 to determine the image transformation and convolution of the noisy image to the desired image. The system is checked to ensure that a static noisy image N exists (702), and if no noisy image exists, such an image is input into the system 703, which image may be sent from the same source as the variable hyperparameters θ, or may be inserted by a separate system, or manually entered by a user on the system itself. After the noise image N is inserted into the system 703, or if the noise image already exists, for example if the noise image remains static and the same in all embodiments of the disclosed invention, an image transformation function f (704) may be executed, however instead of Figure 6In the example of FIG. 3 , no hyperparameters are input into the system, but rather the hyperparameters are input into the system (701), and before output (705), the static noise image 305 is transformed (704) into an image I', which is an approximation of the desired image I 304. In this way, image compression is performed by specifying how to reconstruct the image rather than changing the image data itself, thereby allowing the specification of the reconstructed image to be transmitted rather than the image itself, ensuring excellent compression of the data.

[0071] Figure 8 is a method diagram of high level components used in the operation of a system for lossy image compression utilizing a neural network in accordance with a preferred embodiment. First, a desired image must exist or be input (801) prior to accessing a static noisy image N 305 (802), which is the image that will be attempted to be reconstructed using the image transformation network 410. The noisy image will be used by the image transformation function 410 to attempt to transform it into an approximation of the input desired image I 304 using variable hyperparameters. After receiving or accessing the noisy image (if it is static and unchanging) (304), the image transformation function f 410 is executed (803), utilizing variable hyperparameters θ 302 and fed into a set of machine learning algorithms 250 running on the neural network 260, resulting in varying the variable hyperparameters θ (804, 805) of the image transformation function f to attempt to transform a closer approximation to the input image I 304. When an image is produced from the transformation of the noisy image N 305, the variable hyperparameters θ 302 (806) are output so that a user can utilize the image transformation function 410, the noisy image 305, and the hyperparameters 302 to reconstruct a close approximation of the original desired image 304. In this way, the image may be heavily compressed and may be encrypted by the user without having access to the image transformation function itself. The variable hyperparameters θ 302 may then be sent to another network endpoint 307 (807), such as another laptop, desktop computer or workstation device or a computer capable of running Figure 5 The destination network endpoint 307 may then use the variable hyperparameter θ 302 to execute the image transformation function f (808) with the variable hyperparameter θ as input as a parameter (809), where the neural network 410 transforms the noisy image 305 into an approximation of the desired image I 304 (810).

[0072] Fig. 9is a state diagram for two users using the disclosed compression system, one encoding an image and one decoding an image. In a first state 910, user A has a desired image I 304, while user B wants to receive image I 304 and does not yet have it. This is similar to a user having an image to be compressed and sending it to another user, perhaps a friend, colleague, or third party service, who needs to compress the file first for one of many possible reasons. In a second state 920, the second state necessarily follows the first state, trained with a neural network 410 that allows for tuning of hyperparameters θ 302 to allow the image transformation function to transform a static noise image N 305 into an approximation of the desired image I 304. The third state 930 starting from the second state allows user A to obtain the parameters θ302 as described above from training, and then proceeds to the fourth state 940, where the hyperparameters θ302 are sent from user A to user B, which allows the fifth state (950) to be reached, whereby user B, who now has the hyperparameters θ302, can obtain I304 by transforming the noisy image N 305 into an approximation of the desired image I 304 using the image transformation function f 410.

[0073] Fig.10 1 is a diagram illustrating a method for training a neural network with a single image (as opposed to learning from multiple images), resulting in a non-generic method for image compression. As previously detailed in the disclosed system, a desired image is input into the system (1010), such an image comprising any two-dimensional graphics file including .BMP, .PNG or one of many possible formats. A neural network 410 is trained (1020) using a static noisy image N 305 without changing or swapping out the image 305, resulting in a network trained on a single data point rather than a large data set, unlike a generalized neural network. Rather than feeding a large data set to the network, the edge weights of the network are adjusted (1030) based on how close a given output of the network's image transform is to the desired image I 304. In this manner, training continues (1040) on the network 410 for a single noisy image 305 using the adjusted weights and parameters. Table 1050 illustrates the differences and relationship between a conventional neural network that derives a general solution to the problem and the proposed neural network 410 that utilizes a specialized training method and does not learn a general image transformation technique by iteratively training on a single noisy image, changing the weights and parameters of the network when generating a new transformed image until a close approximation to the desired image is generated.

[0074] Fig.11is a diagram showing an actual image and a noisy image entering an image compression engine, identifying parameters that allow the noisy image to be approximately transformed to the original non-noisy image, and relayed through a network to a second system to regenerate an approximation of the original image from the noisy image. A sample image 1101 and a static noisy image 1102 are input into an image compression engine 303. The hyperparameters 303 used to reconstruct the image 1101 are designed by training a neural network to transform the noisy image 1102 into a close approximation of the desired image 1101, which is possible because the noisy image 1102 is a maximum entropy image that can be transformed into any number of images using the correct transformation steps. The hyperparameters 302 used to reconstruct the image 1101 are sent through a network 306 where they are fed into an image decompression engine 309 along with a copy 1103 of the static noisy image 1102, which outputs a close approximation 1104 of the desired image 1101. It should be noted that this is not an identical copy, but rather a close approximation, thus demonstrating a lossy compression method, and creating a separate image file 1104 that is different from the original image 1101.

[0075] Fig.12 1 is a system diagram of an image compression engine 1200 for decompressing an image according to one embodiment. It should be noted that the image decompression engine 1200 uses the same components as the image compression engine 303, but is configured to receive the noisy image 305 and the hyperparameters 302 (eg, via the processor 1230), and output an approximated image I'310. Fig.12 As shown in the right hand portion of , the noisy image N 305 and the variable hyperparameters 302 of the neural network 1210 are input into the processor 1230. The processor 1230 is a common computing device for processing arithmetic and logical operations and is well known in the art. The processor 1230 has a two-way communication with a volatile memory 1240 (commonly referred to as a random access memory ("RAM")) of at least one bit for short-term storage of data to be accessed by the processor 1230, and the processor changes the data in the RAM and uses the data in the RAM to change its own processes. This is generally understood to be one of the key interactions in modern computing. The processor 1230 also communicates with the neural network 1210, which operates as a function f in the algorithm 220, which operates as a framework for the machine learning algorithm 1220 to process the incoming data from the noisy image 305 and (for the purpose of decompression) the hyperparameters θ302 of the previous neural network compression of the target image 304. After the hyperparameters θ302 are input into the processor 1230, the processor directs the neural network 1210 to incorporate and use these hyperparameters, and the network uses Figure 2The algorithm 220 in guides its operation in transforming the noisy image 305 into an approximation 310 of the input image 304, wherein the approximated image I'310 is output to the network endpoint 307 for further use, including transmission to a destination endpoint, storage on a hard drive or external storage device, or other common uses now for uncompressed images.

[0076] Fig.13 1 is a system diagram of an alternative embodiment of an image decompression engine 1310 arrangement for decompression while utilizing a dedicated 2D convolution application specific integrated circuit ("ASIC") 1311, according to one embodiment. It should be noted that the image decompression engine 1300 uses the same components as the image compression engine 309, but is configured to receive the noisy image 305 and the hyperparameters 302, and output an approximated image I' 310. Figure 3 As shown in the right hand portion of FIG. 1 , the noise image N 305 and the variable hyperparameters 302 of the neural network 1310 are input into the processor 1330. The processor 1330 is a common computing device for processing arithmetic and logical operations and is well known in the art. The processor 1330 has two-way communication with a volatile memory 1340 (commonly referred to as a random access memory ("RAM")) for short-term storage of at least one bit of data to be accessed by the processor 1330, and the processor changes the data in the RAM and uses the data in the RAM to change its own process. This is generally understood to be one of the key interactions in modern computing. The 2D convolution ASIC 1311 also communicates bidirectionally with both the system processor 1330 and the memory 1340 to process the specialized processing of the convolutional neural network 1310 data much faster than a general central processing unit could achieve. Processor 1330 also communicates with neural network 1310, which operates as function f in algorithm 220, operating as a framework for machine learning algorithm 1320 to process incoming data from noisy image 305 and (for decompression purposes) hyperparameters θ302 of the previous neural network compression of target image 304. After hyperparameters θ302 are input into processor 1330, processor directs neural network 1310 to incorporate and use these hyperparameters, and the network uses Figure 2 The algorithm 220 in guides its operation in transforming the noisy image 305 into an approximation 310 of the input image 304, wherein the approximated image I'310 is output to the network endpoint 307 for further use, including transmission to a destination endpoint, storage on a hard drive or external storage device, or other common uses now for uncompressed images.

[0077] Because training a single neural network to efficiently map a known noisy image to a desired image typically requires many hours on a highly optimized graphics processing unit (GPU) based machine to complete for each desired image to be compressed using normal training methods based on (e.g., steepest gradient descent), it is critical to find different ways to efficiently determine the weights f for a given desired image. According to one aspect, the inventors have determined that the weights of f (in the sense of how closely the output image obtained by applying N to f matches the input image used to generate the corresponding weights for f) can be generated with arbitrary precision in a single pass through the meta-network described below. Specifically, the weights of f are obtained in a single pass by passing the desired image I and the known noisy image N as input to a function g using a meta-network, as described below with reference to Figures 14 to 17 More specifically, the inventors have shown that Fig.17 The meta-network of the arrangement shown demonstrates a massive increase in coding efficiency with respect to runtime compression of images. Specifically, training via SGD on a high-end GPU takes more than 6 hours. In contrast, the runtime g is less than a second on similar hardware, and can be less than 100 milliseconds if it is dedicated hardware (such as a hardware-based optimized convolution stage).

[0078] Fig.14 is a system diagram of high-level components used in the operation of a system that uses a meta-network to implement image or video compression and decompression in accordance with a preferred embodiment. Fig.14 As shown on the left side of FIG, a given target image I 304 is provided to a meta-network engine 1410 along with a static (i.e., known and unchanging) noise image N 305. The meta-network engine 1410 comprises a "network of neural networks," including a plurality of single neural networks, each of which performs processing on a portion of the target image I 304, such as a particular feature or color spectrum (e.g., in a three-network meta-network, each network may focus on one of the red, blue, and green color bands within the image). The meta-network engine 1410 uses convolutional training between the individual networks within it to obtain a set of hyperparameters θ 302 suitable for any given image decompression engine 309 to reconstruct the target image 304 (or any similar reconstruction thereof) from the noise image 305. This set of hyperparameters 302 may then be stored for future reuse, e.g., to enable many clients to reconstruct the target image 304 from a single set of calculated hyperparameters (e.g., as may be useful for photo or video streaming services, where the same content is repeatedly sent to a large number of destinations), and may be sent from one endpoint 301 to another endpoint 307 via a network 306. At the network destination, the received set of hyperparameters 308 is then provided as input to an image decompression engine 309 along with the noisy image I 311 so that the image decompression engine 309 can produce a reasonably close reconstruction of the target image, the reconstructed image being denoted I' 310.

[0079] Fig.15 1 is a system overview diagram of a meta-network engine 1410 according to a preferred embodiment, which is used to train a meta-network with a set of training images, and to compress and find filters that can be used to convert a noisy image to an approximation of an input image. According to the embodiment, a given target image I 304 is provided as input to the meta-network engine 1410 along with a static noise image N 305. The meta-network engine 1410 includes a processor 1540 and a memory 1559 for performing machine learning tasks associated with the training and transformation process. The neural network 1510 is trained on a set of training images 1530 using a machine learning algorithm 1520 to obtain a suitable set of hyperparameters for operation, given the inputs of N 305 and I 304, the hyperparameters include a set of hyperparameters required for the meta-network engine 1410 to create a set of variable hyperparameters 302 to store and transmit to the target image encoding network. In other words, the set of hyperparameters determined by training is the set used for the meta-network operation, and the operation (when given the correct, trained set of hyperparameters) then produces the required set of hyperparameters 302 from the noise image 305 and the target image 304.

[0080] Fig.16 16 is a data flow diagram of a system 1610 for lossy compression according to one embodiment, the system utilizing a meta-network 1602 to train, compress, and send data for decompression to a system utilizing a specific neural network configuration 1640. According to the embodiment, the meta-network 1602 instantiates a function g, where g, when given a known noisy image 1603 and an original image 1601 as inputs, determines a specific set of hyperparameters θ 1604 for a given original image I 1601 and a noisy image N 1603 as inputs. The hyperparameters 1604 have the following properties: when inserted as weights into a neural network (e.g., at a destination device) to define a particular instance of a decompression function f 1640, the hyperparameters can be used to map the same known / static noise image 1630 (here, "same" means that the noise image 1630 is the same image as the noise image 1603 used as input to g in the meta-network 1602 to produce the hyperparameters θ 1604. According to a preferred aspect, the hyperparameters θ 1604 can be transmitted to the destination image encoding network 1640, which includes a function f such that f produces a replica of the original image I, called I' 1650, given the noise image N 1630 and the hyperparameters θ 1604 as input. Mathematically, the operation takes the form of a set of functions 1620:

[0081] Given: {I1,...,IK}, N, ε, f

[0082] Found: g

[0083] So that for j = {1, ..., K}: g(Ij, N): = θ'j

[0084] f(g(Ij,N)|N):=I'j

[0085]

[0086] Fig.17 is a system diagram of multiple individual neural networks 1710a communicating with each other within a meta-network 1710 to produce multiple sets of filter weights 1721, 1723, 1725, which can be inserted into neural networks 1722, 1724 and 1726, which together embody a specific instance of a function f that can be used to map a noisy image 1701 to an image 1730 that is (within arbitrary error limits) approximately identical to the input image 1711. As shown, the meta-network 1710 is a network of neural networks, including multiple individual networks, each performing an independent machine learning task focused on a portion of the original image 1711. As shown, one arrangement uses three individual networks (and the inventors have determined through practice that this configuration provides excellent performance for the relative amount of resources utilized, reducing the image compression task from hours to milliseconds), however alternative meta-network arrangements are possible according to embodiments, such as using only two individual networks or using four or more individual networks.

[0087] When the original image 1711 is provided as input to a separate network 1710a (although only the particular separate network is highlighted in the figure, this is done for clarity, and it should be understood that the discussion of the operation of 1710a applies to any and all separate networks 1710 present in the meta-network 1710), a series of convolutional filters 1714a-n are applied, where the output of each filter can be provided as an input to the next filter to implement machine learning. The convolution output of each separate network can also be provided as an input to the next sequential filter in the processing 1712a-n, 1713a-n of another separate network, so that the networks effectively learn from each other's processing and jointly "zero in" on an ideal solution. After the convolution process, multiple non-convolutional final filters 1715a-n, 1716a-n, 1717a-n can be applied to the output of each separate network (and in this case, only the output of the network) using the static noise image 1701 as an additional input.

[0088] This produces a set of filters 1721, 1723, 1725 that are developed based on a combination of convolutions of the original image, common processing, and network-specific processing based on the noisy image, such that each filter is different (based only on the processing of a single network, and therefore differs from each other due to differences in the networks and their corresponding outputs and learning processes), but the combination of filters, when applied sequentially at the target image encoding network 1720, successively reconstructs more accurate representations 1722, 1724, 1726 before achieving a final, arbitrarily similar reconstruction 1730 of the original image 1711.

[0089] exist Fig.17 As will be seen in the diagram, data is passed between layers within a given convolutional neural network (e.g., 1712a-n) as per convention. However, in addition, according to one aspect, meta-network connections are established to link the output of one neural network (one horizontal row) of a meta-network (e.g., 1712a-n) to the input of the next stage of a different neural network (a different horizontal row, such as 1713a-n), as shown in FIG. Fig.17 During some initial testing of multiple meta-networks without these inter-meta-network connections, we encountered training and quality challenges. The training challenge is the vanishing gradient problem. This is when the updates that the learning algorithm wants to make to the weights become very small (10 -8 These small updates cause the network to become "stale" and stop learning; in other words, the functionality stops improving. This problem can be addressed by adding skip connections between layers within the meta-network, as was done in the meta-network connections, to shorten the path that the update must take to acquire its parameters, thereby reducing the magnitude of the update reduction. The described techniques were fully utilized in an exemplary manner, where substantial improvements were noted.

[0090] To further improve the information flow, inter-meta-network paths are added. These act as information injection between meta-networks. Their main role is to update the current meta-network based on the state of the previous meta-network. This further helps the flow of gradient information during training, resulting in a more stable training process while taking less time.

[0091] Another reason to add inter-meta-network paths is that the current meta-network can know the filters (and therefore the transforms) produced by the previous network, and it can produce its filters in a way that will complement the transforms of the previous filters rather than cancel them. This ability to know the state of the previous meta-network is important because in image coding networks according to aspects of the present invention, transforms to the input (noise) are typically applied sequentially. Therefore, if the task of a meta-network is to predict the transform of a particular stage, it should know the previous transforms that have been applied, in other words, the state of the previous meta-network row.

[0092] The task of each meta-network is to predict the filter that will best transform the image encoding network input into a progressively closer approximation to the desired image. The input to the meta-network is simply the desired image (the known noise image is injected late at layers 1715a-n in the first row, as Fig.17 ), so the meta-network sees an incomplete picture. It does not see the "coming" image, but only the "going" image. Therefore, the meta-network is provided with a late fusion of the IEN input (noise). This increases the information that the meta-network can use to make predictions and improves the image quality. Furthermore, this information fusion is done in an "or" manner, so the meta-network can choose whether it wants this information fusion or not.

[0093] In summary, according to one aspect of the method, such as Fig.17 The method shown in can be considered as an ensemble of meta-networks, since the method uses a collection or group of meta-networks. One could argue that this is the same network; however, this is wrong, since all the individual meta-networks are independent of each other in terms of their trainable parameters, inputs, and outputs. It is not even necessary to train them together, but it is convenient.

[0094] The meta-network 1710 must be trained before it can perform its function of determining the weights 1722, 1723, 1725 required for the function f to map the noisy image N 1701 to the target image 1730. Therefore, Fig.18 is a method diagram showing the steps required to train a meta-network 1710 for lossy compression of images and videos using a meta-network. In an initial step 1810, a set of training images j are provided as input to the meta-network. These are used to train the meta-network (1820) based on images j and a given noisy image N. This training, as described above Fig.15As described in , a set of filters is implemented that need to be applied to a noisy image N to achieve a reconstruction of a given original image I (1830). In more detail, this reconstruction process involves multiple independent networks within the meta-network, each of which generates a subset of hyperparameters θ (1840), and each subset is different from the other subsets and focuses on a specific part or attribute of the input image I (1850). Each subsequent individual network takes as input the state of all previous networks 1860, so that their joint machine learning is used as a convolutional neural network built from smaller, more specialized convolutional neural networks (each focusing on a specific attribute of the image, while the entire meta-network focuses on the whole by using each specific network as a convolutional filter in its own processing). In order to reconstruct the target image, the static image N is passed through the continuous filters developed by the network 1870, resulting in a close approximation image I' (1880). When the approximate image is determined to be acceptable, the set of hyperparameters θ is stored and training is considered complete; the hyperparameters θ can then be sent (1890) to any arbitrary number of destination image encoding networks, each of which can then use the hyperparameters θ and the noise image N to quickly produce a reconstructed image I'.

[0095] Fig.19 is a flow chart of the steps taken by a single network within a meta-network to train on a set of images and produce a network that operates as part of a function g for neural network hyperparameter prediction of an image encoding network f, according to one embodiment. In an initial step 1901, a set of training images j are input to the meta-network. At each individual network within the meta-network, a check is performed (1902) to determine if a noisy image is present. If no noisy image is present, one is provided (1903), and then (or if a noisy image is present) the meta-network performs a check on each training image j. n The image transformation function g is performed (1904). After performing the transformation, the results are checked to determine if the output image is acceptably close to the desired image I (1905), and if not, each individual network within the meta-network slightly changes its parameters (1906) and the function g is repeated to iterate closer to the desired output result. When the reconstructed image I' is acceptably close to the original I, the set of hyperparameters is stored as an acceptable set of hyperparameters required for an acceptable image transformation (1907).

[0096] Fig. 20Flowchart of a process for multiple networks communicating within a meta-network for the purpose of cross-training and developing progressive filters to transform static images to help alleviate the gradient vanishing problem according to one embodiment. In an initial step 2010, convolution processing layer i is performed on all individual networks within the meta-network. Then at 2020, a check is performed to determine whether there are additional layers waiting to be executed. If so, each individual network sends the result of its corresponding processing layer i as an output to be used as input for the next convolution layer i+1 of all next networks (2040). Then, i is incremented (2050) and operation continues (2010). When there are no more layers, the convolution process ends, and each individual network then generates its corresponding filter 2030, thereby generating a complete set of filters required to reconstruct the target image from the noisy image.

[0097] Hardware Architecture

[0098] Typically, the techniques disclosed herein can be implemented in hardware or a combination of software and hardware. For example, they can be implemented in an operating system kernel, in a separate user process, in a library package bound to a network application, on a specially constructed machine, on an application specific integrated circuit ("ASIC"), or on a network interface card.

[0099] The software / hardware hybrid implementation of at least some aspects disclosed herein can be implemented on a programmable network-resident machine (which should be understood to include an intermittently connected network-aware machine), which is selectively activated or reconfigured by a computer program stored in a memory. Such a network device can have multiple network interfaces, which can be configured or designed to utilize different types of network communication protocols. The general architecture of some of these machines can be described herein to illustrate one or more exemplary devices that can implement a given functional unit. According to specific aspects, at least some of the features or functions of the various aspects disclosed herein can be implemented on one or more general-purpose computers associated with one or more networks, such as end-user computer systems, client computers, network servers or other server systems, mobile computing devices (e.g., tablet computing devices, mobile phones, smart phones, laptop computers or other suitable computing devices), consumer electronic devices, music players or any other suitable electronic devices, routers, switches or other suitable devices or any combination thereof. In at least some aspects, at least some of the features or functions of the various aspects disclosed herein can be implemented in one or more virtualized computing environments (e.g., network computing clouds, virtual machines hosted on one or more physical computers, or other suitable virtual environments).

[0100] Reference now Fig.21, a block diagram depicting an exemplary computing device 10 suitable for implementing at least a portion of the features or functions disclosed herein is shown. The computing device 10 may be, for example, any of the computing machines listed in the previous paragraph, or indeed any other electronic device capable of executing software- or hardware-based instructions according to one or more programs stored in a memory. The computing device 10 may be configured to communicate with multiple other computing devices such as clients or servers over a communication network such as a wide area network, a metropolitan area network, a local area network, a wireless network, the Internet, or any other communication network, using known protocols for such communications, whether wireless or wired.

[0101] In one embodiment, computing device 10 includes one or more central processing units (CPUs) 12, one or more interfaces 15, and one or more buses 14 (such as a peripheral component interconnect (PCI) bus). When acting under the control of appropriate software or firmware, CPU 12 can be responsible for implementing specific functions associated with the functions of a particular configured computing device or machine. For example, in at least one embodiment, computing device 10 can be configured or designed to be used as a server system utilizing CPU 12, local memory 11 and / or remote memory 16 and interface 15. In at least one embodiment, CPU 12 can be made to perform one or more different types of functions and / or operations under the control of software modules or components, which can include, for example, an operating system and any appropriate application software, drivers, etc.

[0102] The CPU 12 may include one or more processors 13, such as, for example, a processor from one of the Intel, ARM, Qualcomm, and AMD families of microprocessors. In some embodiments, the processor 13 may include specially designed hardware for controlling the operation of the computing device 10, such as an application specific integrated circuit (ASIC), an electrically erasable programmable read-only memory (EEPROM), a field programmable gate array (FPGA), and the like. In a particular embodiment, a local memory 11 (such as a non-volatile random access memory (RAM) and / or a read-only memory (ROM), including, for example, one or more levels of cache memory) may also form part of the CPU 12. However, there are many different ways to connect memory to the system 10. The memory 11 may be used for a variety of purposes, such as caching and / or storing data, programming instructions, and the like. It should also be understood that the CPU 12 may be one of a variety of system-on-chip (SOC) type hardware, which may include additional hardware such as a memory or graphics processing chip, such as the QUALCOMM SNAPDRAGON SoCs that are becoming increasingly common in the art for use in mobile devices or integrated devices. TM or SAMSUNG EXYNOSTM CPU.

[0103] As used herein, the term "processor" is not limited to those integrated circuits known in the art as processors, mobile processors or microprocessors, but refers broadly to microcontrollers, microcomputers, programmable logic controllers, application specific integrated circuits, and any other programmable circuits.

[0104] In one embodiment, the interface 15 is provided as a network interface card (NIC). Typically, a NIC controls the sending and receiving of data packets on a computer network; other types of interfaces 15 may, for example, support other peripheral devices used with the computing device 10. Among the interfaces that may be provided are Ethernet interfaces, frame relay interfaces, cable interfaces, DSL interfaces, token ring interfaces, graphics interfaces, and the like. In addition, various types of interfaces may be provided, such as universal serial bus (USB), serial, Ethernet, FIREWIRE TM , THUNDERBOLT TM , PCI, Parallel, Radio Frequency (RF), Bluetooth TM , near field communication (e.g., using near field magnetic), 802.11 (WiFi), frame relay, TCP / IP, ISDN, Fast Ethernet interface, Gigabit Ethernet interface, Serial ATA (SATA) or external SATA (ESATA) interface, High Definition Multimedia Interface (HDMI), Digital Video Interface (DVI), analog or digital audio interface, Asynchronous Transfer Mode (ATM) interface, High Speed ​​Serial Interface (HSSI) interface, Point of Sale (POS) interface, Fiber Data Distributed Interface (FDDI), etc. Typically, such interface 15 may include a physical port suitable for communicating with the appropriate media. In some cases, they may also include an independent processor (such as a dedicated audio or video processor for high-fidelity A / V hardware interface as is common in the art), and in some cases, include volatile and / or non-volatile memory (e.g., RAM).

[0105] although Fig.21The system shown in illustrates a particular architecture of a computing device 10 for implementing one or more of the inventions described herein, but is by no means the only device architecture on which at least a portion of the features and techniques described herein may be implemented. For example, an architecture having one or any number of processors 13 may be used, and such processors 13 may be present in a single device or distributed among any number of devices. In one embodiment, a single processor 13 handles communications as well as routing calculations, while in other embodiments, a separate dedicated communications processor may be provided. In various embodiments, different types of features or functions may be implemented in a system according to the invention, the system including a client device (such as a tablet device or smartphone running client software) and a server system (such as the server system described in more detail below).

[0106] Regardless of the network device configuration, the system of the present invention may employ one or more memories or memory modules (e.g., such as remote memory block 16 and local memory 11) configured to store data, program instructions for general network operations, or other information related to the functionality of the embodiments described herein (or any combination thereof). For example, program instructions may control the execution of an operating system and / or one or more application programs or include an operating system and / or one or more application programs. Memory 16 or memories 11, 16 may also be configured to store data structures, configuration data, encrypted data, historical system operation information, or any other specific or general non-program information described herein.

[0107] Because such information and program instructions may be used to implement one or more systems or methods described herein, at least some network device embodiments may include non-transitory machine-readable storage media that may, for example, be configured or designed to store program instructions, state information, etc., for performing the various operations described herein. Examples of such non-transitory machine-readable storage media include, but are not limited to, magnetic media such as hard disks, floppy disks, and magnetic tapes; optical media such as CD-ROM disks; magneto-optical media such as optical disks, and hardware devices specifically configured to store and execute program instructions, such as read-only memory devices (ROMs), flash memory (common in mobile devices and integrated systems), solid-state drives (SSDs), and "hybrid SSD" storage drives that may combine the physical components of solid-state and hard disk drives in a single hardware device (increasingly common in the field of personal computers), memristor memory, random access memory (RAM), and the like. It should be understood that such storage devices may be integrated and non-removable (such as RAM hardware modules that may be soldered to a motherboard or otherwise integrated into an electronic device), or they may be removable, such as swappable flash memory modules (such as a "thumb drive" or other removable media designed for rapid exchange of physical storage devices), "hot-swappable" hard drives or solid-state drives, removable optical storage disks, or other such removable media, and that such integrated and removable storage media may be utilized interchangeably. Examples of program instructions include object code (such as may be produced by a compiler), machine code (such as may be produced by an assembler or linker), byte code (such as may be produced by, for example, a JAVA TM A file that contains code that can be executed by a computer using an interpreter (e.g., a script written in Python, Perl, Ruby, Groovy, or any other scripting language).

[0108] In some embodiments, the system according to the present invention can be implemented on a stand-alone computing system. Fig. 22 , shows a block diagram depicting a typical exemplary architecture of one or more embodiments or components thereof on a stand-alone computing system. The computing device 20 includes a processor 21 that can run software that performs one or more functions or applications of embodiments of the present invention, such as, for example, a client application 24. The processor 21 can execute computing instructions under the control of an operating system 22, such as, for example, a MICROSOFT WINDOWS TM Operating system, APPLE OSX TM or iOS TM Operating system versions, some types of Linux operating systems, ANDROID TMoperating system, etc. In many cases, one or more shared services 23 may be operable in the system 20 and may be used to provide common services to the client applications 24. The service 23 may be, for example, a WINDOWS TM Services, user space public services in a Linux environment, or any other type of public service architecture used with the operating system 21. The input device 28 can be of any type suitable for receiving user input, including, for example, a keyboard, a touch screen, a microphone (e.g., for voice input), a mouse, a touchpad, a trackball, or any combination thereof. The output device 27 can be of any type suitable for providing output to one or more users, whether remotely or locally to the system 20, and can include, for example, one or more screens, speakers, printers, or any combination thereof for visual output. The memory 25 can be a random access memory having any structure and architecture known in the art for use by the processor 21, for example, for running software. The storage device 26 can be any magnetic, optical, mechanical, memristor, or electrical storage device for storing data in digital form (such as those described above, with reference to Fig.21 ). Examples of storage device 26 include flash memory, magnetic hard drive, CD-ROM, etc.

[0109] In some embodiments, the system of the present invention may be implemented in a distributed computing network, such as one having any number of clients and / or servers. Fig.23 , a block diagram depicting an exemplary architecture 30 for implementing at least a portion of a system according to an embodiment of the present invention on a distributed computing network is shown. According to an embodiment, any number of clients 33 may be provided. Each client 33 may run software for implementing the client-side portion of the present invention; the client may include software such as Fig. 22 33. In addition, any number of servers 32 may be provided to handle requests received from one or more clients 33. Clients 33 and servers 32 may communicate with each other via one or more electronic networks 31, which in various embodiments may be any of the Internet, a wide area network, a mobile phone network (such as a CDMA or GSM cellular network), a wireless network (such as WiFi, WiMAX, LTE, etc.), or a local area network (or indeed any network topology known in the art; the present invention does not favor any one network topology over any other). Network 31 may be implemented using any known network protocol, including, for example, wired and / or wireless protocols.

[0110] In addition, in some embodiments, the server 32 may call the external service 37 to obtain additional information when needed, or refer to additional data about a particular call. For example, communication with the external service 37 may occur via one or more networks 31. In various embodiments, the external service 37 may include web-enabled services or functions associated with or installed on the hardware device itself. For example, in an embodiment where the client application 24 is implemented on a smartphone or other electronic device, the client application 24 can obtain information on an external service 37 stored in a server system 32 in the cloud or deployed on the premises of one or more specific enterprises or users.

[0111] In some embodiments of the present invention, the client 33 or the server 32 (or both) may utilize one or more specialized services or devices that may be deployed locally or remotely across one or more networks 31. For example, one or more embodiments of the present invention may use or reference one or more databases 34. It should be understood by those of ordinary skill in the art that the database 34 may be arranged in a variety of architectures and using a variety of data access and manipulation means. For example, in various embodiments, one or more databases 34 may include a relational database system using structured query language (SQL), while others may include alternative data storage technologies, such as those referred to in the art as "NoSQL" (e.g., HADOOP CASSANDRA TM 、GOOGLEBIGTABLE TM etc.). In some embodiments, variant database architectures such as column-oriented databases, in-memory databases, clustered databases, distributed databases, or even flat file data repositories may be used in accordance with the present invention. One of ordinary skill in the art will understand that any combination of known or future database technologies may be used as appropriate, unless a particular database technology or a particular arrangement of components is specified for a particular embodiment herein. In addition, it should be understood that the term "database" as used herein may refer to a physical database machine, a cluster of machines acting as a single database system, or a logical database within an entire database management system. Unless a particular meaning is specified for a given usage of the term "database", it should be interpreted as meaning any of these meanings of the term, all of which are understood by one of ordinary skill in the art as the simple meaning of the term "database".

[0112] Similarly, most embodiments of the present invention may utilize one or more security systems 36 and configuration systems 35. Security and configuration management are common information technology (IT) and web functions, and some functions are commonly associated with any IT or web system. It should be understood by those of ordinary skill in the art that any configuration or security subsystem known in the art now or in the future may be used in conjunction with embodiments of the present invention without limitation, unless the description of any particular embodiment specifically requires a particular security 36 or configuration system 35 or method.

[0113] Fig.24 An exemplary overview of a computer system 40 is shown, which can be used in any of a variety of locations throughout the system. It is an example of any computer that can execute code to process data. Various modifications and changes can be made to the computer system 40 without departing from the broader scope of the systems and methods disclosed herein. A central processing unit (CPU) 41 is connected to a bus 42, to which a memory 43, a non-volatile memory 44, a display 47, an input / output (I / O) unit 48, and a network interface card (NIC) 53 are also connected. The I / O unit 48 can typically be connected to a keyboard 49, a pointing device 50, a hard disk 52, and a real-time clock 51. The NIC 53 is connected to a network 54, which can be the Internet or a local network that may or may not have a connection to the Internet. Also shown as part of the system 40 is a power supply unit 45, which in the example is connected to a main alternating current (AC) power source 46. Not shown are batteries that may be present, as well as many other devices and modifications that are well known but not applicable to the specific novel functions of the current systems and methods disclosed herein. It should be understood that some or all of the components shown may be combined, such as in various integrated applications, such as a Qualcomm or Samsung system-on-a-chip (SOC) device, or whenever it may be appropriate to combine multiple capabilities or functions into a single hardware device (e.g., in a mobile device such as a smartphone, a video game console, an in-vehicle computer system such as a navigation or multimedia system in an automobile, or other integrated hardware device).

[0114] In various embodiments, the functions for implementing the system or method of the present invention can be distributed between any number of clients and / or server components. For example, various software modules can be implemented to perform various functions related to the present invention, and such modules can be implemented differently to run on servers and / or client components.

[0115] A skilled person will recognize a range of possible modifications to the various embodiments described above. Therefore, the present invention is defined by the claims and their equivalents.

Claims

1. A system for lossy image and video compression using a meta-network, comprising: A meta-network engine comprising a processor, a memory, and a first plurality of programming instructions stored in the memory, wherein the first plurality of programming instructions, when executed on the processor, causes the processor to: receiving a desired image; receiving a noisy image; receiving a set of training images; training a plurality of neural networks using the set of training images to reconstruct each of the set of training images by mapping the noisy image to each of the set of training images; storing parameters for each of the plurality of neural networks as a set of meta-network hyperparameters; using the set of meta-network hyperparameters as operating parameters for each of the plurality of neural networks; mapping the noisy image to the desired image using the plurality of neural networks, thereby producing a second set of hyperparameters corresponding to a particular filter resulting from operation of each of the plurality of neural networks, such that the second set of hyperparameters, when applied to the noisy image using the neural networks, produces an approximation of the desired image within an error less than a predetermined threshold; as well as The second set of hyperparameters is stored for use in future image mapping operations.

2. The system of claim 1, wherein each of the plurality of neural networks: generating at least one convolution filter, wherein the noisy image can be filtered through all of the convolution filters in sequence to map the noisy image to an approximation of a desired image; and The communication between the multiple neural networks is facilitated to alleviate the vanishing gradient problem.

3. The system of claim 2, wherein the plurality of neural networks can be located on separate computing devices connected across a network.

4. The system of claim 1, wherein the noise image is static and unchanging.

5. A method for lossy image compression using a meta-network, comprising the following steps: receiving a desired image; receiving a noisy image; receiving a set of training images; training a plurality of neural networks using the set of training images to reconstruct each of the set of training images by mapping the noisy image to each of the set of training images; storing parameters for each of the plurality of neural networks as a set of meta-network hyperparameters; using the set of meta-network hyperparameters as operating parameters for each of the plurality of neural networks; mapping the noisy image to the desired image using the plurality of neural networks, thereby producing a second set of hyperparameters corresponding to a particular filter resulting from operation of each of the plurality of neural networks, such that the second set of hyperparameters, when applied to the noisy image using the neural networks, produces an approximation of the desired image within an error less than a predetermined threshold; as well as The second set of hyperparameters is stored for use in future image mapping operations.

6. The method according to claim 5, further comprising the steps of: generating at least one convolution filter at each of the plurality of neural networks, wherein a noisy image can be filtered through all of the convolution filters in sequence, thereby mapping the noisy image to an approximation of a desired image using a plurality of neural networks; as well as The communication between the multiple neural networks is facilitated to alleviate the vanishing gradient problem.

7. The method of claim 6, wherein the plurality of neural networks can be located on separate computing devices connected across a network. The method of claim 5 , wherein the noise image is static and unchanging.

9. A system for lossy image and video compression and transmission using a neural network, comprising: An image compression engine comprising a first processor, a first memory, and a first plurality of programming instructions stored in the first memory, wherein the first plurality of programming instructions, when executed on the first processor, causes the first processor to: receiving a desired image; Get a noisy image; mapping the noisy image to the desired image using a first neural network to find hyperparameters such that the hyperparameters, when applied to the noisy image using the first neural network, produce an approximation of the desired image within an error less than a predetermined threshold; as well as transmitting the hyperparameters; as well as An image decompression engine comprising a second processor, a second memory, and a second plurality of programming instructions stored in the memory, wherein the second plurality of programming instructions, when executed on the second processor, causes the second processor to: receiving the hyperparameters; Acquire the noise image; as well as The hyperparameters are applied to the noisy image using a second neural network to produce an approximation of the desired image within an error less than the predetermined threshold.

10. The system of claim 9, wherein the image compression engine further comprises a dedicated 2D convolution processor to accelerate operation of the first neural network.

11. The system of claim 9, wherein the image decompression engine further comprises a dedicated 2D convolution processor to accelerate operation of the second neural network.

12. A method for lossy image compression and lossy image transmission using a neural network, comprising the following steps: receiving, at a first computing device, a desired image; Acquire a noise image using the first computing device; mapping the noisy image to the desired image using a first neural network with the first computing device to find hyperparameters such that the hyperparameters, when applied to the noisy image using the first neural network, produce an approximation of the desired image within an error less than a predetermined threshold; as well as transmitting the hyperparameters to a second computing device; as well as receiving, at a second computing device, the hyperparameters; acquiring the noise image at the second computing device; as well as The hyperparameters are applied to the noisy image using a second neural network with the second computing device to produce an approximation of the desired image within an error less than the predetermined threshold.

13. The method of claim 12, wherein the first computing device further comprises a dedicated 2D convolution processor to accelerate operation of the first neural network.

14. The method of claim 12, wherein the second computing device further comprises a dedicated 2D convolution processor to accelerate operation of the second neural network.

Citation Information

Patent Citations

  • Image generation method based on a conditional capsule generative adversarial network

    CN109584337A

  • Systems and Methods for Providing Convolutional Neural Network Based Image Synthesis Using Stable and Controllable Parametric Models, a Multiscale Synthesis Framework and Novel Network Architectures

    US20180068463A1