Methods for encoding, transmitting, and decoding lossy images or videos, and data processing systems.
Trained neural networks with region-specific quantization enhance the efficiency of encoding and decoding lossy images and videos, addressing network strain and energy consumption issues by optimizing compression and reducing data transmission needs.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- INTERDIGITAL VC HOLDINGS INC
- Filing Date
- 2026-04-21
- Publication Date
- 2026-07-29
AI Technical Summary
The increasing demand for high-resolution image and video content is straining communication networks, leading to higher data transmission loads and energy consumption, while existing compression methods, including AI-based approaches, often result in suboptimal compression results.
A method involving trained neural networks for encoding, transmitting, and decoding lossy images and videos, utilizing quantization processes with variable bin sizes and region-of-interest identification to minimize information loss and improve compression efficiency.
The method effectively reduces data transmission requirements and energy consumption by generating high-quality approximations of input images and videos, while optimizing neural network parameters through iterative training based on perceptual metrics and region-specific quantization.
Smart Images

Figure 2026123108000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a method and system for encoding, transmitting, and decoding lost images or lost videos. Methods, apparatus, and computer programs for encoding and transmitting lossy images or lossy videos. and computer-readable storage media, as well as receiving and decoding lost images or videos. This relates to methods, apparatus, computer programs, and computer-readable storage media.
[0002] Demand for image and video content from users of communication networks is increasing. In addition to the number of images displayed and the video playback time, the number of images displayed and the number of videos played are also considered for higher resolution content. Demand for this is also increasing. As a result, the load on communication networks is increasing, and the data being transmitted is increasing. As the amount of data increases, the energy consumption of communication networks is also increasing.
[0003] To mitigate the impact of these problems, image and video content is compressed and shared online. It is transmitted over the network. Image and video content compression includes lossless compression and lossy compression. Lossless compression allows all the original information contained in the content to be restored, including images. and compresses video. However, when using lossless compression, the amount of data that can be reduced is... There are limits to reduction. In lossy compression, information is lost from images and videos during the compression process. Known compression techniques result in images and videos after decompression that are not particularly noticeable to the human visual system. The attempt is to minimize apparent information loss by removing information that would cause change. That is the case.
[0004] Artificial intelligence (AI) based compression technology uses trained neural networks in the compression and decompression processes. By using a network, it enables the compression and decompression of images and videos. Generally, during the training of a neural network, the original images and videos are compressed and decompressed. The differences between the images and videos are analyzed, minimizing the data required for content transmission. While keeping this difference to a minimum, the neural network parameters are modified to reduce this discrepancy. However, with AI-based compression methods, the appearance and transmission of compressed images and videos are important. The compression result may be poor in terms of the amount of information required.
[0005] According to the present invention, methods for encoding, transmitting, and decoding lossy images and videos are provided. This method involves the steps of receiving an input image in a first computer system and first The input image is encoded using a trained neural network to generate a latent image representation. The steps involve performing a quantization process on the latent representation to generate the quantized latent. Steps, where the bin size used in the quantization process is based on the input image. The steps include sending the quantized potential to a second computer system and the second training The quantized latent is decoded using the neural network to generate the output image. The process includes the steps of: 1) making an output image an approximation of the input image; and 2) making an output image an approximation of the input image.
[0006] The bin size may differ between at least two pixels of the latent representation.
[0007] The bin sizes may differ between at least two channels of the latent representation.
[0008] The bin size may be assigned to each pixel of the latent representation.
[0009] The quantization process involves assigning a bin to each pixel in the latent representation. This may include performing operations corresponding to the size.
[0010] Quantization may involve subtracting the average value of the latent representation from each pixel of the latent representation. be.
[0011] The quantization process may include rounding functions.
[0012] The size of the bins used to decode the quantized latent is the size of the quantized latent. It may also be based on previously decoded pixels.
[0013] The quantization process may be constructed from a third trained neural network.
[0014] The third trained neural network is a quantized latent with at least one prior The decoded pixels may be received as input.
[0015] This method generates a hyperlatent representation using a fourth trained neural network. The steps involve encoding the latent representation using the 'k' and generating the quantized hyperlatent. The steps include performing a quantization process on the hyperlatent representation, and the quantized hyper The steps involve sending the potential to a second computer system and obtaining the bin size. Next, the quantized hyperlatency is decoded using a fifth trained neural network. The step involves transforming the latent, and the decoding of the quantized latent uses the size obtained from the bins. The steps may further include:
[0016] The output of the fifth trained neural network is further calculated to obtain the bin size. It can also be processed by a function.
[0017] A further function may be a sixth trained neural network.
[0018] The bin sizes used in the quantization process of the hyperlatent representation may be based on the input image. .
[0019] This method involves the steps of identifying at least one region of interest in an input image, and identifying the region of interest For at least one corresponding pixel of the latent representation within the region, used in the quantization process The process may further include the step of reducing the size of the bottle.
[0020] This method involves the steps of identifying at least one region of interest in an input image, and identifying the region of interest Use different quantization processes for at least one corresponding pixel in the latent representation within the region. The following steps may also be included.
[0021] At least one region of interest is identified by the seventh trained neural network. It's okay.
[0022] The locations of one or more regions of interest may be stored in a binary mask, and the binary mask is It may be used to obtain the size of the bin.
[0023] According to the present invention, a method for training one or more neural networks is provided, The neural network above is used for encoding, transmitting, and decoding lossy images or videos. This method is intended for use, and involves receiving an input image in a first computer system. The first step involves encoding the input image using the first neural network to create a latent representation. The steps involve generating the latent representation and then performing a quantization process on the latent representation to generate the quantized latent representation. Steps, where the bin size used in the quantization process is based on the input image. Then, the quantized latent representation is decoded using a second neural network to generate the output image. The step is to make the output image an approximation of the input image, and the output image and input A step of determining an amount based on the difference between the image and the first N Update the parameters of the neural network and the second neural network. Then, using the first set of input images, repeat the above steps to complete the first training. The generated neural network and the second trained neural network Includes "Pu".
[0024] This method encodes the latent representation using a third neural network, and then performs hyperlatency The steps involve generating a representation and then performing a quantization process on the hyperlatent representation to quantize it. The steps involve generating a hyperlatent representation and then processing the quantized hyperlatent representation into a second quantized representation. The steps involve sending the data to the computer system and then using a fourth neural network to perform quantum processing. A step of decoding the converted hyperlatent representation to obtain the bin size, and a quantity Decoding the childed latent involves a step that uses the size of the obtained bin and a third new The parameters of the neural network and the fourth neural network are the same as those of the third neural network that was trained. To obtain the neural network and the fourth trained neural network , and further include a step that is updated based on the determined amount.
[0025] The quantization process may include a first quantization approximation.
[0026] The determined quantity may also be based on a rate associated with the quantized latent, and a second The quantization approximation may be used to determine the rate associated with the quantized latent. The second quantization approximation may differ from the first quantization approximation.
[0027] The determined quantity may include a loss function, which updates the parameters of the neural network. The steps involve evaluating the gradient of the loss function and, via the neural network, The steps may include backpropagating the gradient of the loss function, and the gradient of the loss function A third quantization approximation is used during the backpropagation of the first quantization approximation. This is the same approximation as the quantization approximation.
[0028] The neural network parameters are further updated based on the distribution of bin sizes. It's okay.
[0029] At least one parameter of the distribution may be learned.
[0030] The distribution may also be an inverse gamma distribution.
[0031] The distribution may be determined by a fifth neural network.
[0032] According to the present invention, a method for lossy image or video coding and transmission is provided, The method involves the steps of receiving an input image with a first computer system and a first trained The step of encoding the input image using a neural network to generate a latent representation. The next step was to perform a quantization process on the latent representation to generate the quantized latent. The bin size used in the quantization process is based on the input image, the step, and the quantized result. The process includes the step of transmitting the potential.
[0033] According to the present invention, a method for receiving and decoding lost images or videos is provided, The method involves the second computer system transmitting the quantized data according to the above method. The steps involve receiving the latent data and then using a second trained neural network to process the data. A step of decoding the childized latent to generate an output image, wherein the output image is the input image It includes the step, which is an approximation of .
[0034] According to the present invention, a method for training one or more neural networks is provided, and one of them Multiple neural networks can encode, transmit, and decode lossy images or videos. This is intended for use in chemical processes, and the method involves inputting into a first computer system. The steps involve receiving an image and encoding the input image using a first neural network. The steps are: generating a latent representation and performing a quantization process on the latent representation to create a quantized latent representation. The steps involve generating and decoding the quantization latent using a second neural network. The step of generating an output image, wherein the output image is an approximation of the input image, and The steps include determining an amount based on the difference between the output image and the input image, and determining an amount based on the determined amount Next, the parameters of the first neural network and the second neural network The steps involve updating the input image and repeating the above steps using multiple sets of input images. , 1 trained neural network and 2 trained neural network A step of generating a q, wherein at least one of a set of input images is a specific q Includes a first proportion of images containing the characteristic, and at least one other set of multiple sets of input images. The first part includes a second proportion of the image containing specific features, and this second proportion is different from the first proportion. Includes steps and
[0035] The first proportion may be all of the images in the set of input images.
[0036] The specific feature was one of the following: a human face, an animal face, letters, eyes, lips, a logo, a car, a flower, or a pattern. That's fine.
[0037] Each of the multiple sets of input images is used the same number of times during the repetition of the method step. It may also be used.
[0038] The difference between the output image and the input image is determined by a neural network that acts as a classifier. This may be determined at least partially.
[0039] A separate neural network, acting as a classifier, processes each of the multiple sets of input images. It may be used for the net.
[0040] One or more parameters of a neural network acting as a classifier are the first number of parameters. The training steps may be updated, and the neural network acting as a discriminator One or more other parameters may be updated for a second number of training steps, The number is smaller than the first number.
[0041] The determined quantity may also be based on the rate associated with the quantized latent, and the input image Updating parameters for at least one of multiple sets of images is related to the quantized latent. A first weighting may be used for the successive rates, among multiple sets of input images. Updating the parameters for at least one other set of values is related to the quantized latent. A second weighting may be used for the rate, the second weighting being the same as the first weighting. It is different.
[0042] The difference between the output image and the input image is at least partially calculated using multiple perceptual metrics. It may be determined that the parameter update applies to at least one of multiple sets of input images. This may use a first weighting set for multiple perceptual metrics, and the input image Updating parameters for at least one of the other sets among multiple sets is performed for multiple sets A second weighting set for perceptual metrics may be used, This is different from the first weighting set.
[0043] The input image is recognized by a third trained neural network as having one or more regions of interest. It may be a modified image that is separated and has other areas of the image masked.
[0044] Areas of interest include one or more features of human faces, animal faces, letters, eyes, lips, logos, cars, flowers, and patterns. It may also include the region.
[0045] The locations of one or more areas of interest may be stored in a binary mask.
[0046] The binary mask may also be an additional input to the first neural network.
[0047] According to the present invention, methods for encoding, transmitting, and decoding lossy images and videos are provided. Therefore, this method involves the steps of receiving an input image in a first computer system and latent The first trained neural network is used to encode the input image to generate a representation. The steps involve quantizing the latent representation and then performing a quantization process on the latent representation to generate the quantized latent. The steps to be executed and the step to send the quantized potential to a second computer system Then, using a second trained neural network to generate the output image, A step of decoding the childized latent, wherein the output image is an approximation of the input image. It includes a first trained neural network and a second trained neural network. The network is trained according to the method described above.
[0048] According to the present invention, a method for lossy image or video coding and transmission is provided, The method involves the steps of receiving an input image with a first computer system and a first trained The step of encoding the input image using a neural network and generating a latent representation. The steps include performing a quantization process on the latent representation to generate a quantized latent, and the quantity The first trained neural network includes the step of transmitting the childized latent. It is trained according to the method described above.
[0049] According to the present invention, a method for receiving and decoding lost images or videos is provided, The method involves a second computer system that processes the quantized latent according to the method of claim 48. The receiving step and the second trained neural network to generate the output image A step of decoding a quantized latent using a key, wherein the output image is an approximation of the input image. The second trained neural network includes the steps described above. They are trained accordingly.
[0050] According to the present invention, a method for training one or more neural networks is provided, and one of them Multiple neural networks can encode, transmit, and decode lossy images or videos. This is intended for use in the following way, and the method involves inputting images into a first computer system. The steps involve receiving an image and encoding the input image using a first neural network. The steps involve generating a latent representation and then performing a quantization process on the latent representation to obtain the quantized representation. The steps involve generating a latent representation and then quantizing it using a second neural network. The step involves decoding the latent representation and generating an output image, wherein the output image is an input image The approximation determines the quantity based on the step and the rate associated with the quantized latent. The step involves a step in evaluating the rate, which includes an interpolation step of the discrete probability rate mass function. Based on the determined quantity, the first neural network and the second neural network Steps to update the network parameters and the first trained neural network Multiple input images are used to generate a first and second trained neural network. This includes the step of repeating the above steps using a set.
[0051] At least one parameter of the discrete probability rate mass function is based on the evaluated rate. It may be updated further.
[0052] This method uses a third neural network to encode the latent representation and then hyperrays The steps involve generating a tent representation and then performing a quantization process on the hyperlatent representation to create a quantum representation. The steps involve generating a hyperlatent and then quantizing it using a fourth neural network. The hyperlatency is decoded to obtain at least one parameter of the discrete probability mass function. Step, and the third neural network and the fourth neural network The parameters are further updated based on the determined quantity, and the third trained neural The steps include obtaining the network and a fourth trained neural network, It may also include more.
[0053] Interpolation methods include piecewise constant interpolation, nearest neighbor interpolation, linear interpolation, polynomial interpolation, spline interpolation, and piecewise interpolation. It may include at least one of cubic interpolation, Gaussian processes, and kriging.
[0054] The discrete probability rate mass function can also be a categorical distribution.
[0055] The categorical distribution may be parameterized by at least one vector.
[0056] The categorical distribution may also be obtained by softmax projection of the vector.
[0057] The discrete probability mass function is parameterized by at least the mean parameter and the scale parameter. It may be converted to a 'ta' form.
[0058] The discrete probability mass function may be multivariate.
[0059] The discrete probability mass function may be constructed from multiple points, and the first adjacent of the multiple points A set of points may have a first interval, and a second set of adjacent points among multiple points is a second The interval may be such that the second interval is different from the first interval.
[0060] The discrete probability mass function may be constructed from multiple points, and the first adjacent of the multiple points A set of points may have a first interval, and a second set of adjacent points among multiple points is a second The intervals may be such that the second interval is equal to the first interval.
[0061] At least one of the first interval and the second interval is obtained using a fourth neural network. It's okay if it's done that way.
[0062] At least one of the first interval and the second interval is the value of at least one pixel in the latent representation It may be obtained based on this.
[0063] According to the present invention, methods for encoding, transmitting, and decoding lossy images and videos are provided. This method involves the steps of receiving an input image in a first computer system and first The input image is encoded using a trained neural network to generate a latent representation. The steps involve performing a quantization process on the latent representation to generate the quantized latent. Step 1, and a step 2, which transmits the quantized potential to a second computer system. The quantized latent is decoded using a trained neural network, and the output image A step of generating an output image in which the output image is an approximation of the input image, and a step of The first trained neural network and the second trained neural network The Ku will be trained according to the method described above.
[0064] According to the present invention, a method for encoding and transmitting lossy images or videos is provided, The method includes the steps of receiving an input image in a first computer system and the first A trained neural network is used to encode the input image and generate a latent representation. Step 1: Perform a quantization process on the latent representation to generate the quantized latent. The process includes a step of transmitting a quantized latent, and a first trained neural network The team is trained according to the method described above.
[0065] According to the present invention, a method for receiving and decoding lost images or videos is provided, The method involves, in a second computer system, quantizing the latent according to the above method. The receiving step and the quantized latency using a second trained neural network The step involves decoding the input to generate an output image, wherein the output image is an approximation of the input image. The second trained neural network includes steps and steps, and follows the method described above. To be trained.
[0066] According to the present invention, methods for encoding, transmitting, and decoding lossy images and videos are provided. This method involves the steps of receiving an input image in a first computer system and then... The input image is encoded using a trained neural network to generate a latent representation. The steps are: performing the first operation on the latent representation to obtain the residual latent; The steps include sending the residual latent to a second computer system and obtaining the acquired latent representation. The step is to perform a second operation on the residual potential in order to obtain The steps include performing operations on the previously obtained latent pixels, The latent representation obtained using the second trained neural network is decoded, and A step of generating a force image, wherein the output image is an approximation of the input image, and include.
[0067] The operation on the acquired latent previously acquired pixels is performed on the acquired latent previously acquired pixels. This can be done for each individual pixel.
[0068] At least one of the first and second operations involves solving an implicit system of equations. That's good too.
[0069] The first operation may include a quantization operation.
[0070] The operations performed on the acquired latent previously obtained pixels include matrix operations. That's good too.
[0071] The matrix used to define a matrix operation may be sparse.
[0072] The matrix that defines the matrix operation is the acquired latent matrix that is not obtained when the matrix operation is performed. It may have a zero value corresponding to the current pixel.
[0073] The matrix used to define matrix operations may be a lower triangular matrix.
[0074] The second operation may be constructed using standard forward substitution.
[0075] The operations performed on the acquired latent previously obtained pixels are performed on the third trained It may be constructed from a neural network.
[0076] This method encodes the latent representation using a fourth trained neural network, The steps involve generating a hyperlatent representation and then processing the quantized hyperlatent using a second computer. The steps involve sending data to the system and using a fifth trained neural network to perform quantum operations. A step of decoding the converted hyperpotential, wherein the obtained latent is a previously obtained The operations performed on the pixels are output to the fifth trained neural network. It may further include steps based on, and
[0077] Decoding the quantized hyperlatency using a fifth trained neural network is Furthermore, mean parameters may be generated, and the implicit equation system is further, mean parameters It may include.
[0078] According to the present invention, a method for training one or more neural networks is provided, The neural network above is used for encoding, transmitting, and decoding lossy images or videos. This method is intended for use in which the input image is received by the first computer system. The steps involve encoding the input image using the first neural network and creating a latent representation. The steps are to generate the latent representation and to perform a first operation on the latent representation to obtain the latent residual. Step 1 and then perform a second operation on the latent residual to obtain the acquired latent representation. The first operation is performed on the previously acquired pixels of the acquired latent pixels. The steps include performing the following, and then quantizing using a second neural network. The step involves decoding the latent data and generating an image, wherein the output image is an approximation of the input image. A step of determining a quantity based on the difference between the output image and the input image, Based on the determined quantity, the first neural network and the second neural network Steps to update the parameters of the workpiece and the above steps using the first set of input images. Repeat the process to train the first trained neural network and the second trained neural network. This includes steps for generating a neural network.
[0079] The operations performed on the acquired latent previously obtained pixels include matrix operations. That's good too.
[0080] The parameters of the matrix that define the matrix operation are updated additionally based on the determined quantity. good.
[0081] The operations performed on the acquired latent previously obtained pixels are performed on the third neural It may include a network, and the parameters of the third neural network are the third train To generate a sophisticated neural network, additionally based on the determined amount... It may be replaced.
[0082] This method uses a fourth neural network to encode the latent representation, and hyperlatentiate The steps involve generating a representation and then performing a quantization process on the hyperlatent representation to quantize it. The steps involve generating a hyperlatent representation and then quantizing the hyperlatent representation into a second The steps involve sending data to a computer system and using a fifth neural network to calculate the quantity. A step of decoding the childized hyperpotential, wherein the acquired latent is previously obtained The operations performed on the pixels are the output of the fifth trained neural network. Based on this, the parameters of the fourth and fifth neural networks The data is further updated based on the determined amount, and the fourth trained neural network The steps include generating a twerk and a fifth trained neural network, and further It is included in.
[0083] According to the present invention, a method for lossy image or video coding and transmission is provided, The method involves the steps of receiving an input image with a first computer system and a first trained The step of encoding the input image using a neural network and generating a latent representation. Then, the first operation is performed on the latent representation to obtain the residual latent, and the residual latent is transmitted Includes the step of sending.
[0084] According to the present invention, a method for receiving and decoding lost images or videos is provided, The method involves the second computer system receiving residual submarines transmitted according to the above method. The steps involve receiving the current and performing a second operation on the residual latent to obtain the latent representation. The first step is to obtain the obtained latent previously obtained pixels, and the second operation is to obtain the obtained latent pixels The steps include performing calculations and then a second trained neural network A step of using to decode the acquired latent representation and generate an output image, The output image is an approximation of the input image, including steps.
[0085] According to the present invention, a method for training one or more neural networks is provided, and one of them Multiple neural networks can encode, transmit, and decode lossy images or videos. This is intended for use in chemical processes, and the method involves inputting into a first computer system. The process involves receiving an image and encoding the input image using a first neural network. The steps are: generating a latent representation, entropy coding the latent representation, and The steps include sending the tropy-encoded latent representation to a second computer system, The steps involve entropy decoding the entropy-encoded latent representation and the second NuV The step involved using a multi-network to decode the latent representation and generate an output image. The output image is an approximation of the input image, based on the step, the difference between the output image and the input image. The first neural network is formed based on the determined quantity. The steps include updating the parameters of the second neural network and the first input image. Repeat the above steps using the set to obtain the first trained neural network. The process includes the steps of generating a first and second trained neural network, and the process of generating an second trained neural network. Entropy decoding of the tropy-encoded latent representation is performed pixel by pixel. The decoding order for each log is updated additionally based on the determined amount.
[0086] The order of decoding each pixel may be based on the latent representation.
[0087] Entropy-encoded latent entropy decoding is performed on previously decoded pixels. It may include operations based on it.
[0088] The determination of the order of pixel-by-pixel decoding involves multiple pixels of the latent representation in a directed aperiodic graph. This may include ordering the cells.
[0089] The determination of the order of decoding each pixel operates on the latent representation using multiple adjacency matrices. It may include and.
[0090] Determining the order of pixel-by-pixel decoding involves dividing the latent representation into multiple subimages. But that's fine.
[0091] Multiple subimages are obtained by convolving the latent representation with multiple binary mask kernels. It's okay if it's done that way.
[0092] Determining the order of decoding each pixel is based on the magnitude of the quantity associated with each pixel. This may include ranking multiple pixels of the latent representation.
[0093] The quantities associated with each pixel are the positional parameters or scale parameters associated with that pixel. A meter is also acceptable.
[0094] The amount associated with each pixel may be additionally updated based on the evaluated difference.
[0095] The determination of the order of pixel-by-pixel decoding is performed by wavelet decomposition of multiple pixels in the latent representation. It may include.
[0096] The order of decoding each pixel is determined by the frequency of the wavelet decomposition related to multiple pixels. It may also be based on the ingredients.
[0097] This method uses a fourth trained neural network to encode the latent representation and The steps involve generating a per-latent representation and sending the hyper-latent to a second computer system. The first step is to decode the hyperlatency using a fifth trained neural network. This is a step in which the order of decoding each pixel is the fifth trained neural network It may further include steps based on the output of the network.
[0098] According to the present invention, methods for encoding, transmitting, and decoding lossy images and videos are provided. This method involves the steps of receiving an input image in a first computer system and then... Use a trained neural network to encode the input image and generate a latent. Steps to entropy encode the latent, and the entropy encoded latent. The steps include sending to a second computer system and the entropy-encoded latent The entropy decoding step and the second trained neural network The step involves decoding the latent and generating an output image, wherein the output image is an approximation of the input image. A certain step includes a first trained neural network and a second training The resulting neural network is trained according to the method described above.
[0099] According to the present invention, a method for lossy image or video coding and transmission is provided, The method includes the steps of receiving an input image in a first computer system and a first instruction Using a sophisticated neural network, the input image is encoded to generate a latent representation. Step, a step to entropy encode the latent representation, and the entropy encoded latent The first trained neural network includes the step of transmitting the representation, and the above They are trained according to the method.
[0100] According to the present invention, a method for receiving and decoding lost images or videos is provided, The method involves the input transmitted in the second computer system according to the above method. Steps include receiving a ropy-encoded latent representation and a second trained neural network. A step of decoding the latent representation using a workpiece to generate an output image, wherein the output image This is an approximation of the input image, including a step and a second trained neural network. The Ku will be trained according to the method described above.
[0101] According to the present invention, a method for training one or more neural networks is provided. The one or more neural networks are used for encoding, transmitting, and lossy images or videos. This is for use in decryption, and the method is performed on a first computer system. The steps include receiving an input image and encoding the input image using a first neural network. The steps involve transforming the data and generating a latent representation, and then using a second neural network to generate the latent representation. The step involves decoding the input image to generate an output image, wherein the output image is an approximation of the input image. Based on the steps, the difference between the output image and the input image, and the rate associated with the latent representation This is a step in determining the quantity, in which a first weighting is applied to the output image and the input. The difference from the image is applied, and the second weight is applied to the rate related to the latent representation. Based on the determined quantity, the first neural network and the second neural network The steps involve updating the parameters of the network and using the first set of input images. Repeat the steps described above to train the first neural network and the second training The process includes the step of generating a neural network, and the repetition of the above steps. After at least one of the first weights and at least one of the second weights, It is further updated based on additional quantities, and the additional quantity is the difference between the output image and the input image. and based on at least one of the rates associated with the latent expression.
[0102] At least one of the difference between the output image and the input image and the rate associated with the latent representation is, This may be recorded for each iteration of the step, and further amounts may be recorded between the output image and the input image. Multiple previously recorded rates related to multiple previously recorded differences and latent representations It may be based on at least one of the following.
[0103] Further quantities may be based on the average of multiple previously recorded differences or rates.
[0104] Means include arithmetic mean, median, geometric mean, harmonic mean, exponential moving average, smoothed moving average, and linear mean. It may be at least one of the weighted moving averages.
[0105] Outliers should be removed from previously recorded differences or rates before determining further amounts. It may also be used.
[0106] Outliers may be removed only for the first predetermined number of repetitions of the step.
[0107] The rate associated with the latent expression is calculated using the first method when determining the quantity, and further When determining the quantity, the second method may be used for calculation, and the first method may be used in conjunction with the second method. They are different.
[0108] At least one iteration of the step uses an input image from a second set of input images. It may be executed in this way, and if input images from a second set of input images are used, the first The parameters of the first neural network and the second neural network are further updated. It doesn't have to be done.
[0109] The determined quantity is further based on the output of the neural network that acts as a classifier. That's fine.
[0110] According to the present invention, a method for training one or more neural networks is provided, The neural network above is used for encoding, transmitting, and decoding lossy video. This method involves the step of receiving input video in a first computer system. Then, using the first neural network, multiple frames of the input video are encoded, The steps involve generating latent representations of numbers and using a second neural network to process multiple latent representations. A step of decoding the representation and generating multiple frames of output video, wherein the output video is The input image is approximated by the step, the difference between the output image and the input image, and multiple latent representations. A step of determining an amount based on a rate related to the output image, wherein the first weighting is This is applied to the difference between the image and the input image, and the second weighting is a rate associated with multiple latent representations. Based on the steps and the determined amount applied, the first neural network The steps include further updating the parameters of the second neural network and multiple inputs. By repeating the above steps using video, the first trained neural network The above steps include the step of generating a second trained neural network, and After at least one iteration of the step, the first weight and the second weight are less However, one side is further updated based on a further amount, and the further amount is output video and It is based on the difference between the input video and at least one of the rates associated with multiple latent representations.
[0111] The input video contains at least one I-frame and multiple P-frames.
[0112] The quantity may be based on multiple first or second weightings, each of which is It corresponds to one of multiple frames of the input video.
[0113] After at least one iteration of the step, at least one of the multiple weights, each weight It may be updated additionally based on the additional amount associated with the discovery.
[0114] Each additional amount is a predetermined target value of the difference between the output frame and the input frame, or a latent representation of that difference. It may also be based on relevant rates.
[0115] The additional amount associated with the I-frame may have a first target value, and the smaller amount associated with the P-frame At the very least, one additional quantity may have a second target value, and the second target value is equal to the first target value. It is different.
[0116] Each additional quantity associated with the P-frame may have the same target value.
[0117] Multiple first or second weights may be initially set to zero.
[0118] According to the present invention, the encoding of lossy images and videos is provided with a method for transmission and decoding. The method is provided and comprises the steps of receiving an input image in a first computer system and The input image is encoded using a trained neural network to generate a latent representation. The steps involve performing a quantization process on the latent representation to generate the quantized latent. Step 1, and Step 2, which transmits the quantized potential to a second computer system. The quantized latent is decoded using two trained neural networks, and the output image is generated. A step of generating an image, wherein the output image is an approximation of the input image, and includes the step of The first trained neural network and the second trained neural network Work will be trained according to the method described above.
[0119] According to the present invention, a method for lossy image or video coding and transmission is provided, The method includes the steps of receiving an input image or video in a first computer system, and Using a trained neural network, encode the input image or video, and The process includes the steps of generating an active representation and transmitting a latent representation, and the first trained The neural network is trained according to the method described above.
[0120] According to the present invention, a method for receiving and decoding lost images or videos is provided, The method involves receiving the latent expression in the second computer system according to the method described above. The first step is to decode the latent representation using a second trained neural network. A step of generating an output image or video, wherein the output image or video is an input image Or is an approximation of video, including steps and a second trained neural network The Ku will be trained according to the method described above.
[0121] According to the present invention, a method for encoding, transmitting, and decoding lossy images or videos is provided. Therefore, this method involves the steps of receiving an input image in a first computer system and latent The input image is coded using the first trained neural network to generate a representation. The steps involve transforming the latent and then performing a quantization process on the latent representation to generate the quantized latent. The steps involve performing the following steps, and entropy coding the quantized latent using a probability distribution. The probability distribution is defined using a tensor network, with steps and an entry. Steps to transmit the morphy-encoded quantized latent to a second computer system. Then, the entropy-encoded quantized latent is entropy-decoded using a probability distribution. Then, the step of obtaining the quantized latent and the second trained neural network The steps include: decoding the quantized latent quantity using the quantized latent quantity to generate an output image, and output image The image is an approximation of the input image.
[0122] Even if the probability distribution is defined by a Hermitian operator acting on a quantized latent, Often, Hermitian operators are defined using tensor networks.
[0123] A tensor network includes a non-normal core tensor and one or more orthonormal tensors. But that's fine.
[0124] This method encodes the latent representation using a third trained neural network, The steps involve generating a hyperlatent representation and then performing quantization on the hyperlatent representation. , the step of generating a quantized hyperlatent representation, and the quantized hyperlatent representation The steps include sending the data to a second computer system and a fourth trained neural network. A step of decoding a quantized hyperlatent representation using a twerk, the fourth The output of a trained neural network is one or more parameters of a tensor network. It may also include steps and other elements.
[0125] A tensor network consists of a non-normal core tensor and one or more orthonormal tensors. Alternatively, the output of the fourth trained neural network is 1 of the unnormalized core tensor. There may be more than one parameter.
[0126] One or more parameters of the tensor network may be calculated using one or more pixels of the latent representation and may be calculated using one or more pixels of the latent representation.
[0127] The probability distribution may be related to a subset of pixels of the latent representation.
[0128] The probability distribution may be related to the channels of the latent representation.
[0129] The tensor network may be at least one of a tensor tree, a locally purified state, a bone machine, a matrix product state, and a projected entangled pair state factorization.
[0130] According to the present invention, there is provided a method for training one or more networks, the one or more networks being used for encoding, transmitting, and decoding a lossy image or video, the method including receiving a first input image, encoding the first input image using a first neural network to generate a latent representation, performing quantization processing on the latent representation to generate a quantized latent, entropy encoding the quantized latent using a probability distribution, the probability distribution being defined using a tensor network, entropy decoding the entropy-encoded quantized latent using the probability distribution to obtain the quantized latent, decoding the quantized latent using a second neural network to generate an output image, the output image being an approximation of the input image, determining a quantity based on the difference between the output image and the input image, updating parameters of the first neural network and the second neural network based on the determined quantity, and repeating the steps for a plurality of input images wherein the output image is an approximation of the input image, updating parameters of the first neural network and the second neural network based on the determined quantity, Repeat the above steps using the first trained neural network and the second This includes the step of generating a trained neural network.
[0131] One or more parameters of the tensor network are additionally updated based on the determined quantity. It may also be used.
[0132] A tensor network may include a non-normal core tensor and one or more orthonormal tensors. The parameters of all tensors in a tensor network, excluding the unnormalized core tensor, are It may be updated based on the determined amount.
[0133] Tensor networks may be computed using latent representations.
[0134] The tensor network may be computed based on linear interpolation of the latent representation.
[0135] The determined quantity may also be based on the entropy of the tensor network.
[0136] According to the present invention, a method for lossy image or video coding and transmission is provided, The method involves the steps of receiving an input image with a first computer system and a first trained Steps to encode an input image using a neural network and generate a latent representation The steps are: 1) Performing a quantization process on the latent representation to generate a quantized latent; , a step of entropy coding a quantized latent using a probability distribution, The rate distribution is defined using a tensor network, with step and entropy coding. The process includes the step of transmitting the quantized latent.
[0137] According to the present invention, a method for receiving and decoding lost images or videos is provided, The method involves the input transmitted in the second computer system according to the above method. The steps involve receiving the entropy-coded quantization latent and using the probability distribution to obtain the entropy-coded quantity The steps are to entropy-decode the quantization latent to obtain the quantization latent, and then the second trained The neural network is used to decode the quantization latent and generate the output image. The process includes steps such as a step where the output image is an approximation of the input image.
[0138] According to the present invention, methods for encoding, transmitting, and decoding lossy images and videos are provided. This method involves the steps of receiving an input image in a first computer system and first The input image is encoded using a trained neural network to generate a latent representation. The first step is to encode the latent representation using a second trained neural network. , the step of generating a hyperlatent representation and a third trained neural network The step of using to encode the hyperlatent representation and generate the hyperhyperlatent representation. And, latent representation, hyper-latent representation and hyper-hyper-latent representation are used in a second computer The steps involve sending data to the system and using a fourth trained neural network to perform high-speed processing. Steps to decode the per-hyper latent representation and the fourth trained neural network Using the output of the 5th trained neural network, a hyper-hyper-latent representation is created. The steps to decode the output of the fifth trained neural network and the sixth training The latent representation is decoded using the output of the neural network, and the output image is generated. A step of forming, wherein the output image is an approximation of the input image, and.
[0139] This method may further include a step of determining the rate of the input image, and if the determined rate satisfies a predetermined condition, the steps of encoding the hyper latent representation and the hyper latent representation decoding steps are not executed.
[0140] If the determined rate satisfies a predetermined condition, the steps of encoding the hyper flat representation and decoding the hyper flat representation are not executed.
[0141] According to the present invention, a method for training one or more networks is provided, and the one or more networks are for use in encoding, transmitting, and decoding lossy images or videos, and the method includes receiving an input image at a first computer system, encoding the input image using a first neural network to generate a latent representation, encoding the latent representation using a second neural network to generate a hyper latent representation, encoding the hyper latent representation using a third neural network to generate a hyper - hyper latent representation, decoding the hyper - hyper latent representation using a fourth neural network, decoding the hyper latent representation using the output of the fourth neural network and a fifth neural network, decoding the latent representation using the output of the fifth neural network and a sixth neural network to generate an output image, where the output image is an approximation of the input image, a step, determining a quantity based on the difference between the output image and the input image, determining and determining and a step of generating an output image, where the output image is an approximation of the input image, and a step of determining a quantity based on the difference between the output image and the input image, and determining The parameters of the third and fourth neural networks are updated based on the amount obtained. The third and fourth training steps are performed by repeating the above steps using multiple input images. This includes the step of generating a neural network.
[0142] The parameters of the first, second, fifth, and sixth neural networks are the step repetitions. It is not necessary for at least one of the returns to be updated.
[0143] This method may further include the step of determining the rate of the input image, and the determined rate If the specified conditions are met, the first, second, fifth, and sixth neural networks The parameters are not updated in that iteration of the step.
[0144] A predetermined condition may be that the rate is below a predetermined value.
[0145] The parameters of the first, second, fifth, and sixth neural networks are determined by the step size. It is not necessary to update after the number of repetitions.
[0146] The parameters of the first, second, fifth, and sixth neural networks are determined by the amount. Based on this, the first, second, fifth, and sixth trained neural networks were further updated. You may generate a workpiece.
[0147] Before performing any other steps, for at least one of the multiple input images, Perform at least one of the following processes: sampling, smoothing filter, and random cropping. You may do so.
[0148] According to the present invention, a method for image or video encoding and transmission is provided, and this method This involves the steps of receiving an input image with a first computer system and a first trained N The process involves a step of encoding an input image using a neural network to generate a latent representation, and The latent representation is encoded using a second trained neural network, and the hyperlatency is then processed. The steps involve generating representations and using a third trained neural network to hyper The steps involve encoding the latent expression to generate a hyper-hyper latent expression, and the latent expression, The process includes the steps of transmitting a hyper latent expression and a hyper-hyper latent expression.
[0149] According to the present invention, a method for receiving and decoding lost images or videos is provided. The law applies to latent expressions, hyperlatent expressions and hyperhyper - The steps of receiving the latent expression in a second computer system and a fourth trained nu The steps involve decoding the hyper-hyper-latent representation using a multi-network, and the fourth step The output of the trained neural network and the fifth trained neural network The steps involve decoding the hyperlatent representation using force, and the fifth trained neural network The latent representation is decoded using the twerk and the output of the sixth trained neural network. Steps include generating an output image by converting it, wherein the output image is an approximation of the input image. Includes "Pu".
[0150] According to the present invention, a data processing system constructed to perform any of the above methods It will be provided.
[0151] According to the present invention, a data processing device constructed to perform any of the above methods is provided. To be served.
[0152] According to the present invention, when a program is executed by a computer, the computer A computer program is provided that includes instructions to perform one of the methods described below.
[0153] According to the present invention, when executed by a computer, the computer performs the above method A computer-readable storage medium is provided that contains instructions for executing one of the following. [Brief explanation of the drawing]
[0154] Herein, embodiments of the present invention will be described with reference to the following figures and examples. [Figure 1] This shows an example of an image or video compression, transmission, and decompression pipeline. [Figure 2] Further examples of image or video compression, transmission, and decompression pipelines, including hypernetworks, are presented. [Figure 3] This diagram shows a schematic representation of the coding phase of an example Al-based compression algorithm for video and image compression. [Figure 4] This diagram shows a schematic representation of the decoding phase of an example of an AI-based compression algorithm for video and image compression. [Figure 5] This shows an example of the distribution of quantization bin size values that can be learned during training of an Al-based compression pipeline. [Figure 6] A heatmap is shown illustrating how the size of the learned quantization bins changes across latent channels for a given image. [Figure 7] A schematic diagram illustrating the coding phase of an Al-based compression algorithm utilizing hyperpliers and learned quantization bin sizes is shown. [Figure 8] A schematic diagram illustrating the decryption phase of an Al-based compression algorithm utilizing hyperpliers and learned quantization bin sizes is shown. [Figure 9] Here are some examples of inverse gamma distributions. [Figure 10] This section provides an example overview of a GAN architecture. [Figure 11] An example of a standard generative adversarial compression pipeline is shown. [Figure 12] This shows the failure modes of an architecture combining a GAN and an autoencoder. [Figure 13] This example demonstrates a compressed pipeline using multi-classifier cGAN training with dataset bias. [Figure 14] This shows a comparison of rebuilding the same generative model trained at the same bitrate, with and without using a multi-discrimination dataset biasing scheme. [Figure 15] Here is an example of the results of bitrate-adjusted dataset bias. [Figure 16] An example of an Al-based compression pipeline is shown. [Figure 17] Further examples of Al-based compression pipelines are shown. [Figure 18] This example shows a compression pipeline with quantization using a quantization map. [Figure 19] This example shows the result of implementing a pipeline with a face detector used for identifying regions of interest. [Figure 20] The following shows different quantization functions Qm for the regions identified by the ROI detection network H(x). [Figure 21] Here are three examples of typical one-dimensional distributions that can be used in training an Al-based compression pipeline. [Figure 22] This section compares piecewise linear interpolation with piecewise cubic Hermitian interpolation. [Figure 23] This example demonstrates how an autoregressive structure defined by a sparse context matrix L can be used to parallelize components of a serial decoding path. [Figure 24] This example demonstrates the encoding process using the prediction context matrix Ly in an Al-based compression pipeline. [Figure 25]This example demonstrates the decoding process using the prediction context matrix Ly in an AI-based compression pipeline. [Figure 26] An example of a raster scan order for a single-channel image is shown. [Figure 27] Instead of all preceding variables, the next pixel shows a 3x3 receptive field conditioned on the local variable. [Figure 28] Here is an example of a DAG that describes the joint distribution of four variables {y1, y2, y3, y4}. [Figure 29] Here is an example of a two-step analytical algorithm (AO) where all current variables at each step are conditionally independent and can be evaluated in parallel. [Figure 30] The directed graph corresponding to AO in Figure 29 is shown. [Figure 31] This example shows how a 2x2 binary mask kernel that conforms to constraints (61) and (62) generates four subimages. [Figure 32] An example of an adjacency matrix A that determines the graph connectivity of an AO as defined by the binary mask kernel framework is shown. [Figure 33] This shows the index for the Adam7 interlacing method. [Figure 34] An example of visualizing the scale parameter σ, where AO is defined as y1, y2, ..., y16, is shown. [Figure 35] An example of permutation matrix representation is shown. [Figure 36] This example illustrates the concept of a ranking table as applied to the binary mask kernel framework. [Figure 37] This example demonstrates hierarchical autoregressive ordering based on wavelet transforms. [Figure 38] This diagram illustrates various tensors and their products. [Figure 39] Examples of 3-tensor decomposition and matrix product states are shown graphically. [Figure 40] An example of a Locally Purified State is shown graphically. [Figure 41] An example of a tensor tree is shown graphically. [Figure 42] An example of a 2x2 Projected Entangled Pair State is shown graphically. [Figure 43] An example of the procedure for converting a matrix product state into a canonical form is shown graphically. [Figure 44] This example shows an image or video compression pipeline with a tensor network predicted by a hyperencoder / hyperdecoder. [Figure 45] This shows an image or video decompression pipeline with a tensor network predicted by a hyperdecoder.
[0155] Compression reduces the amount of data needed to store information, i.e., the file size. Therefore, it can be applied to all forms of information. Image and video information is compressed. This is an example of the information being stored. The file size required to store the information is specifically the compressed file size. During the compression process that references a file, this is sometimes called the rate. Generally, compression involves reversible pressure. There are two types of compression: reduction and lossy compression. Both compression formats result in smaller file sizes. However, In lossless compression, information is compressed, and no information is lost when it is subsequently decompressed. In other words, during the decompression process, the original file that stored the information is completely reconstructed. In other words, with lossy compression, information is lost during the compression and decompression process, and the reconstructed file is of the original quality. The file may differ. Image and video files containing image and video data. These are common compression formats: JPEG, JPEG2000, AVC, HEVC, AVI. This is an example of image and video file compression.
[0156] In image compression processes, the input image is sometimes represented as x. The image may be stored in a tensor of dimension H × W × C, where H represents the height of the image. W represents the width of the image, and C represents the number of channels in the image. Each H×W data point in the image is paired with Represents the pixel value of the image at the corresponding position. Each channel C of the image is used when the image file is downloaded. Represents the different components of the image of each pixel, which are combined when displayed by the viewfinder. For example, an image file has three channels, and each channel represents the red, green, and Represents the blue component. In this case, the image information is stored in the RGB color space, and the model or format It is sometimes called a set. Other examples of color spaces and formats include CMKY and Y There is the CbCr color model. However, the channels in an image file are limited to storing color information. Video is not something that is simply displayed; other information may also be represented on the channel. Since these are likely to be a series of images, the compression process applied to the images may also apply to the video. Each image that makes up a video is sometimes called a video frame.
[0157] Video frames may be labeled according to their characteristics. For example, video Deo frames may be labeled as I-frames and P-frames. The frame may be the first frame of a new section of video. For example, a scene transition. The first frame after migration is sometimes labeled as an I-frame. A P-frame is an I-frame. It may be a subsequent frame after the frame. For example, the background or An object may not change until it progresses from the I-frame to the P-frame. The changes in the P-frame compared to the I-frame, which is progressing through the frame, are due to the objects present within the frame. It may be described by the movement of the figure or by the perspective movement of the frame.
[0158] The output image may differ from the input image. JPEG2026123108000002.jpg75 (hereinafter also referred to as "x (hat)"). Hereafter, characters with the symbol ^ above them will be referred to as characters (hat It may also be written as (sometimes written as ) and can be represented as . The difference between the input image and the output image is distortion and This is sometimes called a difference in image quality. Distortion occurs when the input image and output image are received, and the input image and It can be measured using any distortion function that provides an output that numerically represents the difference between the output images. An example of such a method is the mean squared error (MSE) between pixels of the input and output images. There is a method using ), but as is well known to those skilled in the art, there are many other methods for measuring strain. Yes, it exists. The distortion function may include a trained neural network.
[0159] Generally, the compression rate and distortion in lossy compression are related. Increasing the rate reduces distortion. Therefore, a decrease in rate leads to an increase in distortion. Changes in distortion are reflected in the rate in a corresponding manner. This may affect the . The relationship between these quantities with respect to a given compression technique is It can be defined by the T-strain equation.
[0160] Al-based compression processes may involve the use of neural networks. A network is an operation that can be performed on an input and produce an output. A network can be constructed with multiple layers. The first layer of the network is the input layer. It receives the input. One or more operations are performed on the input by each layer, and the output of the first layer is generated. The output of the first layer is then passed to the next layer of the network, and in a similar manner to one or more others. The following calculation is performed. The output of the final layer becomes the output of the neural network.
[0161] Each layer of a neural network may be divided into nodes. Each node is a descendant of the previous layer or It receives at least a portion of the input and provides an output to one or more nodes in the subsequent layer. It is also possible that each node in the layer performs one or more operations on the layer for at least a portion of the inputs to the layer. It may be done. For example, a node can receive input from one or more nodes in the previous layer. Yes, it is possible. One or more operations include convolution, weights, biases, and activation functions. Convolutional operations are used in convolutional neural networks. Alternatively, convolution may be performed across the entire input to the layer. The process may be performed on at least a portion of the input to the layer.
[0162] Each of the one or more operations is defined by one or more parameters associated with each operation. For example, weight calculations are applied to each input from each node in the previous layer to each node in the current layer. The weights may also be defined by a weight matrix. In this example, each value of the weight matrix is nu These are the parameters of a convolutional network. Convolution is also known as the kernel. It can be defined by a matrix. In this example, one or more values of the convolution matrix are nu. The activation function may be a parameter of the neural network. The parameters of the network may also be defined by values that can be parameters of the network. This can be changed during network training.
[0163] Other features of a neural network are predetermined, and therefore, training the network Some things remain unchanged. For example, the number of network layers, the number of network nodes. The one or more operations performed in each layer, and the connections between layers are predetermined, therefore These pre-set features may be fixed before the training process is performed. These are sometimes called hyperparameters. These features are used in the network architecture. It is sometimes called "kucha".
[0164] To train a neural network, the expected output (called ground truth) is A training set of inputs that are known (and may be affected) can be used. The network's initial parameters are randomized, and the first training input is provided to the network. The network output is compared to the expected output, and the difference between the output and the expected output is calculated. Based on this, the network is configured to minimize the difference between the network output and the expected output. The parameters of the network are changed. This process is repeated for multiple training inputs. The network is trained. The difference between the network's output and the expected output is defined by the loss function. This can be done. The result of the loss function is obtained by using the difference between the network output and the expected output. The gradient of the loss function can be calculated and determined. Backpropagation of gradient descent of the loss function Pageation uses the gradient dL / dy of the loss function to measure the parameters of the neural network. It may be used to update the data. Multiple neural networks in the system Each network is simultaneously trained by backpropagating the gradient of the loss function. That's fine.
[0165] For Al-based image or video compression, the loss function is determined by the rate-distortion equation. It can be understood. The rate-distortion equation can be expressed as Loss = D + λ*R, where D is The Lagrange multiplier is a distortion function where λ is the weighting coefficient and R is the rate loss. These are provided as weights for specific terms in the loss function relative to each other term, and during network training. It can be used to control which terms of the loss function are given priority.
[0166] For AI-based image or video compression, a training set of input images can be used. Yes, it's possible. An example of a training set of input images is the Kodak image set (for example, www. cs.albany.edu / xypan / research / snr / Kodak.h There is a tml. An example of a training set of input images is an IMAX image set. An example of an image training set is the ImageNet dataset (e.g., www.imag). (e-net.org / download) is available. An example of a training set of input images is: CLIC Training Dataset P (“professional”) and M("mobile") exists (e.g., http: / / challenge.c (compression.cc / tasks / ).
[0167] An example of Al-based compression 100 is shown in Figure 1. The first step of Al-based compression is Then, input image 5 is provided. Input image 5 is used with the function f acting as an encoder. θ to special Provided to the trained neural network 110 as a characteristic. Encoder neural Network 110 generates an output based on the input image. This output is based on the input image 5. This is called a latent representation. In the second step, a quantification process 1 characterized by operation Q is performed. At step 40, the latent representation is quantized, and as a result, a quantized latent representation is obtained. Quantization processes convert continuous latent representations into discrete quantized latent representations. An example of a quantization process is... There is a rounding function.
[0168] In the third step, the quantized latent quantity is entropy coded using entropy coding process 150. —Encode and generate bitstream 130. Entropy coding is performed, for example, It may be range coding or arithmetic coding. In the fourth step, bitstream 130 may be transmitted over a communication network.
[0169] In the fifth step, the bitstream undergoes entropy decoding in the entropy decoding process 160. Entropy is decoded. The quantized latent quantity is decoded by the quantized latent quantity Function g that functions as a coda θ Another trained neural network characterized by Provided to work 120. The trained neural network 120 is quantized The output is generated based on the latent state. The output is the output image of the Al-based compression process 100. It is also acceptable. An encoder-decoder system may be called an autoencoder.
[0170] The system described above may be distributed across multiple locations and / or devices. For example, Encoder 110 is compatible with laptop computers, desktop computers, and smart computers. It may be located on a device such as a phone or server. The decoder 120 is a receiver It may be placed in another device called a vice. Input image 5 to obtain output image 6. The system used to encode, transmit, and decode data is called a compression pipeline. It can happen.
[0171] Al-based compression processing is a hypernetwork for transmitting metadata that improves compression processing. It may further include hyper-network 105. Hyper-network 105 is a hyper-encoder. A trained neural network 115 functioning as JPEG2026123108000003.jpg78, and a hyperdecoder. It is constructed from 125 trained neural networks that function as JPEG2026123108000004.jpg76. An example of the system is shown in Figure 2. The system components that will not be explained further are those mentioned above. We can assume they are the same. Neural network 1 functioning as a hyperdecoder 15 receives the latent signal, which is the output of encoder 110. Hyperencoder 115, It generates output based on a latent expression sometimes called a hyperlatent expression. Next, high Per potential is Q h In the quantization process 145 characterized by, the quantized and quantized Hyperpotential is generated. Q h The quantization process 145 characterized by the above This may be the same as the quantization process 140 characterized by Q.
[0172] In the same manner as described above for the quantized latent, the quantized hyperlatency is: Then, in the entropy coding process 155, it is entropy coded and bitstream Generate 135. Bitstream 135 is processed in entropy decoding process 165. The entropy-decoded, quantized hyperpotential can be extracted. The hyperlatency then becomes a trained neural network that acts as a hyperdecoder. It is used as input to the twerk 125. However, the compression pipeline 100 is In contrast, the output of the hyperdecoder is not an approximation of the input to the hyperdecoder 115. In some cases, the output of the hyperdecoder is the entropy of the main compression process 100. Parameters for use in the P-coding process 150 and the entropy decoding process 160 It is used to provide the following. For example, the output of the hyperdecoder 125 is the average, standard deviation. Entropy coding process 150 of difference, variance, or latent 10 representation and entropy decoding One or more other parameters used to describe the probabilistic model of the 160 treatment process It can include a single entropy decoding process 1. In the example shown in Figure 2, for simplicity, Only 65 and Hyper Decoder 125 are shown. However, in practice, decompression is usually performed. Since this is performed by a separate device, the parameters used in the entropy coding process 150 are provided. To provide this functionality, copies of these processes will exist on the device used for encoding.
[0173] At any stage of the Al-based compression process 100, at least one of the latent and hyperlatent Further transformations can be applied. For example, at least one of the latent and hyperlatent Even if the entropy coding processes 150 and 155 are performed, the value is converted to a residual value. Good. The residual value is the mean of the distribution of latent or hyperlatent values from each latent or hyperlatent value. The value may be determined by subtraction. Alternatively, the residual value of 20 may be normalized.
[0174] To perform the training of the AI-based compression process described above, the input image is trained as described above. A training set may be used. During the training process, the encoder 110 and decoder 120 Both parameters may be updated simultaneously during each training step. If workpiece 105 also exists, then hyper encoder 115 and hyper decoder 125 Both parameters are updated simultaneously in each training step.
[0175] The training process also involves generative adversarial networks. It can include a GAN (Generative Adversarial Network). Suitable for Al-based compression processing. If used, in addition to the compression pipeline described above, an additional neutral network that functions as a discriminator is used. The twerk is included in the system. The classifier receives input and outputs a score based on the input. The classifier then provides an indicator of whether it considers the input to be ground truth or fake. For example, this metric is a score, where a high score is associated with true input and a low score is associated with false input. It relates to the input. For training the classifier, the input ground truth and input fei A loss function is used that maximizes the difference in output display between the two.
[0176] If the GAN is incorporated into the training of the compression process, the output image 6 may be provided to the classifier. The output of the classifier may be used in the compression loss function as a measure of the distortion of the compression process. Alternatively, the classifier receives both input image 5 and output image 6, and compresses the difference between the output images. It can be used as a measure of distortion in the loss function of the compression process. Training of a neural network and other neural networks in compression processing is done simultaneously. It can be performed. A trained compression program for image or video compression and transmission. While using iPline, the discriminator neural network is removed from the system and compressed pi The output of the plane is output image 6.
[0177] When a GAN is incorporated into the training process, the decoder 120 may perform hallucinations. The process involves adding information that was not present in the input image 5 to the output image 6. For example... These are details that were not present in input image 5 or were not received by decoder 120. Details may be added to output image 6. The hallucination being performed is by decoder 120. It may also be based on the quantized latent information received.
[0178] As mentioned above, video is constructed from a series of images arranged sequentially. Even if you apply compression process 100 multiple times to compress, transmit, and decompress the video, For example, each frame of the video may be compressed, transmitted, and decompressed individually. Then, You can group the captured frames to obtain the original video.
[0179] Here, we will explain many of the concepts related to the Al compression process mentioned above. Each concept is explained individually. As explained below, one or more of the desired concepts are related to the above-mentioned Al-based compression process. It may be applied in the following cases.
[0180] Learned quantization Quantization is a crucial step in Al-based compression pipelines. Generally, Childing is achieved by rounding the data to the nearest integer. This is because images and videos Some domains can tolerate a higher level of information loss, while others require finer detail. Therefore, below, we fix the size of the quantum quantization bin to round to the nearest integer. Instead, we will discuss methods that can be used for learning. Hypernetworks, contextual modal This can be achieved by using a `resource` to predict bin sizes from an additional neural network. Several architectures for this purpose are detailed. Also, A with a learned quantization bin size The necessary loss function and quantization procedure for training an l-based compression pipeline Document any necessary changes and use Bayesian programming to control the distribution of bin sizes learned during training. This shows how to introduce the 'Iya'. The learned quantization bins can be used with or without partition quantization. This demonstrates that it can be used. This technological innovation allows strain gradients to be transmitted via the decoder to the hypernet It also becomes possible to flow into twerks. Finally, generalized quantum mechanics improves performance and execution time. I will explain the compression function in detail. Specifically, this technological innovation will improve the compression pipeline. Decoders can now include context models, but arithmetic (or other non- No runtime penalty occurs from repeatedly executing the loss-decryption algorithm. Our method for learning quantization bins involves hyperpliers, autoregressive models, and implicit models. It is compatible with all methods of transmitting metadata, such as [mention specific methods here].
[0181] The following description, however, is not limited to, the use of AI-based image and video compression. The functions, range, and future prospects of learned quantization bins and generalized quantization functions Let's give an overview.
[0182] Compression algorithms can be divided into two phases: an encoding phase and a decoding phase. Yes, it is possible. In the encoding phase, the input data is smaller (by bits) than the original input variable. It is transformed into a latent variable with a different representation. In the decoding phase, the inverse transformation is applied to the latent variable. Then, the original data (or an approximate value of the original data) is restored.
[0183] Al-based compression systems also need to be trained. This is necessary for good compression results. Parameters for an Al-based compression system that achieves small file size and minimal distortion. This is the procedure for selecting. During training, how to adjust the parameters of the Al-based compression system. To determine what to do, parts of the encoding and decoding algorithms are executed. .
[0184] More precisely, in Al-based compression, encoding generally takes the following form:
number
[0185] Here, x is the data to be compressed (image or video), and f enc It is an encoder. It is typically a neural network with trained parameters θ. The encoder is The input data x is converted into a latent representation y in an improved form to further compress it to a lower dimension. Convert.
[0186] To further compress y and send it as a bit stream, established methods such as arithmetic coding are used. Such lossless coding algorithms can be used. In some cases, it is required that y is discrete rather than continuous, and the latent representation is also important. Knowledge of rate distributions may also be necessary. To achieve this, continuous data can be converted to discrete values y( A quantification function Q (usually rounding to the nearest integer) is used to convert to a hat.
[0187] The required probability distribution p(y(hat)) can be obtained by fitting the probability distribution to the latent space. is required. The probability distribution can be learned directly, but in many cases, it has parameters determined by a hypernetwork consisting of a hyper-encoder and a hyper-decoder, which is a parameter distribution. When using a hypernetwork, an additional bitstream ẑ (also called "side information") may be encoded, transmitted, and decoded:
Equation
[0188] The encoding process (using a hypernetwork) is shown in Figure 3. Figure 3 is a schematic illustration of an example of the encoding phase of an AI-based compression algorithm for video and image compression.
[0189] Decoding is performed as follows:
Equation
[0190] Summary: In an arithmetic decoder (or other lossless decoding algorithm), the latent distribution p( ŷ) is used, and the bitstream is converted to the quantized latent ŷ. Then, the function f converts the quantized latent data into a loss reconstruction of the input dec data denoted by x̂. In AI-based compression, f is usually a neural network according to the learned para dec meters θ.
[0191] When using a hypernetwork, first decode the side information bitstream, then To decode the main bitstream, we need to construct p(y(hat)) Used to obtain the necessary parameters. (Decryption process using a hypernetwork) An example is shown in Figure 4. Figure 4 is an example of the decoding phase of an Al-based compression algorithm. It is for video and image compression.
[0192] Al-based compression uses typical optimization techniques with a "loss function" to encode and This depends on learning the parameters of the decoding neural network. The loss function is: To compress images or videos to a small file size while maximizing reconstruction quality: This is chosen to balance the objective. Therefore, the loss function consists of two terms. :
number
[0193] Here, R determines the cost of encoding the quantified latent according to the distribution p(y(hat)). D measures the quality of the reconstructed image, and λ is the ratio of low file size to reconstruction quality. This parameter determines the trade-off between the two. A typical choice in R is cross-entropy. That is
number
[0194] The selection of JPEG2026123108000010.jpg712 is for quantization, and the latent is rounded to the nearest integer, so p(y(hat) The probability distribution of ) is given by the cumulative distribution function from y(hat)-1 / 2 to y(hat)+1 / It is given by the integral of the (unquantized) latent distribution p(y) up to 2, which is cumulative. It is given by the distribution function p(y(hat)) term.
[0195] Function D can be selected as the mean squared error, but MS-SSIM, LPIPS, and / or (when using adversarial neural networks to enforce image quality) It can also be a combination of other metrics of perceived quality, such as adversarial loss.
[0196] When using a hypernetwork, this represents the cost of transmitting additional side information. Additional terms can be added to R:
number
[0197] In summary, the loss function explicitly depends on the choice of quantization scheme through the R term, and y( Note that (t) is implicitly dependent on the choice of quantization scheme.
[0198] Here, the learned quantization bins are used in how they are used in AI-based image and video compression. We will discuss whether it can be used for this purpose. The steps to explain are as follows: • Architecture required for learning and predicting the size of quantization bins • Standard quantization functions and coding / decoding processes for incorporating learned quantization bins. Change of logic • A method for training a neural network using learned quantization bins.
[0199] A key step in a typical AI-based image and video compression pipeline is "quantum This is called "bit rounding," and pixels in the latent representation are usually rounded to the nearest integer. It is necessary for algorithms that encode streams losslessly. However, the quantization step is This can lead to data loss and affect the quality of reconstruction.
[0200] A neural network is used to predict the size of the quantization bin that should be used for each latent pixel. It is possible to improve the quantization function by training the workpiece. Typically, latent y It is rounded to the nearest integer, which corresponds to a "bin size" of 1. That is, a section of length 1. All possible values of y in between are mapped to the same y (hat).
number
[0201] However, this may not be the optimal choice in terms of information loss. Therefore, it is possible to ignore more information without significantly affecting the reconstruction quality. Yes (equivalently: use a bin greater than 1). Also, for other latent pixels, optimal The bottle size is smaller than 1.
[0202] This problem can be solved by predicting the quantization bin size for each image and each pixel. This is a tensor Perform the operation using JPEG2026123108000013.jpg636, and modify the quantization function as follows.
number
[0203] This is called "quantized latent residue". Let's call it JPEG2026123108000015.jpg87. Therefore, Equation 7 becomes as follows:
number
[0204] Figure 6 shows that the size of the learned quantification bins across latent channels relative to a given image. This is a heatmap that shows how it changes. Different pixels are larger, It is predicted that it will benefit from smaller quantization bins, and that larger or smaller information Address the loss.
[0205] The learned quantization bin size is incorporated into the modification of the quantization function Q, so it is encoded Note that any data you want to send can utilize the size of the learned quantization bins. For example, instead of encoding latent y, we want to use the mean-subtracted latent y-μ. y Encode If you want to do this, you can achieve it.
number
[0206] Similarly, hyper-latency, hyper-hyper-latency, and other things we want to quantize The object is a properly learned Δ with all the corrected quantization functions Q Δ Use It is possible.
[0207] Here, we discuss several architectures for predicting the size of quantization bins. Possible architectures use hypernetworks to quantize bin size. The task is to predict Δ. The bitstream is encoded as follows:
number
[0208] Figure 7 shows an example of modified encoding processing using a hypernetwork, and hyperpliers Encoding phase of an Al-based compression algorithm using learned quantization bin sizes This is a schematic diagram illustrating an example for video and image compression.
[0209] During decoding, the bitstream is decoded in the usual lossless manner. Then, two values are obtained using Δ. Rescale by multiplying the elements ξ y Perform this conversion. Let the result be y (hat). The data is then passed to the decoder network as usual.
number
[0210] Figure 8 shows an example of a modified decoding process using a hypernetwork. Decoding of an Al-based compression algorithm using Lyra and learned quantization bin size This is a schematic diagram illustrating an example of ASE, used for video and image compression.
[0211] Applying the above technology improves the rate of Al-based compression pipelines by 1.5%, and M Distortion may improve by 1.9% when measured with SE. Therefore, Al-based Compression performance will be improved.
[0212] Several variations of the above architecture will be explained in detail. The size of the hyperlatency quantization bin can also be the learned parameter. It is also possible to include hyper-hypernetworks that predict these variables. • Predicting the quantized bin size improves the prediction for a particular pixel using neighboring pixels. This can be improved by using the so-called "context module" feature. • Quantization bin size Δ from hyperdecoder y After obtaining this tensor, further delinear This can be processed using a shape function. This nonlinear function is generally used in neural networks. (However, this is not the only option.)
[0213] Furthermore, our method for learning quantization bins is hyperplier, hyperhyper Compatible with all methods of conveying metadata, including pliers, autoregressive models, and implicit models. Emphasize that it exists.
[0214] To train a neural network with quantization bins learned for AI-based compression Therefore, we can change the loss function. Specifically, the data described in equation 5a The cost of encoding can be changed as follows:
number
[0215] Instead of integrating over an interval of length 1, From JPEG2026123108000023.jpg616 This means that we need to integrate the probability distribution of the materials up to JPEG2026123108000024.jpg516.
[0216] Similarly, when using hypernetworks, the coding costs of hypernetworks must be addressed. Items to do JPEG2026123108000025.jpg537 incorporates the learned quantization bin size, so the encoding cost is exactly the same as Latento's. It will be corrected to "uni".
[0217] Neural networks typically update the parameters of the trained network in order to It is trained by a variation of gradient descent using backpropagation. For this purpose, We need to calculate the gradients of all layers in the network, and to do that, the network The layers must be constructed with differentiable functions. However, the quantization function Q and its learning bin Modification Q Δ It is not differentiable because a rounding function exists. In Al-based compression, Q( To replace y), two differentiable approximations for quantization are used during network training. One of these is used (when the network is trained and used for inference, the approximation is not used). ).
number
[0218] When using the learned quantization bins, the approximation of the training quantization is as follows:
number
[0219] Instead of selecting one of these differentiable approximations during training, Al-based compressed pi When calculating rate loss R, the plan We will use JPEG2026123108000028.jpg613, but the decoder will be used for training. It can also be trained using "partition quantization" by sending JPEG2026123108000029.jpg612. The model can be trained using both partitioned quantization and learned quantization bins.
[0220] First, note that there are two types of partitioned quantization. • Hard partition quantization: Decoder Receive JPEG2026123108000030.jpg612 • Soft partition quantization: The decoder is in the forward pass. I receive JPEG2026123108000031.jpg612, When calculating JPEG2026123108000032.jpg77, in the back pass of backpropagation Use JPEG2026123108000033.jpg613.
[0221] Note that in integer rounding quantization, hard partitioning and soft partitioning quantization are equivalent.
number
[0222] However, when using learned quantization bins, the back pass will be as follows, The quantization of double split and soft split is not equivalent.
number
[0223] In all quantization schemes, the rate gradient wrtΔ is negative.
number
[0224] In any quantization scheme, the rate term is Since it receives JPEG2026123108000037.jpg613, The filename becomes JPEG2026123108000038.jpg712. Next, always Since the file is JPEG2026123108000039.jpg518, we obtain the following:
number
[0225] The strain gradient wrtΔ differs depending on the quantization scheme. The gradients are as follows:
number
[0226] In soft partitioning quantization, JPEG2026123108000042.jpg812 and This results in JPEG2026123108000043.jpg631. This is because the random noise ε in the back pass is in the front pass, This is because it does not rely on STE rounding when performing JPEG2026123108000044.jpg66. This means the following:
number
[0227] Therefore, in soft partition quantization, the rate gradient increases Δ, and the strain gradient is 0 on average. Therefore, the overall Δ→∞, making it impossible to train the network.
[0228] Conversely, in hard partitioning quantization... JPEG2026123108000046.jpg66 is Since it does not depend on JPEG2026123108000047.jpg612, it will be as follows.
number
[0229] In summary, when using partitioned quantization with learned quantization bins, soft partitioned quantization is not the way to go. We will use hard partition quantization.
[0230] A non-trivial strain gradient that can be achieved with or without partition quantization is obtained when the strain gradient passes through the decoder. This means that it flows into a hypernetwork. This is not possible with models that have [a certain characteristic], but our method of learning the size of the quantization bins allows us to [perform a specific action]. This is a feature that has been introduced.
[0231] Some compression pipelines control the distribution of values learned for the quantization bin size. It is important to do this (though it is not always necessary). Add it to the loss function if necessary. This is achieved by introducing a term.
number
[0232] F Δ is distribution p Δ Characterized by the choice of (Δ), this is a term used in Bayesian statistics. Therefore, we call it a "plier" for Δ. There are several options for the plier. • Any parametric probability distribution on positive numbers. Specifically, the inverse gamma distribution, and how many of them An example is shown in Figure 9. The distribution is shown for several values of the parameter. Parameters can be fitted from other models, selected a priori, or learned during training. It can also be learned. • A neural network that learns a prior distribution during training.
[0233] The previous section detailed the simplified quantization function Q that utilizes a tensor of bin size Δ. I gave a detailed explanation.
number
[0234] All of these methods can be extended to more generalized quantization functions. In this case, Q is some invertible function of y and Δ. Then the coding is given by the following equation:
number
number
[0235] This quantization function is more flexible than Equation 24, resulting in improved performance. The quantization function incorporates quantization parameters, for example, when using an autoregressive context model. By including context, we can also enable the system to recognize the context.
[0236] All methods in the previous section are compatible with the framework of generalized quantization functions: • Hypernetworks, context modules, and implicit equations if necessary Predict essential parameters. • Appropriately adjust the loss function. • When using split quantization, hard split quantization should be employed. • If necessary, to control the behavior of the loss function, apply Bayesian pliers to the loss function. We will implement this.
[0237] More flexible generalized quantization functions improve performance. In addition to this, generalized quantities The quantization function can depend on parameters that are determined autoregressively, which is quantization. This means it relies on pixels that have already been encoded / decoded.
number
number
[0238] Generally, using autoregressive context models improves the performance of AI-based compression. ru.
[0239] Autoregressive generalized quantization functions are also beneficial from a runtime perspective, such as PixelCNN. Other standard autoregression models use a context model to decode pixels. Each time, an arithmetic decoder (or other lossless decoder) needs to be run. In real-world applications of image and video compression, serious performance issues arise. This will be a blow to the generalized quantization function framework, such as PixelCNN. Integrating an autoregressive context model into AI-based compression without causing runtime problems. It can be included. This is Q -1 This is the ξ(hat) that has been completely decoded from the bitstream. This is because it acts automatically and regressively on ). Therefore, the arithmetic decoder is executed automatically and regressively. It is not necessary, and the execution time issue will be resolved.
[0240] A generalized quantization function can be any invertible function. For example, • Invertible rational function: Q(·,Δ) = P(·,Δ) / Q(·,Δ), where P and Q are polynomials. formula. • Logarithmic, exponential, and trigonometric functions, with appropriate domain control so that these functions are invertible. Limited edition. • Invertible function of matrix Q(·,Δ) • A reversible function including the context parameter L predicted by the hyperdecoder: Q(·,Δ,L)
[0241] Furthermore, Q does not generally need to be closed or reversible. For example, Q en c (·,Δ) and Q dec (·,Δ) can be defined, and these functions are not necessarily related to each other. It doesn't need to be the inverse function of , and the entire pipeline can be trained end-to-end. In this case, Q enc and Q dec This is a neural network, or Gaussian process, or probabilistic graph. Modeling as a specific process, such as a physical model (a simple example being a Hidden Markov Model). It is possible.
[0242] To train an Al-based compression pipeline using a generalized quantization function, use the same method as described above. Use many tools. • Appropriately adjust the loss function (specifically the rate term). • When using partitioned quantization, use hard partitioned quantization. • Bayesian pliers and / or regularization terms (l1, l2, second moment penalty) Introduce the distribution parameter Δ, and if necessary, the context Control the distribution of the stress parameter L.
[0243] Depending on the choice of generalized quantization function, other tools may be used to train an Al-based compression pipeline. This will be necessary. Reinforcement learning (Q-Learning, Monte Carlo Estimati) Techniques from DQN, PPO, SAC, DDPG, TD3) • General proximity gradient method • Continuous relaxation. Here, discrete quantification residues are approximated by a continuous function. This function is related to It can have hyperparameters to control the smoothness of the numbers, and the network training Furthermore, modifications are made at various points in the training to improve the final performance.
[0244] There are several context modeling approaches that are compatible with the generalized quantization function framework. There is a possibility that... Q is an autoregressive neural network. The most common example that can be adapted to Q is P It's an ixelCNN-style network, but it uses Resnets and Transform ers, Recurrent Neural Networks, Fully-Conn Other neural network building blocks, such as ected networks, are also available. It can be used for the Q-function of the self-regression type. ·Δ can be predicted as a linear combination of previously decoded pixels.
number
[0245] For example, if important metadata is obtained from the attention mechanism / focus mask, this can be used for Δ prediction. It can be incorporated into this. In this case, the size of the quantization bin is the sensitive area of the image and video. To adapt more precisely to the region, knowledge of sensitive areas is stored in this metadata. In this way, Less information is lost from perceptually important areas, while information from unimportant areas is lost. Because it is ignored, compared to Al-based compression pipelines that do not have adaptable bin sizes. And performance is improved in a more enhanced way.
[0246] Furthermore, we will outline the relationship between the learned quantization bins and the variability rate model. One form of Dell uses a free hyperparameter δ to control the bin size, Al Train the base compression pipeline. During inference, δ is sent as metadata, and the transmission rate Controls the cost (cost per bit).
[0247] In a variable rate framework, δ means that the size of all bins is controlled simultaneously. Taste is a global parameter. Our technological innovation predicts the quality of each pixel. We locally determine the size of the tensor Δ. In addition, we use δ to control the transmission rate. The rate model uses global predictions as needed to control the rate during interpolation. This allows us to scale local predictions element by element, which is compatible with our framework. It is interchangeable.
number
[0248] Dataset bias This section describes the training procedure applied to the generative adversarial network framework. This will be explained in detail. This approach allows for generative processing of any type of image data. The compression model can be biased, and the resulting image quality depends on the subject matter in the image. It can be controlled.
[0249] Generalized Adversarial Networks (GANs) are used in various different domains such as images, video, and audio. This approach shows excellent results when applied to the task of production. The reviewers pitted their two models against each other, and as a result, both were made stronger, using game theory. This is the inspiration. The first GAN model takes a noise variable z as input and a synthetic data sample The second model is a generator G that outputs a hat x, and the second model is a sun from the real data distribution. A discriminator D is trained to distinguish between data generated by a pull and a generator. An example of a GAN architecture is shown in Figure 10.
[0250] P x The data distribution on the actual sample x, P z The data distribution on the noise sample z, P g to data Let this be the generator distribution on x.
[0251] The training of the GAN is shown as a minimax game in which the following function is optimized.
number
[0252] Applying a generative adversarial approach to the image compression task, Let's start by considering JPEG2026123108000058.jpg519, where C is the number of channels and H and W are the height and width in pixels.
[0253] The autoencoder-based compression encoder pipeline is constructed from the following: The da function f θ (x)=y represents the latent representation of image x. Encode it as JPEG2026123108000059.jpg721. Q is the quantization function required to transmit y as a bitstream. Decoder function JPEG2026123108000060.jpg616 is a reconstructed image of the quantized latent y (hat). This function decodes the image into JPEG2026123108000061.jpg519.
number
[0254] In this case, encoder f θ , quantization function Q and decoder g θ The combination is, It can be thought of as a network. For simplicity of notation, this generative network is called G This is denoted as (x). This generative network is complemented by the discriminative network D, and is a two-stage process. Train the generative network together with the law.
number
[0255] An example of a standard generative adversarial compression pipeline is shown in Figure 11. Next, a standard compression pipeline Training is performed using the T-strain loss function.
number
[0256] By complementing this trained compression network with a classifier model, the perceived quality of the output image can be improved. This can improve the performance. In this case, the compression encoder decoder network generates the network. It can be considered a twerk, and the two models are trained using a two-level approach in each iteration. This is possible. The discriminator architecture has conditions for generating higher quality reconstructed images. We decided to use a classifier with a hat. In this case, the classifier d(x,y(hat)) is quantized. It is conditioned on the given latent y (hat). First, the classifier is trained with the classifier loss.
number
[0257] (32) To train the generative network, adversarial agents used to train the GAN generator The rate distortion loss in (36) is compensated for by adding a "non-saturated" loss.
number
[0258] By adding adversarial loss to rate distortion loss, the network can reproduce natural patterns and textures. Prompt the system to generate a signal. Use an architecture that combines GANs and autoencoders. This significantly improves the perceptual quality of the reconstructed image and yields excellent results in image compression. We were able to do that. However, even though the overall result of this architecture is great, Nevertheless, there are many noteworthy failure modes. These models are based on human faces and text. We are struggling to compress areas of high visual importance, including but not limited to strikes. This has been observed. An example of such a failure mode is shown in Figure 12. The image on the left is the original image. The image on the right is a reconstructed image synthesized using a generative compression network. It's important to note that what's happening is in the human faces present in the image, that is, the parts with a lot of visual information. Please be aware of this. To address this problem, we will bias the model towards certain types of images, such as faces. We propose a method that can improve the perceived quality of reconstructed images.
[0259] Under this framework, we apply different classifiers to multiple datasets. We will use this to train the network. First, we will use N additional datasets X1 that are biased towards the model. ,…X N Select this. A good example of such a dataset useful for face modeling is a person. This is likely a dataset consisting of portraits. Each dataset X i For identification mode RuD i We will introduce each classifier model D. i Dataset X i It is trained using only data, The encoder-decoder model is trained on images from the entire dataset.
number
[0260] Figure 13 shows the compressed pipeline trained by a multiple classifier cGAN using dataset bias. The figure shows the dataset X. i Image x i It passes through the encoder, is digitized, and is rangecode It is converted into a bitstream using a decoder. Then the decoder decodes the bitstream. Transform, x i We attempt to reconstruct the (hat). Then, the original image x and the reconstructed x (hat) are compared. This is passed to the classifier corresponding to dataset i.
[0261] For illustrative purposes, we will focus on facial failure modes, as shown in Figure 12. The technique described here... All of these are any number of regions of interest in the image (e.g., letters, animal faces, eyes, lips, logos, cars, etc.) Please note that this is applicable to flowers, patterns, etc.
[0262] For example, biasing a dataset by using only one additional dataset. Let's consider this. In this case, X1- is a general training dataset, and X2- is only portrait images. This is a dataset that includes [dataset name]. The case when using a multi-discrimination dataset bias scheme and the case when using [dataset name]. Figure 14 shows a comparison of the reconstruction of the same generative model trained at the same bitrate, but without this step. This shows that the image on the left is a reconstruction of an image synthesized by a standard generative compression network, and the image on the right is... The images were generated using the same generative network trained with a bias on a multi-identifier dataset. Therefore, it is a reconstruction of the same composite image. Using this method, the remaining part of the image The perceived quality of human faces is improved without compromising quality.
[0263] [Table 1]
[0264] [Table 2]
[0265] The above approach is an architecture that uses a single classifier for all datasets. It can also be used. Furthermore, classifier D for specific datasets i More frequently than generators It can be trained to increase the effect of bias on that dataset.
[0266] Given the above generation and compression network, here the image or frame A change in architecture that allows for higher bit allocation, conditional on the text Define. Enhance the effect of dataset bias and differentiate different regions of the image depending on the subject of the image. To alter the perceived quality, Lagrangian coefficients are used to control the bitrate in the dataset. Tox i We propose a different training procedure for each. The generator modifies the loss function of (37) as follows. do.
number
[0267] This approach uses a higher percentage of the bitstream for the face region of the compressed image. Train the model to assign biases to bitrate-adjusted datasets. The results can be observed in Figure 15. Figure 15 shows the same results for the face dataset. While maintaining JPEG2026123108000071.jpg59, the background dataset is different. This shows three images synthesized by a model trained on JPEG2026123108000072.jpg59. The image on the left is background data. The set has a low bitrate, the center image has a medium bitrate, and the right image has a high bitrate. Trained by rate.
[0268] Extending the method proposed above, we use different bias functions d for each dataset used for biasing. I propose using (x, x (hat)). This method allows us to focus the model. It can be adapted to each specific type of data. For example, as a distortion function, MS A linear combination of E, LPIPS, and MS-SSIM metrics can be used.
number
[0269] Coefficients of different components of the distortion function By modifying JPEG2026123108000074.jpg638, the generative compression model can reconstruct different regions of the image in different ways. This allows us to change the perceived quality of the resulting image. Then, equation 37 is given by each Dataset X i To index the distortion function d(x, x (hat)) for a given value. It can be corrected by [method].
number
[0270] Here, in order to bias the generation and compression model towards particularly important regions in the image, salience (Note (i) Discuss the use of masks. These masks are used in conjunction with another pre-trained network. It is generated using the workpiece, and its output is used to further improve the performance of the compression model. Cut.
[0271] Take image x as input and a binary mask Consider network H that outputs JPEG2026123108000076.jpg517. Let's start from there.
number
[0272] X i Prominent pixels are indicated by 1 in m, and 0 indicates areas that the network does not need to pay attention to. This indicates the region. This binary mask further biases the compression network towards that region. It can be used. Examples of such important areas include the eyes and lips, which are features of the human face. These are some examples, but are not limited to them. Given m, these regions are prioritized. The input image x can be modified as shown. Modified image x H Adversarial compression network It is used as input to the compression pipeline. An example of such a compression pipeline is shown in Figure 16.
[0273] Extending the approach proposed above, Salience applies a bias to the compression pipeline. We propose an architecture that utilizes a pre-trained network to generate masks. This approach allows you to change the mask without retraining the compression network. This allows you to change the bitrate allocation to different parts of the reconstructed image. In the example, the mask m in equation 41 is used as an additional input, and m is used as the salient (1). Train the network to allocate more bits to the worked area. During the inference phase after the model has been trained, the bit allocation is adjusted by changing the mask m. Yes, it is possible. An example of such a compression pipeline is shown in Figure 17.
[0274] We further developed a training scheme that ensures the model is exposed to a wide range of natural image examples. I propose that the training dataset is constructed from images from N different classes, and each image is They are labeled accordingly. During training, images are sampled from the dataset according to their class. The model is sampled. By sampling images equally from each class, the model is sufficiently representative. It allows you to encounter classes that haven't been explored before and learn the entire distribution of natural images.
[0275] Area emphasis Modern trained image compression methods, such as VAE and GAN-based architectures, , excellent compression at low bitrates and a significant improvement in the perceptual quality of the reconstructed image. It makes possible. However, the overall result of such an architecture is great, However, as mentioned above, there are many notable failure modes. These models are visual It has been observed that it is difficult to compress the areas of high importance, and this includes the human face. Or it may include text, but is not limited to these. Figure 1 shows an example of such a failure mode. As shown in 2, the image on the left is the original image, and the image on the right is the composite image created using a generative compression network. The reconstructed image is shown. The most distorted areas are the human faces present in the image. Note that this is a visually important part.
[0276] By changing the quantization bin within the ROI, more bits can be allocated to the bitstream. We propose an approach that improves the perceived quality of the region of interest (ROI) by focusing on specific areas.
[0277] To encode the latent y into a bitstream, we first need to know that it is discrete. To guarantee this, it can be quantized. Using the quantization parameter Δ, the domain is divided into sections. We propose controlling the assigned BPP. Δ is the size of the quantization bin or the interval between quantizations. This is a gap, representing the coarseness of quantization in latent and hyperlatent space. The coarser the quantization, the higher the degree of coarseness. The number of bits allocated to the data will decrease.
[0278] The quantization of latent y is achieved as follows:
number
[0279] We propose using spatially varying delta to control the coarseness of quantization within an image. This allows us to control the number of bits allocated and the visual quality of different areas of the image. It can be controlled. This proceeds as follows:
[0280] Let's consider the function H that detects the region of interest, which is typically represented in neural networks. Let's begin. The function H(s) takes image x as input and a binary mask Outputs JPEG2026123108000079.jpg517. A value of 1 for m indicates that the corresponding pixel of image x is within the region of interest, and zero means It corresponds to the pixels outside of it.
[0281] In one example, the network H(x) is trained before the compression pipeline is trained, and in another example... It is trained in combination with an encoder-decoder. Map m is quantized for each pixel. Used to create a quantization map Δ to which a parameter is assigned. If the value in m is 1, the corresponding value in Δ is small. The function Q defined in Equation 42 is as follows: Before encoding into a bitstream, y is quantized to y (hat) using a spatial map Δ. As a result of such a quantization scheme, the bit of the region of interest is compared to the rest of the image. The rate will increase.
[0282] The proposed pipeline is shown in Figure 18. Figure 18 depends on the binary mask of ROI m.. This diagram illustrates a proposed compression pipeline having quantization that utilizes a quantization map Δ. Such pipelines are used for identifying regions of interest, such as face detectors and network H( The results of the implementation using (x) are shown in Figure 19. The image on the left shows the ROI in the quantification map Δ. This is a compressed image synthesized with faces assigned as such, and the image on the right is a standard generated compression model. This is a compressed image by Dell.
[0283] In another example, for the region identified by ROI detection network H(x) from equation (41) , different quantization functions Q m This can be used. An example of such an arrangement is shown in Figure 20. Figure 20 shows the quantization function Q for the entire image and the quantization function Q for the region of interest. m Using A diagram illustrating the proposed compression pipeline is shown.
[0284] Discrete PMFS A story of discrete symbols (such as latent pixels in an Al-based compression pipeline) To encode and decode the data into a binary bitstream, a discrete probability mass function (PM) is used. Access to F) may be required. However, in an Al-based compression pipeline... Training such discrete PMFs is an approach that allows training to access continuous probability distribution functions (PDFs). It is widely believed that this is impossible because it requires a certain amount of energy. Therefore, Al-based compression is not possible. In pipeline training, sequential PDFs are trained, and only after training is complete are the quantization point separations performed. Approximating continuous PDFs with discrete PMFs evaluated in a discrete manner has become the de facto standard. It is.
[0285] The following describes the reverse procedure: interpolating discrete PMFs into a continuous real-valued space. This allows for direct training of Al-based compression pipelines with discrete PMFs. Diverse PMFs can be learned or predicted, and can also be parameterized.
[0286] The following explanation is limited to discrete probability mass functions and Al-based image and video compression. This document outlines the functionality, scope, and future prospects of interpolation for use without modification. Below, we will discuss high-level explanations, inferences, and AI-based compression algorithms for discrete probability mass functions. This section describes training and interpolation methods for functions (such as discrete probability mass functions).
[0287] In the literature on Al-based compression, a standard approach for constructing an entropy model is... This is the continuous probability density function (PDF) p y (y) From (Laplace distribution or Gaussian distribution, etc.) The first step is to begin. Shannon entropy is a discrete variable y (hat) (usually Since it is defined only in JPEG2026123108000080.jpg69), this PDF is a discrete probability mass function (PMF) Convert to JPEG2026123108000081.jpg610 for use with, for example, a reversible arithmetic encoder / decoder. This is because y (hat) is in the middle This can be done by collecting all the (continuous) mass in a central (unit) bottle. .
number
[0288] This approach was developed by Johannes Balle, Valero Laparra, and E. First proposed by Ero P. Simoncelli. End-to-end This is optimized image compression. 5th International Conference ce on Learning Representations, ICLR 201 7, Toulon, France, April 24-26, 2017, Co OpenReview in the inference Track Proceedings .NET, 2017, is incorporated here by reference. This function is an integer. It is not only defined for JPEG2026123108000083.jpg69, but also accepts any real number argument. This is an encoder. A PDF for the latent values of continuous real numbers output by the system is needed during training. This is convenient. Therefore, the "new" function JPEG2026123108000084.jpg523 is defined. By definition, this is a PMF defined on an integer. This function is an exact match to JPEG2026123108000085.jpg55. JPEG2026123108000086.jpg56 is a PMF (Product-Market Fitting) on which an end-to-end AI-based compression algorithm is actually trained.
[0289] In summary, continuous real-valued PDF JPEG2026123108000087.jpg617 is a discrete PMF. Converted to JPEG2026123108000088.jpg517, then during training JPEG2026123108000089.jpg513 contains a continuous PDF It is evaluated as JPEG2026123108000090.jpg56. Figure 21 shows three used to train the Al-based compression pipeline. This shows an example of a typical one-dimensional distribution. In reality, the original PDF p y During training, reasoning is also performed. However, it is never used explicitly. In practice, it is used as follows: JPEG2026123108000091.jpg56 (used for inference) and There are only two functions: JPEG2026123108000092.jpg56 (used for training).
[0290] This thinking model can also be reversed. Instead of starting with PDFs, start with discrete PMFs. By interpolating the PMF, we can recover a continuous PDF (which is used only in training). It is possible to do so.
[0291] PMF Suppose the image JPEG2026123108000093.jpg56 is given. This PMF consists of two vectors of length N, i.e., y i (hat ) and p i It can be expressed using a hat symbol, where i=1... and N represents discrete points. In this way of thinking (when PMF is defined as a function), It is defined as JPEG2026123108000094.jpg517. However, generally speaking... JPEG2026123108000095.jpg58 can be any non-negative vector whose sum is 1. i The results are sorted in ascending order. It does not necessarily have to be an integer value.
[0292] Here is the query point Suppose the image JPEG2026123108000096.jpg619 is given. The query points must be enclosed by discrete extrema. (Approximate) Training PDF To define JPEG2026123108000097.jpg58, an interpolation routine is used.
number
[0293] The following is a non-exhaustive list of possible interpolation routines: • Piecewise constant interpolation (nearest neighbor interpolation) Linear interpolation Polynomial interpolation • Spline interpolation (e.g., tertiary interpolation) • Gaussian process / Kriging
[0294] In general, functions defined by interpolation are not necessarily PDFs. Depending on the method, the interpolated value may be negative or may not have a unit mass. These problems can be mitigated by selecting appropriate routines. For example, Piecewise linear interpolation preserves mass, preserves positivity, and the interpolated function is actually PDF. This guarantees that. The tertiary Hermitian interpolation method is based on Randall L. Dougherty's work. Discussed by Alan S. Edelman and James M. Hyman Thus, if the interpolation point itself is positive, it can be constrained to be positive. Non-negativity, monotonicity. Cubic and quintic Hermitian interpolation that preserves convexity. f Computation, 52(186):471-494, 1989, this is a reference. This is incorporated herein by means of.
[0295] However, piecewise linear interpolation has another problem. Its derivative is piecewise constant, for example, as shown in Figure As shown in the left figure of 22, interpolation errors can be quite severe. For example, if PMF is When generated using Balle's method, the interpolation error is It is defined as JPEG2026123108000099.jpg523. Other interpolation schemes, such as piecewise cubic Hermitian interpolation, are also used. The scheme results in smaller interpolation errors, as shown in the right-hand diagram of Figure 22, for example.
[0296] Discrete PMFs can be directly trained with AI-based image and compression algorithms. During training, the latent probability values of the real numbers output from the encoder are used as discrete values of the PMF model. Interpolation is performed using this method. In training, the gradient is calculated from the rate (bitstream size) loss using PMF. The model parameters are then used to learn PMF (Product-Market Fit).
[0297] PMF models can be learned or predictive. Learning means that P This means that the MF model and its hyperparameters are independent of the input image. This means that the PMF model is conditionally dependent on "side information" such as hyperpotential. This means that in this scenario, the PMF parameters are predicted by the hyperdecoder. Furthermore, PMF can also conditionally depend on adjacent latent pixels. (In this case, PMF is called a discrete PMF context model). Regardless, during training, the PMF values are interpolated to provide an estimate of the probability values of real-valued (non-quantized) points. This can then be input into the rate loss of the training objective function.
[0298] The PMF model can be parameterized in one of the following ways (however, this The list is not exhaustive. • PMF is a categorical distribution, and the probability values of the categorical distribution are the finite quantization points on the solid line. It corresponds to. The probability values of a categorical distribution correspond to a finite number of quantization points on the solid line. The value of can be parameterized by a vector, and that vector is mapped to a probabilistic simplex. The projection is cast. This projection may be a softmax style projection, or a probabilistic simplex projection. Other projections onto S are also acceptable. PMF can be parameterized by several parameters. For example, PMF If it is defined for N points, then to control the PMF value, n (n <N)のパラ A meter can be used. For example, PMF can be controlled by mean and scale parameters. This can be done, for example, by quantizing the mass of a continuous one-dimensional distribution using discrete quantization bins. It can be done by gathering people together. • PMF can be multivariate, in which case PMF is a multidimensional set of quantization points. It is defined over a wide range of periods. • For any of the previous items, the quantization point is, for example, a number of integer values. Or it can be any interval. The interval between quantization bins is an auxiliary net such as a hyperdecoder. It can also be predicted by the work, or from the context (adjacent latent pixels) It can also be predicted. • For any of the previous items, the parameters that control PMF are fixed or variable. Predicted by auxiliary networks such as the PER decoder, or context (neighbors) It is possible to predict from latent pixels.
[0299] This framework can be extended in several ways. For example, if discrete PMFs are numerous... If it is a variable (multidimensional), a multivariate (multidimensional) interpolation scheme is used to realize the PMF value. It is possible to interpolate to vector values. For example, multilinear interpolation (bilinear in 2D, 3D). In 2D, multicubic interpolation (such as trilinear interpolation) can be used. You can also use icubic interpolation (or tricubic interpolation in 3D). .
[0300] This interpolation method is not limited to modeling only discrete PMF values. Al-based compression. Any discrete-value function can be interpolated at any point in the pipeline, as described here. The technology is not strictly limited to modeling the probability mass / density function.
[0301] Context Model In Al-based compression, the autoregressive context model is a powerful entropy modeling machine. It has the capability, but because it must be executed in series, it has the problem of having a very short execution time. Yes. In this paper, we predict the autoregressive model components from the hyperdecoder (and these This explains how to overcome this difficulty by conditioning the components with "side" information. A more advanced modeling system with real-time autoregressive capabilities. The system is obtained. This real-time capability requires an autoregressive system with a lossless decoder. This is achieved by decoupling from the model. Instead, the autoregressive system does not decode at all. This involves solving a quadratic equation, which can be solved extremely quickly using numerical linear algebra techniques. Encoding can also be performed quickly by solving simple implicit equations.
[0302] This document is limited to AI and deep learning-based image and video data compression. Without doing so, the current and future use of autoregressive stochastic models with linear decoding systems This document outlines its functions and scope.
[0303] In AI-based image and video compression, the input image x is mapped to a latent variable y. This latent variable is encoded into a bitstream, and the bitstream is decoded back into the latent variable. It is then set in the receiver. The receiver then uses the recovered latent variables to reconstruct the representation of the original image. Convert to (build) x (hat).
[0304] To perform the step of converting the latent to a bitstream, the latent is given an integer value representation y( It can be quantized into a (hat). This quantized latent y (hat) is arithmetic encoding. Lossless encoder / decoder schemes such as D / decoder or range encoder / decoder It is converted to a bitstream via [a specific method / platform].
[0305] In lossless coding / decoding schemes, for each element of the latent quantization variables, the model's one-dimensionality is used. A scattered probability mass function (PMF) may be required. The optimal bitstream length (phi The size of this model is achieved when the PMF matches the true one-dimensional data distribution of the latent data. It can be done.
[0306] Therefore, file size is closely related to the power with which the model PMF matches the true data distribution. This is linked to the fact that the more powerful PMF model reduces file size and improves performance. This results in compression. In some cases, this can lead to better reconstruction errors (given file) (Because more information can be sent relative to the size in order to reconstruct the original image.) Therefore, in the development of a powerful model PMF (often called the entropy model) A lot of effort has been put in.
[0307] A typical approach to modeling one-dimensional PMF in Al-based compression is... Lametric one-dimensional distribution, The image used is JPEG2026123108000100.jpg519, where θ is a parameter of the one-dimensional PMF. For example, quantization You can use a quantized Laplacian or a quantized Gaussian. In this example, θ constructs the position and scale of the distribution. For example, quantized When Gauss (Laplacian) is used, PMF is expressed as follows:
number
[0308] Parameters such as position μ or scale σ are stored in the bitstream. By "conditioning" the information, we can create more powerful models. In other words, If the PMF parameters are constant across all inputs of the Al-based compression system Rather than being statically fixed, the parameters can respond dynamically to the input. Cut.
[0309] This is generally done in two ways. The first is to add an extra side y (hat) This is a method for sending information z (hat) to the bitstream. This variable z (hat) The (t) variable is often called the hyperlatency. This variable is used before decoding y (hat). Since the entire expression is decoded, it can be used for encoding / decoding y (hat). Next Furthermore, μ and σ can be expressed as functions of z (hat), for example, in a neural network. It returns μ and σ through this. Then, the one-dimensional PMF is conditional on z (hat) I was told, Provided by JPEG2026123108000102.jpg532.
[0310] Another approach is to use an autoregressive stochastic model. For example, Aaron v an den Oord,Nal Kalchbrenner,Lasse Espeh olt, Koray Kavukcuoglu, Oriol Vinyals, and Alex Graves has described PixelCNN. Conditional image generation using a CODA. In Daniel D. Lee, Masashi S. ugiyama,Ulrike von Luxburg,Isabelle Guyo n,and Roman Garnett,editors,Advances in Neural Information Processing Systems 29 :Annual Conference on Neural Information Processing Systems 2016,December 5-10,2 016,Barcelona,Spain,pages 4790-4798,2016 ,which is herein incorporated by refernc e,David Minnen,Johannes Balle,and Joint autoregressive and hierarchical priors f or as done in learned image compression, It is widely used in academic papers on l-based compression. (Samy Bengio, Hanna) M. Wallach, Hugo Larochelle, Kristen Graum an,Nicolo Cesa-Bianchi,and Roman Garnett ,editors,Advances in Neural Information Processing Systems 31:Annual Conference on Neural Information Processing Systems 2018,NeurIPS 2018,December 3-8,2018,Mon This is described in treal, Canada, pages 10794-10803, 2018. In this framework, the context pixel is the P of the current pixel. These are used to condition the position μ and scale σ parameters of the MF. A text pixel is a previously decoded pixel adjacent to the current pixel. Example For example, suppose the previous k pixels have been decoded. Because images have their own unique spatial correlations, These pixels contain relevant information about the currently active pixel. There are many. Therefore, these context pixels are used to determine the current pixel position and The scale prediction can be improved. The current PMF of the pixel is The image is given in JPEG2026123108000103.jpg569, where μ and σ are functions of the previous k variables (usually a convolutional neural network). )
[0311] These two approaches involve conditions using hyperlatent or autoregressive context models. Both options have their advantages and disadvantages.
[0312] One of the main advantages of hyperlatency conditioning is that quantization is not affected by location shift. This means that the quantization bin can be centered around the position parameter μ. Using integer bins, the quantization potential is as follows:
number
[0313] The main advantage of autoregressive context models is that they can utilize contextual information (decoded neighboring pixels). This is because images (and videos) are highly spatially correlated, and these adjacent pins Xel provides extremely accurate and precise predictions about what a pixel should look like at present. It is possible. Most state-of-the-art academic Al-based compression pipelines use bits. Its impressive performance when measured in terms of stream length and reconstruction error. The results indicate that an autoregressive context model is being used. However, its impressive relative Despite its performance, it's plagued by two problems.
[0314] In other words, the PMF of the current pixel depends on all previously decoded pixels. Furthermore, the position function μ(·) and the scale function σ(·) are usually large neural networks. These two facts mean that autoregressive context models can be run in real time. This means that there are no computational requirements that are greater than those required for real-time performance on edge devices. It takes a lot of time. Therefore, currently, autoregressive context models are impressively pressured. Despite the fact that it offers improved performance, it is not commercially viable.
[0315] Secondly, due to the effects of cascade errors, the autoregressive context model is rounded straight. You must use JPEG2026123108000107.jpg515. Position shift rounding JPEG2026123108000108.jpg527 is impossible because a small floating-point error introduced early in the decoding path During the serial decoding path, the predictions are amplified and expanded, resulting in significant differences between the encoder and decoder. This is because the lack of position shift rounding is a problem, and all other construction requirements If the primes are the same, an autoregressive model with position shift rounding (if it can be constructed) is... This is considered superior to the rate-rounded autoregressive model.
[0316] Therefore, the advantages of conditioning for hyperlatencies (fast execution time, position shift rounding) And the excellent performance of the autoregressive model (from pre-decoded context pixels) It is necessary to develop a PMF modeling framework that combines (creating powerful predictions). be.
[0317] The following describes how to use a hyperdecoder to predict additional parameters of an autoregressive model. This is a technique to modify the hyperdecoder in order to achieve this. In other words, autoregressive model Dell's parameters are conditioned on the hyperlatent z (hat). This means that the autoregressive function is static. It is constant and invariant, and does not change depending on the input of the compression pipeline, for autoregressive modeling. This is in contrast to the standard setup.
[0318] Here, we mainly use the following quasi-linear setup (where the decoding path is linear, but the encoding is linear). It deals with (not) the hyperdecoder, in addition to μ and σ predictions, also called the context matrix. A sparse matrix L can be output. This sparse matrix is PMF as follows: Used as an element of autoregressive context modeling. Latent pixel order (raster Given a can order, etc., the previous k latent pixels are encoded / decoded, and It is assumed that this can be used for self-regression context modeling. The approach is as follows: A modified position shift quantization is used. Quantization is performed as follows:
number
[0319] Probability models are, It is given as JPEG2026123108000110.jpg848. In matrix-vector notation, it is as follows:
number
[0320] Note that this is a form of autoregressive context modeling. This is a one-dimensional P. This is because MF depends on previously decoded latent pixels. However, the position parameter Only the data parameter depends on previously decoded latent pixels; the scale parameter does not. .
[0321] The integer values that can actually be encoded by the arithmetic encoder / decoder are the quantization residues. Please note that this is a reserve item.
number
[0322] Therefore, in decoding, the arithmetic encoder returns from the bitstream y( It is not a hat, but ξ (hat). Next, y (hat) is given by the following system of linear equations. It can be reconstructed by solving this equation.
number
[0323] Solving system (50) is separated from the arithmetic decoding process. In other words, arithmetic decoding The conversion process must be performed serially when the bitstream is received. Solving (50) is independent of this process and can be done using any numerical linear algebra algorithm. It can be executed using M. The decoding path for the L context modeling step is Furthermore, the process does not have to be a serial procedure; it can be executed in parallel.
[0324] Another way to look at this result is that, equivalently, the arithmetic encoder / decoder is at position 0. It operates on ξ (hat). In other words, the arithmetic encoder is not latent on y (hat) It operates with respect to residual ξ (hat). Considering this, PMF is This is JPEG2026123108000115.jpg524. Only after ξ (hat) is restored from the bitstream, then y (hat) Recover it. However, recovery from the bitstream may not be autoregressive. (The only dependency is σ, which is not a context / autoregressive dependency), this procedure is It can be very fast. Then, a highly optimized linear to solve (50) We can use an algebraic routine to recover y (the hat).
[0325] In both encoding and training the L context system, (48) can be solved, Since the unknown variable y(hat)-y(hat) is not explicitly given, it must be determined. No. In fact, (48) is an implicit system. Here, y(H) that satisfies (48) This outlines several possible approaches to finding the (bot). The first approach solves (48) sequentially and the dependencies in the autoregressive model This involves manipulating pixels in a specific order. In this configuration, all pixels are self-revolved. Simply iterate in a recursive order, applying (47) at each iteration to obtain the quantized latent in the current iteration. Take out the contents. Since (48) is an implicit equation, the second approach is to use an implicit equation solver. This involves employing an implicit coding solver. This is an iterative solution that finds the fixed point solution of (48). It's a silver. • Finally, in certain special cases, it is defined by the sparse context matrix L. By utilizing an autoregressive structure, the components of the serial decoding path can be parallelized. This approach first defines the dependencies between latent pixels in a dependency graph (D Create an rectified acyclic graph. This dependency graph is L It can be constructed based on the sparse structure of matrices. Then, at the same level of the DAG Note that the pixels are conditionally independent of each other. Therefore, at that level It is possible to compute all pixels in parallel without affecting the calculation of other pixels. Yes, it is possible. Therefore, in coding (and training), the graph starts from the root node. This is iterated by working through the levels of the DAG. At each level, all nodes are This process is performed in parallel. This procedure is performed in a parallel computing environment (graphics processing). If a processing unit or neural processing unit is available, a simple direct This results in a dramatic speed increase compared to column implementation. This procedure is shown in Figure 23. The image on the left shows the L context parameter associated with the i-th pixel in the example. Neighboring context pixels are the pixels directly above the current pixel and the left neighboring pixels. These are cells. The image on the right shows the pixels listed in raster scan order. The image below shows... Given dependencies generated by the L context matrix, directed acyclic graph This shows an example of constructing a DAG (Directed Acyclic Graph). Pixels at the same level are conditionally independent of each other. Therefore, it can be encoded / decoded in parallel.
[0326] Many of the techniques described in the previous section can also be applied during decoding. Specifically, linear equations
number
[0327] The following provides a detailed example of an L context module within an Al-based compression pipeline. I will explain in detail.
[0328] Figure 24 shows the prediction context matrix L in an example of an Al-based compression pipeline. y Using This shows the encoding process. A typical implicit solver is depicted in this figure. So, the input image x (hat) is supplied to an encoder function such as a neural network. The encoder outputs a latent y, and this latent is sent to the hyperencoder, and the hyperlatent y outputs a latent y. Returns the value of z.
[0329] The hyperlatency is the learned position μ z and scale σ z Use ID PMF which depends on The arithmetic encoder, and the lossless encoder, quantize to z (hat), and the bit It is sent to the stream. Optionally (though not shown in Figure 24), it is learned. The L context module can also be used in the entropy model with respect to y (hat). The quantized hyperlatency is then sent to the hyperdecoder, and the entropy with respect to y is calculated. - Output the model parameters. These include position μ y , scale σ y and L-conte Ström matrix L y This includes the following: The residue is obtained by solving the coding equation using one of the methods described above. It is calculated by and . This quantized residual ξ (hat) is zero mean and scale parameter σ y It is sent to the bitstream using an MF that has [a specific parameter].
[0330] Figure 25 shows the prediction context matrix L in an example of an Al-based compression pipeline. y Use This figure shows the decryption process. A linear equation solver is drawn in this figure. Decryption So, the learned position parameter μ z and σ z One-dimensional PMF with scale parameter Using a lossless decoder, the hyperlatent z (hat) is first used in the bitstream. It is restored from the system. Optionally (not shown in the diagram), it is used during encoding. The L context module can also be used. The hyperpotential is supplied to the hyperdecoder. The hyperdecoder is at position μ y , scale σ y and sparse context matrix L y Leave It exerts power. The remaining components are a lossless decoder, zero mean PMF, and scale parameter σ y Use Then it is restored from the bitstream. Next, as described above, the linear decoding system is solved By doing so, the quantized latent is restored. Finally, another neural network By supplying the quantized latent y (hat) through decoder functions such as, The constructed image is restored.
[0331] In the previous sections, we assumed that L is a lower triangle with respect to the pixel decoding order. As a generalization, we can relax this assumption and assume a general matrix A that is not necessarily a lower triangular matrix. Determined. In this case, the coding equation is,
number
number
number
[0332] In general, context functions can be nonlinear. For example, in coding problems,
number
number
number
[0333] One interpretation of this latter extension is implicit PixelCNN. For example, f(·) If we have a triangular Jacobian (matrix of the first derivative), (55) models an autoregressive system. However, (55) is more general than this interpretation, and in fact, not only autoregressive systems, but also P This model describes a stochastic system whose order has both forward and backward conditional dependence. It is possible.
[0334] Learned AR sequence In AI-based image and video compression, autoregressive modeling is used for latent space entry A powerful technique for morphic modeling, conditioned on previously decoded pixels. The context model is used in the latest AI-based compression pipelines. However, However, the autoregressive order in context models is often the raster scan order. These are predefined, elementary elements that could potentially impose undesirable biases on learning. Yes. Therefore, as an autoregressive order in the context model, raster scans and We propose external fixed orders, conditional orders, learning orders, or direct optimization orders.
[0335] In mathematical terms, the goal of loss-based Al compression is to make possible the latent distribution that generates the observed data. The goal is to infer prior probability distributions and entropy models that are as consistent as possible. This can be achieved by training neural networks through optimization frameworks such as descent control. It is possible. Entropy modeling underpins the entire Al-based compression pipeline. Therefore, better distribution matching is achieved by lower reconstruction loss and bitrate. It is characterized by better compression performance.
[0336] In image and video data exhibiting large spatial and temporal redundancy, the entropy model is To utilize this redundancy in the process, an autoregressive model called context modeling is used. The process is very useful. At a high level, the general idea is to use existing available information. This is to condition the explanation of subsequent information. The previous variable is used to realize the next variable. A conditioning process refers to an autoregressive information retrieval structure with a specific order. This concept is A It has proven to be very powerful in L-based image and video compression, and It is commonly used as part of advanced neural compression architectures.
[0337] However, ordering of autoregressive structures in AI-based image and video compression, Regressive ordering (AO for short) is predetermined This can happen. In such context models, for example, the image data type (3D; RG This is the so-called raster scan order, which naturally follows a data sequence of length x width x channels (e.g., B). — is often adopted. Figure 26 shows an example of the raster scan order for a single-channel image. This illustrates the following: Gray squares represent conditional variables, and white squares represent unconditional variables. However, adopting the raster scan order as the basic AO is arbitrary, and currently Because it is not possible to use information from below and to the right of a variable (or pixel) as a condition, it is harmful. This may be due to inefficiencies or potential in neural network training. This could potentially lead to undesirable biases.
[0338] Below, we will discuss many fixed or learnable AOs, and many different frameworks that can be used to formulate them. We will explain this along with the framework. AO in context modeling is these frameworks This can be generalized through the concept of "potential change." (a) Theoretical aspects of AI-based image and video compression and contextual models This section details the purpose of autoregressive modeling. (b) The number of conventional and unconventional AOs that can be assigned to context modeling. Explain and illustrate each point. (c) The AO can be trained by network optimization using gradient descent or reinforcement learning. This explains many of the frameworks.
[0339] Al-based image and video compression pipelines typically follow an autoencoder structure. This is a convolutional neural network that constructs encoding and decoding modules. Constructed by a CNN (Convolutional Neural Network), its parameters are based on natural image and video data. Optimization can be achieved by training with sets. (Observation) data is generally denoted as p, and the data is divided into It is assumed that the distribution follows the cloth p(x). The feature representation after the encoder module is latent. It is called "present" and is denoted as y. In encoding, this is the entropy code of the bitstream. It is converted into a number, and the reverse happens during decoding.
[0340] The true distribution of the latent space p(y|x) cannot be practically obtained. This is because the data distribution Because it is difficult to marginalize the joint distribution for y and x in order to calculate JPEG2026123108000126.jpg637. Therefore, we can only find an approximate representation of this distribution, and That is precisely what entropy modeling does.
[0341] The true latent distribution of JPEG2026123108000127.jpg612 can be expressed as a joint probability distribution with the conditional dependent variable without loss of generality. can
Mathematics
Mathematics
[0342] Applying either of the two concepts imposes constraints that invalidate the equivalence of the joint probability and factorization into conditional components as described in Equation (59), but this is often done to trade off modeling complexity. The first concept is actually mostly done for high-dimensional data, such as modeling context based on PixelCNN where only local receptive fields are considered. An example of this process is shown in Figure 27, which shows a 3×3 receptive field where the next pixel conditions on local variables (where the arrow occurs) instead of all preceding variables. An example of this process is shown in Figure 27, which shows a 3×3 receptive field where the next pixel conditions on local variables (where the arrow occurs) instead of all preceding variables. The condition is (place). However, in this paper, in order to generalize the innovation discussed here Therefore, we will not consider imposing this constraint.
[0343] The second concept is the factorization entropy model (which does not impose conditions on random variables and does not use deterministic patterns). (Conditioning only on the lameter) and hyperprientropy model (the latent variable is high By conditioning the set of per latent z, we can assume that they are all conditionally independent. Including the case, both of these cases have a 1-step AO, i.e., the joint distribution The inference is performed in a single step.
[0344] The following describes applications for AI-based image and video compression. This document describes three different frameworks for specifying AO (Autoregressive Autoregression) for serial execution. I will explain this. This includes entropy modeling using context models. , but not limited to this. Each framework has (1) a way to define AO and (2) AO This provides a method for formulating optimization techniques.
[0345] The data is in a 2D format with dimensionality M = H × W (where H is the height dimension and W is the width dimension) (single Assuming it is arranged as a channel image or single frame, single channel video This is possible. The concept shown here also applies to data with multiple channels or multiple frames. Applicable to various situations.
[0346] Graphical models, more specifically directed acyclic graphs (DAGs), are probability distributions and their corresponding patterns. It is very useful for describing dependency structures. The graph shows nodes corresponding to the variables of the distribution and , indicating conditional dependency (the variable at the beginning of the arrow is conditioned on the variable at the end of the arrow) It is constructed by directional links (arrows). As a visual example, Figure 28 illustrates the collaborative structure. The fabric will look like this.
number
[0347] The main constraint for a directed graph to correctly describe the probability of connections is that it must not contain directed cycles. This means that there is no path that starts from any node on the path and ends at the same node. This means it will not be a directed, uncycled graph. Raster scan order Figure 2 shows an example of a DAG describing the joint distribution of four variables {y1, y2, y3, y4}. Equation (60) shows that when the variables are organized from 1 to N in the raster scan pattern, It follows exactly the same structure. Considering our assumptions, this is M step AO and complete To evaluate the coupling distribution, M passes or executions of the context model are required. It means that.
[0348] Another AO with fewer than M steps is checkerboard ordering. Figure 29 shows each step. We demonstrate a two-step algorithm (AO) where all current variables are conditionally independent and can be evaluated in parallel. In Figure 29, in step 1, the current pixel distribution is in parallel without any conditions. Inferred. In step 2, the current pixel distribution is inferred in parallel with the previous step. It is conditioned on all pixels. At each step, all current variables are conditioned on The steps are independent. The corresponding directed graph is shown in Figure 30. As shown in Figure 30, the steps Step 1 evaluates the node in the top row, and Step 2 evaluates the node in the bottom row with an arrow indicating a condition. It is represented by a code.
[0349] The binary mask kernel for autoregressive modeling is N< <MのNステップのAOを It is a convenient framework for specifying. Given data y, a binary mask kar The Nell method extracts it into N low-resolution subimages {y1,…,y N It is necessary to split it into} Previously, each pixel was variable y i Although defined as such, here, JPEG2026123108000131.jpg513 has K conditionally independent (conditioned to be combined for future steps) pixels or It is defined as a group of variables.
[0350] Sub-image y i is, k H ×k W Element M having a stride of i,pq Binary mask carne Ru The data is extracted by convolving it with JPEG2026123108000132.jpg651. In order to define a valid AO, Inari mask kernels must adhere to the following constraints:
number
number
[0351] Here, 1 kH,kW is size k H ×k W It is a single matrix. In other words, each mask is unique. And it must have only a single entry of 1 (the remaining elements are 0). This is a sufficient condition to ensure that AO is exhaustive and follows a logical order of conditions. To enforce constraints (61) and (62) while establishing AO, kH k W Learn k logits and learn one for each position (p, q) in the kernel, for example ranking the logits from high to low to order the autoregressive process. Also, the trick of Gumbel-Softmax can be applied to finally enforce one-hot encoding while decreasing the temperature Figure 31 shows an example where a 2×2 binary mask kernel according to constraints (61) and (62) generates four sub-images and is conditioned according to the graphical model (but vector variables instead of scalar) described in Figure 28. This example is very similar to the checkerboard AO described in the previous section, but some intermediate steps are added. In fact, by not conditioning between y1 and y2, and between y3 and y4, it is possible to define AO with a binary mask kernel so as to exactly reflect the previous checkerboard AO The choice of whether to condition on the previous sub-image y
[0352] can be determined by an adjacency matrix A that is strictly lower triangular and a threshold condition T. In this case, A should be of a more practical dimension k × k Following the previous example, Figure 32 shows how this is possible. Figure 32 shows an example of an adjacency matrix A that determines the graph connectivity of AO defined in the binary mask kernel framework. If there is a link where the associated adjacent term is less than the threshold T, that link is ignored. When a link is ignored, a conditional independent structure appears and the autoregressive process is constructed in fewer steps i can be determined by an adjacency matrix A that is strictly lower triangular and a threshold condition T. In this case, A should be of a more practical dimension k × H k W × k H ...... k W Following the previous example, Figure 32 shows how this is possible. Figure 32 shows an example of an adjacency matrix A that determines the graph connectivity of AO defined in the binary mask kernel framework. If there is a link where the associated adjacent term is less than the threshold T, that link is ignored. When a link is ignored, a conditional independent structure appears and the autoregressive process is constructed in fewer steps Figure 32 shows an example of an adjacency matrix A that determines the graph connectivity of AO defined in the binary mask kernel framework. If there is a link where the associated adjacent term is less than the threshold T, that link is ignored. When a link is ignored, a conditional independent structure appears and the autoregressive process is constructed in fewer steps If there is a link where the associated adjacent term is less than the threshold T, that link is ignored. When a link is ignored, a conditional independent structure appears and the autoregressive process is constructed in fewer steps If there is a link where the associated adjacent term is less than the threshold T, that link is ignored. When a link is ignored, a conditional independent structure appears and the autoregressive process is constructed in fewer steps If there is a link where the associated adjacent term is less than the threshold T, that link is ignored. When a link is ignored, a conditional independent structure appears and the autoregressive process is constructed in fewer steps
[0353] In binary mask kernels such as Adam7 used in PNG, conventional interlacing It is also possible to represent the method. Figure 33 shows the index of the Adam7 interlacing method. This indicates the kernel layout. This notation shows how the kernel is arranged and grouped. Please note that this is used as a reference and is not actually used directly within the model. Please note. For example, 1 corresponds to one mask in the top left position. 2 corresponds to the same position Corresponds to another single mask. 3's correspond to two masks grouped together. Correspondingly, each has one corresponding element at the same position as the 3's in the index. The same principle applies to the remaining elements. This also applies to the index. The kernel size becomes 8x8, and there are 64 mask kernels. A group is generated, and the grouping is indexed as shown in Figure 33, resulting in 7 Two sub-images are generated (thus resulting in a 7-step AO). This is binary mass The kernel framework utilizes low-resolution information to generate high-resolution data. This demonstrates that it is adapted to handle common interlacing algorithms.
[0354] The raster scan order is within the framework of a binary mask kernel with kernel size H x W. It can also be defined within the unit. In this case, H × W = N steps of AO, and H × W N binary mask kernels are arranged so that they are in order during a raster scan.
[0355] In summary, binary mask kernels are suitable for gradient descent-based learning methods, and the following applies: As explained below, this relates to further concepts concerning autoregressive ordering in frequency space. It is.
[0356] The ranking table is a third framework that characterizes AO, and under fixed rankings This allows for describing M-step AO without the representational complexity of binary mask kernels. It is particularly effective. The concept of a ranking table is simple, Given JPEG2026123108000135.jpg513 (flattened and corresponding to the total number of variables), each AO ranks the elements of q. System q,q i It is determined based on the following. Indexing is done using the argsort operator This can be done using q, and the ranking will be either in descending or ascending order, depending on the interpretation of q. It is possible. The largest index is assigned as y1, and the second largest q i of The index will be assigned as y2. Indexing is done using the argsort operation. This is done using children, and the ranking will be either descending or ascending depending on the interpretation of q.
[0357] q is the entropy parameter of y (learned or hyperpreferenced). Therefore, the predicted parameters, such as the scale parameter σ, relating to the source data y. This can be an existing quantity that conveys specific information. This is a large scale parameter σ ij The region of high uncertainty associated with a variable having this property can be easily obtained through the context. Since it contains missing information, an example of an unprocessed visualization is shown in Figure 34. In the formula, AO is y1, y2, ... y 16 It is defined as follows. Figure 34 shows the result after being flattened to q and then sorted by the argsort operator. Define a ranking table and sort the variables in descending order based on the magnitude of each scale parameter. This shows an example of visualizing scale parameters, organized in a specific way.
[0358] q can also be derived from existing quantities, such as the first or second derivative of the position parameter μ. Both of these can be obtained using either a gradient vector (for the first derivative) or a Hessian matrix (for the second derivative). In the case of q and ar, this can be obtained by applying the finite difference method, and q and ar This is obtained before calculating gsort(q). Then the ranking is obtained by the norm of the gradient vector. The norm of the eigenvalues of the Hessian matrix, or any measure of the curvature of the latent image, is established. Alternatively, in the case of the second derivative, the ranking is equivalent to the trace of the Hessian matrix. This can be based on the magnitude of the Laplacian.
[0359] Finally, q can be an entirely different entity. Fixed q can be optionally pre-trained. They can be defined, and like hyperparameters, they remain static or dynamic during training. Alternatively, learning can be done using gradient descent to optimize the model. It can be transformed, or parameterized by a hypernetwork. .
[0360] The method of accessing elements of y depends on whether you want to pass gradients through the ranking operator. do: · When a gradient is not needed Sort :q in ascending / descending order to access elements, and based on that order Next, access the elements of y. · If a gradient is required : The order is a discrete permutation matrix P or a continuous relaxation matrix The result is represented as JPEG2026123108000136.jpg55, and the sorted result y is multiplied by a matrix: y sort =Py
[0361] If the ranking table is optimized based on gradient descent, then argsort or argmax Some index operators may not be differentiable. Therefore, permutation matrices. The continuous relaxation of JPEG2026123108000137.jpg55 needs to be used, which can be implemented with the soft sort operator.
number
[0362] The concept of ranking tables will be extended to work with binary mask kernels as well. This is possible. The matrix q will have the same dimensions as the mask kernel itself, and AO will have the same rank as the elements of q. It is specified based on the following. Figure 36 shows the order in which it is applied to the binary mask kernel framework. This is a visual example of the concept of a rank table.
[0363] Another possible autoregressive model is defined by a hierarchical transformation of the latent space. From this perspective, The potential is converted into a hierarchy of variables, and the lower hierarchical levels are conditioned on the higher hierarchical levels This is the case
[0364] This concept can be best explained using wavelet decomposition. In wavelet decomposition , the signal is decomposed into high-frequency and low-frequency components. This is done through the wavelet operator W Let the potential of size H×W pixels be y 0 The superscript 0 indicates that the potential is at the lowest (or root) level of the hierarchy. Applying the wavelet transform once converts the potential into a set of four small images y 1 ll , y 1 lh , y 1 hl , and y 1 hh , each of size is JPEG2026123108000142.jpg5\2\1. H and L represent the high-frequency and low-frequency components respectively. The first letter of the tuple corresponds to the first spatial dimension (e.g., height) of the image , and the second letter corresponds to the second dimension (e.g., width) . Thus, for example, y 1 hl is the wavelet component of the potential image y 0 corresponding to the high frequency in the height dimension and the low frequency in the width dimension 0 .
[0365] In matrix notation, it is as follows
Equation
[0366] Applying this procedure recursively to the low-frequency blocks constructs a hierarchical tree of decomposition. Figure 3 Figure 7 is an example of a procedure with two hierarchies. Figure 37 shows a hierarchical procedure based on wavelet transform. The autoregressive order is shown. The top figure shows that the forward wavelet transform creates a hierarchy of variables. This shows that there are two levels in this case. The middle image shows the low frequency of the previous level. This shows that the transformation can be reversed to recover several elements. The image below shows one of the hierarchies. This shows an autoregressive model defined by an example of a DAG between elements at a level. Thus, wavelet transforms can be used to create multi-level trees.
[0367] The important thing is that the transformation matrix W is an invertible matrix (in fact, in the case of wavelet transforms, W -1 =W T ) All procedures are reversible. Given the last level of the hierarchy, By simply applying the inverse transform to the last level, the low-frequency components of the previous level can be easily restored. This can be done. Next, after restoring the low-frequency components of the next layer, the inverse transform is applied to the second layer. Then, repeat this process to restore the original image.
[0368] Now, how can we construct an autoregressive order using this hierarchical structure? At the layer level, an autoregressive order is defined between the elements at that level. For example, in Figure 37 Referring to the bottom image, the low-frequency components are at the root of the DAG at that level. Autoregressive model Note that Dell can also be applied to the building elements (pixels) of each variable within the hierarchy. The remaining variables are conditioned on the preceding elements of the level. And all conditions of the level After describing the existence, the lowest frequency component of the preceding level is obtained using the inverse wavelet transform. Restore.
[0369] Another DAG is defined between the next lowest level elements, and the original latent variables are restored until they are recovered. The self-regression process is applied recursively.
[0370] Thus, the autoregressive order is the DAG and inverse wavelet transform between the elements at the tree level. It is defined on a variable given by the level of the wavelet transform of the image.
[0371] This can be generalized in several ways. It does not have to be a wavelet transform; any reversible transform can be used. This includes the following: -Wavelets- For example, permutation matrices defined by a binary mask. - Other orthonormal transforms such as the Fast Fourier Transform - Learned invertible matrices - Learned invertible matrix - For example, a reversible matrix predicted by hyperpreference. • Hierarchical decomposition can also be applied to video. In this case, each level of the tree is lll, hll, It has eight components corresponding to lhl, hhl, llh, lhh, hhh, and hlh, and the first letter in the formula represents the time component. It represents.
[0372] Extended Lagrangian Examples of constrained optimization and rate distortion annealing techniques are described in international patent application P This is described in CT / GB2021 / 052770, which is incorporated herein by reference. It gets included.
[0373] Al-based compression pipelines attempt to minimize rate (R) and strain (D). The target function is as follows:
number
[0374] International patent application PCT / GB2021 / 052770 addresses this problem as a constrained optimization problem. The problem is reformulated as a title. The method for solving this constrained optimization problem is the international patent application PCT. As described in issue / GB2021 / 052770, Augmented Lag This is the rangian method. A constrained optimization problem involves solving the following: Min D (66) R=c etc. (67) In the formula, c is the target compression ratio. Note that D and R are the mean of the entire data distribution. Also, Inequality constraints can also be used. Furthermore, the roles of R and D can be reversed. Rather, the rate can be minimized according to the strain constraint (which may also be a system of constraint equations). ru.
[0375] Typically, constrained optimization problems are solved using stochastic linear optimization. That is, the objective function is The gradient is calculated for a small number of training samples (not the entire dataset), and then the gradient is calculated. The calculation is performed. Then, an update step is executed, and the parameters of the compression algorithm and, In some cases, other parameters related to constrained optimization, such as the Lagrange multiplier, may be changed. This process is repeated thousands of times until an appropriate convergence criterion is reached. Then, perform the following steps. • SGD, mini-batch SGD, Adam (or used for training neural networks) Other optimization algorithms used, etc., Run the optimization step once on JPEG2026123108000145.jpg747. • Update the Lagrange multiplier: JPEG2026123108000146.jpg530, choose a small value for ε in the formula. Repeat the above two steps until the Lagrange multiplier converges based on the target rate r0. return.
[0376] However, this is encountered during training of constrained optimization problems in a stochastic small-batch linear optimization setting. There are several problems. First of all, the constraints on the entire dataset in each iteration It cannot be calculated and is typically only calculated for the small number of training samples used in each iteration. This is done. Using such a small number of training samples in each batch allows for constrained optimization parameters. Lagrangian Multi (Augmented Lagrangian) Updates to pliers, etc., have become extremely dependent on the current batch, resulting in variations in training updates. The number of steps increases, making it impossible to obtain the optimal solution, or the optimization routine becomes unstable. There is a possibility that it will happen.
[0377] The average constraint value (such as the average rate) can be calculated over the N previous iteration steps. Yes, it is possible. This is especially true for parameter updates related to constraint optimization algorithms (extended ragdra). Regarding updating the Lagrange multiplier in the genian, etc., this can be applied to the current optimization step. This has the beneficial effect of extending the constraint information for the last many optimization iteration steps. There is a non-exhaustive method of calculating the average of the past N iterations and applying it to the optimization algorithm. The list is as follows: • It holds a buffer of constraint values for the most recent N training samples. This buffer is, for example, Ba, The extended Lagrange multiplier λ can be updated in iteration t as JPEG2026123108000147.jpg637. In the formula, avg is a general square. This is the average operator. An example of the average operator is as follows: - Arithmetic mean (often simply called "average") - Median -geometric mean -harmonic mean - Exponential moving average -Smoothed moving average -Linear weighted moving average However, any average operator may be used.
[0378] • Instead of calculating the gradient at each training step and immediately applying it to the model weights, N The gradient is accumulated over N iterations, and after N iterations, one training gradient is obtained using the averaging function. Perform the step.
[0379] This averaged constraint value is calculated over the previous N iterations, and includes Lagrange multipliers, etc. Used to update the parameters of the training optimization algorithm.
[0380] The second problem with stochastic linear optimization algorithms is that the constraints on the dataset are extremely large. This involves the inclusion of images or extremely small images (for example, if the rate R is extremely small). (or extremely large values, etc.). If outliers exist in the dataset, the function being trained will be affected. This will involve considering the values, which may result in a worse fit for more general samples. For example, updating the parameters of an optimization algorithm (such as the Lagrange multiplier) can be inaccurate. When there are many values, the variance becomes large, which can lead to suboptimal training.
[0381] Some of these outliers can be removed from the above constraint mean calculation. Several possible methods for filtering (removing) the values are as follows: • When accumulating training samples, the mean can be replaced with a trimmed mean. The difference between a trimmed average and a regular average is that in the trimmed version, x% of the upper and lower values are included in the average. This is something that is not taken into account in the calculation. For example, the top and bottom samples (in sorted order) You can trim 5%. • There is no need to trim throughout the entire training procedure; the normal average is taken in the latter half of the training procedure. You can also use (without trimming). For example, the trimmed average is 1 million It can be turned off after a certain number of iterations. • Fit outlier detectors every N iterations so that the model is based on the mean. The samples can be completely removed or weighted.
[0382] Constraints such as Augmented Lagrangian By using an optimization algorithm with specific mean constraints (R) on the training dataset, Objective c (such as T) can be targeted. However, in the training set, the constraints of that objective can be met. Even if convergence occurs, there is no guarantee that the same constraint values will be obtained in the validation set. This is because, for example, training... The change between the quantization function used in the initial analysis and the quantization function used in the inference (test / verification) results in... This can be caused by, for example, quantization using uniform noise during training, Inference commonly uses rounding (called "STE"). Ideally, While it is desirable for constraints to be met in an argument, achieving this is difficult.
[0383] The following methods are possible: • In each training step, the constraint target value is calculated using the detach operator. Here, the forward For the first pass, the inferred value is used, but for the second pass (gradient calculation), the training gradient is used. If we are going to discuss rate constraints, The value JPEG2026123108000148.jpg548 is used. Here, 'detach' separates the value from the auto-differential graph. It represents that. • An image showing the evaluation of the model using the same constraint algorithm parameters as inference. Update using the "Hold Set".
[0384] As mentioned above, the technology described in international patent application PCT / GB2021 / 052770 is, It can also be applied to AI-based video compression. In this case, it is used in each training step. The Lagrangian multiplier can be applied to the rate and distortion associated with each frame of the video. It is possible to optimize one or more of these Lagrange multipliers using the techniques described above. Yes, it is possible. Alternatively, the multiplier can be averaged across multiple frames during the training process. stomach.
[0385] The target value of the Lagrange multiplier is for each frame of the input video used in the training step. They may be set to equal values. Alternatively, different values may be used. For example, video Different target values can be used for the I-frame and the P-frame. A higher target rate can be used compared to frames. Also, B-frames The same method can be applied to this as well.
[0386] Similar to image compression, for one or more frames of the video used for training, the target The rate may be initially set to zero. If a target value is set for the distortion, then the rate will be set for the distortion. The target value may be set so that the initial weighting is maximized (for example, the target rate is (It may be set to 1.)
[0387] Tensor Network Al-based compression relies on modeling discrete stochastic mass functions (PMFs). These PMFs may seem simple at first glance. Our usual mental model is one Starting with a discrete variable X, it has D possible values X1, ..., X D It is possible to take PMF. The construction of P(X) is P i =P(X i This can be easily done by creating a table defined as follows: It is possible. Of course, Pi must be non-negative and the sum must be 1, but this is for example soft Tomax function This can be achieved by using JPEG2026123108000149.jpg1631. For modeling purposes, it can be fitted to a specific data distribution. Each P in this table i Learning it doesn't seem to be that difficult.
[0388] What about the PMF for two variables, X and Y? Entry P ij =P(X i ,Y j )of This also seems manageable, as it requires a 2D table. It's a bit more complicated. Currently, the table has D 2 There is an entry, but even so, D is large. As long as it's not too much, JPEG2026123108000150.jpg2646 is manageable. Next, using three variables, a 3D table is needed, and NtriP ijkIt is indexed by a 3-tuple.
[0389] However, this naive table-creation "approach" models more than a handful of discrete variables. If you try to do that, it quickly becomes unmanageable. For example, an RGB 1024x1024 image Let's consider modeling PMF on the interval. Each is 256 3 Takes the possible values (Each color channel has 256 possible values, and we have 3 color channels) (to have). Next, the necessary lookup table is This will result in the entry JPEG2026123108000151.jpg615. Calculating in decimal gives approximately The result will be JPEG2026123108000152.jpg68. There are many approaches to dealing with this problem, but the textbook app for discrete modeling... Roach's approach involves using probabilistic graphical models.
[0390] As an alternative approach, PMF can be modeled as a tensor. This simply represents a huge table (which, although not explained here, possesses algebraic properties). This is a statement. Discrete PMFs can always be described as tensors. For example, 2-T An array (also called a matrix) is an array with two indices, that is, a two-dimensional table. Therefore, the above PMF P for two discrete variables X and Y ij =P(X i ,Y j ) T is a 2-tensor. i1 ,…,T iN This is a distribution with N indices. If it is a column and the entries in T are positive and the sum is 1, then this is the PMF for N discrete variables. Yes. Table 1 shows the standard PMF perspective and the tensor perspective for several probabilistic concepts. This is a comparison of the two.
[0391] The main appeal of this perspective is the ability to model huge tensors using the framework of tensor networks. The ability to simplify. Tensor networks use very high-dimensional tensors to create several It is used to approximate by contracting low-dimensional (i.e., easy-to-handle) tensors. Sol networks are used to approximate tensors, which are otherwise difficult to handle, at a low rank. It can be done. [Table 3]
[0392] For example, if we consider a matrix as a 2-tensor, we can use a standard low-rank approximation (singular value decomposition (SVD)). Tensor network analysis (PCA, etc.) is a factorization of a tensor network. Networks are a generalization of low-rank approximations used in linear algebra to multi-linear maps. Examples of using tensor networks in probabilistic modeling in machine learning include "Ivan Glasser, Ryan Sweke, Nicola Pancotti, Jens Eisert,and J Ignacio Cirac.Expressive p owner of tensor-network factorizations fo r probabilistic modeling, with applicatio ns from hidden markov models to quantum machine learning.arXiv preprint,arXiv:19 This is shown in “07.03741,2019”, which is incorporated herein by reference. ru.
[0393] Tensor networks can be considered an alternative to graphical models. There is a correspondence between networks and graphical models, and probabilistic graphical models It can be reconstructed as a tensor network, but the converse is not true. Probabilistic graph It cannot be recast as a physical model, but it has strong performance guarantees and is computationally available. There are tensor networks that are capable of coupling density modeling. In many situations, Sol-networks have higher expressive power than conventional probabilistic graphical models such as HMMs. . Given a fixed number of parameters, tensor networks can experimentally perform HMMs. To surpass. Furthermore, under a certain low-rank approximation, the tensor network can theoretically be re-HMM It has the potential to surpass it.
[0394] Assuming all other modeling assumptions remain the same, tensor networks are preferable to HMMs. It is rare.
[0395] An intuitive explanation for this result is that the probabilistic graph factorizes the combination via conditional probability. However, usually, an index map By considering only JPEG2026123108000154.jpg642, the junctions are modeled as a Boltzmann / Gibbs distribution. This may actually be a restrictive modeling assumption, provided by tensor networks. A completely different approach is to model the combination as an inner product: a certain Hermitian positive ( For a semi-definite operator H, JPEG2026123108000155.jpg527 (This modeling approach is inspired by Born's law for quantum systems). The operator H is a huge ten It can be written as a sol (or tensor network). What is important is that the entries of H are complex. It is not at all clear how (or even if) this can be converted into a graph model. However, it presents a completely different modeling point of view that cannot be obtained elsewhere.
[0396] Let's explain what tensor network decomposition is with a simple example. There is a ij large D×D matrix T (a 2-tensor) with entries T, and we want to make a low-rank approximation (rank-r approximation with r < D) of T. One way to do this is to find an approximate T (hat). That is, in other words, T(hat) = AB, where A is a D×r matrix and B is an r×D matrix. We have introduced hidden dimensions shared between A and B, which are summed up. By setting r to be very small compared to dealing with the huge D×D matrix, we can significantly save computation time and power by reducing the number of parameters from D² to 2Dr. Furthermore, in many modeling
Number
[0397] scenarios, even if r is made very small, a "sufficient" approximation of T can be obtained. Here, let's try to model a 3-tensor with the same approach. Suppose we are given a D×D×D tensor T with entries T. Here A and C are low-rank matrices, and B is a low-rank 3-tensor. Between A and B and between B 2 parameters from D² to 2Dr parameters, we can significantly save computation time and power. Moreover, in many modeling scenarios, even if r is made very small, a "sufficient" approximation of T can be obtained. scenarios, even if r is made very small, a "sufficient" approximation of T can be obtained.
[0398] Now, let's try to model a 3-tensor using the same approach. Suppose we have a D×D×D tensor T with entries T ijk as entries and D ×D×D. There are two hidden dimensions between C and A. One is between A and B, and the other is between B and C. In NS-C network terminology, these hidden dimensions may also be called join dimensions. (Sum of dimensions) This could also be called a contraction.
[0400] Continuing with this example, a 4-tensor can be approximated as a product of lower-dimensional tensors, but The notation for the index quickly becomes cumbersome to write. Instead, the same calculation can be conveyed diagrammatically. We use a tensor network diagram, which is a simple method.
[0401] In a tensor network diagram, tensors are represented by blocks, and as shown in Figure 38, each a The dextension dimension is represented as an arm. The dimensionality of a tensor is simply empty (bra This can be determined by counting the number of arms that are lowered. The top row of Figure 38 shows vectors from left to right. This shows matrices and N-tensors. Tensor product (summation along a specific index dimension) The (abbreviated) tensor arm is represented by connecting two tensor arms. See Figure 38 for a schematic representation. As you can see, there is an arm hanging from the matrix-vector product in the lower left, so the resulting product is a tensor. , that is, it turns out to be a Bekult, as expected. Similarly, the matrix in the lower right and the matrix Since the product has two hanging arms, the result is a matrix, as expected.
[0402] The tensor decomposition of the three-tensor T (hat) given by equation (69) is illustrated in Figure 39. It can be represented as shown in the upper part. Here, a specific element T of T (hat) ijk (Hah) Suppose you want to access (T). Fix the free index to the desired value and perform the necessary contraction. Yes.
[0403] Using this notation, it is possible to factorize tensor networks used in probabilistic modeling. It is possible to explore the nature of this. The key idea is that the true binding distribution of high-dimensional PMFs is not feasible. This means that we must approximate it, and the tensor network factor We approximate it using tensor decomposition. These tensor network factorizations are suitable for the training data. It can learn to match. All tensor network factorizations are appropriate. This is not always the case. The entries in the tensor network are non-negative and the sum is controlled to be 1. It may need to be shortened.
[0404] An example of this approach is Matrix Product State (MPS). One method is to use PMF (also called tensor training). N )of tensor Let's say we want to model it as JPEG2026123108000158.jpg611.
number
[0405] To ensure that the sum of the entries is 1, a normalization constant is passed through all possible states. This is calculated by summing them up. For a general N tensor, this normalization constant is It may be impractical to calculate, but in the case of MPS, due to its linear nature, O( The normalization constant can be calculated in N) time. Here, "linearity" refers to the tensor principle. This means that by manipulating the rows of the tensor, we can perform tensor products one by one sequentially. (Tensor and Both of these tensor network approximations are multilinear functions.
[0406] MPS is very similar to a Hidden Markov Model (HMM). In fact, there is a correspondence: positive terms The MPS, which has eyes, accurately responds to HMM.
[0407] Further examples of tensor network models include Born Machines and Loc There is also the Allied Purified States (LPS). Both are quantum systems. It is inspired by this. Quantum systems assume Born's law. Born's law states that if an event X occurs... The probability of it occurring is given by the inner product <·,H·> with a certain positive (semi) definite volume Hermitian operator H, This is proportional to the square of the norm. This is a powerful feature of probabilistic modeling. It is a homework project, but there is no clear connection to graphical models.
[0408] The locally purified state (LPS) takes the form shown in Figure 40. In LPS, the constructs are A is k There are no restrictions on the sign of a tensor; it can be positive or negative. In fact, A k is multiple It can have a prime number value. In this case, JPEG2026123108000160.jpg54 is a tensor obtained by taking the complex conjugate of entries A. α k The dimension is called the combined dimension. There is, β k Dimensions are sometimes called refined dimensions.
[0409] The element of T (Hat) is determined by the fact that contraction along the purification dimension yields a positive value. It is guaranteed to be positive (for a complex number z, JPEG2026123108000161.jpg511). {i1,…,i N If we consider} as one huge multi-exponential I, then LPS is the diagonal of a huge matrix ( (After collapsing all hidden dimensions), evaluating the LPS operates on the state space. It can be seen that this is equivalent to the dot product.
[0410] Similar to MPS, the calculation of the normalization constant for LPS is fast and can be done in O(N) time. Born Machine is a special case of LPS, and the size of its purification dimension is 1.
[0411] Tensor trees are another example of tensor networks. In the leaves of the tree, dang The ring arm is contracted with data. However, the hidden dimension is placed within the tree, Lie nodes store tensors. Tree edges are the dimensions of the tensor being reduced. Yes, there is. A simple tensor tree is shown in Figure 41. The nodes of the tree store tensors, and The `di` represents contraction between tensors. The leaves of the tree have indices that are contracted by data. Tensor trees are used for modeling probability distributions with multiple resolutions and / or multiple scales. It can be used.
[0412] Each tensor node has a refined dimension added so that it is contracted with the complex conjugate of that node. This allows us to obtain the Elmy given by the tensor tree and its complex conjugate. The inner product is defined according to the T operator.
[0413] As another example of a tensor network, Projected Entangled Pa There is the ir States (PEPS). In this tensor network, tensor no The nodes are arranged in a regular grid and contract with the node immediately next to them. Each tensor is , additional dangling arms (free) contracted with data (such as latent index values) It has an index. In a specific sense, PEPS is a Markov random field and an Isingmo It resembles Dell. A simple example of PEPS in a 2x2 image patch is shown in Figure 42.
[0414] Calculation of tensor networks (calculation of PMF connection probabilities, conditional probabilities, marginal probabilities, and (For example, the calculation of the entropy of PMF) canonically evaluates tensors, as will be explained in detail below. By adopting a formal format, the process can be significantly simplified and sped up considerably. All tensor networks can be placed in canonical form.
[0415] Since the basis on which the hidden dimensions are represented is not fixed (so-called gauge degrees of freedom), these The basis on which a tensor is represented can be simply changed. For example, tensor networks If we set the tensor to a canonical form, then almost all tensors can be converted into orthonormal matrices (unitary matrices). It can be converted.
[0416] This is possible by sequentially performing a series of decompositions into the tensors of the tensor network. These decompositions include QR decomposition (and its variations, RQ, QL, and LQ), SV D decomposition, spectral decomposition (if available), Schur decomposition, QZ decomposition, Takagi decomposition, etc. There are such procedures. The procedure for describing a tensor network in canonical form is to define each tensor in canonical form. It functions by decomposing into orthogonal (unitary) components and other factors. The other factors are adjacent to the It is contracted to the tensor and modifies the neighboring tensor. Then the same procedure is performed on the neighboring tensor and its neighbor. When applied to tensors, all tensors except one are orthonormal (unitary). Repeat until it reaches the desired result.
[0417] The remaining tensors that are not orthonormal (unitary) are called core tensors. Core tensors are S It is similar to the diagonal matrix of singular values in VD decomposition, and is the spectrum of a tensor network. It contains information. The core tensor is, for example, the normalization constant of a tensor network, or It can be used to calculate the entropy of a tensor network.
[0418] Figure 43 shows, from top to bottom, an example of the procedure for converting MPS to canonical form. Sequential core tensor. We perform a QR decomposition. The R tensor is contracted with the next tensor in the chain. This procedure is a core tensor. This process is repeated until all forms except Sol C are in canonical form.
[0419] Here, for probabilistic modeling in AI-based image and video compression, Let's explain the use of Sol Networks in more detail. As mentioned above, Al-based In a compression pipeline, the input image (or video) x is encoded using an encoding function (typically New It is mapped to the latent variable y via a multi-level network. The latent variable y is quantized. It is quantized to an integer value y (hat) using the number Q. These quantized latent variables are above As described above, the data is converted to a bitstream using a reversible coding method such as entropy coding. Arithmetic coding or decoding is one example of such coding processes, and further discussion is needed. It is used as an example in [this context].
[0420] This lossless coding process requires a probabilistic model: the arithmetic encoder / decoder is an integer A probability mass function q(y(hat)) is required to convert the value into a bitstream. During decoding, PMF is similarly used to return the bitstream to its quantized latent, and This is then passed through a decoder function (generally a neural network) to reconstruct the image x Return (hat)
[0421] The size of the bitstream (compression rate) depends on the quality of the probability (entropy) model. It is greatly influenced by this. A better, more powerful probabilistic model will be more effective when used with reconstructed images of the same quality. This results in a smaller bitstream.
[0422] Arithmetic encoders typically operate in one-dimensional PMF. To incorporate this modeling constraint... Typically, the combined PMF q(y(hat)) is assumed to be independent, and the pixels Each of the JPEG2026123108000162.jpg55 images is a one-dimensional probability distribution. It is modeled by JPEG2026123108000163.jpg615. The binding density is then modeled as follows:
number
[0423] In any case, this modeling approach is basically, Assume a one-dimensional distribution for each of the 55 pixels in JPEG2026123108000165.jpg. This may be restrictive. A better approach... The approach is to completely model the connection distribution. Then, the bitstream is coded When encoding or decoding, the one-dimensional distribution required for the arithmetic encoder / decoder is conditional. It can be calculated as a probability.
[0424] Tensor networks can be used to model the connection distribution. This is as follows: Yes, it's possible. Quantized latent Suppose we are given JPEG2026123108000166.jpg630. Each latent pixel is embedded (or lifted) into a higher-dimensional space. In this higher-dimensional space, integers are represented by vectors lying at the vertices of a stochastic simplex. For example, y i D possible integer values Let's call it JPEG2026123108000167.jpg677. This embedding is Map JPEG2026123108000168.jpg55 to a D-dimensional one-hot vector, where slots corresponding to integer values have 1, and so Zero exists in all other locations.
[0425] For example, each JPEG2026123108000169.jpg55 takes values of {-3,-2,-1,0,1,2,3}. Let's assume it's JPEG2026123108000170.jpg612. Next, the embedding is The file will be named JPEG2026123108000171.jpg536.
[0426] Therefore, the embedding is JPEG2026123108000172.jpg649 This maps to JPEG2026123108000173.jpg633. In effect, this maps y (hat) existing in M-dimensional space to D M dimensional space Map to.
[0427] Here, each of these items in the embedding is a dimension that indexes a higher-dimensional tensor and It can be considered as such. Therefore, the approach we take is a tensor network T( This involves modeling the connection probability density through the hat symbol.
number
[0428] During encoding / decoding, arithmetic encoders / decoders cannot use coupling probabilities. Instead Therefore, a one-dimensional distribution must be used. Conditional probability is used to calculate the one-dimensional distribution. It is possible.
[0429] Conveniently, conditional probability marginalizes the hidden variable and fixes the precondition variable, These can be easily calculated by normalization. All of these can be easily done using tensor networks. It is possible to do so.
[0430] For example, suppose we encode / decode in raster scan order. Then, for each pixel, Conditional probability is required: JPEG2026123108000175.jpg565. Each of these conditional probabilities is a hidden variable (unseen variable) in the tensor network. By reducing the expression, fixing the index of the condition variable, and normalizing it with an appropriate normalization constant, It can be easily calculated.
[0431] If the tensor network is in canonical form, this is a particularly fast procedure, and in this case, the hidden This is because contraction along a given dimension is equivalent to multiplication by a unit.
[0432] The tensor network spans all latent pixels or patches of latent pixels. Probabilistic modeling of PMFs, or modeling of binding probabilities across channels of latent representations. It can be applied to G, or any combination thereof.
[0433] Tensor network-based connection stochastic modeling is performed as follows: It can be easily integrated into iPline. Tensor networks are end-to-end It is learned during training and fixed after training. Alternatively, a tensor network or The components can be predicted by a hypernetwork. The network is an entropy coding of hyperlatency in hypernetworks and It can be used additionally or alternatively for decoding. In this case, hyperlatency The parameters of the tensor network used for entropy coding and decoding are: It can be learned during end-to-end training and solidified after the training.
[0434] For example, a hypernetwork uses the core tensors of a tensor network in a patch-by-patch manner. It can be predicted by the patch. In this scenario, the core tensor is between the pixel patches. It changes, but the rest of the tensor is learned and fixed between pixel patches. For example, high A tensor network with predictions made by a per-encoder / hyperdecoder Figure 44 shows an L-based compression encoder and a T&T in an Al-based compression pipeline. Tensor network predicted by hyperdecoder for use of sol network See Figure 45, which shows an Al-based compression decoder with a marker. See Figures 1 and 2. As mentioned above, we can assume that the features corresponding to the features shown are the same. In these examples, quantization, encoding, and decoding are performed using a tensor network probabilistic model. The residual ξ = y - μ. In this case, the tensor network parameter is T y Represented by In the examples of Figures 44 and 45, the quantized hyperlatency z (hat) is further T z Encoding and decoding are performed using a ssol network probabilistic model with parameters represented by It can be done.
[0435] Using hypernetworks to predict tensor network components Rather than (or in some cases combined with) one of the tensor networks The part uses a context module that uses previously decoded latent pixels. It may be measured.
[0436] During training of an AI-based compression pipeline using a tensor network probabilistic model, The Sol Network has latent values of non-integer values (y(hat) = Q(y) not y, and here It can be trained on Q (where Q is a quantization function). To do this, embedding functions e can be defined as a non-integer value. For example, the embedding function can be defined as an appropriate integer value for I. It takes one integer value and becomes zero for all other integer values, and is constructed using a tent function that linearly interpolates between them. This allows for multilinear interpolation. Other real numbers for the embedding scheme Value extensions are permitted as long as they match the original embeddings for integer points.
[0437] The performance of a tensor network entropy model is improved by some form of regularization during training. It is possible to increase the value. For example, entropy regularization can be used. In this case, The entropy H(q) of the ensemble network is calculated, and its multiples are added to or subtracted from the training loss function. It is possible. The entropy of a normal form tensor network is the entropy of the core tensor. Note that this can be easily calculated by calculating the tropy.
[0438] Hyper-Hyper Network For use in compressing image and video data based on AI and deep learning. However, training techniques for auxiliary hyperhyper The current functionality and scope of use are described below.
[0439] The network configuration commonly used in AI-based image and video compression is O It is a pixel encoder. This encodes the input data, which is an alternative representation of the input data, and pixels. Encoder module that converts to a "latent" (y) which is often modeled as a set of values. It takes a set of latent data and transforms it into input data (or something as close to it as possible). It is built with a decoder module intended to convert and return. Each latent pixel is one The "latent space" representing dimensions possesses higher-dimensional properties, hence the so-called "entropy model." The parametric distribution is "fitted" to the latent space using p(y(hat)). The Ropy model uses a lossless arithmetic encoder to convert y (hat) into a bitstream. Used to do so. Parameters of the entropy model ("entropy parameters") The entropy model is learned within the network. This can be predicted by the hyperplier structure. A diagram of this structure is shown in Figure 1.
[0440] The entropy parameter is most commonly the position parameter and the scale parameter (positive It is constructed by (often expressed as a real number) (however, it is not limited to this) (There is none). Naturally, whether parametric or nonparametric, there are many more There are different fabric types, and a wide variety of parameter types.
[0441] The hyperplier structure (Figure 2) is obtained through a set of transformations that predict the entropy parameter. Next, we introduce an additional set of "latent" (z). We assume that y is modeled by a Gaussian distribution. Then, the hyperplier model can be defined by the following equation.
number
[0442] Hyperpliers are added to models that already have trained hyperpliers. This can be done, and is sometimes called an auxiliary hyperplier. This technique is particularly useful at low frequencies. To improve the model's ability to capture features and to improve performance on "low-rate" images This applies, but is not limited to, low-frequency characteristics, sharp along its axis. It is present in images when there are no drastic color changes. Therefore, images with many low-frequency features are This means that the entire image will contain only one color. This involves extracting the power spectrum of the image. This allows us to identify the low-frequency features of the image.
[0443] Hyper-Hyperpliers are trained in conjunction with Hyperpliers and entropy parameters. It is possible. However, it is not possible to use hyper-hyperpliers on all images. This could potentially increase computational costs. Maintaining performance with images other than low-rate images, In order to give the network the ability to model low-frequency features, the image is low bitrate. By employing auxiliary hyper pliers, which are used when the product meets certain characteristics such as [specific characteristics]. This is possible. An example of a low-rate image is one where the bits per pixel (bpp) is approximately 0. This applies to cases where the value is less than 1. An example of this is shown in Algorithm 3.
[0444] The auxiliary hyper-hyper-plier framework allows you to adjust the model only when necessary. It is possible. Once trained, this particular image will no longer require hyper-hyperpliers. A flag indicating that can be encoded into a bitstream can be used. This generalizes to the infinite construction elements of entropy models such as hyperhyperpliers. It is possible.
[0445] [Table 4]
[0446] The most direct way to train a hyper-hyper-plier is to use an encoder and decoder. This includes "freezing" existing trained hyperplier networks and hyperhyper - This involves optimizing only the module weights. In this book, "freeze" refers to free The weights of modules that are frozen are not trained, and the weights of modules that are not frozen are not trained. This means that gradients are not accumulated. By freezing the existing entropy model, Hyperpliers are parameters of hyperpliers, such as μ and σ in the case of a normal distribution. The data can be modified to favor lower-rate images.
[0447] Using this training scheme offers several advantages. • Because gradients need to be calculated and saved with fewer parameters, training time and memory Mori consumption can be scaled better with respect to image size. • Hyperpliers training can be repeated infinitely, and Hyperpliers can be free You can then slide it and train another hyperplier on top of it.
[0448] One approach is to first train the hyperplier network through N iterations. Once it reaches the end, the entropy model is frozen, and if the image is low rate, hyperp It can be switched to a lyre. This means that the hyperplier model is already good at Hyperpliers specialize in images that are being processed, and function as intended. Algorithm 4 This shows the training scheme. This training involves dividing the image into K blocks of size NxN. Then, by applying this scheme to that block, the image contains low-frequency regions. In such cases, it is also possible to train the system only in that specific domain.
[0449] Another possibility is to start training the hyperpliers N times, as shown in Algorithm 5. The key is not to wait for the next repetition.
[0450] [Table 5]
[0451] [Table 6]
[0452] The rate calculated using the distribution selected as the prior distribution is the mean value of the image's power spectrum. Alternatively, you can use the median, or the median or mean of the frequencies obtained by the Fast Fourier Transform. To classify an image as low-rate, you can choose from a variety of criteria.
[0453] Data augmentation is achieved by increasing the number of samples that have low-frequency features associated with low-rate images. This allows for the creation of sufficient data. There are various methods for modifying images. • Upsample all images to a fixed size NxN. • Use a constant upsampling factor. • Randomly sample upsampling coefficients from the distribution. Uniform distribution, Gaug S distribution, Gumbel distribution, Laplacian distribution, Gaussian mixture distribution, geometric distribution, Student distribution t-distribution, nonparametric distribution, CHI 2 Distribution, beta distribution, gamma distribution, Pareto You can choose any distribution from the Cauchy distribution. • Apply smoothing filters to the image. (Average, Weighted Average, Median, Gaussian, Bi-Smoothing) Any lateral filter will work. • Use the actual image size unless it is smaller than a specific threshold N, and then use a fixed size. Upsample and use a constant upsampling factor, or as described above. The upsampling coefficient is sampled from the distribution.
[0454] In addition to image upsampling or blurring, random cropping may also be performed.
Claims
1. A method for encoding, transmitting, and decoding lossy images and videos, wherein the method 、 The first computer system receives an input image, To generate a latent representation, the first trained neural network is used with the input The steps to encode the image, The step of performing a quantization process on the latent representation to generate a quantized latent. The size of the bins used in the quantization process is based on the input image, Steps and The steps include transmitting the quantized latent to a second computer system, The quantized latent is decoded using a second trained neural network. A step in which the output image is an approximation of the input image, Methods that include...
2. The size of the bin differs between at least two of the pixels of the latent representation. The method according to claim 1.
3. The size of the bin differs between at least two channels of the latent representation, claim The method described in 1 or 2.
4. Any of claims 1 to 3, the bin size is assigned to each pixel of the latent representation. The method described in paragraph 1.
5. The quantization process assigns the value of each pixel in the latent representation to the pixel. The method according to claim 4, comprising performing a calculation corresponding to the bin size.
6. The quantification process involves subtracting the average value of the latent representation from each pixel of the latent representation. A method according to any one of claims 1 to 5, including the method described in any one of claims 1 to 5.
7. The method according to any one of claims 1 to 6, wherein the quantification process includes a rounding function.
8. The size of the bin used to decode the quantized latent is the quantum Based on previously decoded pixels of the converted latent, as described in any one of claims 1 to 7 Method of loading.
9. The quantization process comprises a third trained neural network, according to claims 1 to 8. The method described in any one of the items.
10. The third trained neural network is at least one of the quantized latents The method according to claim 9, which takes one previously decoded pixel as input.
11. To generate the hyperlatent representation, a fourth trained neural network is used. The steps include encoding the latent representation and Quantization is performed on the hyperlatent representation to generate the quantized hyperlatent. The steps to take, The steps include: transmitting the quantized hyperpotential to the second computer system; 、 To obtain the aforementioned bin of the aforementioned size, a fifth trained neural network is used. The method further includes the step of decoding the quantized hyperlatency using the method, The decoding of the quantized latent uses the obtained size of the bin. 、 The method according to any one of claims 1 to 10.
12. The output of the fifth trained neural network takes the size of the bin. The method according to claim 11, wherein the data is processed by a further function to obtain the result.
13. The further function is a sixth trained neural network, as described in claim 12. Method of loading.
14. The size of the bin used in the quantization process of the hyperlatent representation is the input A method according to any one of claims 11 to 13, based on an image.
15. The method includes the step of identifying at least one region of interest in the input image, Before at least one corresponding pixel of the latent representation in the identified region of interest The process further includes the step of reducing the size of the bin used in the quantification process, The method according to any one of claims 1 to 14.
16. The method further includes the step of identifying at least one region of interest of the input image. 、 For at least one corresponding pixel of the latent representation within the identified region of interest Different quantification processes are used. The method according to any one of claims 1 to 15.
17. The aforementioned at least one region of interest is recognized by the seventh trained neural network. The method according to claim 15 or 16, as otherwise provided.
18. The positions of the one or more regions of interest are stored in the binary mask. The binary mask is used to obtain the size of the bins. The method according to any one of claims 15 to 17.
19. A method for training one or more neural networks, wherein the one or more neural The network is used for encoding, transmitting, and decoding lost images or videos. The method described above is The first computer system receives an input image, The input image is encoded using the first neural network to generate a latent representation. Steps and The step involves performing a quantization process on the latent representation to generate a quantized latent. The size of the bins used in the quantization process is based on the input image. Top, Using a second neural network, the quantized latent is decoded to produce the output image. A step of generating an output image, wherein the output image is an approximation of the input image, A step of determining a quantity based on the difference between the output image and the input image, Based on the determined quantity, the first neural network and the second neural network Steps to update the parameters of the holographic network, The above steps are repeated using the first set of input images to obtain the first trained N Steps to generate a neural network and a second trained neural network P and, Methods that include...
20. The latent representation is encoded using a third neural network to obtain a hyperlatent representation. The steps to generate, Quantization processes are performed on the aforementioned hyperlatent representation to obtain the quantized hyperlatent. The steps to generate, The steps include: transmitting the quantized hyperpotential to the second computer system; 、 The quantized hyperlatency is decoded using a fourth neural network, A step of obtaining the size of the record bin, The decoding of the quantized latent uses the obtained bin size, The parameters of the third neural network and the fourth neural network The meter is further updated based on the determined amount, and a third trained neural Steps to obtain a new neural network and a fourth trained neural network, and 、 The method according to claim 19, further comprising:
21. The method according to claim 19 to 20, wherein the quantization process includes a first quantization approximation.
22. The determined quantity is, in addition to the rate associated with the quantized latent, A second quantization approximation is used to determine the rate associated with the quantized latent. And so, The second quantization approximation is different from the first quantization approximation. The method according to claim 21.
23. The determined quantity includes the loss function, and further modifies the parameters of the neural network. The new step described above is The steps include evaluating the gradient of the loss function, The gradient of the loss function is backpropagated through the neural network. Steps, including, During the backpropagation of the gradient of the loss function, a third quantization approximation is used. The third quantization approximation is the same as the first quantization approximation. The method according to claim 21 or 22.
24. The parameters of the neural network are based on the distribution of the sizes of the bins. The method according to any one of claims 19 to 23, which is further updated accordingly.
25. The method according to claim 24, wherein at least one parameter of the distribution is learned.
26. The method according to claim 24 or 25, wherein the distribution is an inverse gamma distribution.
27. The distribution is determined by a fifth neural network, according to claim 24 or 25. Method of description.
28. A method for encoding and transmitting lossy images or videos, wherein the method is The first computer system receives an input image, The input image is encoded using the first trained neural network to form a latent representation. The steps to generate, The step involves performing a quantization process on the latent representation to generate a quantized latent. The size of the bins used in the quantization process is based on the input image. Top, The step of transmitting the quantized latent, Methods that include...
29. A method for receiving and decoding a lost image or video, wherein the method is: The quantized latent transmitted according to the method described in claim 28 is used in a second computer Steps to receive data in the data system, The quantization is performed using a second trained neural network to generate the output image. The steps include decoding the latent, wherein the output image is an approximation of the input image. 、 method.
30. Data processing configured to perform the method described in any one of claims 1 to 31 system.
31. A data processing device configured to perform the method described in claim 28 or 29.
32. When the program is executed by a computer, the program according to claim 28 or 29 A computer program that includes instructions to cause the computer to perform the method described above.
33. When executed by a computer, the method according to claim 28 or 29 is performed by the computer A computer-readable storage medium containing instructions to be executed by a computer.
34. A method for training one or more neural networks, wherein the one or more neural The network is used for encoding, transmitting, and decoding lost images or videos. The method described above is The first computer system receives an input image, The input image is encoded using the first neural network to generate a latent representation. Steps and The steps include: performing a quantization process on the latent representation to generate a quantized latent; The quantized latent is decoded using a second neural network, and the output image is generated. A step of generating, wherein the output image is an approximation of the input image, A step of determining a quantity based on the difference between the output image and the input image, Based on the determined quantity, the first neural network and the second neural network The steps include updating the parameters of the network, The above steps are repeated using multiple input image sets to obtain the first trained nucleus. Steps to generate a first neural network and a second trained neural network. And, At least one of the plurality of sets of the input images is a first image containing a specific feature Including the proportion, At least one of the multiple sets of input images has the specific feature Includes the second proportion of the included images, The second ratio is different from the first ratio, in steps, Methods that include...
35. The first proportion is all of the images in the set of input images, as described in claim 34. The method.
36. The aforementioned specific features include human faces, animal faces, text, eyes, lips, logos, cars, flowers, and patterns. The method according to claim 34 or 35, which is one of the methods.
37. Each of the plurality of sets of input images is an equal number of times during the iteration of the method step. The method according to any one of claims 34 to 36, as used.
38. The difference between the output image and the input image acts as a discriminator in the neural network. As determined at least partially by the workpiece, according to any one of claims 34 to 37 The method.
39. Separate neural networks acting as classifiers analyze the aforementioned set of input images. The method according to any one of claims 34 to 38, used for each set.
40. One or more of the parameters of the neural network acting as a classifier are trained The first number of steps that have been performed is updated, One or more other neural networks of the neural network acting as a discriminator The mark is updated for the trained steps of the second number, and the second number is the first Lower than the number The method according to claim 39.
41. The determined quantity is, in addition to the rate associated with the quantized latent, The update of the parameter for at least one of the multiple sets of the input images is performed as follows: Using a first weighting for the rate associated with the quantized latent, Before the parameter for at least one other set of the plurality of sets of input images The update uses a second weighting for the rate associated with the quantized latent, and The second weighting is different from the first weighting. The method according to any one of claims 34 to 40.
42. The difference between the output image and the input image is at least partially a plurality of perceptual metrics Determined using riks, The update of the parameter for at least one of the sets of the plurality of input images However, using the first set of weights for the aforementioned multiple perceptual metrics, The parameter for at least one other set of the plurality of sets of the input images The aforementioned update of the data uses a second set of weights for the plurality of perceptual metrics, The second set of weights is different from the first set of weights. The method according to any one of claims 34 to 41.
43. The aforementioned input image has one or more regions of interest, which are analyzed by a third trained neural network. A modified image identified by and with other areas of the aforementioned image masked, according to claims 34 to 42. The method described in item 1.
44. The aforementioned area of interest is one or more of the following: human face, animal face, letters, eyes, lips, logo, car, flower, pattern. The method according to claim 43, which is a region including the above features.
45. The location of the area of the one or more regions of interest is stored in a binary mask, claim The method described in 43 or 44.
46. The binary mask is an additional input to the first neural network, according to the claim. The method described in 45.
47. A method for encoding, transmitting, and decoding lossy images and videos, wherein the method is The first computer system receives an input image, The input image is encoded using the first trained neural network to form a latent representation. The steps to generate, The steps include: performing a quantization process on the latent representation to generate a quantized latent; The steps include transmitting the quantized latent to a second computer system, The quantized latent is decoded using a second trained neural network. The steps include: generating an output image, wherein the output image is an approximation of the input image; The first trained neural network and the second trained neural network The work is trained according to the method described in any one of claims 34 to 46. That is the case. method.
48. A method for encoding and transmitting lossy images or videos, wherein the method is The first computer system receives an input image, The input image is encoded using the first trained neural network to form a latent representation. The steps to generate, The steps include: performing a quantization process on the latent representation to generate a quantized latent; The step of transmitting the quantized latent, The first trained neural network is described in any one of claims 34 to 46. It is trained in accordance with the aforementioned method. method.
49. A method for receiving and decoding a lost image or video, wherein the method is The quantized latent according to the method described in claim 48 is used in a second computer system Steps to receive via Tem, The quantized latent is decoded using a second trained neural network. A step of generating an output image, wherein the output image is an approximation of the input image. Includes, The second trained neural network is described in any one of claims 34 to 46. Trained according to the aforementioned method, method.
50. A data processing system configured to perform the method described in any one of claims 34 to 46 A rational system.
51. A data processing device configured to perform the method described in claim 48 or 49.
52. When the program is executed by a computer, the program described in claim 48 or 49 A computer program that includes instructions to cause the computer to perform the method described above.
53. When executed by a computer, the method described in claim 48 or 49 is performed by the computer A computer-readable storage medium containing instructions to be executed by a computer.
54. A method for training one or more neural networks, wherein the one or more neural The network is used for encoding, transmitting, and decoding lost images or videos. The method described above is The first computer system receives an input image, The input image is encoded using the first neural network to generate a latent representation. Steps and The steps include: performing a quantization process on the latent representation to generate a quantized latent; The quantized latent is decoded using a second neural network, and the output image is generated. A step of generating an image wherein the output image is an approximation of the input image, A step of determining a quantity based on a rate associated with the quantized potential, The evaluation of the rate includes the step of interpolating a discrete probability mass function, Based on the determined quantity, the first neural network and the second neural network The steps include updating the parameters of the network, The above steps are repeated using multiple sets of input images to obtain the first trained N Steps to generate a neural network and a second trained neural network P and, Methods that include...
55. At least one parameter of the discrete probability mass function is based on the evaluated rate. The method according to claim 54, further updated.
56. A third neural network is used to encode the latent representation and generate a hyperlatent representation. The steps to take, quantization is performed on the aforementioned hyperlatent representation to obtain the quantized hyperlatent representation. The steps to generate, The quantized hyperlatent representation is decoded using a fourth neural network. , a step of obtaining at least one parameter of the discrete probability mass function, and further Including, The parameters of the third neural network and the fourth neural network The meter is further updated based on the determined amount, and a third trained neural Obtain a new network and a fourth trained neural network. The method according to claim 54 or 55.
57. The interpolation methods include piecewise constant interpolation, nearest neighbor interpolation, linear interpolation, polynomial interpolation, spline interpolation, and section interpolation. Claim 54, comprising at least one of segmental cubic interpolation, Gaussian process, and kriging. The method described in any one of paragraphs 56 to 56.
58. The discrete probability mass function is a categorical distribution, according to any one of claims 54 to 57. The aforementioned method.
59. Claim 5, wherein the categorical distribution is parameterized by at least one vector. The method described in 8.
60. The category distribution is obtained by softmax projection of a vector, as described in claim 59. The aforementioned method.
61. The discrete probability mass function is determined by at least the mean parameter and the scale parameter. The method according to any one of claims 54 to 57, which is parameterized.
62. The method according to any one of claims 54 to 61, wherein the discrete probability mass function is multivariate. Law.
63. The discrete probability mass function includes multiple points, The first pair of adjacent points of the plurality of points have a first interval between them. The second pair of adjacent points of the plurality of points has a second interval, and the second interval is the Different from the interval of 1, A discrete probability mass function according to any one of claims 54 to 62.
64. The discrete probability mass function includes multiple points, The first pair of adjacent points of the plurality of points have a first interval between them. The second pair of adjacent points of the plurality of points has a second interval, and the second interval is the Equal to the interval of 1, The method according to any one of claims 54 to 62.
65. At least one of the first interval and the second interval is a fourth neural network The method according to claim 63 or 64, obtained using the k.
66. At least one of the first interval and the second interval is at least one of the latent expressions The method according to claim 63 or 64, which is obtained based on the value of one pixel.
67. A method for encoding, transmitting, and decoding lossy images and videos, wherein the method is The first computer system receives an input image, The input image is encoded using the first trained neural network to form a latent representation. The steps to generate and The steps include: performing a quantization process on the latent representation to generate a quantized latent; The steps include transmitting the quantized latent to a second computer system, The quantized latent is decoded using a second trained neural network. A step of generating a force image, wherein the output image is an approximation of the input image. Includes, The first trained neural network and the second trained neural network The work is trained according to the method described in any one of claims 54 to 66. That is the case. method.
68. A method for encoding and transmitting lossy images or videos, wherein the method is The first computer system receives an input image, The input image is encoded using the first trained neural network to form a latent representation. The steps to generate and The steps include: performing a quantization process on the latent representation to generate a quantized latent; The step of transmitting the quantized latent, The first trained neural network is described in any one of claims 54 to 66. Trained according to the aforementioned method, method. .
69. A method for receiving and decoding a lost image or video, wherein the method is The quantized latent is converted to a second computer system according to the method described in claim 68. The steps to receive via M, The quantized latent is decoded using a second trained neural network. A step of generating an output image, wherein the output image is an approximation of the input image. Includes, The second trained neural network is described in any one of claims 54 to 66. Trained according to the aforementioned method, method.
70. A data processor configured to perform the method described in any one of claims 54 to 67 A rational system.
71. A data processing device configured to perform the method described in claim 68 or 69.
72. When the program is executed by a computer, the program described in claim 68 or 69 A computer program that includes instructions to cause the computer to perform the method described above.
73. When executed by a computer, the computer is described in claim 68 or 69 A computer-readable storage medium including an instruction to perform the method described above.
74. A method for encoding, transmitting, and decoding lossy images and videos, wherein the method is The first computer system receives an input image. The input image is encoded using the first trained neural network to form a latent representation. The steps to generate, The steps include: performing a first operation on the latent representation to obtain the residual latent; The steps include transmitting the residual potential to a second computer system, A second operation is performed on the residual latent to obtain the acquired latent representation, The second operation then performs the operation on the acquired latent previously obtained pixels. Steps and Using a second trained neural network, the acquired latent representation is decoded. A step of generating the output image, wherein the output image is an approximation of the input image. Steps and Methods that include...
75. The operation on the previously obtained latent pixels is performed on the previously obtained pixels The cell is performed on each of the acquired latent pixels as described in claim 74. The method.
76. At least one of the first and second operations includes a method for solving the implicit system of equations. The method according to claim 74 or 75.
77. The method according to any one of claims 74 to 76, wherein the first operation includes a quantization operation.
78. The operation performed on the acquired latent previously obtained pixels is a matrix operation. The method according to any one of claims 74 to 77, including the method described in any one of claims 74 to 77.
79. The method according to claim 78, wherein the matrix defining the matrix operation is sparse.
80. The matrix that defines the matrix operation is not obtained when the matrix operation is performed. The method according to claim 78 or 79, which has zero values corresponding to the acquired latent pixels. method.
81. The matrix defining the matrix operation is a lower triangular matrix, as described in any one of claims 78 to 80. Method of loading.
82. The method according to claim 81, wherein the second operation includes a standard forward substitution.
83. The operation performed on the acquired latent previously obtained pixels is the third lesson A method according to any one of claims 74 to 82, comprising a refined neural network. 。
84. The representation is encoded using a fourth trained neural network, and the hyperlatency Steps to generate the representation, Step to transmit the quantized hyperlatent representation to the second computer system P and, The quantized hyperlatency is decoded using a fifth trained neural network. A step of converting, which is performed on the acquired latent previously obtained pixels. The operation performed is based on the output of the fifth trained neural network, step P and, The method according to any one of claims 74 to 83, further comprising:
85. If dependent on claim 76, the fifth trained neural network is used The decoding of the quantized hyperlatency additionally generates average parameters, The implicit equation system further includes the mean parameter, The method according to claim 84.
86. A method for training one or more neural networks, wherein the one or more neural The network is used for encoding, transmitting, and decoding lossy images or videos. The method described above is The first computer system receives an input image, The input image is encoded using the first neural network to generate a latent representation. Steps and The first step is to perform an operation on the latent representation to obtain the residual latent, A step of performing a second operation on the residual latent to obtain the acquired latent representation, The second operation then performs the operation on the acquired latent previously obtained pixels. Steps that include, The quantized latent is decoded using a second neural network to generate an output image. A step in which the output image approximates the input image, A step of determining a quantity based on the difference between the output image and the input image, Based on the determined quantity, the first neural network and the second neural network The steps include updating the parameters of the network, The above steps are repeated using the first set of input images to obtain the first trained N Steps to generate a neural network and a second trained neural network P and, Methods that include...
87. The calculation performed on the acquired latent previously obtained pixels is: The method according to claim 86, comprising matrix operations.
88. The parameters of the matrix defining the matrix operation are additionally determined based on the determined quantity. The method according to claim 87, as updated.
89. The operation performed on the acquired latent previously obtained pixels is a third Including neural networks, The parameters of the third neural network are adjusted based on the determined quantity. It is additively updated to generate a third trained neural network. The method according to any one of claims 86 to 88.
90. The latent representation is encoded using a fourth neural network to obtain a hyperlatent representation. The steps to generate, The hyperlatent representation is subjected to quantization to generate a quantized hyperlatent. The steps, The steps include: transmitting the quantized hyperpotential to the second computer system; 、 A fifth neural network is used to decode the quantized hyperlatency. A step which is performed on the acquired latent previously obtained pixels. The calculation is based on the output of the fifth trained neural network, and , further including, The parameters of the fourth neural network and the fifth neural network The meter is further updated based on the determined amount, and a fourth trained neural Generates a neural network and a fifth trained neural network. The method according to any one of claims 86 to 89.
91. A method for encoding and transmitting lossy images or videos, wherein the method is The first computer system receives an input image. The input image is encoded using the first trained neural network to form a latent representation. The steps to generate, The steps include: performing a first operation on the latent representation to obtain the residual latent; The step of transmitting the residual potential, Methods that include...
92. A method for receiving and decoding a lost image or video, wherein the method is The residual potential transmitted according to the method described in claim 91 is sent to a second computer system Steps to receive via Tem, A step in which a second operation is performed on the residual latent to obtain the acquired latent representation. The second operation then performs the operation on the acquired latent previously obtained pixels. The steps to do, The acquired latent representation is decoded using a second trained neural network. A step of generating an output image, wherein the output image is an approximation of the input image. Step and, Methods that include...
93. A data processor configured to perform the method described in any one of claims 74 to 90 A rational system.
94. A data processing device configured to perform the method described in claim 91 or 92.
95. When the program is executed by a computer, the program according to claim 91 or 92 A computer program that includes instructions to cause the computer to perform the method described above.
96. When executed by a computer, the computer is described in claim 91 or 92. A computer-readable storage medium containing instructions for performing the aforementioned method.
97. A method for training one or more neural networks, wherein the one or more neural The network is used for encoding, transmitting, and decoding lost images or videos. The method described above is The first computer system receives an input image, The input image is encoded using the first neural network to generate a latent representation. Steps and The steps include entropy encoding the aforementioned latent representation, Steps to transmit the entropy-encoded latent representation to a second computer system P and, The steps include entropy decoding the entropy-encoded latent representation, The latent representation is decoded using a second neural network to generate an output image. A step in which the output image is an approximation of the input image, A step of determining a quantity based on the difference between the output image and the input image, Based on the determined quantity, the first neural network and the second neural network The steps include updating the parameters of the network, The above steps are repeated using the first set of input images to obtain the first trained N Steps to generate a neural network and a second trained neural network Includes, The entropy decoding of the entropy-encoded latent representation is performed entropy by entropy It was executed, The order of decoding for each entropy is further updated based on the determined amount. to be done, method.
98. The order of decoding each pixel is based on the latent representation, as described in claim 97. Law.
99. The entropy-encoded latent, the entropy-decoded, is previously decoded The method according to claim 97 or claim 98, comprising calculations based on Xels.
100. The determination of the order of the pixel-by-pixel decoding in the directed aperiodic graph Any of claims 97 to 99, comprising ordering a plurality of the aforementioned pixels of the representation. The method described in paragraph 1.
101. The determination of the order of decoding for each pixel is performed using a plurality of adjacency matrices to determine the latent table The method according to any one of claims 97 to 100, including operating on the current state.
102. The determination of the order of the pixel-by-pixel decoding divides the latent representation into a plurality of sub-images. The method according to any one of claims 97 to 101, comprising dividing.
103. The aforementioned sub-images are used to convolve the latent representation with multiple binary mask kernels. The method described in claim 102, obtained thereby.
104. The determination of the order of decoding for each pixel depends on the magnitude of the amount associated with each pixel. Claims 97 to 97 include ranking a plurality of pixels of the latent representation based on the The method described in any one of paragraphs 103.
105. The quantity associated with each pixel is the position or scale parameter associated with that pixel. The method according to claim 104, wherein the data is a data.
106. The amount associated with each pixel is further updated based on the evaluated difference, The method described in paragraph 104 or 105.
107. The determination of the order of decoding for each pixel is the order of the multiple pixels of the latent representation The method according to any one of claims 97 to 106, including beamlet decomposition.
108. The sequence of decoding for each pixel is related to the wave number associated with the plurality of pixels. The method according to claim 107, based on the frequency components of the decomposition.
109. The latent representation is encoded using a fourth trained neural network, Steps to generate latent representations, The steps include transmitting the hyperpotential to the second computer system, Step 5: Decode the hyperlatency using a trained neural network. The sequence of the pixel-by-pixel decoding is the fifth trained neural The steps, based on the output of the network, further include: The method according to any one of claims 97 to 108.
110. A method for encoding, transmitting, and decoding lossy images and videos, wherein the method is The first computer system receives an input image, The input image is encoded using the first trained neural network to form a latent representation. The steps to generate, The steps include entropy encoding the aforementioned latent representation, Steps to transmit the entropy-encoded latent representation to a second computer system P and, The steps of entropy decoding an entropy-encoded latent representation and The latent is decoded using a second trained neural network to produce an output image. A step comprising the steps of: fruit, The first trained neural network and the second trained neural network The work is trained according to the method described in any one of claims 97 to 109. It is method.
111. A method for encoding and transmitting lossy images or videos, wherein the method is The first computer system receives an input image, The input image is encoded using the first trained neural network to form a latent representation. The steps to generate, The steps of entropy encoding the latent representation and The step of transmitting the entropy-encoded latent representation includes, The first trained neural network is described in any one of claims 97 to 109. Trained according to the method described above, method.
112. A method for receiving and decoding a lost image or video, wherein the method is The entropy-encoded latent table transmitted according to a method of claim 111 The steps include receiving the data in a second computer system, The latent representation is decoded using a second trained neural network, and the output image is generated. A step of generating an output image, wherein the output image is an approximation of the input image, Includes, The second trained neural network is described in any one of claims 97 to 109. Trained according to the method described above, method.
113. Data configured to perform the method described in any one of claims 97 to 110 Processing system.
114. A data processing device configured to perform the method described in claim 111 or 112 。
115. When the program is executed by a computer, the following applies: A computer program that includes instructions to cause the computer to execute the method described above.
116. When executed by a computer, the method according to claim 111 or 112 A computer-readable storage medium containing instructions to be executed by a computer.
117. A method for training one or more neural networks, wherein the one or more neural The network is used for encoding, transmitting, and decoding lost images or videos. The method described above is The first computer system receives an input image, The input image is encoded using the first neural network to generate a latent representation. Steps and The latent representation is decoded using a second neural network to generate an output image. A step in which the output image is an approximation of the input image, Based on the difference between the output image and the input image and the rate associated with the latent representation A step in which the quantity is determined, wherein when the quantity is determined, the output image and the input image and A first weight is applied to the difference between them, and a second weight is applied to the rate associated with the latent representation. The steps to which the attachment is applied, Based on the determined quantity, the first neural network and the second neural network The steps include updating the parameters of the network, The above steps are repeated using the first set of input images to obtain the first trained N Steps to generate a neural network and a second trained neural network Includes, After at least one iteration of the above step, the first weighting and before At least one of the second weightings is further updated based on a further amount, and the further The quantity is related to the difference between the output image and the input image and the latent representation. Based on at least one of the rates, method.
118. The difference between the output image and the input image and the rate associated with the latent representation At least one of them is recorded for each 20 repetitions of the above step, The further amount is the multiple previously recorded differences between the output image and the input image. minutes and at least of the previously recorded rates related to the aforementioned potential expression Based on one side, The method according to claim 117.
119. The further amount is based on the average of a plurality of previously recorded differences or rates, according to the claim. The method described in 118.
120. The mean is the arithmetic mean, the median, the geometric mean, the harmonic mean, the exponential moving mean. The claim is that at least one of the mean, the smoothed moving average, and the linear weighted moving average. The method described in 119.
121. Before determining the aforementioned further amount, the quantity anomaly is determined by the difference or rate of the multiple previously recorded values. The method according to any one of claims 118 to 120, wherein the removal is performed from.
122. The abnormal value is removed only for the first predetermined number of repetitions of the step, The method described in item 121.
123. The rate associated with the latent expression is calculated using the first method when determining the quantity. And when determining further amounts, the second method is used for calculation, and the first method is used for calculation. The method according to any one of claims 117 to 122, which differs from the second method.
124. At least one repetition of the above step involves selecting an input image from a second set of input images. It is used and executed, The parameters of the first neural network and the second neural network The meter is not updated when an input image from a second set of input images is used. The method according to any one of claims 117 to 123.
125. The determined quantity is added to the output of the neural network acting as a discriminator. The method according to any one of claims 117 to 124, based on the present invention.
126. A method for training one or more neural networks, wherein the one or more neural The network is intended for use in lossy video coding, transmission, and decoding. The method described above is A first computer system receives input video, The first neural network is used to encode multiple frames of the input video, Steps to generate multiple latent representations, The plurality of latent representations are decoded using a second neural network, and the output video A step of generating multiple frames, wherein the output image is an approximation of the input image. , steps and, The difference between the output video and the input video and the rates associated with the plurality of latent representations. A step of determining an amount based on the first weighting of the output video and the input The difference between the video and the second weighting is applied to the plurality of latent representations The steps applied to the rate, Based on the amount determined above, the first neural network and the second neural network The steps include updating the parameters of the neural network, The above steps are repeated using multiple input videos to obtain the first trained neural network. The process includes the steps of generating a network and a second trained neural network. 、 After at least one iteration of the above step, the first weighting and before At least one of the second weightings is further updated based on a further amount, and the further The quantity is related to the difference between the output image and the input image and the latent representation. Based on at least one of the rates, method.
127. The input video comprises at least one I-frame and a plurality of P-frames, The method described in paragraph 126.
128. The aforementioned quantity is based on a plurality of first weightings or second weightings, and each of the weightings is , corresponding to one of the plurality of frames of the input video, claim 126 or 1 The method described in 27.
129. After at least one of the iterations of the above step, of the plurality of weights At least one of the bills is updated based on the additional amount associated with each weighting. The method described in paragraph 128.
130. Each of the aforementioned additional amounts is the difference between the output frame and the input frame or the latent representation. The method according to claim 129, based on a predetermined target value of the relevant rate.
131. The additional amount related to the I-frame has a first target value, and the smaller amount related to the P-frame has At least one additional amount has a second target value, and the second target value is different from the first target value. The method according to claim 130.
132. The method described in claim 130 or 131, wherein each additional amount related to the P-frame has the same target value. The aforementioned method.
133. Claim 1, wherein the plurality of first weights or second weights are initially set to zero. The method described in any one of paragraphs 28 to 132.
134. A method for encoding, transmitting, and decoding lossy images and videos, wherein the method is The first computer system receives an input image, The input image is encoded using the first trained neural network to form a latent representation. The steps to generate, The steps include: performing a quantization process on the latent representation to generate a quantized latent; The steps include transmitting the quantized latent to a second computer system, The quantized latent is decoded using a second trained neural network. A step of generating an output image, wherein the output image is an approximation of the input image. Includes, The first trained neural network and the second trained neural network The work is trained according to the method described in any one of claims 117 to 133. It was, method.
135. A method for encoding and transmitting lossy images or videos, wherein the method is The steps include receiving an input image or video in a first computer system, The input image or video is encoded using a first trained neural network. Then, the step of generating a latent expression, The step of transmitting the latent expression, The first trained neural network is any one of claims 117 to 133. Trained according to the method described above, method.
136. A method for receiving and decoding a lost image or video, wherein the method is The latent representation is received by a second computer system according to the method described in claim 135. Steps to believe, The latent representation is decoded using a second trained neural network, and the output image is also The step of generating an image, wherein the output image or video is the input image or video It includes a step, which is an approximation of deo. The second trained neural network is any one of claims 117 to 133. Trained according to the method described above, method.
137. Data configured to perform the method described in any one of claims 117 to 134 Processing system.
138. A data processing device configured to perform the method described in claim 135 or 136 。
139. When the program is executed by a computer, the following applies: A computer program that includes instructions to cause the computer to execute the method described above.
140. When executed by a computer, the method according to claim 135 or 136 A computer-readable storage medium containing instructions to be executed by a computer.
141. A method for encoding, transmitting, and decoding lossy images or videos, wherein the method is The first computer system receives an input image, The input image is encoded using the first trained neural network to generate a latent representation. The steps to take, The steps include: performing a quantization process on the latent representation to generate a quantized latent; A step of entropy coding the quantized latent using a probability distribution, wherein The following probability distributions are defined using tensor networks, with steps and The entropy-encoded quantized latent is transmitted to a second computer system. The steps, The entropy-encoded quantized latent is entropy-recovered using the probability distribution The steps include: quantizing the latent to obtain the quantized latent, The quantized latent is decoded using a second trained neural network. A step of generating an output image, wherein the output image is an approximation of the input image. Top and, including, method.
142. The probability distribution is defined by a Hermitian operator operating on the quantized latent. Claim 14, wherein the Hermitian operator is defined by the tensor network. The method described in 1.
143. The tensor network comprises a non-normal orthonormal core tensor and one or more orthonormal tensors. The method according to claim 141 or claim 142, including the ru.
144. The latent representation is encoded using a third trained neural network, Steps to generate latent representations, Quantization is performed on the hyperlatent representation to generate the quantized hyperlatent. The steps to take, The steps include: transmitting the quantized hyperpotential to the second computer system; 、 Using a fourth trained neural network, the quantized hyperlatency is restored. This further includes the step of assigning a number, The output of the fourth trained neural network is the tensor network One or more parameters of The method according to any one of claims 141 to 143.
145. When it conforms to claim 143, the tensor network is a non-normalized core tensor and It contains one or more orthonormal tensors, The output of the fourth trained neural network is the unnormalized core tensor One or more parameters, The method according to claim 144.
146. One or more parameters of the tensor network correspond to one or more pixels of the latent representation. The method according to any one of claims 141 to 143, which is calculated using
147. Claims 141 to 141, wherein the probability distribution relates to a subset of the pixels of the latent representation. The method described in any one of paragraphs 146.
148. Any of claims 141 to 146, wherein the probability distribution relates to the channel of the latent representation. The method described in paragraph 1.
149. The aforementioned tensor network is a tensor tree, a locally refined state, a bone machine. , the state of the matrix product, and at least one of the factorizations of the state of the projected entangled pair The method according to any one of claims 141 to 148.
150. A method for training one or more networks, wherein the one or more networks lose For use in encoding, transmitting and decoding images or videos, the method 、 The steps include receiving a first input image, The first input image is encoded using the first neural network to generate a latent representation. The steps to accomplish, The steps include: performing a quantization process on the latent representation to generate a quantized latent; A step of entropy coding the quantized latent using a probability distribution, wherein The following probability distributions are defined using tensor networks, with steps and Using the aforementioned probability distribution, the entropy-encoded quantized latent is entropy-recovered The steps include: quantizing the latent to obtain the quantized latent, The quantized latent is decoded using a second neural network, and the output image is generated. A step of generating an image wherein the output image is an approximation of the input image, A step of determining a quantity based on the difference between the output image and the input image, Based on the determined quantity, the first neural network and the second neural network The steps include updating the parameters of the network, The above steps are repeated using multiple input images to obtain a first trained neural network. The steps include generating a network and a second trained neural network, include, method.
151. One or more of the method parameters of the tensor network are based on the determined quantity. The method according to claim 150, further updated thereto.
152. The tensor network comprises a non-normal orthonormal core tensor and one or more orthonormal tensors. Includes, All of the tensor networks except for the orthonormal core tensor The parameters of the twerk are updated based on the determined amount. The method according to claim 151.
153. Claims 150-152, wherein the tensor network is computed using the latent representation. The method described in any one of the paragraphs.
154. Claim 1, wherein the tensor network is computed based on linear interpolation of the latent representation. The method described in 53.
155. The determined quantity is additionally based on the entropy of the tensor network. The method according to any one of claims 150 to 154, which is determined.
156. A method for encoding and transmitting lossy images or videos, wherein the method is The first computer system receives an input image, The input image is encoded using the first trained neural network to form a latent representation. The steps to generate, The steps include: performing a quantization process on the latent representation to generate a quantized latent; A step of entropy coding the quantized latent using a probability distribution, wherein The specified probability distribution is defined using a tensor network, and the steps and The step of transmitting the entropy-encoded quantized latent, Methods that include...
157. A method for receiving and decoding irreversible images or videos, wherein the method is Entropy-encoded quantized transmitted according to the method described in claim 156 The steps include receiving the potential data in a second computer system, Using the aforementioned probability distribution, the entropy-decoded quantized latent is entropy-decoded Steps to obtain the quantized potential by converting it to a number, The quantized latent is decoded using a second trained neural network. A step of generating an output image, wherein the output image is an approximation of the input image. Includes, method.
158. A data configured to carry out the method described in any one of claims 141 to 155 Processing system.
159. A data processing device configured to perform the method described in claim 156 or 157 。
160. When the program is executed by a computer, the following applies: A computer program that includes instructions to cause the computer to execute the method described above.
161. When executed by a computer, the method according to claim 156 or 157 A computer-readable storage medium containing instructions to be executed by a computer.
162. A method for encoding, transmitting, and decoding lossy images and videos, wherein the method is The first computer system receives an input image, The latent representation is encoded using the first trained neural network, and the latent representation The steps to generate, The latent representation is encoded using a second trained neural network, Steps to generate latent representations, The hyperlatent representation is encoded using a third trained neural network, Steps to generate hyperhyperlatent representations, The aforementioned latent expression, hyper latent expression, and hyperhyper latent expression are used in a second computer Steps to send to the data system, Using the output of the fourth trained neural network, The steps to decode the representation, The fourth trained neural network and the fifth trained neural network The steps include: decoding the hyperlatent representation using the output of the workpiece; The fifth trained neural network and the sixth trained neural network The step involves decoding the latent representation using the output of the workpiece to generate an output image. This includes a step in which the output image is an approximation of the input image, method.
163. The process further includes the step of determining the rate of the input image, wherein the determined rate is If certain conditions are met, the hyperlatent representation is encoded, and the hyperhyperlatent representation is decoded. The method according to claim 162, wherein the step of transforming is not performed.
164. The predetermined condition is that the rate is less than a predetermined value, as described in claim 163. How to write.
165. A method for training one or more networks, wherein the one or more networks However, it is intended for use in encoding, transmitting, and decoding lossy images or videos, and prior The notation method is The first computer system receives an input image, The input image is encoded using the first neural network to generate a latent representation. Steps and The latent representation is encoded using a second neural network to obtain a hyperlatent representation. The steps to generate, The hyper latent representation is encoded using a third neural network, The steps include generating a latent representation and The fourth neural network is used to decode the hyperhyperlatent representation. Top, Using the outputs of the fourth neural network and the fifth neural network The steps include decoding the aforementioned hyperlatent representation, Using the outputs of the fifth neural network and the sixth neural network The step of decoding the latent representation and generating an output image, wherein the output image is The approximation of the input image is a step, A step of determining a quantity based on the difference between the output image and the input image, Based on the determined quantity, the parameters of the third and fourth neural networks Steps to update the meter, The above steps are repeated using multiple input images to obtain a third and fourth trained The steps include generating a neural network, method.
166. The method parameters of the first, second, fifth, and sixth neural networks are, The update is not performed in at least one of the iterations of the step as described in claim 165. The aforementioned method.
167. The step of determining the rate of the input image further includes, If the rate determined above satisfies the predetermined conditions, then the first, second, fifth and sixth of the above The parameters of the neural network are updated in the iteration of the step. The method according to claim 165, which is not applicable.
168. The predetermined condition is that the rate is less than a predetermined value, as described in claim 167. The aforementioned method.
169. The parameters of the first, second, fifth, and sixth neural networks are the same as the state Any one of claims 166 to 168, the update is not performed after a predetermined number of repetitions of the update. The method described above.
170. The parameters of the first, second, fifth, and sixth neural networks are determined Based on the amount obtained, the first, second, fifth, and sixth trained nu The method according to any one of claims 165 to 169, which generates a holographic network. 。
171. Before performing the other steps described above, for at least one of the plurality of input images, At least one of the operations of psample, smoothing filter, and random crop is implemented. The method according to any one of claims 165 to 170, which is carried out.
172. A method for encoding and transmitting lossy images or videos, wherein the method is The first computer system receives an input image, The input image is encoded using the first trained neural network to form a latent representation. The steps to generate, The latent representation is encoded using a second trained neural network, Steps to generate latent representations, The hyperlatent representation is encoded using a third trained neural network, Steps to generate hyperhyperlatent representations, The step includes transmitting the latent, hyperlatent, and hyperhyperlatent representations. ,method.
173. A method for receiving and decoding a lost image or video, wherein the method is The latent, hyper-latent and hyper-latent transmitted according to the method described in claim 172 - The step of receiving the hyperlatent expression in a second computer system, The hyperhyperlatent representation is decoded using a fourth trained neural network. The steps to transform, The output of the fourth trained neural network and the fifth trained neural network The steps include decoding the hyperlatent representation using a network, The output of the fifth trained neural network and the sixth trained neural network The step involves decoding the latent representation using a network to generate an output image. This includes a step in which the output image is an approximation of the input image, method.
174. Data configured to perform the method according to any one of claims 162 to 171 Processing system.
175. A data processing device configured to perform the method described in claim 172 or 173 。
176. When the program is executed by a computer, the following applies: A computer program that includes instructions to cause the computer to execute the method described above.
177. When executed by a computer, the method according to claim 172 or 173 A computer-readable storage medium containing instructions to be executed by a computer.