Method for encoding, decoding, training and related device

The neural network-based encoding and decoding method addresses inefficiencies in existing data compression techniques by learning optimal encoding processes for latent space features, resulting in efficient data compression and reduced bandwidth requirements.

WO2025127957A1PCT designated stage expired Publication Date: 2025-06-19HUAWEI TECH CO LTD +1

Patent Information

Application Number
PCT/RU2023/000384
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-12
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Existing data compression techniques, particularly entropy-based coding, face inefficiencies in encoding large datasets such as video and image data, leading to high bandwidth and storage requirements.

Method used

A neural network (NN) based encoding and decoding method that compresses different features of input data to varying numbers of bits, allowing the NN to learn optimal encoding and decoding processes for latent space features, and is tolerant to errors and losses through end-to-end training with noise simulation.

Benefits of technology

The NN-based solution reduces encoding and decoding complexity and time due to parallel processing, achieves efficient data compression with lower bit usage, and is compatible with existing digital communication infrastructure, making it suitable for real-time communication and video streaming applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure RU2023000384_19062025_PF_FP_ABST
    Figure RU2023000384_19062025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a method for coding, a method for decoding, a method for training, and related devices. The decoding method may include acquiring a bitstream D0; determining a reconstructed feature tensor ẑ0 according to the bitstream D0; performing a decoding operation on the reconstructed feature tensor ẑ0 to obtain a reconstructed latent representation ŷ0 by using a neural network; performing a synthesis transform operation based on the reconstructed latent representation ŷ0 to generate a reconstructed signal ẋ. Based on the method provided by the embodiments of the present application, different features of input data are compressed to different number of bits.
Need to check novelty before this filing date? Find Prior Art

Description

METHOD FOR ENCODING, DECODING, TRAINING AND RELATEDDEVICETECHNICAL FIELD

[0001] Embodiments of the present application relate to the field of signal processing technologies, and more specifically, to a method for encoding, a method for decoding, a method for training, and related devices.BACKGROUND

[0002] With the exponential growth of data, the demand for bandwidth and storage space has become a significant concern. For example, video data typically consists of numerous consecutive frames, each comprising millions of pixels. Consequently, processing and transmitting video data require substantial bandwidth and storage capacity. Similarly, image data can also be sizable, especially in the case of high-resolution images. These datasets necessitate considerable bandwidth for transmission and occupy significant storage space. One solution to address this challenge is through data compression techniques. By compressing the data, it becomes possible to reduce the required bandwidth for transmission and save storage space.

[0003] In the field of data compression, encoding plays a crucial role, and one commonly used method is entropy-based coding. However, the efficiency of entropy-based coding is not very high. Therefore, improving encoding efficiency is an important problem that needs to be addressed.SUMMARY

[0004] Embodiments of the present application provide a method for encoding, a method for decoding, a method for training, and related devices. Based on the method provided by theembodiments of the present application, different features of input data are compressed to different number of bits.

[0005] According to a first aspect, an embodiment of the present application provides a decoding method. The method may include acquiring a bitstream D0. The method may include determining a reconstructed feature tensor ẑ0according to the bitstream D0. The method may include performing a decoding operation on the reconstructed feature tensor ẑ0to obtain a reconstructed latent representation ẏo by using a neural network. The method may include performing a synthesis transform operation based on the reconstructed latent representation ẏo to generate a reconstructed signal ẋ.

[0006] The present application provides a neural network (NN) based video, image, audio codecs - source codecs (which are focused only on data compression), channel codecs ((which are focused only on data transmission) and joint source channel coding (which are focused both on data compression and transmission). The NN is used to encode and decode the input data into bitstream. Different features of the input data are compressed to different number of bits(‘words’). The NN may learn their own encoding / decoding way of latent space features represented as floats or integers to binary representation for sending via communication channel, find some correlations, and use less bits to encode data. Further, the ENC NN and the DEC NN may be end to end train together with noise - simulating both bit errors and packet losses - thus, becoming tolerant to errors, distortions and losses in communication channel.

[0007] Compared with conventional entropy encoding solutions, the NN based encoding solution provided by the present application have less complexity and less encoding / decoding time due to massive parallel processing on NN supporting hardware. Finally, the present application works in digital domain and do not making any constrains to physical layer having friendly design to existing widely deployed digital communication infrastructure. Therefore, the present application may be deployed with existing hardware (such as routers, Wi-Fi access points, 3G / 4G / 5G wireless communication devices).

[0008] In some embodiments, the decoding operation may be a decoding operation that is able to pass a loss gradient.

[0009] In a possible implementation of the first aspect, the method further includes: acquiring a decoding information, the determining a reconstructed feature tensor ẑ0accordingto the bitstream D0, includes: determining the reconstructed feature tensor ẑ0according to the decoding information and the bitstream D0, wherein a size of the reconstructed feature tensor ẑ0is larger than a size of the bitstream D0.

[0010] In a possible implementation of the first aspect, the decoding information is obtained based on information signaled with the bitstream D0.

[0011] In a possible implementation of the first aspect, the decoding information includes a length tensor Nbitso.

[0012] In a possible implementation of the first aspect, the determining the reconstructed feature tensor ẑ0according to the decoding information, includes: determining a mask tensor MSK0according to the length tensor Nbitso and an index tensor idxo, wherein the bitstream D0includes X1bit(s), the mask tensor MSK0includes X1first element(s) with a first value and X2second element(s) with a second value, X1, a sum of X1and X2is the size of the reconstructed feature tensor ẑ0, X1is a positive integer, X2is a positive integer; determining the reconstructed feature tensor ẑ0according to the mask tensor MSK0and the bitstream D0, wherein the reconstructed feature tensor ẑ0includes X1third element(s) and X2fourth element(s), a location of each element in the reconstructed feature tensor zo is a same as a location of a corresponding element in the mask tensor MSK0, a value of each of the X1third element(s) is a same as a value of a corresponding bit in the bitstream D0, and a value of each of the X2fourth element(s) is a preset value that is not equals to 1.

[0013] In a possible implementation of the first aspect, the method further includes: acquiring M bitstream(s), wherein M is a positive integer; determining M reconstructed feature tensor(s) according to the M bitstream(s) respectively; performing M decoding operation(s) on the M reconstructed feature tensor(s) to obtain M reconstructed latent representation(s); the performing a synthesis transform operation based on the reconstructed latent representation ŷ0to generate a reconstructed signal ẋ, includes: performing the synthesis transform operation based on the reconstructed latent representation ŷ0and the M reconstructed latent representation(s) to generate the reconstructed signal ẋ.

[0014] In a possible implementation of the first aspect, the performing the synthesis transform operation based on the reconstructed latent representation ŷ0and the M reconstructed latent representation(s) to generate the reconstructed signal ẋ, includes: performing thesynthesis transform operation on each of the reconstructed latent representation ẏo and the M reconstructed latent representation(s) to obtain M+1 synthesis transform results; generating the reconstructed signal ẋ according to the M+1 synthesis transformation results.

[0015] possible implementation of the first aspect, the performing the synthesis transform operation base on the reconstructed latent representation y0and the M reconstructed latent representation(s) to generate the reconstructed signal ẋ, includes: determining a reference reconstructed latent representation, wherein the reference reconstructed latent representation is a sum of the reconstructed latent representation ẏo and the M reconstructed latent representation(s); performing the synthesis transform operation on the reference reconstructed latent representation to generate the reconstructed signal ẋ

[0016] In a possible implementation of the first aspect, the performing the synthesis transform operation based on the reconstructed latent representation ẏo and the M reconstructed latent representation(s) to generate the reconstructed signal ẋ, includes: performing M+1 synthesis transform operations on M+1 reference information to generate the reconstructed signal ẋ, wherein a first reference information among the M+1 reference information is the reconstructed latent representation ẏo, a m+1threference information is determined according to a mthtransformed information and a mthreconstructed latent representation mthtransformed information is determined according to a synthesis transformation result of a mthsynthesis transform operation among the M+1 synthesis transform operations.

[0017] In a possible implementation of the first aspect, the method further includes: performing a variation operation on a reference reconstructed latent representation, wherein the reference reconstructed latent representation is a reconstructed latent representation determined by a reference decoding operation, the reference decoding operation is the decoding operation or one of the M decoding operation(s), the variation operation includes at least one of the followings: a sampling variation operation, or, a linear transformation, or an analysis transform operation, and the reconstructed signal ẋ is determined according to the varied reconstructed latent representation.

[0018] In a possible implementation of the first aspect, the acquiring a bitstream D0, includes: performing an entropy decoding operation on a received bitstream to obtain the bitstream D

[0019] In a possible implementation of the first aspect, the method further includes: obtaining an optimized parameter; the performing an entropy decoding operation on a received bitstream to obtain the bitstream D0, includes: performing, according to the optimized parameter, the entropy decoding operation on the received bitstream to obtain the bitstream D0.[ ] According to a second aspect, an embodiment of the present application provides an encoding method. The method may include acquiring an input signal. The method may include performing an analysis transform operation based on the input signal to obtain a latent representation y0. The method may include performing an encoding operation based on the latent representation y0to obtain a feature tensor z0by using a neural network; generating a bitstream D0according to the feature tensor z0.

[0021] The present application provides a neural network (NN) based video, image, audio codecs - source codecs (which are focused only on data compression), channel codecs ((which are focused only on data transmission) and joint source channel coding (which are focused both on data compression and transmission). The NN is used to encode and decode the input data into bitstream. Different features of the input data are compressed to different number of bits(‘words’). The NN may learn their own encoding / decoding way of latent space features represented as floats or integers to binary representation for sending via communication channel, find some correlations, and use less bits to encode data. Further, the ENC NN and the DEC NN may be end to end train together with noise - simulating both bit errors and packet losses - thus, becoming tolerant to errors, distortions and losses in communication channel.

[0022] Compared with conventional entropy encoding solutions, the NN based encoding solution provided by the present application have less complexity and less encoding / decoding time due to massive parallel processing on NN supporting hardware. Finally, the present application works in digital domain and do not making any constrains to physical layer having friendly design to existing widely deployed digital communication infrastructure. Therefore, the present application may be deployed with existing hardware (such as routers, Wi-Fi access points, 3G / 4G / 5G wireless communication devices).

[0023] In some embodiments, the encoding operation may be a decoding operation that is able to pass a loss gradient.

[0002] In a possible implementation of the second aspect, the method further includes:acquiring an encoding information; the generating a bitstream D0according to the feature tensor z0, includes: generating the bitstream D0according to the encoding information and the feature tensor z0, wherein a size of the feature tensor z0is larger than a size of the bitstream D0.

[0025] In a possible implementation of the second aspect, the encoding information is signaled with the bitstream DO.

[0026] In a possible implementation of the second aspect, the encoding information is a length tensor NbitsO.

[0027] In a possible implementation of the second aspect, the length tensor Nbisto is obtained by the encoding operation using the neural network.

[0028] In a possible implementation of the second aspect, the generating the bitstream D0according to the encoding information and the feature tensor z0, includes: determining a mask tensor MSK0according to the length tensor NbitsO and an index tensor idxo, wherein the mask tensor MSK0includes X1first element(s) with a first value and X2second element(s) with a second value, a sum of X1and X2is the size of the feature tensor z0, X1is a positive integer, X2is a positive integer; determining the bitstream D0according to the mask tensor MSK0and the feature tensor z0, wherein the bitstream D0includes X1bit(s), the feature tensor z0includes X1third element(s) and X2fourth element(s), a location of each elements in the feature tensor z0is a same of a location of a corresponding element in the mask tensor MSK0, a value of each of the X1third element(s) is configured to obtain a value of a corresponding bit in the bitstream D0.

[0029] In a possible implementation of the second aspect, the method further includes: performing M encoding operation(s) and M decoding operation(s) according to the bitstream D0to obtain M bitstream(s), wherein a mthbitstream among the M bitstream(s) is generated by performing a mthencoding operation among the M encoding operation(s) on an input data INPm, and the input data INPmis determined according to m reconstructed latent representation(s) obtained according to a first m-1 decoding operation(s) of the M decoding operation(s), m=1 , ... ,M.

[0030] In a possible implementation of the second aspect, the input data INPmand the m reconstructed latent representation(s) satisfy:

[0031] INPm= AT^x - £ ST(yi)

[0032] wherein INPmis the input data INPm, AT ( ) denotes the analysis transform operation, ST ( ) denotes a synthesis transform operation, x is the input signal, and ẏi is an ithreconstructed latent representation among the m reconstructed latent representation(s).

[0033] In a possible implementation of the second aspect, the input data INPmand the m reconstructed latent representation(s) satisfy:

[0035] wherein INPmis the input data INPm, y0is the latent representation y0, and ẏi is an ithreconstructed latent representation among the m reconstructed latent representation(s).

[0036] In a possible implementation of the second aspect, the input data INPmand the m reconstructed latent representation(s) satisfy:

[0037]

[0038] wherein INPmis the input data INPm, ATK-mis a result of a (K-m)thanalysis transform operation among K analysis transform operation(s); STmis a result of a mthsynthesis transform operation among K synthesis transform operation(s), and the mthsynthesis transform operation is performed on a target data determined according to the m reconstructed latent representation(s)

[0039] In a possible implementation of the second aspect, the method further includes: performing a first variation operation on a reference input data, wherein the reference input data is an input data of a reference encoding operation, the reference encoding operation is the encoding operation or one of the M encoding operation(s), the first variation operation includes at least one of the followings: a sampling variation operation, or, a linear transformation, or an analysis transform operation, and the reference encoding operation is performed on the varied first input data; performing a second variation operation on a result of a reference decoding operation, wherein the second variation operation is a variation operation corresponding to the first the variation operation.

[0040] In a possible implementation of the second aspect, the method further includes: performing an entropy coding operation on the bitstream D0.

[0041] In a possible implementation of the second aspect, the performing an entropy codingoperation on the encoded bitstream D0, includes: determining an optimized parameter according to the bitstream D0by an optimization model; performing, according to the optimized parameter, the entropy coding operation on the bitstream D0.

[0042] According to a third aspect, an embodiment of the present application provides a method for training a neural network. The method may include including: acquiring an input signal x. The method may include performing an analysis transform operation based on the input signal to obtain a latent representation y0. The method may include performing an encoding operation based on the latent representation y0to obtain a feature tensor z0containing values in the range of 0 to 1 by using a first neural network. The method may include obtaining a reconstructed feature tensor ẑ0based on the feature tensor z0. The method may include performing a decoding operation on the reconstructed feature tensor ẑ0to obtain a reconstructed latent representation ẏo by using a second neural network. The method may include performing a synthesis transform operation based on the reconstructed latent representation ẏo to generate a reconstructed signal ẋ. The method may include determining a loss function value according to the input signal x and the reconstructed signal ẋ. The method may include determining a first target neural network according to the first neural network and the loss function value. The method may include determining a second target neural network according to the second neural network and the loss function value.

[0043] The present application provides a neural network (NN) based video, image, audio codecs - source codecs (which are focused only on data compression), channel codecs ((which are focused only on data transmission) and joint source channel coding (which are focused both on data compression and transmission). The NN is used to encode and decode the input data into bitstream. Different features of the input data are compressed to different number of bits(‘words’). The NN may learn their own encoding / decoding way of latent space features represented as floats or integers to binary representation for sending via communication channel, find some correlations, and use less bits to encode data. Further, the ENC NN and the DEC NN may be end to end train together with noise - simulating both bit errors and packet losses - thus, becoming tolerant to errors, distortions and losses in communication channel.

[0044] Compared with conventional entropy encoding solutions, the NN based encoding solution provided by the present application have less complexity and less encoding / decodingtime due to massive parallel processing on NN supporting hardware. Finally, the present application works in digital domain and do not making any constrains to physical layer having friendly design to existing widely deployed digital communication infrastructure. Therefore, the present application may be deployed with existing hardware (such as routers, Wi-Fi access points, 3G / 4G / 5G wireless communication devices).[ ] In a possible implementation of the third aspect, the obtaining a reconstructed feature tensor ẑ0based on the feature tensor z0includes: assigning the reconstructed feature tensor ẑ0equal to the feature tensor z0.In a possible implementation of the third aspect, the obtaining a reconstructed feature tensor ẑ0based on the feature tensor z0includes: distorting at least one element of the feature tensor z0to obtain the reconstructed feature tensor ẑ0.

[0047] According to the above-mentioned technical solution, the present application allows to get trainable error tolerant data coding, or NN based data encryption system. The most obvious applications may include real time communication and video streaming systems, such as online meetings and video transmission, such as from phone to TV, from computer to projector, from drone to VR glasses and so on.In a possible implementation of the third aspect, the distorting at least one element of the feature tensor z0includes inverting 0 to 1 and 1 t

[0049] According to a fourth aspect, an embodiment of the present application provides an electronic device, and the electronic device has a function of implementing the method in the first aspect. The function may be implemented by hardware, or may be implemented by hardware executing corresponding software. The hardware of the software includes one or more units corresponding to the function.

[0050] According to a fifth aspect, an embodiment of the present application provides an electronic device, and the electronic device has a function of implementing the method in the second aspect. The function may be implemented by hardware, or may be implemented by hardware executing corresponding software. The hardware of the software includes one or more units corresponding to the function.

[0051] According to a sixth aspect, an embodiment of the present application provides an electronic device, and the electronic device has a function of implementing the method in thethird aspect. The function may be implemented by hardware, or may be implemented by hardware executing corresponding software. The hardware of the software includes one or more units corresponding to the function.

[0052] According to a seventh aspect, an embodiment of the present application provides a computer readable storage medium including instructions. When the instructions run on an electronic device, the electronic device is enabled to perform the method in the first aspect or any possible implementation of the first aspect.

[0053] According to an eighth aspect, an embodiment of the present application provides a computer readable storage medium including instructions. When the instructions run on an electronic device, the electronic device is enabled to perform the method in the second aspect or any possible implementation of the second aspect.

[0054] According to a ninth aspect, an embodiment of the present application provides a computer readable storage medium including instructions. When the instructions run on an electronic device, the electronic device is enabled to perform the method in the third aspect or any possible implementation of the third aspect.

[0055] According to a tenth aspect, an embodiment of the present application provides an electronic device, including a processor and a memory. The processor is connected to the memory. The memory is configured to store instructions, and the processor is configured to execute the instructions. When the processor executes the instructions stored in the memory, the processor is enabled to perform the method in the first aspect or any possible implementation of the first aspect.

[0056] According to an eleventh aspect, an embodiment of the present application provides an electronic device, including a processor and a memory. The processor is connected to the memory. The memory is configured to store instructions, and the processor is configured to execute the instructions. When the processor executes the instructions stored in the memory, the processor is enabled to perform the method in the second aspect or any possible implementation of the second aspect.

[0057] According to a twelfth aspect, an embodiment of the present application provides an electronic device, including a processor and a memory. The processor is connected to the memory. The memory is configured to store instructions, and the processor is configured toexecute the instructions. When the processor executes the instructions stored in the memory, the processor is enabled to perform the method in the third aspect or any possible implementation of the third aspect.

[0058] According to a thirteenth aspect, an embodiment of the present application provides a chip system, where the chip system includes a memory and a processor, and the memory is configured to store a computer program, and the processor is configured to invoke the computer program from the memory and run the computer program, so that an electronic device on which the chip system is disposed performs the method in the first aspect or any possible implementation of the first aspect.

[0059] According to a fourteenth aspect, an embodiment of the present application provides a chip system, where the chip system includes a memory and a processor, and the memory is configured to store a computer program, and the processor is configured to invoke the computer program from the memory and run the computer program, so that an electronic device on which the chip system is disposed performs the method in the second aspect or any possible implementation of the second aspect.

[0060] According to a fifteenth aspect, an embodiment of the present application provides a chip system, where the chip system includes a memory and a processor, and the memory is configured to store a computer program, and the processor is configured to invoke the computer program from the memory and run the computer program, so that an electronic device on which the chip system is disposed performs the method in the third aspect or any possible implementation of the third aspect.

[0061] According to a sixteenth aspect, an embodiment of the present application provides a computer program product, where when the computer program product runs on an electronic device, the electronic device is enabled to perform the method in the first aspect or any possible implementation of the first aspect.

[0062] According to a seventeenth aspect, an embodiment of the present application provides a computer program product, where when the computer program product runs on an electronic device, the electronic device is enabled to perform the method in the second aspect or any possible implementation of the second aspect.

[0063] According to an eighteenth aspect, an embodiment of the present applicationprovides a computer program product, where when the computer program product runs on an electronic device, the electronic device is enabled to perform the method in the third aspect or any possible implementation of the third aspect.

[0064] According to a nineteenth aspect, an embodiment of the present application provides computer readable storage medium, where the computer readable storage medium stores a bitstream, the bitstream is obtained by using the method according to the second aspect or any possible implementation of the second aspect.DESCRIPTION OF DRAWINGS

[0065] FIG. 1 is a schematic block diagram illustrating a coding system according to some embodiments of the present application.

[0066] FIG. 2 illustrates a network architecture of a compression network.

[0067] FIG. 3 illustrates a network architecture of a compression network in accordance with some embodiments of the present application.

[0068] FIG. 4 illustrates an encoding method according to some embodiments of the present application.

[0069] FIG. 5 illustrates an encoding method according to some embodiments of the present application.

[0070] FIG. 6 illustrates another encoding method according to some embodiments of the present application.

[0071] FIG. 7 illustrates an encoding method according some embodiments of the present application.

[0072] FIG. 8 illustrates an encoding-decoding procedure according to some embodiments of the present application.

[0073] FIG. 9 illustrates an encoding-decoding procedure according to some embodiments of the present application.

[0074] FIG. 10 illustrates a decoding method according to some embodiments of the present application.

[0075] FIG. 11 illustrates a decoding procedure according to some embodiments of thepresent application.

[0076] FIG. 12 illustrates 5 DECs and 5 synthesis transform operations.

[0077] FIG. 13 illustrates an encoding procedure and a corresponding decoding procedure according to some embodiments of the present application.

[0078] FIG. 14 illustrates an encoding procedure and a corresponding decoding procedure according to some embodiments of the present application.

[0079] FIG. 15 illustrates a system architecture according to some embodiments of the present application.

[0080] FIG. 16 illustrates a training method according to some embodiments of the present application.

[0081] FIG. 17 is a schematic block diagram of an electronic device according to some embodiments of the present application.

[0082] FIG. 18 is a schematic block diagram of an electronic device according to some embodiments of the present application.

[0083] FIG. 19 is a schematic block diagram of an electronic device according to some embodiments of the present application.

[0084] FIG. 20 is a schematic block diagram of an electronic device according to some embodiments of the present application.DESCRIPTION OF EMBODIMENTS

[0085] The following describes the technical solutions in the present application with reference to the accompanying drawings.

[0086] FIG. 1 is a schematic block diagram illustrating a coding system according to some embodiments of the present application. Referring to FIG. 1, the coding system 100 includes a source device 110 configured to provide encoded picture data to a destination device 120 for decoding the encoded picture data. For convenience, it is assumed that the coding system 100 is a picture coding system. As will be apparent for the skilled person, the picture coding system is just example embodiments of the invention and embodiments of the invention are not limited thereto.The source device 110 may include an encoding unit 111. Optionally, the source device 110 may further include a picture source unit 112, a pre-processing unit 113, and a communication unit 114.

[0087] The picture source unit 112 may include or be any kind of picture capturing device, for example for capturing a real-word picture, and / or any kind of a picture generating device, for example a computer- graphics processor for generating a computer animated picture, or any kind of device for obtaining and / or providing a real-word picture, a computer animated picture(e.g., a screen content, a virtual reality (VR) picture) and / or any combination thereof (e.g., an augmented reality (AR) picture). In the following, all these kinds of pictures and any other kind of picture will be referred p

[0088] A (digital) picture is or can be regarded as a two-dimensional array or matrix of samples with intensity values. A sample in the array may also be referred to as pixel (short form of picture element) or a pel. The number of samples in horizontal and vertical direction (or axis) of the array or picture define the size and / or resolution of the picture. For representation of color, typically three-color components are employed, i.e. the picture may be represented or include three sample arrays. In RBG format or color space a picture comprises a corresponding red, green and blue sample array. However, in video coding each pixel is typically represented in a luminance / chrominance format or color space, e.g. YCbCr, which comprises a luminance component indicated by Υ (sometimes also L is used instead) and two chrominance components indicated by Cb and Cr. The luminance (or short luma) component Y represents the brightness or grey level intensity (e.g. like in a grey-scale picture), while the two chrominance (or short chroma) components Cb and Cr represent the chromaticity or color information components.Accordingly, a picture in YCbCr format comprises a luminance sample array of luminance sample values (Y), and two chrominance sample arrays of chrominance values (Cb and Cr).Pictures in RGB format may be converted or transformed into YCbCr format and vice versa, the process is also known as color transformation or conversion. If a picture is monochrome, the picture may comprise only a luminance sample array.

[0089] The picture source unit 112 may be, for example a camera for capturing a picture, a memory, e.g. a picture memory, comprising or storing a previously captured or generated picture, and / or any kind of interface (internal or external) to obtain or receive a picture. Thecamera may be, for example, a local or integrated camera integrated in the source device, the memory may be a local or integrated memory, e.g. integrated in the source device. The interface may be, for example, an external interface to receive a picture from an external video source, for example an external picture capturing device like a camera, an external memory, or an external picture generating device, for example an external computer- graphics processor, computer or server. The interface can be any kind of interface, e.g. a wired or wireless interface, an optical interface, according to any proprietary or standardized interface protocol. The interface for obtaining the picture data 312 may be the same interface as or a part of theCommunication unit 114.

[0090] In distinction to the pre-processing unit 113 and the processing performed by the pre- processing unit 113, the picture or picture data 131 may also be referred to as raw picture or raw picture data 131.

[0091] The pre-processing unit 113 is configured to receive the (raw) picture data 131 and to perform pre-processing on the picture data 131 to obtain a pre-processed picture 132 or pre- processed picture data.

[0092] The pre-processing performed by the pre-processing unit 113 may, e.g., comprise trimming, color format conversion (e.g. from RGB to YCbCr), color correction, or de-noising.The encoding unit 111 is configured to receive the pre-processed picture data 132 and provide encoded picture data 133.

[0093] The communication unit 114 of the source device 110 may be configured to receive the encoded picture data 133 and to directly transmit it to another device, e.g. the destination device 120 or any other device, for storage or direct reconstruction, or to process the encoded picture data 133 for respectively before storing the encoded picture data 133 and / or transmitting the encoded picture data 133 to another device, e.g. the destination device 120 or any other device for decoding or storing.

[0094] The destination device 120 comprises a decoding unit 121, and may additionally, i.e. optionally, comprise a communication unit 124, a post-processing unit 123 and a display unit122.

[0095] The communication unit 124 of the destination device 120 is configured receive the encoded picture data 133, e.g. directly from the source device 110 or from any other source, e.g.a memory, e.g. an encoded picture data memory.

[0096] The communication unit 114 and the communication unit 124 may be configured to transmit respectively receive the encoded picture data 133 via a direct communication link between the source device 110 and the destination device 120, e.g. a direct wired or wireless connection, or via any kind of network, e.g. a wired or wireless network or any combination thereof, or any kind of private and public network, or any kind of combination thereof.

[0097] The communication unit 114 may be, e.g., configured to package the encoded picture data 133 into an appropriate format, e.g. packets, for transmission over a communication link or communication network, and may further comprise data loss protection and data loss recovery.

[0098] The communication unit 124, forming the counterpart of the communication unit 114, may be, e.g., configured to de-package the packets to obtain the encoded picture data 133 and may further be configured to perform data loss protection and data loss recovery, e.g. comprising error concealment.

[0099] The communication unit 124, forming the counterpart of the communication unit 114, may further be configured to perform communication without or with limited data loss protection, and without re-transmission of lost or corrupted data to minimize communication delay and end-to-end latency between source and destination device. In such configuration the encoded picture data 133 may contain errors after receiving by communication unit 124. For binary represented signals the communication errors will lead to inverting 0 to 1 and vice versa.

[0100] Both, the communication unit 114 and the communication unit 124 may be configured as unidirectional communication interfaces as indicated by the arrow for the encoded picture data 133 in FIG. 1 pointing from the source device 110 to the destination device 120, or bi- directional.

[0101] The communication units, and may be configured, e.g. to send and receive messages, e.g. to set up a connection, to acknowledge and / or re-send lost or delayed data including picture data, and exchange any other information related to the communication link and / or data transmission, e.g. encoded picture data transmission.

[0102] The decoding unit 121 is configured to receive the encoded picture data 133 and provide decoded picture data 134.

[0103] The post-processing unit 123 of destination device 120 is configured to post-process the decoded picture data 134 to obtain post-processed picture data 135. The post-processing performed by the post-processing unit 123 may comprise, e.g. color format conversion (e.g. from YCbCr to RGB), color correction, trimming, or re-sampling, or any other processing, e.g. for preparing the decoded picture data 134 for display, e.g. by display unit 122.

[0104] The display unit 122 of the destination device 120 is configured to receive the post- processed picture data 135 for displaying the picture, e.g. to a user or viewer. The display unit122 may be or comprise any kind of display for representing the reconstructed picture, e.g. an integrated or external display or monitor. The displays may, e.g. comprise cathode ray tubes(CRT), liquid crystal displays (LCD), plasma displays, organic light emitting diodes (OLED) displays or any kind of other display, such as beamer, hologram (3D), or the like.

[0105] Although FIG. 1 depicts the source device 110 and the destination device 120 as separate devices, embodiments of devices may also comprise both or both functionalities, the source device 110 or corresponding functionality and the destination device 120 or corresponding functionality. In such embodiments the source device 110 or corresponding functionality and the destination device 120 or corresponding functionality may be implemented using the same hardware and / or software or by separate hardware and / or software or any combination thereof.

[0106] As will be apparent for the skilled person based on the description, the existence and(exact) split of functionalities of the different units or functionalities within the source device110 and / or destination device 120 as shown in FIG. 1 may vary depending on the actual device and application.

[0107] Therefore, the source device 110 and the destination device 120 as shown in FIG. 1 are just example embodiments of the invention and embodiments of the invention are not limited to those shown in

[0108] The source device 110 and the destination device 120 may comprise any of a wide range of devices, including any kind of handheld or stationary devices, e.g. notebook or laptop computers, mobile phones, smart phones, tablets or tablet computers, cameras, desktop computers, set-top boxes, televisions, display devices, digital media players, video gaming consoles, video streaming devices, broadcast receiver device, or the like and may use no or anykind of operating system.

[0109] The embodiments of this application relate to application of a large quantity of neural networks. Therefore, for ease of understanding, related terms and related concepts such as the neural network in the embodiments of this application are first described below.

[0110] Neural Network

[0111] The neural network may include neurons. The neuron may be an operation unit that uses xsand an intercept 1 as inputs, and an output of the operation unit may be as follows:

[0112]

[0113] Herein, s=1, 2, . . . , or n, n is a natural number greater than 1, Ws is a weight of xs, and b is bias of the neuron, f is an activation function of the neuron, and the activation function is used to introduce a non-linear feature into the neural network, to convert an input signal in the neuron into an output signal. The output signal of the activation function may be used as an input of a next convolutional layer. The activation function may be a sigmoid function. The neural network is a network formed by connecting many single neurons together. To be specific, an output of a neuron may be an input of another neuron. An input of each neuron may be connected to a local receptive field of a previous layer to extract a feature of the local receptive field. The local receptive field may be a region including several neurons.

[0114] (2) Deep Neural Network

[0115] The deep neural network (DNN), also referred to as a multi-layer neural network, may be understood as a neural network having many hidden layers. The “many” herein does not have a special measurement standard. The DNN is divided based on locations of different layers, and a neural network in the DNN may be divided into three types: an input layer, a hidden layer, and an output layer. Generally, the first layer is the input layer, the last layer is the output layer, and the middle layer is the hidden layer. Layers are fully connected. To be specific, any neuron at the ith layer is certainly connected to any neuron at the (i+l)th layer. Although the DNN looks to be complex, the DNN is actually not complex in terms of work at each layer, and is simply expressed as the following linear relationship expression: where is aninput vector, s an output vector, is a bias vector, W is a weight matrix (also referred toas a coefficient), and α( ) is an activation function. At each layer, the output vector isobtained by performing such a simple operation on the input vector Because there are many layers in the DNN, there are also many coefficients W and bias vectors Definitions of theseparameters in the DNN are as follows: The coefficient W is used as an example. It is assumed that in a DNN having three layers, a linear coefficient from the fourth neuron at the second layer to the second neuron at the third layer is defined as. The superscript 3 represents a layer at which the coefficient W is located, and the subscript corresponds to an output third- layer index 2 and an input second-layer index 4. In conclusion, a coefficient from the kthneuron at the (L-1 )thlayer to the jthneuron at the Lthlayer is defined as. It should be noted that there is no parameter W at the input layer. In the deep neural network, more hidden layers make the network more capable of describing a complex case in the real world.

[0116] Theoretically, a model with a larger quantity of parameters indicates higher complexity and a larger “capacity”, and indicates that the model can complete a more complex learning task. Training the deep neural network is a process of learning a weight matrix, and a final objective of the training is to obtain a weight matrix of all layers of the trained deep neural network (a weight matrix including vectors W at many layers).

[0117] (3) Convolutional Neural Network

[0118] The convolutional neural network (CNN) is a deep neural network having a convolutional structure. The convolutional neural network includes a feature extractor including a convolutional layer and an optional sub sampling layer. The feature extractor may be considered as a filter. A convolution process may be considered as using a trainable filter to perform convolution on an input image or a convolutional feature plane (feature map). The convolutional layer is a neuron layer that is in the convolutional neural network and at which convolution processing is performed on an input signal. At the convolutional layer of the convolutional neural network, one neuron may be connected only to some adjacent-layer neurons. A convolutional layer usually includes a plurality of feature planes, and each feature plane may include some neurons arranged in a rectangular form. Neurons in a same feature plane share a weight. The shared weight herein is a convolution kernel. Weight sharing may beunderstood as that an image information extraction manner is irrelevant to a location. A principle implied herein is that statistical information of a part of an image is the same as that of another part. This means that image information learned in a part can also be used in another part. Therefore, image information obtained through same learning can be used for all locations in the image. At a same convolutional layer, a plurality of convolution kernels may be used to extract different image information. Usually, a larger quantity of convolution kernels indicates richer image information reflected by a convolution operation.

[0119] The convolution kernel may be initialized in a form of a random-size matrix. In a process of training the convolutional neural network, the convolution kernel may obtain an appropriate weight through learning. In addition, a direct benefit brought by weight sharing is that connections between layers of the convolutional neural network are reduced and an overfitting risk is lowered.

[0120] (4) Recurrent Neural Network

[0121] A recurrent neural network (RNN) is used to process sequence data. In a conventional neural network model, from an input layer to a hidden layer and then to an output layer, the layers are fully connected, and nodes at each layer are not connected. Such a common neural network resolves many difficult problems, but is still incapable of resolving many other problems. For example, if a word in a sentence is to be predicted, a previous word usually needs to be used, because adjacent words in the sentence are not independent. A reason why the RNN is referred to as the recurrent neural network is that a current output of a sequence is also related to a previous output of the sequence. A specific representation form is that the network memorizes previous information and applies the previous information to calculation of the current output. To be specific, nodes at the hidden layer are connected, and an input of the hidden layer not only includes an output of the input layer, but also includes an output of the hidden layer at a previous moment. Theoretically, the RNN can process sequence data of any length. Training for the RNN is the same as training for a conventional CNN or DNN. An error back propagation algorithm is also used, but there is a difference: If the RNN is expanded, a parameter such as W of the RNN is shared. This is different from the conventional neural network described in the foregoing example. In addition, during use of a gradient descent algorithm, an output in each step depends not only on a network in the current step, but also ona network status in several previous steps. The learning algorithm is refened to as a back propagation through time (BPTT) algorithm.

[0122] Now that there is a convolutional neural network, why is the recunent neural network required? A reason is simple. In the convolutional neural network, it is assumed that elements are independent of each other, and an input and an output are also independent, such as a cat and a dog. However, in the real world, many elements are interconnected. For example, stocks change with time. For another example, a person says: I like traveling, and my favorite place isYunnan. I will go if there is a chance. If there is bank filling, people should know that “Yunnan” will be filled in the blank. A reason is that the people can deduce the answer based on content of the context. However, how can a machine do this? The RNN emerges. The RNN is intended to make the machine capable of memorizing like a human. Therefore, an output of the RNN needs to depend on current input information and historical memorized information.

[0123] (5) Loss Function

[0124] In a process of training the deep neural network, because it is expected that an output of the deep neural network is as much as possible close to a predicted value that is actually expected, a predicted value of a cunent network and a target value that is actually expected may be compared, and then a weight vector of each layer of the neural network is updated based on a difference between the predicted value and the target value (certainly, there is usually an initialization process before the first update, to be specific, parameters are preconfigured for all layers of the deep neural network). For example, if the predicted value of the network is large, the weight vector is adjusted to decrease the predicted value, and adjustment is continuously performed until the deep neural network can predict the target value that is actually expected or a value that is very close to the target value that is actually expected. Therefore, “how to obtain, through comparison, a difference between a predicted value and a target value” needs to be predefined. This is the loss function or an objective function. The loss function and the objective function are important equations used to measure the difference between the predicted value and the target value. The loss function is used as an example. A higher output value (loss) of the loss function indicates a larger difference. Therefore, training of the deep neural network is a process of minimizing the loss as much as possible.

[0125] (6) Transformer

[0126] A transformer is a deep learning architecture that relies on the parallel multi-head attention mechanism. The transformer includes two main components: an encoder and a decoder. The encoder may include encoding layers that sequentially process the input tokens, and similarly, the decoder may include decoding layers that iteratively process the encoder's output and the decoder's generated tokens. Each encoder layer is responsible for producing contextualized representations of the tokens. These representations capture information from other input tokens through the self-attention mechanism. In other words, each representation combines information from different input tokens to create a comprehensive understanding. On the other hand, each decoder layer contains two attention sublayers. The first sublayer, called cross-attention, incorporates the contextualized representations generated by the encoder. It enables the decoder to leverage the knowledge learned from the input tokens during the generation of output tokens. The second sublayer, known as self-attention, facilitates the integration of information among the input tokens within the decoder. This self-attention mechanism allows the decoder to consider the tokens that have already been generated, enabling it to automatically learn which tokens are important and need to be predicted accurately. In other words, through self-attention, the decoder can attend to and analyze the previously generated tokens during the sequence generation process, enhancing its understanding of context and making more precise predictions.To further enhance the processing of outputs, both the encoder and decoder layers include a feed-forward neural network. This network performs additional computations on the outputs. Additionally, residual connections and layer normalization steps are applied in both the encoder and decoder layers to improve the model's stability and performance. The transformer has found widespread applications in natural language processing (NLP) and computer vision tasks, including machine translation, text classification, image classification, speech recognition, and more.

[0127] FIG. 2 illustrates a network architecture of a compression network. In FIG. 2, the following abbreviations and notations may be used.

[0128] gamay denote an analysis transform (e.g., an analysis transform operation) for transforming x into a latent representation y. gsmay denote a synthesis transform (e.g., synthesis transform operation) for generating a reconstructed picture ẋ.

[0129] Rectangles marked with “conv” may denote convolutional layers. A convolutionallayer may be represented by “the number of channels” x “kernel filter width” x “kernel filter height” and “down-scaling or up-scaling factor”, “↑” and “↓”respectively denote up-scaling and down-scaling through transposed convolutions.

[0130] An input picture input may be normalized to fit a scale between -1 and 1.

[0131] In a convolutional layer, “N” may indicate the number of feature map channels.

[0132] “GDN” may denote generalized divisive normalization (GDN). “IGDN” may denote inverse generalized divisive normalization (IGDN). “Q” may denote quantization e.g. uniform(rounding-off), or adaptive step size quantization. “EC” may denote an entropy-encoding process, “ED” may denote an entropy-decoding process.

[0133] Referring to FIG. 2, the analysis transform operation gais used to transform an input picture x into a latent representation y. Next, y may be quantized into ẏ, and the entropy- encoding process may be performed to generate an encoded bitstream. The encoded bitstream may be decoded through the entropy-decoding process to obtain ẏ. The synthesis transform operation gsis performed on ẏ to generate a reconstructed picture.

[0134] FIG. 3 illustrates a network architecture of a compression network in accordance with some embodiments of the present application. In FIG. 3, the following abbreviations and notations may be used.

[0135] “Down Sample” may denote a down-scaling process, and “Up Sample” may denote an up-scaling process.

[0136] Rectangles marked with “conv” may denote convolutional layers. A convolutional layer may be represented by “kernel filter height” x”kernel filter width”, “the number of input ” channels” x”the number of output channels” and “down-scaling or up-scaling factor” and “↓” respectively denote up-scaling and down-scaling through transposed convolutions.

[0137] In block 311, “N” may indicate the number of input channels. “L” may indicate the number of output channels of the layer. In block 321, L may indicate the number of input channels, and N may indicate the number of output channels of the layer.

[0138] “GDN” may denote generalized divisive normalization (GDN). “IGDN” may denote inverse generalized divisive normalization (IGDN).

[0139] “Flatten” may denote a flatten process. The flatten process may be used to convert multidimensional data into one-dimensional format. For example, video frames or pictures aretypically represented as two-dimensional arrays of pixel values. The flatten process may rearrange the pixel values into a single continuous line.

[0140] “Position Encoding” may denote a position encoding process that is used to encode spatial information or location within a video frame or a picture.

[0141] “Add beta token” may denote adding a weighted beta coefficient (controls the trade- off between rate and distortion) to the latent vector, in order to indicate how much information should be sent.

[0142] “Trans. Layer El” may denote a Transformer encoder layer number 1, that relies on the parallel multi-head attention mechanism, which allows it to capture long-range dependencies.

[0143] “Trans. Layer E2” may denote a Transformer encoder layer number 2.

[0144] “Trans. Layer D1” may denote a Transformer decoder layer number 1.

[0145] “Trans. Layer D2” may denote a Transformer decoder layer number 2.

[0146] “Normalize” may denote a normalization layer, that calculates the mean and standard- deviation over the number of embedded dimensions and applies normalization to a given vector with learnable per-element affine parameters initialized to ones (for weights) and zeros (for biases).

[0147] “Reshape” may denote the procedure of reshaping latent vector in order to “anti-flatten” it - to get back two-dimensional arrays of pixel values instead of a single continuous line.

[0148] Referring to FIG. 3, the compression network includes two networks. A first network310 is configured to derive an encoded bitstream according to an input image, and a second network 320 is configured to derive a reconstructed image corresponding to the encoded bitstream. Therefore, the first network 310 may be referred to as an encoding network, while the second network 320 may be referred to as a decoding network.

[0149] The encoding procedure performed by the first network may include two operations, where one of the two operations is an analysis transform operation 311, while the other operation may be referred to as a datastream deriving operation 312.

[0150] A neural network may be configured to implement the analysis transform operation311. Referring to FIG. 3, the analysis transform operation 311 may include a convolutional layer and a nonlinear transformation unit. The nonlinear transformation unit is a network unitthat contains nonlinear operations (such as ReLU, Sigmoid, Tanh, PWL, etc.). The overall computation of the nonlinear transformation unit does not conform to linear characteristics. The analysis transform operation 311 is configured to transform the input image from a signal space to a latent space. In other words, an output of the analysis transform operation is a latent representation of the input image.

[0151] The datastream deriving operation 312 may be configured to obtain the latent representation of the input image and derive a bitstream corresponding to the obtained latent representation. The datastream deriving operation 312 includes a bitstream deriving operation314 and an encoding operation 313. The encoding operation 313 may be configured to obtain a feature tensor and relevant parameters from the latent representation of the input image. The bitstream deriving operation 314 is configured to derive a bitstream according to the feature tensor and the parameters obtained by the encoding operation 313.

[0152] The second network 320 may be configured to perform a decoding procedure, and the decoding procedure may include two operations, where one of the two operations is a synthesis transform operation 321, while the other operation may be referred to as a latent reconstructing operation 322. The latent reconstructing operation 322 may be configured to decode the encoded bitstream to obtain an output data in the latent space. The synthesis transform operation321 is configured to transform the output data in the latent space to the reconstructed image in the signal space. A neural network may be configured to perform the synthesis transform operation 321.

[0153] The latent reconstructing operation 322 may include a decoding operation 323 and a tensor reconstructing operation 324. The tensor reconstructing operation 324 is configured to obtain a reconstructed feature tensor from the encoded stream, and the decoding operation 323 may be used to determine a reconstructed latent representation according to the reconstructed feature tensor. The reconstructed latent representation may be forwarded to the synthesis transform operation 321, and the synthesis transform operation 321 may determine a reconstructed image according to the reconstructed latent representation.

[0154] In general terms, the second network 320 may be regarded as an inverse operation of the first network 310. The latent reconstructing operation 322 may be regarded as an inverse operation of the datastream deriving operation 312, while the synthesis transform operation 321may be regarded as an inverse operation of the analysis transform operation 311.Correspondingly, the decoding operation 323 may be regarded as an inverse operation of the encoding operation 313, and a tensor reconstructing operation 324 may be regarded as an inverse operation of the bitstream deriving operation 314.

[0155] Details about the first network 310 may refer to FIG. 4, and details about the second network may refer to FIG. 10.

[0156] Although input data of the most embodiments of the present application is a video frame / picture / image, a type of the input data is not limited by the present application. The input data of the encoding device may be an image, a video, an audio, a 3D information, a text document, or the like.

[0157] The encoding operation 313 and the decoding operation 323 are based on transformer architecture. The encoding operation 313 and the decoding operation 323 may be implemented by other neural networks. For example, the encoding operation 313 and the decoding operation323 may be based on a CNN, a RNN, and so forth.

[0158] FIG. 4 illustrates an encoding method according to some embodiments of the present application. For convenience, it is assumed that the operation illustrated in FIG. 4 is performed by an encoder or an encoding device. However, the operation illustrated in FIG. 4 may also be performed by a component of the encoder or the encoding device, such as a circuit in the encoder or the encoding device, or a chip or a system on chip (SoC) in the encoder / encoding device.

[0159] 401, The encoding device acquire an input signal.

[0160] The input signal may be an image, a video frame, an audio or the like. For convenience, in the following embodiments, it is assumed that the input is an image. In other words, the input signal is an input image.

[0161] 402, The encoding device performs an analysis transform operation based on the input image to obtain a latent representation of the input image. For convenience, letter “x” is used to symbolize the input image, and “y0” is used to symbolize the latent representation of the input image. Therefore, the input image and the latent representation satisfies:

[0162] y0= AT(x), (equation 1

[0163] where AT ( ) represents the analysis transform operation.

[0164] 403, The encoding device performs an encoding operation based on the latent representation y0to obtain a feature tensor z0by using a neural network.

[0165] The present application involves multiple neural networks. In order to distinguish these neural networks, the neural network that is used for performing the encoding operation may be referred to as an encoding neural network (ENC NN). As previously mentioned, the analysis transform operation may be performed by using a neural network, or another type of transforms like discrete cosine transform (DCT), wavelet, fast Fourier transform (FFT) etc. For convenience, the neural network which is used to perform the analysis transform operation may be referred to as an analysis transform neural network (AT NN).

[0166] 404, The encoding device generates a bitstream D0according to the feature tensor z0by means of bitstream deriving operation 314.

[0167] In some embodiments, the encoding device may acquire an encoding information.Then the encoding device may generate the bitstream D0according to the encoding information and the feature tensor z0. The size of the feature tensor z0is larger than a size of the bitstream D0. In some embodiments, the encoding information may also be referred to as a decoding information obtained by training or other optimization process, and available for encoding and decoding device.[ ] In some embodiments, the encoding information is singled with the bitstream D0. In some embodiments, the encoding information may also be referred to as a decoding information.

[0169] In some embodiments, the encoding information is a length tensor NbitsO. In some embodiments, the length tensor NbitsO may be obtained according to the latent representation y0by using the ENC NN, more specifically as an output of encoding operation network 313.

[0170] In some embodiments, the encoding device may determine a mask tensor MSK0according to the length tensor NbitsO and an index tensor idxo. Then, the encoding device determines the bitstream D0according to the mask tensor MSK0and the feature tensor z0.

[0171] In some other embodiments, the encoding information (or the decoding information) may be fixed and obtained during a training process.

[0172] The mask tensor MSK0comprises X1first element(s) with a first value and X2second element(s) with a second value, X1, a sum of X1and X2is the size of the feature tensor z0, X1is a positive integer, X2is a positive integer. In other words, the number of the first element(s),the number of the second element(s), and the size of the feature tensor z0satisfy:

[0175] For example, if the number of the first element(s) is 5 and the number of the second element(s) is 4, then the size of the feature tensor is 9.

[0176] In some embodiments, the first value may be 1, while the second value may be 0. In some other embodiments, the first value may be 0, while the second value may be 1. For convenience, in the following embodiments, it is assumed that the first value is 1, and the second value is

[0177] The bitstream D0may include X1bit(s), and the feature tensor z0may include X1third element(s) and X2fourth element(s), where a location of each elements in the feature tensor z0is a same of a location of a corresponding element in the mask tensor MSK0, a value of each of the X1third element(s) is the same as a value of a corresponding bit in the bitstream D0.

[0178] In some embodiment, a binarization transform operation may be performed on the feature tensor z0. An output of the binarization transform operation of the feature tensor z0may be referred to as a binary feature tensor ẑ0. A function B(y) may be used to transform the feature tensor z0into the binary feature tensor z0. The function B(y) may also be referred to as a binarization function. The binarization function B(y) may include the differentiable function.The feature tensor z0and the binary feature tensor z0may satisfy:

[0179] z0= B( z0), (equation )

[0180] In some embodiments, the binarization function B(y) may include a sigmoid function and a mathematical rounding function. The feature tensor z0is an input of the sigmoid function, an input of the mathematical rounding function is an output of the sigmoid function, and an output of the mathematical rounding function is the binary feature tensor z0. The mathematical rounding function may be a round function, a ceil function, or a floor function. For example, when the mathematical rounding function is the round function, the binary feature tensor z0and the feature tensor z0may satisfy:

[0181] z0= round (sigmoid(z0)), (equation 1.4)

[0182] where round ( ) denotes the round function, and sigmoid ( ) denotes the sigmoidfunction.

[0183] When the binarization transform operation is performed, the bitstream D0may be determined according to the binary feature tensor zo. The binary feature tensor zo may also include X hird element(s) and X fourth element(s), where a value of the third element maybe the same as the value of the first element, and a value of the fourth element may be the same as the value of the second element.

[0184] For example, the feature tensor may be: , and each row of thefeature tensor z0is a feature of the input image. The binary feature tensor zo determined according to the above-mentioned equation (1.4) is In this example, the thirdelements of the binary feature tensorinclude the elements with a value of 1 , and the fourth elements include the element with a value of 0.

[0185] The binarization function B(y) may make all output values set to 0 or 1 preserving gradients of the feature tensorThe sigmoid function is an example function which is used to convert the fundamental feature tensor z0into the fundamental binary feature tensor z0, and the sigmoid function in the above-mentioned equations may be replaced by any other differentiable function f(z0) with values range between 0 and 1. For example, f, where tanh( ) denotes a hyperbolic tangent function, and softsing ( )denotes a softsign function.

[0186] The index tenor idxo is a tensor of the same shape as the binary feature tensor z0. For example, if binary feature tensor , the index tensor idxo is a 3><3 matrix. Valuesof the index tenor idxo is determined according to a dimensional parameter Ho. For example, the index tensor idxo may be filled continuously with values from 0 to , where N isdimension along which number of bits should be cut. A relationship between the index tenor idxo and the dimensional parameter N may be:

[0187] i ([0, 1, 2, 3, ..., ), (equation 1.5)

[0188] where tensor ( ) denotes a tensor of the same shape as the binary feature tensor z0. Ifthe binary feature tensor the dimensional parameter Nmax_o may be 2, and the index tensor idxo may be .

[0189] As previous mentioned, the encoding device may determine the mask tensor MAK0according to the length tensor Nbitso and an index tensor idxo. The mask tensor MSK0is a tensor of the same shape as the index tensor idxo, and values of the tensor (that is, MSK0) may be determined according to the length tensor Nbitso and the index tensor idx0. When the size of the input image is WxH, the size of the feature tensor (or the binary feature tensor) may be wxhxc, where W and w are the number of columns, H and h are the number of rows, and c is the number of channels or features for each latent space, w is smaller than W, h is smaller than H, and all of these elements (that is, W, H, w, h, and c) are positive integer. In addition, the size of the length tensor Nbits0may be 1 xhxc. The length tensor Nbitso specifies how may features of the binary feature tensor z0need to be transferred from every latent space position of the binary feature tensor z0. In other words, the length tensor Nbitso shows how much values (features of the latent space) will be taken in each rows of the binary feature tensor z0, and the remaining values are discarded, and therefore the binary feature tensor z0is compressed. In addition, the index tensor idxo and the mask tensor MAK0may be the same size as the binary feature tensor z0. Each element of a xthrow in the h rows of the index tensor idxo may be compared with an element of a xthrow in the h rows of the length tensor Nbits0. For convenience, V may be referred to as a value of a ythelement in the xthrow of the index tensor idxo,may be referred to as a value of the element in the xthrow of the length tensor Nbits_o, and y maybe referred to as a value of ythelement in the xthrow of the mask tensor MAK0. Whenis larger than or equal to is set to 0; V y is less than,isset to 1 .

[0190] For example, i then

[0191] The mask tensor MAK0may be configured to extract bit(s) from the binary feature tensor z0to construct the encoded bitstream D0. As previous mentioned, the size of the mask tensor MAK0may be the same as the size of the binary feature tensor z0. Elements of the masktensor MAK0may be in one-to-one correspondence with elements of the binary feature tensor z0. When a value of an element of the mask tensor MAK0corresponding to an element of the binary feature tensor zo is equal to 1 , the element of the binary feature tensor zo may be extracted as a bit in the encoded bitstream, otherwise the element may be ignored. In other words, the size of the encoded bitstream D0may be equal to the number of elements in MAK0that have a value of 1.

[0192] For example, if and M , an element in the first row and ,first column of zo corresponds to an element in the first row and first column of MAK0. Since the value of the element in the first row and first column MAK0is 1 , the element in the first row and first column of zo may be reserved as a bit in the encoded bitstream D0. Similarly, since a value of an element in a first row and second column of MAK0is 1 , an element in the first row and second column of zo may be reserved as a bit in the encoded bitstream D0. However, a value of an element in the first row and third column of MAK0is 0, then an element in the first row and third column of zo may be ignored. Based on the same principle, the fundamental encoded bitstream D0may be obtained, and D0=01100.

[0193] In some other embodiments, the binarization transform operation may be performed after the element extraction procedure mentioned in the previous paragraph. Similarly, when a value of an element of the mask tensor MSKO corresponding to an element of the feature tensor zo is equal to 1, the element of the feature tensor zo may be extracted as a bit in the encoded bitstream, otherwise the element may be ignored. For example, may be usedto extract elements from the feature tensor . Elements extracted from zomay be (-0.1, 1.2, 1.3, -2.3, -0.9). The binarization transform operation may be performed on the extracted elements to obtain then encoded bitstream D0.

[0194] In some embodiments, the encoding device may transmit the encoded bitstream D0to a decoding device. Correspondingly, the decoding device may receive the encoded bitstream D0from the encoding device and decode the bitstream D0to obtain the reconstructed image.The decoding device may determine a reconstructed feature tensor zo and perform a decodingoperation on the reconstructed feature tensor ẑ0to obtain a reconstructed latent representation ŷ0by using a neural network. For convenience, the neural network that is used to obtain the reconstructed latent representation ŷ0may be referred to as a DEC NN. After obtaining the reconstructed latent representation ŷ0, the decoding device may perform a synthesis transform operation based on the reconstructed latent representation ŷ0to generate a reconstructed signal ẋ

[0195] In some embodiments, in addition to the encoded bitstream D0, the decoding device need additional information to decode the encoded bitstream D0. For convenience, the additional information that is sued to decode the encoded bitstream D0may be referred to as a decoding information. Correspondingly, the decoding information can be known on decoding device in advance, or, alternatively, the encoding device may transmit the decoding information to the decoding device.

[0196] In some embodiments, the decoding information may include the length tensor Nbitso-The decoding device may determine the mask tensor MSK0according to the index tensor idxo and the length tensor Nbitso and determine the reconstructed feature tensor ẑ0according to the encoded bitstream D0and the mask tensor MSK0. In some other embodiments, the decoding information may further include the index tensor idxo. In some other embodiments, the decoding device may determine the index tensor idxo according to the dimensional parameterNmax_0. Details about how to determine the index tensor according to the dimensional parameter may refer to the above-mentioned embodiments and will not be described here.

[0197] In some other embodiments, the decoding information may include the mask tensor MSK0. The decoding device may determine the reconstructed feature tensor ẑ0according to the encoded bitstream D0and the mask tensor MSK0.

[0198] As previously described, in some embodiments, the decoding information may include the mask tensor MSK0. In some other embodiments, the decoding information may include the length tensor Nbitso. In some embodiments, the decoding information including the mask tensor MSK0, or, the length tensor Nbits_0may be referred to as a fundamental decoding information.Further, for convenience, the term “datastream information” is used to refer to the information that is sent from the encoding device to the decoding device. As discussed previously, in some embodiments, the datastream information may include the encoded bitstream and the masktensor MSK0. In some other embodiments, the datastream information may include the encoded bitstream and the length tensor Nbits_0. In other words, the datastream information may include the encoded bitstream and the decoding information. In some embodiments, a datastream information including the encoded bitstream D0and the fundamental decoding information may be referred to as a fundamental datastream information.

[0199] In some embodiments, the encoding device may determine at least one additional encoded bitstream and transmit the at least one additional encoded bitstreams to the decoding device. The decoding device may determine the reconstructed image according to the encoded bitstream D0and the at least one additional encoded bitstream. Correspondingly, decoding information corresponding to the at least one additional encoded bitstream may also be transmitted to the decoding device. For convenience, it is assumed that, in addition to the fundamental datastream information, the encoding device may transmit M piece(s) of datastream information to the decoding device. Corresponding, the decoding device may receive the M piece(s) of datastream information. M is a positive integer. Each of the M piece(s) of datastream information includes an encoded bitstream and a decoding information corresponding to the encoded bitstream. The M piece(s) of datastream information may be obtained through M encoding operation(s) and M decoding operation(s). For convenience, an encoding operation which is used for obtaining the fundamental datastream information may be referred to as a fundamental encoding operation, while, unless otherwise indicated, the term“encoding operation” in the following is referred to as the encoding operation for obtaining one of the M piece(s) of datastream information. In other words, unless otherwise indicated, the term “encoding operation” is one of the M encoding operation(s) mentioned in the following context. Similarly, the term “datastream information” is referred to as one of the M piece(s) of datastream information unless otherwise indicated.

[0200] In some embodiments, M may be equal to one. In this case, the encoding device may perform one decoding operation on the fundamental datastream information to obtain a reconstructed image and one encoding operation based on the reconstructed image. For convenience, the image that is used to determine the fundamental encoding information may be referred to as a pre-coding data. The decoding operation which is performed on the fundamental datastream information to obtain a reconstructed image may be referred to as a decodingoperation #1. The encoding operation for obtaining the fundamental datastream information may be referred to as an encoding operation #0. The encoding operation based on the reconstructed image may be referred to as an encoding operation #1. In conclusion, the encoding device may obtain the pre-coding data (hereinafter, referred to as “x”) and perform the encoding operation #0 on the pre-coding data to obtain the fundamental datastream information (hereinafter, referred to as “bits0”) including the encoded bitstream D0and the fundamental decoding information. Then, the encoding device may perform the decoding operation #1 according to bits0 to obtain a reconstructed data (hereinafter, referred to as “rec1”).In addition, the encoding device may perform the encoding operation #1 on [x- rec]1to obtain a datastream information (hereinafter, referred to as “bits1”) including an encoded bitstream and a decoding information corresponding to the encoded bitstream. Details of the encoding operation #1 may be the same as that of the encoding operation #0. For example, the encoding operation #1 may include an encoding operation #1 performed on [x-rec1] by using the ENCNN. Details of the encoding operation #1 may refer to FIG. 3 and FIG. 4 and will not be described here.

[0201] The encoding device may transmit the fundamental datastream information bitso and the datastream information bits1to the decoding device. Corresponding, the decoding device may receive the fundamental datastream information bitso and the datastream information bits1from the encoding device and decode the fundamental datastream information bitso and the datastream information bits1to obtain the reconstructed image.

[0202] In some embodiments, M may be greater than one. In this case, the encoding device may perform the decoding operation M times and perform the encoding operation M times, andM encoding operations do not include the fundamental encoding operation for obtaining the fundamental datastream information. For convenience, a mthdecoding operation among the M decoding operations may be referred to a decoding operation #m, and m=1, ..., M. A mthencoding operation among the M encoding operation may be referred to an encoding operation#m. The fundamental encoding operation for obtaining the fundamental datastream information may be referred to as an encoding operation #0.

[0203] FIG. 5 illustrates an encoding method according to some embodiments of the present application. Referring to FIG. 5, the encoding method performed by an encoding deviceincludes the fundamental encoding procedure, the M decoding procedures and the M encoding procedures. For convenience, for the encoding procedure illustrated in FIG. 5, it is assumed thatM is a positive integer greater than 4.

[0204] As previously mentioned, after performing the encoding operation, the encoding device may perform the bitstream deriving operation to obtain the encoded bitstream. Meanwhile, before performing the decoding operation, the encoding device may perform the tensor reconstructing operation, and the decoding operation is performed on the result of the tensor reconstructing operation. In other words, the encoding operation has a corresponding bitstream deriving operation, and the decoding operation has a corresponding tensor reconstructing operation. For convenience, “ENCj” illustrated in FIG. 5 refer to the encoding operation and the bitstream deriving operation corresponding to the encoding operation, and “ENCj” illustrated in FIG. 5 refers to the decoding operation and the tensor reconstructing operation corresponding to the decoding operation, where j = 1, ..., M. In other words, ENCjillustrated in FIG. 5 corresponds to the block 312 illustrated in FIG. 3, and ENCjincludes the encoding operation #j and the bitstream deriving operation corresponding to the encoding operation #j (hereinafter referred to as “bitstream deriving operation #j”). Similarly, ENCjillustrated in FIG. 5 illustrated the block 322 illustrated in FIG. 3, and ENCjincludes the decoding operation #j and the tensor reconstructing operation corresponding to the decoding operation #j (hereinafter referred to as “tensor reconstructing operation #j”).

[0205] Referring to FIG. 5, bitso is obtained by ENC1, bits1is the datastream information obtained by ENC2, bitS2is the datastream information obtained by ENC3, and so forth; rec1is a decoding result obtained by DEC1and a synthesis transform corresponding to DEC1(that is S1), rec2is a decoding result obtained by DEC2and a synthesis transform corresponding to DEC2(that is S2), and so forth. Ao refers to the analysis transform operation corresponding to ENC0. In other words, ENC0(that is, the encoding operation #0 and the bitstream deriving operation #0) is performed on a result of A0. Similarly, A1refers to the analysis transform operation corresponding to ENC1, A2refers to the analysis transform operation corresponding to ENC2, and so forth. S1refers to the synthesis transform operation corresponding to DEC1. In other words, S1is performed on a result of the decoding operation #1, and the transformed result of S1is rec1. Similarly, S2refers to the synthesis transform operation corresponding to DEC2, S3refers to the synthesis transform operation corresponding to DEC3, and so forth. For convenience, the ENC0and its corresponding analysis transform operation (that is, Ao) may be referred to as an encoding procedure #0, ENC1and its corresponding analysis transform operation (that is, A1) may be referred to as an encoding procedure #1, ENC2and its corresponding analysis transform operation (that is, A2) may be referred to as an encoding procedure #2, and so on. Similarly, DEC1and its corresponding synthesis transform operation(that is, S1) may be referred to as a decoding procedure #1, DEC2and its corresponding synthesis transform operation (that is, S2) may be referred to as a decoding procedure #2, and so on.

[0206] Referring to FIG. 5, the encoding device performs the encoding procedure #1 on [x- rec1]. In other words, the encoding procedure # 1 is performed based on the pre-coding data(that is “x” in FIG. 5) and rec1. Since the encoding procedure #1 is performed based on the decoding result obtained by the decoding procedure #1, the decoding procedure #1 may be referred to as a decoding procedure corresponding to the encoding procedure #1. The encoding device performs the encoding procedure #2 on [x-rec1-rec2]. In other words, the encoding procedure # 2 is performed based on the pre-coding data, rec1, and rec2. Similarly, the decoding procedure #2 is a decoding procedure corresponding to the encoding procedure #2. As the encoding procedure #1 includes the encoding operation #1 and the decoding procedure #1 includes the decoding operation #1, the decoding operation #1 may be referred to as a decoding operation corresponding to the encoding operation #1. Similarly, the decoding operation #2 may be referred to as a decoding operation corresponding to the encoding operation #2, and so on.Therefore, for the mthencoding operation, the encoding device performs the encoding operation#m based on , and the decoding operation #m is a decoding operation correspondingto the encoding operation #m. More specifically, the encoding device performs the encoding operation #m on AT where AT ( ) denotes the analysis transform operation. If yis the reconstructed latent representation obtained by decoding operation, then rec =Therefore, the input of the encoding operation #m may also be expressed aswhere AT ( ) denotes the analysis transform operation, ST ( ) denotes the synthesis transform operation, x is the input signal, and reconstructed latent representation among the mreconstructed latent representation(s). Details of the encoding operation #m may be the same as that of the encoding operation #0. For example, the encoding operation #m may be performed by using the ENC NN. Details of the encoding operation #1 may refer to FIG. 3 and FIG. 4 and will not be described here.

[0207] FIG. 6 illustrates another encoding method according to some embodiments of the present application. Referring to FIG. 6, the encoding method performed by an encoding device includes a fundamental encoding operation, M decoding operations and M encoding operations.Similar to FIG. 5, “ENCj” illustrated in FIG. 6 refer to the encoding operation and the bitstream deriving operation corresponding to the encoding operation, and “ENCj” illustrated in FIG. 6 refers to the decoding operation and the tensor reconstructing operation corresponding to the decoding operation,

[0208] Similar to FIG. 5, bitso is obtained by ENC1, bits1is the datastream information obtained by ENC2, bitS2is the datastream information obtained by ENC3, and so forth; rec1is the decoding result obtained by DEC1, rec2is the decoding result obtained by DEC2, and so forth. Ao refers to the analysis transform operation corresponding to the fundamental encoding operation (that is, the encoding operation #0 included

[0209] Referring to FIG. 6 only ENC0has a corresponding analysis transform operation (that is, Ao) and other ENCs does not have the corresponding analysis transform operation. Referring to FIG. 6, each of the M ENCs corresponds to a DEC. When the ENC does not have the analysis transform operation, the corresponding DEC does not have a corresponding synthesis transform operation.

[0210] For example, ENC1corresponds to DEC1, ENC2corresponds to DEC2.

[0211] Referring to FIG. 5 and FIG. 6, each of the DECs has a corresponding ENC. Since each of the DECs includes a decoding operation and each of the ENCs includes an encoding operation, it may be asserted that each of the decoding operations included in the DECs has a corresponding encoding operation, and the decoding operation may be performed on an input data determined according to the bitstream generated by the corresponding encoding operation.For example, in FIG. 6, DEC1decodes the bitstream generated by ENC0, DEC2#2 decodes the bitstream generated by ENC1. y_rec1is an output of DEC1, and y_rec1may be a reconstructed latent representation corresponding to the encoded bitstream carried by bitso. In someembodiments, the reconstructed latent representation corresponding to the encoded bitstream carried by bitso may also be referred to as a reconstructed latent representation ŷ0. Similarly, y_rec2is an output of DEC2, and y_rec2 may also be referred to as a reconstructed latentmentioned, ENCm includes the encoding operation #m and the bitstream deriving operation #m, and the bitstream deriving operation #m is performed according to the output of the encoding operation #m. Therefore, it ma be asserted that the encoding device performs the encoding

[0212] For convenience, the term “CODEC” may be referred to a DEC and an ENC corresponding to the DEC. For example, CODEC0may include the ENC0and the DEC1,CODEC1may include the ENC1and the DEC2, CODEC2may include the ENC2and the DEC3, and so on. When the encoding procedure includes M decoding operations and M encoding operations, there are M CODECs, that is the CODEC0to CODECM-1, and CODECjincludes the ENCjand the ENCj+1, where j=1, ..., M-1. For the embodiments illustrated in FIG. 5 and FIG.6, except for the CODEC0, input data of each of the CODEC1to the CODECM-1is derived according to the decoded information of the previous CODEC(s) and the pre-coding data (that is “x”)

[0213] In some embodiments, the encoding device may perform one or more the analysis transform operations. When the encoding device performs a plurality of the analysis transform operations, input data of the first analysis transform operation of the plurality of the analysis transform operations is the pre-coding data (that is the input image), and input data of the second analysis transform operation to the last analysis transform operation is output data of the previous analysis transformation. For example, it is assumed that the encoding device may perform the analysis transform operation K time(s). When K is equal to one (in other words, the encoding device only performs the analysis transform operation once), the input data of the analysis transformation is the pre-coding data. When K is a positive integer larger than 1 , the input data of the first analysis transform operation is the pre-coding data, and the input data of the second analysis transform operation is the output data of the first analysist transformoperation, and so on.

[0214] FIG. 7 illustrates an encoding method according some embodiments of the present application. Referring to FIG. 7, the encoding device performs 5 analysis transform operations, and the 5 analysis transform operations may be referred to as the analysis transform operation#1 to #5 respectively, where A1in FIG. 7 denotes the analysis transform operation #1, A2denotes the analysis transform operation #2, and so on. “x” in FIG. 7 is the pre-coding data,ATi is the output data of the analysis transform operation #1, AT2is the output data of the analysis transform operation #2, and so on. As illustrated in FIG. 7, the input data of the analysis transform operation #1 (that is, A1) is the pre-coding data (that is, x), the input data of the analysis transform operation #2 (that is, A2) is AT 1 (that is, the output data of the analysis transform operation #1), the input data of the analysis transform operation #3 (that is, A3) is A2(that is, the output data of the analysis transform operation #2), and so forth.

[0215] The encoding device may perform an encoding operation on the output data of the last of the K analysis transform operation (in other words, the Kthanalysis transform operation, or, the analysis transform operation #K) to generate the fundamental datastream information. For convenience, the encoding operation for generating the fundament datastream information may be referred to as a fundamental encoding operation. In addition to the fundamental encoding operation, the encoding device may perform M encoding operation(s) and M decoding operation(s) to generate M datastream information, where M is a positive integer. For convenience, unless otherwise indicated, the term “encoding operation” in the following is referred to as one of the M encoding operation(s). Similarly, the term “datastream information” is referred to as one of the M datastream information unless otherwise indicated. As previously mentioned, each encoding operation has a corresponding bitstream deriving operation that is used to generate the encoded bitstream, and each decoding operation has a corresponding tensor reconstructing operation that is used to generate the reconstructed tensor feature according to the encoded bitstream. Therefore, similar to FIG. 5, “ENCj” illustrated in FIG. 7 refer to the encoding operation and the bitstream deriving operation corresponding to the encoding operation, and “ENCj” illustrated in FIG. 7 refers to the decoding operation and the tensor reconstructing operation corresponding to the decoding operation, where j = 1, ..., M.

[0216] Referring to FIG. 7, the encoding device may perform ENC0on AT5(that is, the outputdata of the analysis transform operation #5) to generates the fundament datastream information(that is, bitso). Further, the encoding device performs 4 decoding operations and 4 encoding operations to generates 4 datastream information (that is, bits1to bitS4). In order to generate the4 datastream information, the encoding device further performs 4 synthesis transform operations. For convenience, the 4 decoding operations may be referred to as the decoding operation #1 to #4 respectively, the 4 encoding operations may be referred to as the encoding operation #1 to #4 respectively, and the 4 synthesis transform operations may be referred to as the synthesis transform operation #1 to #4. The encoding device performs the decoding operation #1 based on bitso to obtain a reconstructed latent representation #1 (that is, y_rec1) and performs the synthesis transform operation #1 on the reconstructed latent representation #1.The output data of the synthesis transform operation #1 may be referred to as ST1., Similarly, the output data of the synthesis transform operation #2 to #4 may be referred to as ST2to ST4respectively.

[0217] After obtaining ST1,, the encoding device may determine that the input data of ENC1is determined according to AT4and ST1,. INP1denotes the input data of the encoding operation#1. As illustrated in FIG. 7, INP1, ST1, and AT4satisfy: INP1= AT4- ST1. After determining INP1, the encoding device may perform ENC1on INP1to obtain a datastream information #1(that is, bits1). In addition, the encoding device may perform DEC2on bits1to obtain a reconstructed latent representation #2 (that is, y_rec2). According to y_rec2and ST1, the encoding device performs the synthesis transform operation #2 (that is, S2). Input data of the synthesis transform operation #2 is the sum of y_rec2and ST1,. Then, the encoding device may determine the input data of ENC2based on the output of the synthesis transformation #2 (that is, ST2) and the output of the analysis transform operation #3 (that is, AT3) and perform ENC2on the determined input data. As previously mentioned, ENC includes the encoding operation and the bitstream deriving operation, and the bitstream deriving operation #m is performed according to the output of the encoding operation. Therefore, the input data of the ENC may be regarded as the input data of the encoding operation included in the ENC. For example, the input data of the ENCmmay be regarded as the input data of the encoding operation #m.

[0218] In summary, input data of the mthencoding operation (in other words, input data of the encoding operation #m) satisfy:

[0219] INPm= ATK-m-STm, (2.1)

[0220] where INPmis the input data of the mthencoding operation, ATK-mis a result of a (K- m)thanalysis transform operation among the K analysis transform operations; STmis a result ofsynthesis transform operation is performed on a target data determined according to the m decoded bitstream(s). m is a positive integer and less than or equal to M. M is equal to K-l.

[0221] A target data of a synthesis transform operation may also be referred to as an input data of the synthesis transform operation. For example, referring to FIG. 7, the synthesis transform operation #1 is performed on the reconstructed latent representation #1 (that is, y_reci). So, the reconstructed latent representation #1 is the target data (or the input data) of the synthesis transform operation #1, and the reconstructed latent representation #1 is determined according to the first decoding operation (that is, the decoding operation included in DECi). For the target data of the synthesis transform operation #2 (that is, S2), ST1and y_rec2, y_rec2is determined according to the second decoding operation (that is, the decoding operation included in DEC2),STi is determined according to the reconstructed latent representation #1 determined according to the first decoding operation (that is, the decoding operation included in DECi). Therefore, the target data of the synthesis transform operation #2 is determined according to the first two decoding operations (that is, the decoding operations included in DECi and DEC2). For the target data of the synthesis transform operation #3 (that is, S3), ST2 and y_rec3, y_rec3is determined according to the third decoding operation (that is, the decoding operation included in DEC3), ST2 is determined according STi and y_rec2, where y_rec2 is determined according to the second decoding operation (that is, the decoding operation included in DEC2), STi is determined according to the reconstructed latent representation #1 determined according to the first decoding operation (that is, the decoding operation included in DECi). Therefore, the target data of the synthesis transform operation #3 is determined according to the first three decoding operations (that is, the decoding operations included in DECi, DEC2, and DEC3). Therefore, the mthsynthesis transform operation is performed on a target data determined according to the m reconstructed latent representation(s).

[0222] In some embodiments, the encoding device may perform a first variation operation on the input data of at least one of the fundamental encoding operation and the M encodingoperation(s). The first variation operation may include at least one of a sampling variation operation, a linear transformation, or an analysis transformation, and the sampling variation operation may be an up-sampling operation or a down-sampling operation.

[0223] Depending on whether the variation operation is performed, the M+1 encoding operation (that is, the fundamental encoding operation and the M encoding operation) may be referred to as the first encoding operation and the second encoding operation, where the first encoding operation is the encoding operation (or the fundamental encoding operation) on which the variation operation is performed, and the second encoding operation is the encoding operation (or the fundamental encoding operation) on which the variation operation is not performed. Further, an input data of the first encoding operation may be referred to as the first input data, while an input data of the second encoding operation may be referred to as a second input data. Therefore, the encoding device may perform the variation operation on the first input data, then the encoding device may perform the rest process of the encoding operation based on the varied first input data.

[0224] As previously mentioned, each decoding operation has a corresponding encoding operation, and the decoding operation determines a reconstructed latent representation according to the bitstream determined according to the corresponding encoding operation. When the encoding operation of the corresponding encoding operation is the first encoding operation, a variation operation corresponding to the first variation operation may be performed on a result of the decoding operation in the decoding operation. For convenience, the variation operation corresponding to the first variation operation may be referred to as a second variation operation. Similarly, depending on whether the second variation operation is performed, the M decoding operation may be referred to as the first decoding operation and the second decoding operation, where the first decoding operation is the decoding operation on which the second variation operation is performed, and the second decoding operation is the decoding operation on which the second variation operation is not performed. When the encoding operation of a CODEC is the first encoding operation, the decoding operation of the CODEC is the first decoding operation; when the encoding operation of a CODEC is the second encoding operation, the decoding operation of the CODEC is the second decoding operation.

[0225] Thank a down-sampling operation and an up-sampling operation as an example, thedown-sampling operation and the up-sampling operation may be performed with one or more layers of the neural network for performing the decoding / encoding operation (that is, the ENCNN or the DEC NN). When the Enc / DEC NN comprises one or more layers which perform the down-sampling operation or the up-sampling operation, then it may be regarded that a variation operation is performed on the encoding / decoding operation performed by using the Enc / DECNN. When the first variation operation is a down-sampling operation, the second variation operation comprises an up-sampling operation; when the first variation operation is the up- sampling operation, the second variation operation is the down-sampling operation. For example, referring to FIG. 5, when a down-sampling operation (e.g., downscaling 8 times) is performed on the input data of the encoding operation belonging to ENC0, an up-sampling operation (e.g., upscaling 8 times) is performed on the output data of the decoding operation belonging t

[0226] When the first variation operation is a linear transform operation, the second variation operation is a linear transform operation. Referring to FIG. 6, when a linear transform operation is performed on the input data of the encoding operation belonging to ENC1, another linear transform operation may be performed on the output data of the decoding operation belonging to DEC2. The linear transform operation performed on the encoding side may allow to not only transform input tensor from one dimension to another if necessary, but also reorder input channel by their importance, which becomes very useful when finetuning model with preloaded analysis and synthesis transform operations checkpoints is applied.

[0227] When the variation operation is an analysis transform operation, the corresponding variation operation is a synthesis transform operation.

[0228] In some embodiments, the encoding device may further perform an entropy coding operation on the fundamental decoded bitstream. When the encoding device determines M encoded bitstream(s), the entropy coding operation may also be performed on the M encoded bitstream(s). Correspondingly, the decoding device may perform an entropy decoding operation first and then perform the decoding operation. The entropy coding operation may be an arithmetic coding, or, other types of entropy coding like a range coding, Huffman coding,Asymmetric numeral systems (ANS) can be used instead of arithmetic coding.

[0229] FIG. 8 illustrates an encoding-decoding procedure according to some embodiments ofthe present application.

[0230] Referring to FIG. 8, the encoding device may perform context-adaptive binary arithmetic coding (CAB AC) on the encoded bitstream to obtain an arithmetic encoded bitstream. The decoding device may receive the arithmetic encoded bitstream, performs context-adaptive binary arithmetic decoding (CAB AC) on the arithmetic encoded bitstream to obtain the encoded bitstream, and performs the decoding operation on the fundamental encoded bitstream.

[0231] As previously mentioned, the encoding device may further transmit some decoding information to the decoding device. In some embodiments, the encoding device may also perform the arithmetic coding operation on the decoding information and transmit the encoded decoding information to the decoding device. Correspondingly, the decoding device may perform the arithmetic decoding operation on the received encoded decoding information.

[0232] In some embodiments, before performing the entropy coding operation, the encoding device may first input the fundamental encoded bitstream to an optimization model to obtain an optimized parameter. When the encoding device generates M encoded bitstream(s), the M encoded bitstream(s) may also be inputted to the optimization model. Then, the entropy coding operation may be performed on the encoded bitstream based on the optimized parameter.

[0233] FIG. 9 illustrates an encoding-decoding procedure according to some embodiments of the present application.

[0234] Referring to FIG. 9, the encoded bitstream may be inputted to a hyperpior model (or hyperprior network), output of the hyper prior model is inputted to an entropy model, and output of the entropy model is an entropy parameter (that is, the optimized parameter). Then, the encoding device performs a binary arithmetic coding (BAC) on the encoded bitstream according to the entropy parameter to obtain an arithmetic encoded bitstream and transmits the arithmetic encoded bitstream to the decoding device. The entropy parameter may be transmitted to the decoding device. The decoding device receives the arithmetic encoded bitstream and the entropy parameter, performs binary arithmetic decoding on the arithmetic encoded bit stream according to the entropy parameter to obtain encoding bitstream, and performs the decoding operation on the encoding bitstream.

[0235] FIG. 9 illustrates a procedure for optimizing the fundamental encoded bitstream. The procedure for optimizing the M encoded bitstream(s) is similar to FIG. 9 and will refrain fromelaborating further for the sake of brevity.

[0236] FIG. 10 illustrates a decoding method according to some embodiments of the present application. For convenience, it is assumed that the operation illustrated in FIG. 10 is performed by a decoder or a decoding device. However, it should be understood that the operation illustrated in FIG. 10 may also be performed by a component of the decoder / decoding device, such as a circuit in the decoder / decoding device, or a chip or a SoC in the decoder / decoding device. As previously mentioned, the procedure for obtaining the encoded bitstream may also include the decoding operation. Therefore, the method illustrated in FIG. 10 may also be performed by the encoder / encoding device or a component in the encoder / encoding device.

[0237] 1001, The decoding device acquiring a bitstream D0.

[0238] 1002, The encoding determines a reconstructed feature tensor ẑ0according to the bitstream D0.

[0239] 1003, The encoding performs a decoding operation on the reconstructed feature tensor ẑ0to obtain a reconstructed latent representation ŷ0by using a neural network. The neural network may be the previously mentioned DEC NN.

[0240] 1004, The encoding performs a synthesis transform operation based on the reconstructed latent representation ŷ0to generate a reconstructed signal ẋ.

[0241] In some other embodiments, the decoding device may receive a decoding information from the encoding device, and the reconstructed feature tensor ẑ0may be determined according to the bitstream D0and the decoding information.

[0242] In some embodiments, the decoding information may include a length tensor Nbitso.The decoding device may determine an index tensor idxo according to the dimensional parameter Nmax_0. For example, the index tensor idxo may be filled continuously with values from 0 to Nmax_0. Then, the decoding device may determine a mask tensor MSK0according to the length tensor Nbitso and the index tensor idxo. The decoding device employs the same approach as the encoding device to determine the mask tensor MSK0according to the length tensor Nbitso and the index tensor idxo. For the sake of brevity, it will not be retreated here.

[0243] In some other embodiments, the decoding information may include the length tensorNbitso and the index tensor idxo. Under this condition, the decoding device do not need to determine the index tensor idxo by itself.

[0244] In some other embodiments, the decoding information may include the mask tensor MSK0. In other words, the decoding device may directly receive the mask tensor MSK0from the encoding device.

[0245] The bitstream D0may include X1bits, and the mask tensor MSK0includes X elements, where the X elements include X1first element with a first value and X2second element with a second value. X1and X2are positive integer, and a sum of X1and X2is equal to X.

[0246] The size of the encoded bitstream D0is less than a size of the reconstructed feature tensor ẑ0, and the size of the reconstructed feature tensor ẑ0is the same as the size of the mask tensor MSK0. As previously mentioned, if the size of the bitstream D0is X1and the size of the mask tensor MSK0is X, then the size of the reconstructed feature tensor ẑ0may be X. In other words, the reconstructed feature tensor ẑ0may include X elements. The X elements includes X1third element and X2fourth elements. The X1third elements are in one-to-one correspondence with the X1first elements in the mask tensor MSK0. Meanwhile, the X2fourth elements are in one-to-one correspondence with the X2second elements in the mask tensor MSK0. A location of each of the X1third elements is the same as a location of a corresponding first elements, and a location of each of the X2fourth elements is the same as a location of a corresponding second element. For example, if the first value is 1 , the second value is 0, and thenwhere a denotes the third element, and β denotes the fourth element. Further,the X1third elements are in one-to-one correspondence with X1bits in the encoded bitstream, and a value of each of the X1third elements is the same as a value of the corresponding bit. A value of each of the fourth elements is a present value which is different from the first value.For example, if D0=01100 and the preset value is 0, then The preset value 0 ismerely an example of the first value. In some other embodiments, the preset value may be any other value. Further, in some embodiments, values of different fourth elements may be different.

[0247] After obtaining the reconstructed feature tensor ẑ0, the encoding device may determinea reconstructed latent representation ŷ0corresponding to the reconstructed feature tensor ẑ0, and then perform the synthesis transform on the reconstructed latent representation ŷ0to obtain the reconstructed signal ẋ.

[0248] As previously mentioned, in some embodiments, in addition to the fundamental datastream information (that is the datastream information including the fundamental encoded bitstream D0and the fundamental decoding information), the encoding device may determineM piece(s) of datastream information and transmit the M piece(s) of datastream information to the decoding device. Correspondingly, the decoding device may receive the M piece(s) of datastream information and determine the reconstructed signal ẋ. according to the M piece(s) of datastream information and the fundamental datastream information.

[0249] Each of the M piece(s) of datastream information may include an encoded bitstream.Therefore, the decoding device may obtain M encoded bitstream(s).

[0250] The decoding device may perform M decoding operation(s) on the M encoded bitstream(s) to obtain M reconstructed latent representation. Then the decoding device may determine the reconstructed signal ẋ. according to the reconstructed latent representation ŷ0and the M reconstructed latent representation.

[0251] Similarly, each of the M piece(s) of datastream information may further include a decoding information corresponding to the encoded bitstream. The decoding information corresponding to the encoded bitstream include information for decoding the encoded bitstream.Details of the relationship between the encoded bitstream and the decoding information included in the datastream information may referred to the aforementioned encoded bitstream and the decoding information, and details for performing each of the M decoding operation(s) may refer to previously mentioned embodiments and it will not be retreated here.

[0252] In some embodiments, the decoding device may perform a synthesis transform operation on each of the reconstructed latent representation ŷ0and the M reconstructed latent representation to obtain M+1 synthesis transformation results and determine the reconstructed signal ẋ. according to the M+1 synthesis transformation results.

[0253] FIG. 11 illustrates a decoding procedure according to some embodiments of the present application.

[0254] Referring to FIG. 11, the fundamental datastream information obtained from theencoding device is denoted by bitso, a first datastream information among the M piece(s) of datastream information is denoted by bits1, a second datastream information among the M piece(s) of datastream information is denoted by bitS2, and so forth. A decoding procedure performed on bitso may be referred to as a decoding procedure #0, a decoding procedure performed on bits1may be referred to as a decoding procedure #1, and so forth. In FIG. 11, the decoding procedure #0 includes a decoding operation #0 denoted by DECo and a corresponding synthesis transform #0 denoted by So, the decoding procedure #1 includes a decoding operation#1 denoted by DEC1and a corresponding synthesis transform #1 denoted by S1, and so forth.As previously mentioned, each of the decoding operation has a corresponding tensor reconstructing operation. The decoding operation is performed on the result of the corresponding tensor reconstructing operation. However, FIG. 11 does not show the tensor reconstructing operation for brevity.

[0255] In FIG. 11, output of the decoding procedure #0 is denoted by reco, output of the decoding procedure #1 is denoted by rec1, and so forth. The reconstructed signal ẋ is a sum of outputs of the M+1 decoding procedure. In other words, the reconstructed signal ẋ and the output of the M+1 decoding procedures satisfy:

[0256] rec = rec0+ rec1+ rec2+ ... +recM, (3.1)

[0257] where rec = ẋ, that is the reconstructed data.

[0258] The decoding procedure illustrated in FIG. 11 corresponds to the encoding procedure illustrated in FIG. 5. In other words, the datastream information (that is, bitso to bitSM) are results of the encoding procedure illustrated in FIG. 5.

[0259] In some other embodiments, the decoding device may perform M+1 decoding operations and M+1 synthesis transform operations. The M+1 synthesis transform operations are performed on M+1 reference information and the reconstructed data is an output of theM+lthsynthesis transformation. A first reference information among the M+l reference information is the reconstructed latent representation ŷ0, a m+lthreference information is determined according to a mthtransformed information and a mthreconstructed latent representation ym, the mthtransformed information is determined according to a synthesis transformation result of a mthsynthesis transform operation among the M+l synthesis transform operations, and the reconstructed signal ẋ.

[0260] FIG. 12 illustrates a decoding procedure according to some embodiments of the present application. Similar to the abovementioned embodiments, ENC illustrated in FIG. 12 represents an encoding operation and a bitstream deriving operation corresponding to the encoding operation, and DEC illustrated in FIG. 12 represents a decoding operation and a tensor reconstructing operation corresponding to the decoding operation.

[0261] FIG. 12 illustrates 5 DECs (that is, DECo to DEC4 illustrated in FIG. 12) and 5 synthesis transform operations (that is, So to S4 illustrates in FIG. 12). DECo is performed based on the fundamental datastream information (denoted by bitso in FIG. 12), and ŷ0is the reconstructed latent representation determined by DECo. In FIG. 12, four datastream information obtained from the encoding device are denoted by bits1to bitS4, and yi to y4 are four reconstructed latent representations determined by DEC1to DEC4 respectively. The synthesis transformation results of So to S3 are denoted by STo to ST3.

[0262] Referring to FIG. 12, the first reference information is the input of So, which is ŷ0. The second reference information is the input of S1, and the second reference information is determined according to STo and yi. For example, the second reference information is a sum ofSTo and yi. Similarly, the second reference information, the input of S1, is determined according to ST1, and y2, and so forth. The synthesis transformation result of S4 is the reconstructed signal(denoted by rec i

[0263] In some embodiments, bitso to bitS4 illustrated in FIG. 12 may be determined according to the encoding procedure illustrated in FIG. 7. In other words, bitso to bitS4 illustrated in FIG.12 are bitso to bitS4 illustrated in

[0264] In some embodiments, a variation operation may be performed on the reconstructed latent representation ŷ0and / or at least one of the M reconstructed latent representation. For convenience, a term “first reconstructed latent representation” may be referred to a reconstructed latent representation among M+1 reconstructed latent representation (that is, the reconstructed latent representation ŷ0and the M reconstructed latent representation) on which the variation operation is performed, and a term “second reconstructed latent representation” may be referred to a reconstructed latent representation among M+1 reconstructed latent representation on which the variation operation is not performed. Correspondingly, the decoding operation that is used to generate the first reconstructed latent representation may bereferred to as “first decoding operation”, and the term “second decoding operation” may be referred to the decoding operation on which the corresponding variation operation is not performed.

[0265] As previously mentioned, the synthesis transform operation may be performed on the reconstructed latent representation ŷ0and / or the M reconstructed latent representation. When a reconstructed latent representation is the first reconstructed latent representation, the synthesis transform operation may be performed a varied reconstructed latent representation, the varied reconstructed latent representation is a result of the variation operation performed on the first reconstructed latent representation. When a reconstructed latent representation is the second reconstructed latent representation, the synthesis transform operation is performed on the second reconstructed latent representation.

[0266] The reconstructed latent representation is obtained by decoding the encoded bitstream by the decoding operation. If an encoded bitstream among M+1 encoded bitstream (that is, the encoded bitstream D0and the M encoded bitstream(s)) is performed by the first decoding operation, the encoded bitstream may be referred to as a first encoded bitstream; if an encoded bitstream among the M+1 encoded bitstream is not performed by the first decoding operation, the encoded bitstream may be referred to as a second encoded bitstream. Correspondingly, an encoding operation for determining the first encoded bitstream may be referred to as a first encoding operation, and an encoding operation for determining the second encoded bitstream may be referred to as a second encoding operation. Further, an input data of the first encoding operation may be referred to as a first input data, and an input data of the second encoding operation may be referred to as a second input data. The encoding device performs a corresponding variation operation on the first input data to obtain a varied input data, and then performs the first encoding on the varied input data. As previously mentioned, the encoding device may perform the first variation and the second variation. Details about the first variation and the second variation may refer to the previous embodiments, and will not be described here.For convenience, the variation operation performed by the decoding device may be referred to as a third variation operation. The third variation operation is an inverse operation of the first variation operation. For example, when the first variation operation is a down-sampling operation, the third variation operation is an up-sampling operation; when the first variationoperation is an up-sampling operation, the third variation operation is n down-sampling operation.

[0267] In some other embodiments, the decoding device may determine a reference reconstructed latent representation, and then perform a synthesis transform operation on the reference reconstructed latent representation to obtain the reconstructed signal, where the reconstructed signal is a result of the synthesis transform operation. The reference reconstructed latent representation may a sum of the reconstructed latent representation ŷ0and the M reconstructed latent representation. In other words, the reference reconstructed latent representation, the reconstructed latent representation ŷ0and the M reconstructed latent representation satisfy:

[0268]

[0269] where yrefis the reference reconstructed latent representation, ŷ0is the reconstructed latent representation, ymis a mthreconstructed latent representation among the M reconstructed latent representation, m

[0270] FIG. 13 illustrates an encoding procedure and a corresponding decoding procedure according to some embodiments of the present application. The definition of the abbreviations and notations in FIG. 13 may be referred to the aforementioned embodiments. Similar to the previously mentioned embodiments, FIG. 13 omits the bitstream deriving operation and the tensor reconstructing operation. Similar to the abovementioned embodiments, ENC illustrated in FIG. 13 represents an encoding operation and a bitstream deriving operation corresponding to the encoding operation, and DEC illustrated in FIG. 13 represents a decoding operation and a tensor reconstructing operation corresponding to the decoding operation.

[0271] Referring to FIG. 13, in order to obtain bitso, the encoding device performs a downing sampling operation before performing the encoding operation that is used to obtain bitso.Similarly, the encoding device performs a down-sampling operation before performing the encoding operation that is used to obtain bits1. For bitS2to bitSM, the encoding device do not perform a variation operation. For the encoding procedure, the encoding device performs an up-sampling operation on an output of a decoding operation used to decode bitso and performs an up-sampling operation on an output of a decoding operation used to decode bits1. For decoding operation for decoding bitS2to bitSM, the encoding device does not perform a variationoperation on outputs of these operations.

[0272] For the decoding procedure, the decoding device performs an up-sampling operation on a decoding operation that is used to decode bitso and performs an up-sampling operation on a decoding operation that is used to decode bits1. For decoding operation for decoding bitS2to bitSM, the decoding device does not perform a variation operation on outputs of these operations. In addition, the decoding device determines the reference reconstructed latent representation yrefaccording to ŷ0to yMand performs a synthesis transform operation on the reference reconstructed latent representation yrefto obtain a reconstructed signal ẋ (e.g., a reconstructed image or a reconstructed image).

[0273] FIG. 14 illustrates an encoding procedure and a corresponding decoding procedure according to some embodiments of the present application. The definition of the abbreviations and notations in FIG. 14 may be referred to the aforementioned embodiments. “L” in FIG. 14 denotes a linear transformation.

[0274] Similar to the embodiment illustrated in FIG. 13, an encoding device performs a down- sampling operation before performing an encoding operation for obtaining bitso and bits1and performs an up-sampling operation on results of a decoding operation performed on bitso and bitS2. In addition, the encoding device performs a linear transformation. For bitso and bits1, the linear transformation is performed before the down-sampling operation and after the up- sampling operation. For bitS2to bitSM, the linear transformation is performed before the encoding operation and after the decoding operation. Similarly, the decoding device further performs a linear transformation. For bitso and bits1, the linear transformation is performed after the up-sampling operation. For bitS2to bitSM, the linear transformation is performed after the decoding operation.

[0275] A loss function according to some embodiments of the present application may be the

[0276]

[0277] where WxH is a shape of the pre-coding data (that is, a input image of the encoding device), Nbits_idenotes to the number of bits (channels to use) in i-th row of the latent vector y (which size is equal to batches_number x rows_number x channels_number) , β value denotesdesired ration between both losses. The larger is β, the larger is distortion importance, and by thus, more bits will be used. For example, during a training procedure β may be changing from 02to 802(float number), and from 102, 202, 302, ..., 802(integer number) on an evaluation procedure. A distortion metrics of the loss function in equation 4.1 is mean square error (MSE), and MSE(ẋ,x) is an average of squares of the differences between a predicted value x and a true value x, where the predicted value ẋ. is a reconstructed signal determined by a decoding procedure according to the encoded bitstream determined by the encoding device, and the true value x is the input data of the encoding procedure.

[0278] For the embodiments illustrated in FIG. 11, the loss function may be:

[0279]

[0280] where bitSi(β) denotes number of bits of the encoded bitstream obtained by Bdei, and rec1(β) denotes number of bits of the output of the decoding operation #i. β represents a balance between rate and distortion in loss function. distortion( is a distortion metricsbetween x and , and x is the pre-coding data. The distortion metrics may be MSE,multi-scale structural similarity (MS-SSIM), structural similarity (SSIM) or the like.

[0281] Equation 4.3 shows MSE distortion with weights 8:1 :1 over a YUV image:

[0282] Distortion(A, B) = 0.8*MSE(A.Y, B.Y)+0.1*MSE(A.U, B.U)+0.1MSE(A.V, B.V)(4.3)

[0283] A, B are samples for the distortion evaluation.

[0284] The above-mentioned loss function may be used to train the ENC NN and the DECNN. FIG. 15 and FIG. 16 are used to depict the training procedure of the ENC NN and the DECNN.

[0285] Referring to FIG. 15, an embodiment of the present application provides a system architecture 1500. As shown in the system architecture 1500, a data collection device 1560 is configured to collect training data. The training data may be stored into a database 1530. A training device 1520 may obtain a target model / rule 1501 through training based on the training data maintained in the database 1530. In some embodiments, the target model / rule 1501 can be used to implement the ENC NN to obtain the feature tensor. In some other embodiments, the target model / rule 1501 can be used to implement the DEC NN to obtain the reconstructed latentrepresentation. The target model / rule 1501 in this embodiment of this application may specifically be a neural network obtained through training. In this embodiment provided in this application, the neural network is obtained by training an initialized neural network. It should be noted that, in actual application, the training data maintained in the database 1530 is not necessarily all collected by the data collection device 1560, and may be received from another device. In addition, it should be noted that the training device 1520 does not necessarily perform training completely based on the training data maintained in the database 1530 to obtain the target model / rule 1501, and may obtain training data from a cloud or another place to perform model training. The foregoing description shall not be construed as a limitation on this embodiment of this application.

[0286] The target model / rule 1501 obtained by the training device 1520 through training may be applied to different systems or devices, for example, applied to an execution device 1510 shown in FIG. 15. The execution device 1510 may be a terminal, such as a mobile phone terminal, a tablet computer, a notebook computer, an augmented reality (AR) device, a virtual reality (VR) device, or a vehicle-mounted terminal, or may be a server or the like. In FIG. 15, an I / O interface 1512 is configured on the execution device 1510 and is configured to exchange data with an external device. A user may input data into the I / O interface 1512 by using a customer device 1540. The input data may be data collected by the execution device 1510 by using the data collection device 1560, may be data in the database 1530, or may be data from the customer device 1540. The input data may be an image, a video frame, an audio data, a text data or the like, and the present application is not limited thereto.

[0287] A preprocessing module 1513 is configured to perform preprocessing based on the input data (for example, the input image) received by the I / O interface 1512. For example, the preprocessing module 1513 may be configured to implement one or more of the following operations: flipping, rotation, scaling, grayscale conversion, demozaicking, color correction, lens shading correction, auto white balance and the like; and is further configured to implement another preprocessing operation. This is not limited in this application.

[0288] In a related processing procedure in which the execution device 1510 preprocesses the input data or a calculation module 1511 of the execution device 1510 performs calculation, the execution device 1510 may invoke data, code, and the like in a data storage system 1550 toimplement corresponding processing, and may also store, into the data storage system 1550, data, an instruction, and the like obtained through corresponding processing.

[0289] Finally, the I / O interface 1512 returns a processing result, for example, the foregoing obtained image processing result, to the customer device 1540, to provide the processing result for the user.

[0290] It should be noted that the training device 1520 may obtain, through training based on different training data, corresponding target models / rules 1501 for different targets that are alternatively referred to as different tasks. The corresponding target models / rules 1501 may be used to implement the foregoing targets or complete the foregoing tasks, to provide a required result for the user.

[0291] In a case shown in FIG. 15, the user may manually provide the input data. The manually providing may be performed by using a screen provided on the I / O interface 1512. In another case, the customer device 1540 may automatically send the input data to the I / O interface 1512. If it is required that the customer device 1540 needs to obtain authorization from the user to automatically send the input data, the user may set corresponding permission on the customer device 1540. The user may view, on the customer device 1540, a result output by the execution device 1510. Specifically, the result may be displayed or may be presented in a form of sound, an action, or the like. The customer device 1540 may also be used as a data collection end to collect the input data that is input into the I / O interface 1512 and an output result that is output from the I / O interface 1512, as shown in the figure, use the input data and the output result as new sample data, and store the new sample data into the database 1530. Certainly, alternatively, the customer device 1540 may not perform collection, and the I / O interface 1512 directly stores, into the database 1530 as new sample data, the input data that is input into theI / O interface 1512 and an output result that is output from the I / O interface 1512, as shown in the figure.

[0292] It should be noted that FIG. 15 is merely a schematic diagram of a system architecture provided in an embodiment of the present application. A location relationship between a device, a component, a module, and the like shown in the figure constitutes no limitation. For example, in FIG. 15, the data storage system 1550 is an external memory relative to the execution device1510. In another case, the data storage system 1550 may be alternatively disposed in theexecution device 1510. In this application, the target model / rule 1501 obtained through training based on the training data may be a neural network used for performing the aforementioned encoding / decoding operation.

[0293] FIG. 16 illustrates a training method according to some embodiments of the present application. The method illustrated in FIG. 16 is used for training the ENC NN and the DECNN.

[0294] 1601, A training device acquires an input signal x.

[0295] 1602, The training device performs an analysis transform operation based on the input signal to obtain a latent representation ŷ0.

[0296] 1603, The training device performs an encoding operation based on the latent representation y0to obtain a feature tensor z0containing values in the range of 0 to 1 by using a first neural network.

[0297] 1604, The training device obtains a reconstructed feature tensor ẑ0based on the feature tensor z0.

[0298] 1605, The training device performs a decoding operation on the reconstructed feature tensor ẑ0to obtain a reconstructed latent representation ŷ0by using a second neural network.

[0299] 1606, The training device performs a synthesis transform operation based on the reconstructed latent representation ŷ0to generate a reconstructed signal ẋ.

[0300] 1607, The training device determines a loss function value according to the input signal x and the reconstructed signal ẋ.

[0301] 1608, The training device determines a first target neural network according to the first neural network and the loss function value.

[0302] 1609, The training device determines a second target neural network according to the second neural network and the loss function value.

[0303] The first target neural network is the ENC NN, and the second target neural network is the DEC NN.

[0304] In some embodiments, the training device may assign the reconstructed feature tensor ẑ0equal to the feature tensor z0.

[0305] In some embodiments, the training device may distort at least one element of the feature tensor z0before obtaining the reconstructed feature tensor ẑ0. For example, in someembodiment, the training device may invert 0 to 1 and 1 to 0. Inverting 0 to 1 and 1 to 0 simulates bit errors and / or packet losses. Therefore, the first neural network and the second neural network are trained with noise. Consequently, the first neural and the second neural may be tolerant to errors, distortions and losses in communication channel.

[0306] Equation 5.1 is used to guarantee that the dropping of bits with index larger than Nbitso.

[0307]

[0308] where B( ) represents a binarization function, y is hyper parameter responsible for the binarization function B sharpness. For example, in the beginning of training the binarization function may be softer that can be achieved by making y is smaller, e.g. 0.01. With increasing number of training epochs y may be increased e.g. up to 1 or 2. For example, in some embodiments, the binarization function B may satisfy:

[0309]

[0310] Further, during the training process, the binary function for obtaining the binary feature tensor z0and the feature tensor z0may satisfy:

[0311]

[0312] where z0’ satisfies equation 5.4, and 5 satisfies equation 5.5.

[0313]

[0315] where 9 is an optional noise addition. For example, δ = 0.5 * U(-l, 1), where U(-l,l) means random variable uniformly distributed from -1 to 1, emulating effect of round operation during training stage. By default δ = 0. Addition of δ allows to keep gradients unaffected, thought values in tensor z0are strictly binary, detach(z0’) refers to obtain a new tensor detached from z0’

[0316] The present application provides a NN based video, image, audio codecs - source codecs (which are focused only on data compression), channel codecs ((which are focused only on data transmission) and joint source channel coding (which are focused both on data compression and transmission). The NN is used to encode and decode the input data into bitstream. Different features of the input data are compressed to different number of bits(‘words’)

[0317] The present application allows to get trainable error tolerant data coding, or NN baseddata encryption system. The most obvious applications may include real time communication and video streaming systems, such as online meetings and video transmission, such as from phone to TV, from computer to projector, from drone to VR glasses and so on.

[0318] The NN may learn their own encoding / decoding way of latent space features represented as floats or integers to binary representation for sending via communication channel, find some correlations, and use less bits to encode data. Further, the ENC NN and the DEC NN may be end to end train together with noise - simulating both bit errors and packet losses - thus, becoming tolerant to errors, distortions and losses in communication channel.

[0319] Compared with conventional entropy encoding solutions, the NN based encoding solution provided by the present application have less complexity and less encoding / decoding time due to massive parallel processing on NN supporting hardware. Finally, the present application works in digital domain and do not making any constrains to physical layer having friendly design to existing widely deployed digital communication infrastructure. Therefore, the present application may be deployed with existing hardware (such as routers, Wi-Fi access points, 3G / 4G / 5G wireless communication devices).

[0320] FIG. 17 is a schematic block diagram of an electronic device 1700 according to some embodiments of the present application. The electronic device 1700 may be the aforementioned decoding device. Referring to FIG. 17, the electronic device 1700 includes an acquiring unit 1701, a determining unit 1702, a decoding unit 1703, and a transforming unit 1704.

[0321] The acquiring unit 1701 may be configured to acquire a bitstream D0.

[0322] The determining unit 1702 may be configured to determine a reconstructed feature tensor ẑ0according to the bitstream D0.

[0323] The decoding unit 1703 may be configured to perform a decoding operation on the reconstructed feature tensor ẑ0to obtain a reconstructed latent representation ŷ0by using a neural network.

[0324] The transforming unit 1704 may be configured to performing a synthesis transform operation based on the reconstructed latent representation ŷ0to generate a reconstructed signal ẋ

[0325] The determining unit 1702, the decoding unit 1703, and the transforming unit 1704 may be implemented by a processor. In some embodiments, the determining unit 1702, thedecoding unit 1703 and the transforming unit 1704 may be implemented by the one processor.In some other embodiments, the determining unit 1702, the decoding unit 1703, and the transforming unit 1704 may be implemented by two or more processors. In some embodiments, the acquiring unit 1701 may be implemented by a receiver or a receiving circuit of the processor.

[0326] Details on how to decode the acquired bitstream D0may refer to the above-mentioned embodiments and will not be described here.

[0327] FIG. 18 is a schematic block diagram of an electronic device 1800 according to some embodiments of the present application. The electronic device 1800 may be the aforementioned encoding device. Referring to FIG. 18, the electronic device 1800 includes an acquiring unit 1801, a transforming unit 1802, an encoding unit 1803, and a processing unit 1804.

[0328] The acquiring unit 1801 may be configured to acquire an input signal.

[0329] The transforming unit 1802 may be configured to perform an analysis transform operation based on the input signal to obtain a latent representation y0.

[0330] The encoding unit 1803 may be configured to perform an encoding operation based on the latent representation y0to obtain a feature tensor zO by using a neural network.

[0331] The processing unit 1804 may be configured to generating a bitstream D0according to the feature tenso

[0332] The transforming unit 1802, the encoding unit 1803, and the processing unit 1804 may be implemented by a processor. In some embodiments, the transforming unit 1802, the encoding unit 1803, and the processing unit 1804 may be implemented by the one processor. In some other embodiments, the transforming unit 1802, the encoding unit 1803, and the processing unit 1804 may be implemented by two or more processors. The acquiring unit 1801 may be implemented by a receiver or a receiving circuit of the processor.

[0333] Details on how to encode the input signal may refer to the above-mentioned embodiments and will not be described here.

[0334] FIG. 19 is a schematic block diagram of an electronic device 1900 according to some embodiments of the present application. The electronic device 1900 may be aforementioned training device. Referring to FIG. 19, the electronic device 1900 includes an acquiring unit 1901, a transforming unit 1902, an encoding unit 1903, an obtaining unit 1904, a decoding unit1905, and a processing unit 1906.

[0335] The acquiring unit 1901 may be configured to acquire an input signal x.

[0336] The transforming unit 1902, may be configured to perform an analysis transform operation based on the input signal to obtain a latent representation y0.

[0337] The encoding unit 1903 may be configured to perform an encoding operation based on the latent representation y0to obtain a feature tensor z0containing values in the range of 0 to 1 by using a first neural network.

[0338] The obtaining unit 1904 may be configured to obtain a reconstructed feature tensor ẑ0based on the feature tensor z0.

[0339] The decoding unit 1905 may be configured to perform a decoding operation on the reconstructed feature tensor ẑ0to obtain a reconstructed latent representation ŷ0by using a second neural network.

[0340] The transforming unit 1902 may be further configured to perform a synthesis transform operation based on the reconstructed latent representation ŷ0to generate a reconstructed signal

[0341] The processing unit 1906 may be configured to determine a loss function value according to the input signal x and the reconstructed signal x;

[0342] The processing unit 1906 may be further configured to determine a first target neural network according to the first neural network and the loss function value;

[0343] The processing unit 1906 may be further configured to determine a second target neural network according to the second neural network and the loss function value.

[0344] The transforming unit 1902, the encoding unit 1903, the obtaining unit 1904, the decoding unit 1905, and the processing unit 1906 may be implemented by a processor. In some embodiments, the transforming unit 1902, the encoding unit 1903, the obtaining unit 1904, the decoding unit 1905, and the processing unit 1906 may be implemented by the one processor. In some other embodiments, the transforming unit 1902, the encoding unit 1903, the obtaining unit 1904, the decoding unit 1905, and the processing unit 1906 may be implemented by two or more processors. The acquiring unit 1001 may be implemented by a receiver or a receiving circuit of the processor.

[0345] Details on how to train the first neural network and the second network may refer to the above-mentioned embodiments and will not be described

[0346] As shown in FIG. 20, an electronic device 2000 may include a receiver 2001, a processor 2002, and a memory 2003. The memory 2003 may be configured to store code, instructions, and the like executed by the processor 2002.

[0347] It should be understood that the processor 2002 may be an integrated circuit chip and has a signal processing capability. In an implementation process, steps of the foregoing method embodiments may be completed by using a hardware integrated logic circuit in the processor, or by using instructions in a form of software. The processor may be a general purpose processor, a central processing unit (CPU), a graphics processing unit (GPU), a neural processing unit(NPU), a system on chip (SoC) or another programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component. The processor may implement or perform the methods, the steps, and the logical block diagrams that are disclosed in the embodiments of the present application. The general purpose processor may be a microprocessor, or the processor may be any conventional processor or the like. The steps of the methods disclosed with reference to the embodiments of the present application may be directly performed and completed by the processor, or may be performed and completed by using a combination of hardware in the processor and a software module. The software module may be located in a mature storage medium in the art, such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, or a register. The storage medium is located in the memory, and the processor reads information in the memory and completes the steps of the foregoing methods performed by the encoding device or the decoding device in combination with hardware in the processor.

[0348] It may be understood that the memory 2003 in the embodiments of the present application may be a volatile memory or a nonvolatile memory, or may include both a volatile memory and a nonvolatile memory. The nonvolatile memory may be a read-only memory(Read-Only Memory, ROM), a programmable read-only memory (Programmable ROM,PROM), an erasable programmable read-only memory (Erasable PROM, EPROM), an electrically erasable programmable read-only memory (Electrically EPROM, EEPROM), or a flash memory. The volatile memory may be a random access memory (Random Access Memory,RAM) and is used as an external cache. By way of example rather than limitation, many formsof RAMs may be used, and are, for example, a static random access memory (Static RAM,SRAM), a dynamic random access memory (Dynamic RAM, DRAM), a synchronous dynamic random access memory (Synchronous DRAM, SDRAM), a double data rate synchronous dynamic random access memory (Double Data Rate SDRAM, DDR SDRAM), an enhanced synchronous dynamic random access memory (Enhanced SDRAM, ESDRAM), a synchronous link dynamic random access memory (Synchronous link DRAM, SLDRAM), and a direct rambus random access memory (Direct Rambus RAM, DR RAM).

[0349] It should be noted that the memory in the electronic device and the methods described in this specification includes but is not limited to these memories and a memory of any other appropriate type.

[0350] The present application provides a computer readable storage medium including instructions. When the instructions run on an electronic device, the electronic device is enabled to perform the aforementioned method.

[0351] The present application provides a computer readable storage medium. The computer readable storage medium stores the bitstream obtained by the aforementioned method.

[0352] The present application provides a chip system. The chip system includes a memory and a processor, and the memory is configured to store a computer program, and the processor is configured to invoke the computer program from the memory and run the computer program, so that an electronic device on which the chip system is disposed performs the aforementioned method.

[0353] The present application provides a computer program product. When the computer program product runs on an electronic device, the electronic device is enabled to perform the aforementioned method.

[0354] In the embodiments of the present application, “at least one” means one or more, and“a plurality of’ means two or more. The term “and / or” describes an association relationship between associated objects and represents that three relationships may exist. For example, A and / or B may represent the following three cases: only A exists, both A and B exist, and only B exists, where A and B may be singular or plural. The character “ / ” generally indicates an “or” relationship between the associated objects. “At least one of the following” and a similar expression thereof refer to any combination of these items, including any combination of oneitem or a plurality of items. For example, at least one of a, b, and c may indicate: a, b, c, a and b, a and c, b and c, or a, b, and c, where a, b, and c may be singular or plural.

[0355] A person of ordinary skill in the art may be aware that, in combination with the examples described in the embodiments disclosed in this specification, units and algorithm steps can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are performed by hardware or software depends on particular applications and design constraints of the technical solutions. A person skilled in the art may use different methods to implement the described functions for each particular application, but it should not be considered that the implementation goes beyond the scope of this application.

[0356] It may be clearly understood by a person skilled in the art that, for the purpose of convenient and brief description, for a detailed working process of the foregoing system, apparatus, and unit, refer to a corresponding process in the foregoing method embodiment.Details are not described herein again.

[0357] In the several embodiments provided in this application, it should be understood that the disclosed system, apparatus, and method may be implemented in other manners. For example, the described apparatus embodiment is merely an example. For example, the unit division is merely logical function division and may be other division in actual implementation.For example, a plurality of units or components may be combined or integrated into another system, or some features may be ignored or not performed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections may be implemented through some interfaces. The indirect couplings or communication connections between the apparatuses or units may be implemented in electronic, mechanical, or other forms.

[0358] The units described as separate parts may be or may not be physically separate, and parts displayed as units may be or may not be physical units, may be located in one position, or may be distributed on a plurality of network units. Some or all of the units may be selected based on actual requirements to achieve the objectives of the solutions of the embodiments.

[0359] In addition, functional units in the embodiments of this application may be integrated into one processing unit, or each of the units may exist alone physically, or two or more units are integrated into one unit.

[0360] When the functions are implemented in a form of a software functional unit and sold or used as an independent product, the functions may be stored in a computer readable storage medium. Based on such an understanding, the technical solutions in this application essentially, or the part contributing to the prior art, or some of the technical solutions may be implemented in a form of a software product. The computer software product is stored in a storage medium, and includes several instructions for instructing a computer device (which may be a personal computer, a server, a network device, or the like) to perform all or some of the steps of the methods described in the embodiments of this application. The foregoing storage medium includes: any medium that can store program code, such as a USB flash drive, a removable hard disk, a read-only memory (Read-Only Memory, ROM), a random access memory (RandomAccess Memory, RAM), a magnetic disk, or an optical disc.

[0361] The foregoing descriptions are merely specific implementations of this application, but are not intended to limit the protection scope of this application. Any variation or replacement readily figured out by a person skilled in the art within the technical scope disclosed in this application shall fall within the protection scope of this application. Therefore, the protection scope of this application shall be subject to the protection scope of the claims.

Claims

CLAIMSWhat is claimed is:

1. A decoding method, comprising: acquiring a bitstream D0; determining a reconstructed feature tensor zo according to the bitstream D0; performing a decoding operation on the reconstructed feature tensor zo to obtain a reconstructed latent representation yo by using a neural network; performing a synthesis transform operation based on the reconstructed latent representation yo to generate a reconstructed signal x.

2. The method according to claim 1, wherein the method further comprises: acquiring a decoding information, the determining a reconstructed feature tensor zo according to the bitstream D0, comprises: determining the reconstructed feature tensor zo according to the decoding information and the bitstream D0, wherein a size of the reconstructed feature tensor zo is larger than a size of the bitstream D0.3 The method according to claim 2, wherein the decoding information is obtained based on information signaled with the bitstream D0.

4. The method according to claims 2 or 3, wherein the decoding information comprises a length tensor Nbitso.

5. The method according to claim 4, wherein the determining the reconstructed feature tensor zo according to the decoding information, comprises: determining a mask tensor MSK0according to the length tensor Nbitso and an index tensor idxo, wherein the bitstream D0comprises X1bit(s), the mask tensor MSK0comprises X1first element(s) with a first value and X2 second element(s) with a second value, X1, a sum of X1and X2 is the size of the reconstructed feature tensor zo, X1is a positive integer, X2is a positive integer; determining the reconstructed feature tensor zo according to the mask tensor MSK0and the bitstream D0, wherein the reconstructed feature tensor zo comprises X1third element(s) andX2fourth element(s), a location of each element in the reconstructed feature tensor ẑ0is a same as a location of a corresponding element in the mask tensor MSK0, a value of each of the X1third element(s) is a same as a value of a corresponding bit in the bitstream D0, and a value of each of the X2fourth element(s) is a preset value that is not equals to 1.

6. The method according to any one of claims 1 to 5, wherein the method further comprises: acquiring M bitstream(s), wherein M is a positive integer; determining M reconstructed feature tensor(s) according to the M bitstream(s) respectively; performing M decoding operation(s) on the M reconstructed feature tensor(s) to obtain M reconstructed latent representation(s); the performing a synthesis transform operation based on the reconstructed latent representation ŷ0to generate a reconstructed signal ẋ, comprises: performing the synthesis transform operation based on the reconstructed latent representation y0and the M reconstructed latent representation(s) to generate the reconstructed signal ẋ.

7. The method according to claim 6, wherein the performing the synthesis transform operation based on the reconstructed latent representation ŷ0and the M reconstructed latent representation(s) to generate the reconstructed signal ẋ, comprises: performing the synthesis transform operation on each of the reconstructed latent representation ŷ0and the M reconstructed latent representation(s) to obtain M+1 synthesis transform results; generating the reconstructed signal ẋ according to the M+1 synthesis transformation results.

8. The method according to claim 6, wherein the performing the synthesis transform operation base on the reconstructed latent representation ŷ0and the M reconstructed latent representation(s) to generate the reconstructed signal ẋ, comprises: determining a reference reconstructed latent representation, wherein the reference reconstructed latent representation is a sum of the reconstructed latent representation ŷ0and theM reconstructed latent representation(s); performing the synthesis transform operation on the reference reconstructed latent representation to generate the reconstructed signal ẋ.

9. The method according to claim 6, wherein the performing the synthesis transform operation based on the reconstructed latent representation ŷ0and the M reconstructed latent representation(s) to generate the reconstructed signal ẋ, comprises: performing M+1 synthesis transform operations on M+1 reference information to generate the reconstructed signal ẋ, wherein a first reference information among the M+1 reference information is the reconstructed latent representation ŷ0, a m+lthreference information is determined according to a ,th transformed information and a mthreconstructed latent representation, the mthtransformed information is determined according to a synthesis transformation result of a mthsynthesis transform operation among the M+1 synthesis transform operations.

10. The method according to any one of claims 7 to 9, wherein the method further comprises: performing a variation operation on a reference reconstructed latent representation, wherein the reference reconstructed latent representation is a reconstructed latent representation determined by a reference decoding operation, the reference decoding operation is the decoding operation or one of the M decoding operation(s), the variation operation comprises at least one of the followings: a sampling variation operation, or, a linear transformation, or an analysis transform operation, and the reconstructed signal ẋ is determined according to the varied reconstructed latent representation.

11. The method according to any one of claims 1 to 10, wherein the acquiring a bitstream D0, comprises: performing an entropy decoding operation on a received bitstream to obtain the bitstream D0.

12. The method according to claim 11, wherein the method further comprises: obtaining an optimized parameter; the performing an entropy decoding operation on a received bitstream to obtain the bitstream D0, comprises: performing, according to the optimized parameter, the entropy decoding operation on the received bitstream to obtain the bitstream D0.

13. An encoding method, comprising:acquiring an input signal; performing an analysis transform operation based on the input signal to obtain a latent representation y performing an encoding operation based on the latent representation y0to obtain a feature tensor z0by using a neural network; generating a bitstream D0according to the feature tensor z0.

14. The method according to claim 13, wherein the method further comprises: acquiring an encoding information; the generating a bitstream D0according to the feature tensor z0, comprises: generating the bitstream D0according to the encoding information and the feature tensor z0, wherein a size of the feature tensor z0is larger than a size of the bitstream D0.15 The method according to claim 14, wherein the encoding information is signaled with the bitstream16. The method according to claims 14 or 15, wherein the encoding information is a length tensor Nbitso.

17. The method according to claim 16, wherein the length tensor Nbisto is obtained by the encoding operation using the neural network.

18. The method according to claim 16 or 17, wherein the generating the bitstream D0according to the encoding information and the feature tensor z0, comprises: determining a mask tensor MSK0according to the length tensor Nbitso and an index tensor idxo, wherein the mask tensor MSK0comprises X1first element(s) with a first value and X2second element(s) with a second value, a sum of X1and X2is the size of the feature tensor z0, X1is a positive integer, X2is a positive integer; determining the bitstream D0according to the mask tensor MSK0and the feature tensor z0, wherein the bitstream D0comprises X1bit(s), the feature tensor z0comprises X1third element(s) and X2fourth element(s), a location of each elements in the feature tensor z0is a same of a location of a corresponding element in the mask tensor MSK0, a value of each of the X1third element(s) is configured to obtain a value of a corresponding bit in the bitstream D0.

19. The method according to any one of claims 13 to 18, wherein the method further cperforming M encoding operation(s) and M decoding operation(s) according to the bitstream D0to obtain M bitstream(s), wherein a mthbitstream among the M bitstream(s) is generated by performing a mthencoding operation among the M encoding operation(s) on an input data INPm, and the input data INPmis determined according to m reconstructed latent representation(s) obtained according to a first m-1 decoding operation(s) of the M decoding operation(s)20. The method according to claim 19, wherein the input data INPmand the m reconstructed latent representation(s) satisfy:wherein INPmis the input data INPm, AT ( ) denotes the analysis transform operation, ST ( ) denotes a synthesis transform operation, x is the input signal, and yi is an ithreconstructed latent representation among the m reconstructed latent representation(s).

21. The method according to claim 19, wherein the input data INPm and the m reconstructed latent representation(s) satisfy:wherein INPmis the input data INPm, y0is the latent representation y0, and y; is an ; it'h reconstructed latent representation among the m reconstructed latent representation(s).

22. The method according to claim 19, wherein the input data INPmand the m reconstructed latent representation(s) satisfy:wherein INPm is the input data INPm,( ) analysis transform operation among K analysis transform operation(s);STmis a result of a mthsynthesis transform operation among K synthesis transform operation(s), and the mthsynthesis transform operation is performed on a target data determined according to the m reconstructed latent representation(s).

23. The method according to any one of claims 19 to 22, wherein the method further comprises: performing a first variation operation on a reference input data, wherein the reference inputdata is an input data of a reference encoding operation, the reference encoding operation is the encoding operation or one of the M encoding operation(s), the first variation operation comprises at least one of the followings: a sampling variation operation, or, a linear transformation, or an analysis transform operation, and the reference encoding operation is performed on the varied first input data; performing a second variation operation on a result of a reference decoding operation, wherein the second variation operation is a variation operation corresponding to the first the variation operation.

24. The method according to any one of claims 13 to 23, wherein the method further comprises: performing an entropy coding operation on the bitstream D0.

25. The method according to claim 24, wherein the performing an entropy coding operation on the encoded bitstream D0, comprises: determining an optimized parameter according to the bitstream D0by an optimization model; performing, according to the optimized parameter, the entropy coding operation on the bitstream26. A method for training a neural network, comprising: acquiring an input signal x; performing an analysis transform operation based on the input signal to obtain a latent representation y ; performing an encoding operation based on the latent representation y0to obtain a feature tensor z0containing values in the range of 0 to 1 by using a first neural network; obtaining a reconstructed feature tensor ẑ0based on the feature tensor z0; performing a decoding operation on the reconstructed feature tensor ẑ0to obtain a reconstructed latent representation ŷ0by using a second neural network; performing a synthesis transform operation based on the reconstructed latent representation ŷ0to generate a reconstructed signal ẋ; determining a loss function value according to the input signal x and the reconstructed signal ẋ;determining a first target neural network according to the first neural network and the loss function value; determining a second target neural network according to the second neural network and the loss function value.

27. The method according to claim 26, wherein the obtaining a reconstructed feature tensor ẑ0based on the feature tensor z0comprises: assigning the reconstructed feature tensor ẑ0equal to the feature tensor z0.

28. The method according to claim 26, wherein the obtaining a reconstructed feature tensor ẑ0based on the feature tensor z0comprises: distorting at least one element of the feature tensor z0to obtain the reconstructed feature tenso29. The method according to claim 28, wherein the distorting at least one element of the feature tensor z0comprises inverting 0 to 1 and 1 t30. A decoding device, comprising: an acquiring unit, configured to acquire a bitstream D0; a determining unit, configured to determine a reconstructed feature tensor ẑ0according to the bitstreama decoding unit, configured to perform a decoding operation on the reconstructed feature tensor ẑ0to obtain a reconstructed latent representation ŷ0by using a neural network; a transforming unit, configured to performing a synthesis transform operation based on the reconstructed latent representation ŷ0to generate a reconstructed signal ẋ31. The decoding device according to claim 30, wherein the acquiring unit, further configured to acquire a decoding information, the determining unit is further configured to determining the reconstructed feature tensor ẑ0according to the decoding information and the bitstream D0, wherein a size of the reconstructed feature tensor ẑ0is larger than a size of the bitstream D0.32 The decoding device according to claim 31, wherein the decoding information is obtained based on information signaled with the bitstream D0.

33. The decoding device according to claims 31 or 32, wherein the decoding information comprises a length tensor34. The decoding device according to claim 33, wherein the determining unit is further configured to: determining a mask tensor MSK0according to the length tensor NbitsO and an index tensor idxo, wherein the bitstream D0comprises X1bit(s), the mask tensor MSK0comprises X1first element(s) with a first value and X2second element(s) with a second value, X1, a sum of X1and X2is the size of the reconstructed feature tensor ẑ0, X1is a positive integer, X2is a positive integer; determining the reconstructed feature tensor ẑ0according to the mask tensor MSK0and the bitstream D0, wherein the reconstructed feature tensor ẑ0comprises X1third element(s) and X2fourth element(s), a location of each element in the reconstructed feature tensor ẑ0is a same as a location of a corresponding element in the mask tensor MSK0, a value of each of the X1third element(s) is a same as a value of a corresponding bit in the bitstream D0, and a value of each of the X2fourth element(s) is a preset value that is not equals to 1.

35. The decoding device according to any one of claims 30 to 34, wherein the acquiring unit is further configured to acquire M bitstream(s), wherein M is a positive integer; the determining unit is further configured to determine M reconstructed feature tensor(s) according to the M bitstream(s) respectively; the decoding unit is further configured to perform M decoding operation(s) on the M reconstructed feature tensor(s) to obtain M reconstructed latent representation(s); the performing unit, is further configured to perform the synthesis transform operation based on the reconstructed latent representation y0and the M reconstructed latent representation(s) to generate the reconstructed signal ẋ.

36. The decoding device according to claim 35, wherein the performing unit is further configured to perform the synthesis transform operation on each of the reconstructed latent representation ŷ0and the M reconstructed latent representation(s) to obtain M+1 synthesis transform results; generating the reconstructed signal ẋ according to the M+1 synthesis transformation results.

37. The decoding device according to claim 35, wherein the performing unit is further configureddetermining a reference reconstructed latent representation, wherein the reference reconstructed latent representation is a sum of the reconstructed latent representation ŷ0and theM reconstructed latent representation(s); performing the synthesis transform operation on the reference reconstructed latent representation to generate the reconstructed signal38. The decoding device according to claim 35, wherein the performing t unit is further configured to perform M+1 synthesis transform operations on M+1 reference information to generate the reconstructed signal ẋ, wherein a first reference information among the M+1 reference information is the reconstructed latent representation ŷ0, a m+lthreference information is determined according to a mthtransformed information and a mthreconstructed latent representation, the 111thtransformed information is determined according to a synthesis transformation result of a mthsynthesis transform operation among the M+1 synthesis transform operations.

39. The decoding device according to any one of claims 36 to 38, wherein the performing unit is further configured to perform a variation operation on a reference reconstructed latent representation, wherein the reference reconstructed latent representation is a reconstructed latent representation determined by a reference decoding operation, the reference decoding operation is the decoding operation or one of the M decoding operation(s), the variation operation comprises at least one of the followings: a sampling variation operation, or, a linear transformation, or an analysis transform operation, and the reconstructed signal ẋ is determined according to the varied reconstructed latent representation.

40. The decoding device according to any one of claims 30 to 39, wherein the acquiring unit is further configured to perform an entropy decoding operation on a received bitstream to obtain the bitstream41. The decoding device according to claim 40, wherein the acquiring unit is further configured to obtain an optimized parameter; performing, according to the optimized parameter, the entropy decoding operation on the received bitstream to obtain the bitstream D0.

42. An encoding device, comprising: an acquiring unit, configured to acquire an input signal;a transforming unit, configured to perform an analysis transform operation based on the input signal to obtain a latent representation y0; an encoding unit, configured to perform an encoding operation based on the latent representation y0to obtain a feature tensor z0by using a neural network; a processing unit, configured to generate a bitstream D0according to the feature tensor z0.

43. The encoding device according to claim 42, wherein the acquiring unit is further configured to acquire an encoding information; the processing unit is further configured to generate the bitstream D0according to the encoding information and the feature tensor z0, wherein a size of the feature tensor z0is larger than a size of the bitstream D0.44 The encoding device according to claim 43, wherein the encoding information is signaled with the bitstream D0.

45. The encoding device according to claims 43 or 44, wherein the encoding information is a length tensor Nbitso.

46. The encoding device according to claim 45, wherein the length tensor Nbisto is obtained by the encoding operation using the neural network.

47. The encoding device according to claim 45 or 46, wherein the processing unit is further configured to determine a mask tensor MSK0according to the length tensor Nbitso and an index tensor idxo, wherein the mask tensor MSK0comprises X1first element(s) with a first value and X2second element(s) with a second value, a sum of X1and X2is the size of the feature tensor z0, X1is a positive integer, X2is a positive integer; determine the bitstream D0according to the mask tensor MSK0and the feature tensor z0, wherein the bitstream D0comprises X1bit(s), the feature tensor z0comprises X1third element(s) and X2fourth element(s), a location of each elements in the feature tensor z0is a same of a location of a corresponding element in the mask tensor MSK0, a value of each of the X1third element(s) is configured to obtain a value of a corresponding bit in the bitstream D0.

48. The encoding device according to any one of claims 42 to 47, wherein the encoding unit is further configured to performing M encoding operation (s) and M decoding operation(s) according to the bitstream D0to obtain M bitstream(s), wherein a mthbitstream among the M bitstream(s) is generated by performing a ithencoding operation among the M encodingoperation(s) on an input data INPm, and the input data INPm is determined according to m reconstructed latent representation(s) obtained according to a first m-1 decoding operation(s) of the M decoding operation(s), m49. The encoding device according to claim 48, wherein the input data INPmand the m reconstructed latent representation(s) satisfy:wherein INPmis the input data INPm, AT ( ) denotes the analysis transform operation, ST( ) denotes a synthesis transform operation, x is the input signal, and y; is an ithreconstructed latent representation among the m reconstructed latent representation(s).50 The encoding device according to claim 48, wherein the input data INPmand the m reconstructed latent representation(s) satisfy:wherein INPmis the input data INPm, y0is the latent representation y0, and y; is an ithreconstructed latent representation among the m reconstructed latent representation(s).

51. The encoding device according to claim 48, wherein the input data INPmand the m reconstructed latent representation(s) satisfy:wherein INPmis the input data INPmanalysis transform operation among K analysis transform operation(s);STmis a result of a mthsynthesis transform operation among K synthesis transform operation(s), and the mthsynthesis transform operation is performed on a target data determined according to the m reconstructed latent representation(s).

52. The encoding device according to any one of claims 48 to 51, wherein the processing unit is further configured to: perform a first variation operation on a reference input data, wherein the reference input data is an input data of a reference encoding operation, the reference encoding operation is the encoding operation or one of the M encoding operation(s), the first variation operation comprises at least one of the followings: a sampling variation operation, or, a lineartransformation, or an analysis transform operation, and the reference encoding operation is performed on the varied first input data; perform a second variation operation on a result of a reference decoding operation, wherein the second variation operation is a variation operation corresponding to the first the variation operation.

53. The encoding device according to any one of claims 42 to 52, wherein the processing unit is further configured to perform an entropy coding operation on the bitstream D0.

54. The encoding device according to claim 53, wherein the processing unit is further configured to: determine an optimized parameter according to the bitstream D0by an optimization model; perform, according to the optimized parameter, the entropy coding operation on the bitstream55. A training device, comprising: an acquiring unit, configured to acquire an input signal x; a transforming unit, configured to perform an analysis transform operation based on the input signal to obtain a latent representation y0; an encoding unit, configured to perform an encoding operation based on the latent representation y0to obtain a feature tensor z0containing values in the range of 0 to 1 by using a first neural network; an obtaining unit, configured to obtain a reconstructed feature tensor ẑ0based on the feature tensor z a decoding unit, configured to perform a decoding operation on the reconstructed feature tensor ẑ0to obtain a reconstructed latent representation ŷ0by using a second neural network; the transforming unit, further configured to perform a synthesis transform operation based on the reconstructed latent representation ŷ0to generate a reconstructed signal ẋ; a processing unit, configured to determine a loss function value according to the input signal x and the reconstructed signal ẋ; the processing unit, further configured to determine a first target neural network according to the first neural network and the loss function value; the processing unit, further configured to determine a second target neural networkaccording to the second neural network and the loss function value.

56. The training device according to claim 55, wherein the obtaining unit is further configured to assign the reconstructed feature tensor ẑ0equal to the feature tensor z0.

57. The training device according to claim 55, wherein the obtaining unit is further configured to distort at least one element of the feature tensor z0to obtain the reconstructed feature tensor ẑ0.

58. The training device according to claim 57, wherein the obtaining unit is further configured to invert 0 to 1 and 1 to 0.

59. A computer readable storage medium, wherein the computer readable storage medium stores instructions, and when the instructions run on an electronic device, the electronic device is enabled to perform the method according to any one of claims 1 to 12.

60. A computer readable storage medium, wherein the computer readable storage medium stores instructions, and when the instructions run on an electronic device, the electronic device is enabled to perform the method according to any one of claims 13 to 25.

61. A computer readable storage medium, wherein the computer readable storage medium stores instructions, and when the instructions run on an electronic device, the electronic device is enabled to perform the method according to any one of claims 26 to 29.

62. A computer readable storage medium, wherein the computer readable storage medium stores a bitstream, wherein the bitstream is obtained by using the method according to any one of claims 13 to 25.

63. An electronic device, comprising a memory and a processor, wherein the memory is configured to store a computer program, and the processor is configured to invoke the computer program from the memory and run the computer program, so that the electronic device performs the method according to any one of claims 1 to 12.

64. An electronic device, comprising a memory and a processor, wherein the memory is configured to store a computer program, and the processor is configured to invoke the computer program from the memory and run the computer program, so that the electronic device performs the method according to any one of claims 13 to 25.

65. An electronic device, comprising a memory and a processor, wherein the memory is configured to store a computer program, and the processor is configured to invoke the computerprogram from the memory and run the computer program, so that the electronic device performs the method according to any one of claims 25 to 29.

66. A chip system, comprising a memory and a processor, wherein the memory is configured to store a computer program, and the processor is configured to invoke the computer program from the memory and run the computer program, so that an electronic device on which the chip system is disposed performs the method according to any one of claims 1 to 12.

67. A chip system, comprising a memory and a processor, wherein the memory is configured to store a computer program, and the processor is configured to invoke the computer program from the memory and run the computer program, so that an electronic device on which the chip system is disposed performs the method according to any one of claims 13 to 25.

68. A chip system, comprising a memory and a processor, wherein the memory is configured to store a computer program, and the processor is configured to invoke the computer program from the memory and run the computer program, so that an electronic device on which the chip system is disposed performs the method according to any one of claims 26 to 29.

69. A computer program product, wherein when the computer program product runs on an electronic device, the electronic device is enabled to perform the method according to any one of claims 1 to 12.

70. A computer program product, wherein when the computer program product runs on an electronic device, the electronic device is enabled to perform the method according to any one of claims 13 to 25.

71. A computer program product, wherein when the computer program product runs on an electronic device, the electronic device is enabled to perform the method according to any one of claims 26 to 29.

Citation Information

Patent Citations

  • Enhanced coding efficiency with progressive representation

    US10977553B2

  • Method and apparatus for multi-rate neural image compression with micro-structured masks

    US20210406691A1

Cited By

  • Context-aware error concealment to improve inference accuracy

    US12641301B2

  • Context-aware error concealment to improve inference accuracy

    US20260059148A1