Decoding method and coding method based on neural network and electronic equipment

By using a neural network-based decoding method in the audio encoding and decoding scheme and using context information to predict entropy decoding parameters, the problems of high entropy decoding code rate consumption and high complexity in the prior art are solved, and more efficient decoding performance is achieved.

CN120071944APending Publication Date: 2025-05-30TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311615622.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-28
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing end-to-end audio encoding and decoding scheme based on deep learning has the problem of high code rate consumption and high complexity of solution methods during the entropy decoding process, which affects the decoding efficiency and performance.

Method used

By using a neural network-based decoding method, by obtaining the context information of the elements to be decoded, using the parameter prediction model to predict the entropy decoding parameters, entropy decoding is directly performed, avoiding the introduction of a code stream specifically used to determine the decoding parameters.

Benefits of technology

Reduces the complexity of bit rate consumption and solution methods, and improves decoding efficiency and performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071944A_ABST
    Figure CN120071944A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a decoding method and a coding method based on a neural network and electronic equipment, and relates to the technical field of coding and decoding, and the method comprises the steps: obtaining a first to-be-decoded element based on a first code stream; obtaining context information of the first to-be-decoded element based on at least one decoded element decoded before the first to-be-decoded element; based on the context information, predicting a first decoding parameter for performing entropy decoding on the first to-be-decoded element by using a parameter prediction model based on a neural network; and performing entropy decoding on the first to-be-decoded element based on the first decoding parameter to obtain decoding data of the first to-be-decoded element. According to the method, the code rate consumption and the decoding complexity can be reduced, so that the decoding efficiency and performance can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of coding and decoding technologies, and more specifically, to a decoding method, an encoding method, and an electronic device based on a neural network. Background Art

[0002] In recent years, deep learning solutions have been widely applied to the processing of signals in different dimensions (such as audio, images, and videos).

[0003] In an end-to-end audio coding and decoding solution based on deep learning, an encoder-quantization-decoder structure is usually adopted. During encoding, the encoder first inputs the input audio signal into an encoding network for non-linear transformation; then, a quantization operation is performed based on the latent variables obtained after the transformation; then, entropy coding is performed on the quantization result, and a binary code stream is generated. During decoding, the decoder first performs entropy decoding on the code stream to obtain latent variables, then performs an inverse quantization operation on the recovered latent variables, and inputs them into a decoding network for non-linear transformation to obtain the reconstructed audio.

[0004] However, with the development of technology, there is still a need to pursue better probability modeling techniques to improve the entropy decoding performance, so as to improve the decoding efficiency and performance. Summary of the Invention

[0005] Embodiments of the present application provide a decoding method, an encoding method, and an electronic device based on a neural network, which can reduce the code rate consumption and decoding complexity, and thus can improve the decoding efficiency and performance.

[0006] In a first aspect, an embodiment of the present application provides a decoding method based on a neural network, including:

[0007] Obtaining a first element to be decoded based on a first code stream;

[0008] Obtaining context information of the first element to be decoded based on at least one decoded element decoded before the first element to be decoded;

[0009] Predicting a first decoding parameter for performing entropy decoding on the first element to be decoded based on the context information by using a parameter prediction model based on a neural network;

[0010] Performing entropy decoding on the first element to be decoded based on the first decoding parameter to obtain decoded data of the first element to be decoded.

[0011] In a second aspect, an embodiment of the present application provides an encoding method based on a neural network, including:

[0012] Obtaining a first element to be encoded;

[0013] Obtain context information of the first element to be encoded based on at least one encoded element that has been encoded before the first element to be encoded;

[0014] Based on the context information, use a neural network-based parameter prediction model to predict a first encoding parameter for entropy encoding the first element to be encoded;

[0015] Entropy encode the first element to be encoded based on the first encoding parameter to obtain encoded data of the first element to be encoded.

[0016] In a third aspect, an embodiment of the present application provides a decoder, including:

[0017] A first acquisition unit, configured to obtain a first element to be decoded based on a first bitstream;

[0018] A second acquisition unit, configured to obtain context information of the first element to be decoded based on at least one decoded element that has been decoded before the first element to be decoded;

[0019] A determination unit, configured to predict a first decoding parameter for entropy decoding the first element to be decoded based on the context information by using a neural network-based parameter prediction model;

[0020] An entropy decoding unit, configured to perform entropy decoding on the first element to be decoded based on the first decoding parameter to obtain decoded data of the first element to be decoded.

[0021] In a fourth aspect, an embodiment of the present application provides an encoder, including:

[0022] A first acquisition unit, configured to obtain a first element to be encoded;

[0023] A second acquisition unit, configured to obtain context information of the first element to be encoded based on at least one encoded element that has been encoded before the first element to be encoded;

[0024] A determination unit, configured to predict a first encoding parameter for entropy encoding the first element to be encoded based on the context information by using a neural network-based parameter prediction model;

[0025] An entropy encoding unit, configured to perform entropy encoding on the first element to be encoded based on the first encoding parameter to obtain encoded data of the first element to be encoded.

[0026] In a fifth aspect, an embodiment of the present application provides an electronic device, including:

[0027] A processor, adapted to implement computer instructions; and,

[0028] A computer-readable storage medium stores computer instructions, and the computer instructions are adapted to be loaded and executed by a processor to perform the decoding method described in the first aspect or the encoding method described in the second aspect above.

[0029] In one implementation, the processor is one or more, and the memory is one or more.

[0030] In one implementation, the computer-readable storage medium may be integrated with the processor, or the computer-readable storage medium is separately provided from the processor.

[0031] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium storing computer instructions, and when the computer instructions are read and executed by a processor of a computer device, the computer device is caused to perform the decoding method described in the first aspect above or the encoding method described in the second aspect above.

[0032] In a seventh aspect, an embodiment of the present application provides a computer program product or a computer program, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the decoding method described in the first aspect above or the encoding method described in the second aspect above.

[0033] In a ninth aspect, an embodiment of the present application provides a bitstream, which may be a bitstream decoded by using the decoding method described in the first aspect above, or the bitstream may be a bitstream generated by using the encoding method described in the second aspect above.

[0034] For the decoding method based on a neural network provided in the present application, the method includes: obtaining a first element to be decoded based on a first bitstream; obtaining context information of the first element to be decoded based on at least one decoded element decoded before the first element to be decoded; predicting a first decoding parameter for entropy decoding the first element to be decoded by using a parameter prediction model based on a neural network based on the context information; and performing entropy decoding on the first element to be decoded based on the first decoding parameter to obtain decoded data of the first element to be decoded. That is, when predicting the first decoding parameter of the first element to be decoded, by considering the context information of the first element to be decoded, the accuracy of the first decoding parameter can be improved, and further the bitrate consumption and decoding complexity can be reduced, and further the decoding efficiency and performance can be improved. Description of the Drawings

[0035] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments of the present application. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0036] Figure 1 It is a schematic block diagram of an audio encoding and decoding system provided by an embodiment of the present application.

[0037] Figure 2 It is a schematic block diagram of an audio encoding and decoding system provided by an embodiment of the present application.

[0038] Figure 3 It is a schematic block diagram of a decoding method based on a neural network provided by an embodiment of the present application.

[0039] Figure 4 It is an example of an element for determining context information provided by an embodiment of the present application.

[0040] Figure 5 It is another example of an element for determining context information provided by an embodiment of the present application.

[0041] Figure 6 It is a schematic block diagram of another encoding and decoding system provided by an embodiment of the present application.

[0042] Figure 7 It is a schematic block diagram of another encoding and decoding system provided by an embodiment of the present application.

[0043] Figure 8 It is a schematic block diagram of an encoding method based on a neural network provided by an embodiment of the present application.

[0044] Figure 9 It is a schematic block diagram of a decoder provided by an embodiment of the present application.

[0045] Figure 10 It is a schematic block diagram of an encoder provided by an embodiment of the present application.

[0046] Figure 11 It is a schematic block diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0047] The following will clearly describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application belong to the scope of protection of the present application.

[0048] It should be noted that the terms used in the implementation part of this application are only used to explain the embodiments of this application, and are not intended to limit this application.

[0049] For example, the term "and / or" in this text is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The term "at least one" is merely a description of the combination relationship of listed objects, indicating that there can be one or more. For example, at least one of the following: A, B, C can represent the following combination situations: A exists alone, B exists alone, C exists alone, A and B exist simultaneously, A and C exist simultaneously, B and C exist simultaneously, and A, B, and C exist simultaneously. The term "plural" means two or more. The character " / " generally indicates that the associated objects before and after are in an "or" relationship.

[0050] Again, the term "correspond" can indicate a direct or indirect corresponding relationship between two parties, can also indicate an association relationship between two parties, or can be a relationship such as indication and being indicated, configuration and being configured, etc. The term "indicate" can be a direct indication, an indirect indication, or can also indicate an association relationship. For example, A indicates B, which can mean that A directly indicates B. For example, B can be obtained through A; it can also mean that A indirectly indicates B. For example, A indicates C, and B can be obtained through C; it can also mean that there is an association relationship between A and B. The term "predefined" or "preconfigured" can pre-save corresponding codes, tables, or other relevant information that can be used for indication in a device, or can also refer to being agreed upon by a protocol. The "protocol" can refer to a standard protocol in this field. The term "when..." can be interpreted as "if" or "when" or "when...and" or "in response to" and other similar descriptions. Similarly, depending on the context, the phrase "if it is determined" or "if it is detected (the stated condition or event)" can be interpreted as "when it is determined" or "in response to the determination" or "when it is detected (the stated condition or event)" or "in response to the detection (the stated condition or event)" and other similar descriptions. The terms "first", "second", "third", "fourth", "Ath", "Bth", etc. are used to distinguish different objects, rather than to describe a specific order. The terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion.

[0051] The solution provided by this application relates to the field of digital compression technology.

[0052] Among them, digital compression technology mainly compresses huge audio data for the convenience of transmission, storage, etc.

[0053] The solution provided by this application can be applied to the field of digital audio coding technology.

[0054] Among them, the digital audio coding technology field includes, but is not limited to, at least one of the following: the audio codec field, the hardware audio codec field, the dedicated circuit audio codec field, and the real-time audio codec field. In addition, the solution provided in this application can be combined with the following standards: the Audio Video coding Standard (AVS), the second-generation AVS standard (AVS2), or the third-generation AVS standard (AVS3). For example, including but not limited to: the H.264 / AudioVideo coding (AVC) standard.

[0055] For the convenience of understanding, first, in combination with Figure 1 the codec system involved in the embodiments of this application is introduced.

[0056] Figure 1 is a schematic block diagram of a codec system 100 provided by the embodiments of this application.

[0057] As Figure 1 shown, the codec system 100 includes an encoding device 110 and a decoding device 120.

[0058] Among them, the encoding device 110 is used to encode (which can be understood as compressing) audio data to generate a bitstream, and transmit the bitstream to the decoding device 120. The decoding device 120 decodes the bitstream generated by the encoding device 110 to obtain the decoded audio data.

[0059] The encoding device 110 can be understood as a device with audio encoding function, and the decoding device 120 can be understood as a device with audio decoding function, that is, the embodiments of this application include a wider range of devices for the encoding device 110 and the decoding device 120, such as including smart phones, desktop computers, mobile computing devices, notebooks (e.g., laptops) computers, tablet computers, set-top boxes, TVs, cameras, display devices, digital media players, audio game consoles, in-vehicle computers, etc.

[0060] The encoding device 110 can transmit the encoded audio data (such as a bitstream) to the decoding device 120 via the channel 130.

[0061] The channel 130 can include one or more media and / or devices capable of transmitting the encoded audio data from the encoding device 110 to the decoding device 120.

[0062] Channel 130 may include one or more communication media that enable the encoding device 110 to transmit the encoded audio data directly to the decoding device 120 in real time. The encoding device 110 may modulate the encoded audio data according to a communication standard and transmit the modulated audio data to the decoding device 120. The communication media includes wireless communication media, such as radio frequency spectrum. The communication media may also include wired communication media, such as one or more physical transmission lines.

[0063] Channel 130 may include a storage medium that can store the audio data encoded by the encoding device 110. The storage medium includes various locally accessible data storage media, such as optical discs, DVDs, flash memories, etc. In this example, the decoding device 120 can obtain the encoded audio data from the storage medium.

[0064] Channel 130 may include a storage server that can store the audio data encoded by the encoding device 110. In this example, the decoding device 120 can download the stored encoded audio data from the storage server. Optionally, the storage server can store the encoded audio data and can transmit the encoded audio data to the decoding device 120, such as a web server (e.g., for a website), a File Transfer Protocol (FTP) server, etc.

[0065] The encoding device 110 includes an audio encoder 112 and an output interface 113.

[0066] Among them, the output interface 113 may include a modulator / demodulator (modem) and / or a transmitter. The audio encoder 112 directly transmits the encoded audio data to the decoding device 120 via the output interface 113. The encoded audio data can also be stored on a storage medium or a storage server for subsequent reading by the decoding device 120.

[0067] In addition to the audio encoder 112 and the input interface 113, the encoding device 110 may further include an audio source 111.

[0068] The audio source 111 may include at least one of an audio acquisition device, an audio archive, an audio input interface, and a computer graphics system. Among them, the audio input interface is used to receive audio data from an audio content provider, and the computer graphics system is used to generate audio data. The audio encoder 112 encodes the audio data from the audio source 111 to generate a bitstream. The audio data may include one or more audio signals, and the bitstream obtained by encoding the audio contains the encoded data of the audio in the form of a bitstream.

[0069] The decoding device 120 includes an input interface 121 and an audio decoder 122. The input interface 121 may include a receiver and / or a modem.

[0070] In addition to including an input interface 121 and an audio decoder 122, the decoding device 120 may further include a playback device 123.

[0071] Among them, the input interface 121 can receive the encoded audio data through the channel 130. The audio decoder 122 is used to decode the encoded audio data to obtain the decoded audio data, and transmit the decoded audio data to the playback device 123. The playback device 123 plays the decoded audio data. The playback device 123 can be integrated with the decoding device 120 or outside the decoding device 120. The playback device 123 can include various playback devices, such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or other types of playback devices.

[0072] It should be understood that Figure 1 This is only an example of the present application and should not be construed as a display of the present application. That is to say, the technical solutions of the embodiments of the present application are not limited to Figure 1 the system framework shown. For example, the technology of the present application can also be applied to unilateral audio encoding or unilateral audio decoding.

[0073] The audio encoding framework related to the embodiments of the present application will be introduced below.

[0074] Figure 2 is an example of the audio encoding framework provided by the embodiments of the present application.

[0075] As Figure 2 shown, the audio encoding and decoding framework may include an encoder, a decoder, a hyperencoder, a hyper decoder, a factorized entropy model, and an Entropy Parameters model. Among them, the encoder, the decoder, the hyperencoder, and the hyper decoder can all be implemented as neural networks.

[0076] In the encoding process, for the input audio signal x, first send x (here, x may be normalized, such as dividing the signal value by the loudness for normalization, and the loudness can refer to the overall intensity of the signal) into the first-layer neural network (i.e., the encoder) and output the first-layer feature y. Then, further input y into the second-layer neural network (i.e., the hyperencoder) and output the second-layer feature z. Then, first perform a quantization operation (quantify, Q) on z to obtain the quantization result Then, through the trained neural network-based model (i.e., the factorized entropy model) for perform probability modeling, and based on the probability distribution parameters obtained from the modeling for Perform entropy encoding and write it into the bitstream; meanwhile, Input a neural network for decoding (i.e., the hyperprior decoder) and output the transformed feature ψ; then send ψ into a neural network for prediction (i.e., the entropy parameter model) to predict (The probability distribution parameters of y after quantization) (assuming obeys a Gaussian distribution, then N(μ, θ) needs to be predicted, where μ represents the mean and θ represents the standard deviation); then according to the predicted μ and θ, Perform entropy encoding and write it into the bitstream. It should be understood that entropy encoding is a lossless compression technology, which encodes based on the principle of information entropy. During the entropy encoding process, the entropy encoder converts the input data into symbols and encodes them using the probability distribution of the symbols. Among them, entropy encoding includes Shannon encoding, Huffman encoding, exponential Golomb encoding, and arithmetic encoding (Arithmetic Encoding), etc. In this embodiment, an arithmetic encoder (Arithmetic Encoder, AE) can be used to perform and The arithmetic encoding performed.

[0077] During the decoding process, first decode the bitstream on the right side to obtain The information of, and then Input a neural network for decoding (i.e., the hyperprior decoder) and output the transformed feature ψ; then send ψ into a neural network for prediction (i.e., the entropy parameter model) to predict (The probability distribution parameters of y after quantization) (assuming obeys a Gaussian distribution, then N(μ, θ) needs to be predicted, where μ represents the mean and θ represents the standard deviation); then according to the predicted μ and θ, Perform entropy decoding (that is, perform entropy decoding on the bits in the bitstream to obtain and ). Finally, Input into the neural network (i.e., the decoder) and output the reconstructed signal (After that, inverse normalization processing can also be performed on ). Similar to the encoding process, in this embodiment, arithmetic decoding (Arithmetic Decoder, AD) can be used to perform arithmetic decoding on the bits in the bitstream to obtain and

[0078] For the above audio codec framework, during the decoding process, although ψ is used to predict The probability distribution parameters of, but the bitstream on the right side is additionally introduced, increasing the bitrate consumption and decoding complexity, and thus affecting the decoding efficiency and performance.

[0079] In view of this, embodiments of the present application provide a decoding method, an encoding method, and an electronic device based on a neural network. By introducing context information to determine decoding parameters for entropy decoding of elements to be decoded, it is possible to avoid introducing a code stream specifically for determining decoding parameters, reduce code rate consumption and decoding complexity, and thus improve decoding efficiency and performance. The decoding method, encoding method, decoder, and encoder provided by the present application can be applied to any end-to-end audio encoding and decoding framework or system based on deep learning.

[0080] The decoding method provided by the present application will be described below with reference to the accompanying drawings.

[0081] Figure 3 It is a schematic flowchart of a decoding method 200 provided by an embodiment of the present application. It should be understood that the decoding method 200 can be executed by any device or apparatus with decoding capabilities. For example, the decoding method 200 can be executed by Figure 1 the decoding device 120 or the audio decoder 122 shown in the figure. For another example, the decoding method 200 can be executed by Figure 2 the AD in the encoding and decoding framework shown in the figure. For ease of description, the decoder will be used as an example below for explanation.

[0082] As Figure 3 shown, the decoding method 200 may include:

[0083] S210, the decoder obtains a first element to be decoded based on the first code stream.

[0084] Exemplarily, the first element to be decoded may be a data unit extracted from the first code stream and ready for decoding.

[0085] Exemplarily, the first element to be decoded is any one of the following: an audio frame to be decoded, a channel to be decoded, a feature value to be decoded.

[0086] Exemplarily, the first element to be decoded is any one channel to be decoded or any one feature value to be decoded.

[0087] For example, if the format of the decoded feature obtained after entropy decoding of a to-be-decoded audio frame is C×H; correspondingly, the first to-be-decoded element can be a to-be-decoded feature channel, that is, the decoded data obtained after entropy decoding of the first to-be-decoded element can be all the feature values on a certain feature channel in C (for example, the decoder performs a for loop on C, and then restores the decoded feature in the format of C×H); or, the first to-be-decoded element can be a to-be-decoded feature value, that is, the decoded data obtained after entropy decoding of the first to-be-decoded element can be a certain feature value on a certain feature channel (for example, the decoder performs a for loop on H, and then restores the decoded feature in the format of C×H) and then obtains the binary data after entropy encoding.

[0088] In combination with the encoding process, for an audio frame, after the audio frame is encoded by an encoder, the feature y is obtained, and after the feature y is quantized, the feature feature has a feature format of C×H, where C represents the number of channels, and H represents the number of feature values on each feature channel (which can also be called the dimension of the feature on each channel). The feature includes the feature which can be the feature on the i-th feature channel or the i-th feature value; the entropy encoding of the feature obtains the feature corresponding binary data. In this embodiment, the feature obtained after entropy decoding of the first to-be-decoded element is

[0089] Exemplarily, the first to-be-decoded element is any one of the to-be-decoded audio frames, any one of the to-be-decoded channels, or any one of the to-be-decoded feature values among multiple to-be-decoded audio frames.

[0090] For example, if the format of the decoded features obtained after entropy decoding of N (N>1) audio frames to be decoded is N×C×H; correspondingly, the first element to be decoded can be the audio frame to be decoded, that is, the decoded data obtained after entropy decoding the first element to be decoded can be the features of a certain audio frame in N (for example, the decoder performs a for loop on N, and then restores the decoded features in the format of N×C×H); or, the first element to be decoded can be the feature channel to be decoded, that is, the decoded data obtained after entropy decoding the first element to be decoded can be all the feature values on a certain feature channel in C (for example, the decoder performs a for loop on C, and then restores the decoded features in the format of N×C×H); or, the first element to be decoded can be the feature value to be decoded, that is, the decoded data obtained after entropy decoding the first element to be decoded can be a certain feature value on a certain feature channel (for example, the decoder performs a for loop on H, and then restores the decoded features in the format of N×C×H) and then obtains the binary data after entropy encoding.

[0091] In combination with the encoding process, for N (N>1) audio frames, these N audio frames are encoded by the encoder to obtain the feature y, and after quantifying the feature y, the feature feature The feature format of is N×C×H, where N represents the number of audio frames, C represents the number of channels, and H represents the number of feature values on each feature channel (which can also be called the dimension of the features on each channel). The feature includes the feature can be the feature value of the i-th audio frame among the N audio frames, or the feature on the i-th feature channel (i.e., C i ), or the i-th feature value (i.e., H i ); performing entropy encoding on the feature to obtain the binary data corresponding to the feature In this embodiment, the feature obtained after entropy decoding the first element to be decoded is

[0092] Of course, in other alternative embodiments, the first element to be decoded can be some of the feature values to be decoded in the channel to be decoded, and the present application does not make specific limitations on this.

[0093] S220. The decoder obtains the context information of the first element to be decoded based on at least one decoded element that has been decoded before the first element to be decoded.

[0094] Exemplarily, the context information can provide reference information for the decoder to help the decoder understand the meaning of the first element to be decoded.

[0095] Exemplarily, the at least one decoded element may be some or all of the decoded elements.

[0096] For example, when the first element to be decoded is an audio frame to be decoded, the at least one decoded element may include some or all of the audio frames decoded before the audio frame to be decoded; for example, the decoder determines the context information of the currently to-be-decoded audio frame based on some or all of the decoded audio frames.

[0097] Exemplarily, when the first element to be decoded is a channel to be decoded, the at least one decoded element may include some or all of the decoded feature channels of the audio frame to be decoded to which the first element to be decoded belongs. Optionally, it may also include some or all of the feature channels of the decoded audio frame. For example, the decoder may determine the context information of the currently to-be-decoded feature channel based on some or all of the decoded feature channels of the audio frame to be decoded to which the first element to be decoded belongs.

[0098] Exemplarily, when the first element to be decoded is a feature value to be decoded, the at least one decoded element may include some or all of the decoded feature values of the audio frame to be decoded to which the first element to be decoded belongs. Optionally, it may also include some or all of the decoded feature values of the decoded audio frame. For example, the decoder may determine the context information of the currently to-be-decoded feature value based on some or all of the decoded feature values of the audio frame to be decoded to which the first element to be decoded belongs.

[0099] Exemplarily, the number of the at least one decoded element may be a predefined value, which may be agreed upon through a standard protocol or be information stored in the decoder.

[0100] Of course, in other alternative embodiments, the context information may also be determined based on the coding rule or algorithm of the first element to be decoded.

[0101] S230. The decoder predicts a first decoding parameter for entropy decoding the first element to be decoded by using a parameter prediction model based on a neural network based on the context information.

[0102] Exemplarily, the parameter prediction model based on a neural network may be any network model capable of performing probability modeling. Alternatively, the parameter prediction model based on a neural network may be any model capable of predicting the decoding parameters of entropy decoding.

[0103] In some embodiments, the first decoding parameter includes probability values corresponding to respective candidate values (i.e., possible values) of the first element to be decoded, that is, the probability values corresponding to the respective candidate values refer to the probability values when the respective candidate values are used as the decoded data of the first element to be decoded. In other words, the neural network-based parameter prediction model can be any model that can predict the probability values corresponding to the respective candidate values (i.e., possible values) of the first element to be decoded.

[0104] Exemplarily, when the first element to be decoded is an audio frame to be decoded, the respective candidate values of the first element to be decoded include the candidate values of each to-be-decoded feature value of the audio frame to be decoded, that is, the candidate values of each to-be-decoded feature value among C×H to-be-decoded feature values, where C represents the number of channels and H represents the number of feature values on each feature channel (which can also be referred to as the dimension of the feature on each channel).

[0105] Exemplarily, when the first element to be decoded is a feature channel to be decoded, the respective candidate values of the first element to be decoded include the candidate values of each to-be-decoded feature value on the feature channel, that is, the candidate values of each to-be-decoded feature value among H to-be-decoded feature values, where H represents the number of feature values on each feature channel (which can also be referred to as the dimension of the feature on each channel).

[0106] Exemplarily, when the first element to be decoded is a to-be-decoded feature value, the respective candidate values of the first element to be decoded include the candidate values of this to-be-decoded feature value.

[0107] It should be understood that for any to-be-decoded feature value, the specific value of its candidate value is related to the position of the any to-be-decoded feature value. For example, the specific value of the candidate value of the any to-be-decoded feature value can be determined by querying a mapping relation table formed by positions and candidate values.

[0108] In some embodiments, the first decoding parameter includes probability distribution parameters, such as the mean and standard deviation of a Gaussian distribution. In other words, the neural network-based parameter prediction model can be any network model that can perform probability modeling and predict probability distribution parameters. Among them, probability modeling can refer to determining the probability distribution of the first element to be decoded, and predicting probability distribution parameters means predicting the parameters of the probability distribution of the first element to be decoded, that is, probability distribution parameters.

[0109] Exemplarily, the Gaussian distribution is also known as the Normal distribution. Its curve presents a bell shape, low at both ends, high in the middle, and symmetric on both sides; because its curve is bell-shaped, it is also called a bell-shaped curve. The mean of the Gaussian distribution determines the central position of the distribution. The standard deviation of the Gaussian distribution determines the degree of dispersion of the distribution.

[0110] Certainly, in other alternative embodiments, the first decoding parameter may also be a parameter of the Laplace distribution or a parameter of other probability distributions, and the present application does not make specific limitations thereto.

[0111] S240. The decoder performs entropy decoding on the first element to be decoded based on the first decoding parameter to obtain the decoded data of the first element to be decoded.

[0112] In this embodiment, when the decoder predicts the first decoding parameter of the first element to be decoded, by considering the context information of the first element to be decoded, the accuracy of the first decoding parameter can be improved, thereby reducing the code rate consumption and decoding complexity, and further improving the decoding efficiency and performance.

[0113] In some embodiments, S220 may include:

[0114] The decoder uses a context model to obtain the context information.

[0115] Among them, the context model may be a masked CNN based on autoregressive. Among them, the autoregressive model is characterized in that the currently predicted element depends on all the elements predicted before it. Masked-CNN is a variant of the convolutional neural network (CNN). The standard CNN performs a convolution operation on the input and then passes it to the fully connected layer for classification or regression. While masked-CNN, after performing a convolution operation on the input, sets the weights of some regions to 0 (these regions are called masks or masks) to prevent the features of these regions from being used for subsequent predictions.

[0116] In some embodiments, S220 may include:

[0117] Use a convolutional layer with a convolution kernel of 1×k to perform convolution on K elements to obtain the context information; where the K elements include the at least one decoded element, and 1≤k≤K.

[0118] Exemplarily, the K elements only include the at least one decoded element.

[0119] Exemplarily, when the number of decoded elements before the first element to be decoded is greater than or equal to K, the decoder may also select the K elements from the decoded elements before the first element to be decoded.

[0120] Exemplarily, the K elements include the at least one decoded element and undecoded elements.

[0121] Exemplarily, the undecoded element includes at least one of the following: the first element to be decoded, an element located after the first element to be decoded.

[0122] Exemplarily, when the number of decoded elements before the first element to be decoded is less than K, the decoder can also select the K elements from the decoded elements and undecoded elements before the first element to be decoded.

[0123] Exemplarily, the K can be a fixed value, which can be agreed upon by protocol or can be information stored in the decoder.

[0124] Exemplarily, the K increases as the position of the first element to be decoded moves. For example, when the first element to be decoded is the i-th element, the K elements include the first i - 1 elements.

[0125] Exemplarily, k is odd.

[0126] In some embodiments, when the K elements include undecoded elements, the weight corresponding to the undecoded element is configured to 0, or the value of the undecoded element is configured to 0.

[0127] Exemplarily, the weight corresponding to the undecoded element can be understood as the weight used by the undecoded element. When performing convolution on the K elements using a convolutional layer with a convolution kernel of 1×k, each weight in the convolution kernel will perform a multiplication operation with the element at the corresponding position in the K elements, and then the results of these multiplications are added together to obtain an output feature. This output feature will reflect the sensitivity of the convolution kernel to a certain feature of the K elements.

[0128] Exemplarily, the decoder can perform one or more convolutions on the K elements using a convolutional layer with a convolution kernel of 1×k to obtain the context information. For example, when the difference obtained by subtracting k from K is greater than or equal to k, the decoder can perform multiple convolutions on the K elements using a convolutional layer with a convolution kernel of 1×k to obtain the context information.

[0129] In this embodiment, when the K elements include undecoded elements, the weight corresponding to the undecoded element is configured to 0, or the value of the undecoded element is configured to 0; equivalently, only the decoded data of the decoded elements among the K elements is concerned, while the undecoded elements are ignored, which improves the reference value of the context information, and further improves the accuracy of the first decoding parameter, and can reduce the code rate consumption and decoding complexity, and further improve the decoding efficiency and performance.

[0130] In some embodiments, the first element to be decoded is the element located in the middle position among the K elements.

[0131] Exemplarily, if K is odd, the first element to be decoded is the element in the middle position among the K elements.

[0132] Exemplarily, if K is odd and the number of decoded elements before the first element to be decoded is greater than or equal to (K - 1) / 2, it is the number of decoded elements before the first element to be decoded. For example, assume K is 5, then if the number of decoded elements before the first element to be decoded is greater than or equal to 2, the first element to be decoded is the element in the middle position among the 5 elements.

[0133] Of course, even if K is odd, if the number of decoded elements before the first element to be decoded is less than (K - 1) / 2, it means that the decoded elements before the first element to be decoded are not sufficient to support setting the first element to be decoded in the middle position of the K elements. At this time, the elements in the first K positions can be defaulted as the K elements, or other methods can be used to determine the context information. For example, a smaller-sized convolutional kernel determines the context information, or the context information agreed upon by the protocol is used, or the decoding parameters agreed upon by the protocol are directly used as the first decoding parameter. The embodiments of the present application do not make specific limitations on this.

[0134] Figure 4 It is an example of the element for determining context information provided by the embodiments of the present application.

[0135] As Figure 4 shown, for n one-dimensional elements, that is, x 1 ~x n , the first element to be decoded is denoted as x i . The decoder can use a convolutional layer with a convolutional kernel of 1×k to perform convolution on the K elements to obtain the context information. For example, when K = 3, these 3 elements can include the elements shown in the gray area in the figure.

[0136] Of course, in other alternative embodiments, the decoder can use a convolutional layer with a convolutional kernel of s×k to perform convolution on S×K elements to obtain the context information; where the S×K elements include the at least one decoded element, 1 ≤ s ≤ S, 1 ≤ k ≤ K. Optionally, when the S×K elements include undecoded elements, the weight configuration corresponding to the undecoded elements is 0, or the value configuration of the undecoded elements is 0. Optionally, both s and k are odd.

[0137] Figure 5 It is another example of the element for determining context information provided by the embodiments of the present application.

[0138] As Figure 5 shown, for n one-dimensional 2 elements, that is the decoder can perform convolution on the one-dimensional n2 The elements are reorganized to obtain the two-dimensional n 2 elements, that is, x 1×1 ~x n×n The decoder can then use a convolutional layer with a kernel of s×k to 1×1 ~x n×n The context information is obtained by convolving the S×K elements in . For example, when S=K=3, the 9 elements may include the elements shown in the gray area in the figure.

[0139] In some embodiments, S220 may include:

[0140] The decoder determines a decoding parameter used for entropy decoding the at least one decoded element as the context information.

[0141] Exemplarily, the decoding parameter used for entropy decoding the at least one decoded element may be a decoding parameter used by the at least one decoded element.

[0142] Exemplarily, when any one of the at least one decoded elements is a decoded audio frame, the candidate values ​​of the any one decoded element include the candidate value of each decoded feature value of the decoded audio frame, that is, the candidate value of each decoded feature value among C×H decoded feature values, where C represents the number of channels and H represents the number of feature values ​​on each feature channel (also referred to as the dimension of the feature on each channel).

[0143] Exemplarily, when any one of the at least one decoded elements is a decoded feature channel, the candidate values ​​of the any one decoded element include the candidate value of each decoded feature value on the feature channel, that is, the candidate value of each decoded feature value among the H decoded feature values, where H represents the number of feature values ​​on each feature channel (also referred to as the dimension of the feature on each channel).

[0144] Exemplarily, when any one of the at least one decoded element is a decoded feature value, each candidate value of the any one decoded element includes a candidate value of the decoded feature value.

[0145] It should be understood that for any decoded feature value, the specific value of its candidate value is related to the location of the decoded feature value. For example, the specific value of the candidate value of any decoded feature value can be determined by querying a mapping relationship table formed by the location and the candidate value.

[0146] Exemplarily, for any one of the at least one decoded element, the decoding parameters used may include probability distribution parameters, or may include probability values corresponding to respective candidate values (i.e., possible values) of the decoded element, or may include cumulative probability values corresponding to respective candidate values of the decoded element. The probability value corresponding to each candidate value refers to the probability value when taking the respective candidate value as the decoded data. The cumulative probability value corresponding to each candidate value refers to the cumulative probability value when taking the respective candidate value as the decoded data calculated according to a preset arrangement order of each candidate value. For example, assume that any one of the decoded elements includes a first decoded feature value, and its candidate values include a0, a1, a2, a3, a4. If the probability of taking a0 as the decoded data of the first decoded feature value is b0, the probability of taking a1 as the decoded data of the first decoded feature value is b1, the probability of taking a2 as the decoded data of the first decoded feature value is b2, the probability of taking a3 as the decoded data of the first decoded feature value is b3, and the probability of taking a4 as the decoded data of the first decoded feature value is b4; then: the probability values corresponding to each candidate value include: the probability value corresponding to a0 (i.e., b0), the probability value corresponding to a1 (i.e., b1), the probability value corresponding to a2 (i.e., b2), the probability value corresponding to a3 (i.e., b3), the probability value corresponding to a4 (i.e., b4); according to the arrangement order of a0, a1, a2, a3, a4, the cumulative probability values corresponding to each candidate value include: the cumulative probability value corresponding to a0, the cumulative probability value corresponding to a1, the cumulative probability value corresponding to a2, the cumulative probability value corresponding to a3, the cumulative probability value corresponding to a4; among them, the cumulative probability value corresponding to a1 is equal to the probability value corresponding to a0 (i.e., b0), the cumulative probability value corresponding to a2 is equal to the sum of the probability value corresponding to a0 and the probability value corresponding to a1 (i.e., b0 + b1), and so on, the cumulative probability value corresponding to a4 is equal to the sum of the probability value corresponding to a0, the probability value corresponding to a1, the probability value corresponding to a2, the probability value corresponding to a3, and the probability value corresponding to a4 (i.e., b0 + b1 + b2 + b3 + b4).

[0147] Exemplarily, the decoding parameters for performing entropy decoding on the at least one decoded element may be implemented as a parameter table with a size of M×L.

[0148] Wherein, M represents the number of decoded eigenvalue in the at least one decoded element. For example, when the decoded element is a decoded audio frame, the number of its eigenvalues is C×H; L represents the number of decoding parameters used by the decoded eigenvalue or a value determined according to the number of decoding parameters used by the decoded eigenvalue. For example, L represents the number of probability distribution parameters used by the decoded eigenvalue. For another example, L represents the number of probability values corresponding to each candidate value of the decoded eigenvalue. For another example, L represents the number of cumulative probability values corresponding to each candidate value of the decoded eigenvalue. For another example, L represents the number of dividing points of a plurality of probability intervals when a preset interval is divided into a plurality of probability intervals based on the cumulative probability values corresponding to each candidate value of the decoded eigenvalue. For example, when the preset interval is [0,1], the first dividing point among the dividing points of the plurality of probability intervals is 0, and the last dividing point of the dividing points of the plurality of probability intervals is 1. For example, when the preset interval is [0,1], if the first dividing point among the dividing points of the plurality of probability intervals is 0 and the cumulative probability value corresponding to the last candidate value of the decoded eigenvalue is 1, the number of dividing points of the plurality of probability intervals is equal to the number of cumulative probability values corresponding to each candidate value of the decoded eigenvalue plus 1.

[0149] In some embodiments, the S220 may include:

[0150] Determine the context information based on at least one of the following:

[0151] Decoded elements adjacent to the position of the first element to be decoded, decoded elements whose distance from the first element to be decoded is less than or equal to a preset threshold, and decoded elements at a preset position.

[0152] Exemplarily, the decoded element adjacent to the position of the first element to be decoded may include any one of the following:

[0153] The decoded element adjacent to the position of the first element to be decoded in the decoding order of the elements; the decoded element adjacent to the position of the first element to be decoded after element recombination. For example, as Figure 5 shown, the decoded element adjacent to the position of the first element to be decoded after element recombination may include at least one of the following: decoded elements adjacent in the horizontal direction, decoded elements adjacent in the vertical direction, and decoded elements adjacent in the diagonal direction, etc.

[0154] Exemplarily, the distance from the first element to be decoded may be the number of elements spaced from the first element to be decoded.

[0155] Exemplarily, the decoded element whose distance from the first element to be decoded is less than or equal to a preset threshold may include any one of the following:

[0156] Decoded elements whose distance from the first element to be decoded is less than or equal to a preset threshold according to the decoding order of the elements; decoded elements whose distance from the first element to be decoded is less than or equal to a preset threshold after element recombination. For example, as Figure 5 shown, the decoded elements whose distance from the first element to be decoded is less than or equal to a preset threshold after element recombination may include at least one of the following: decoded elements whose distance from the first element to be decoded in the horizontal direction is less than or equal to the preset threshold, decoded elements whose distance from the first element to be decoded in the vertical direction is less than or equal to the preset threshold, and decoded elements whose distance from the first element to be decoded in the diagonal direction is less than or equal to the preset threshold, etc.

[0157] Exemplarily, the preset threshold can be agreed upon through a standard protocol or can be information stored on the decoder.

[0158] Exemplarily, the decoded elements at the preset position may include any one of the following:

[0159] Decoded elements at the preset position according to the decoding order of the elements; decoded elements at the preset position after element recombination. For example, the decoded elements at the preset position according to the decoding order of the elements include but are not limited to at least one of the following: the first position, one or more positions at the front, the middle position, one or more positions at the back. For example, as Figure 5 shown, the decoded elements at the preset position after element recombination may include at least one of the following: decoded elements at the preset position in the vertical direction, decoded elements at the preset position in the horizontal direction. For example, the decoded elements at the preset position after element recombination may include at least one of the following: decoded elements at the central position, decoded elements at the upper right corner, decoded elements at the lower right corner, decoded elements at the upper left corner, decoded elements at the lower left corner, etc.

[0160] Exemplarily, the preset position can be agreed upon through a standard protocol or can be information stored in the decoder.

[0161] Of course, in other alternative embodiments, the preset position can also be a position determined based on the position of the first element to be decoded using a preset rule, and the preset rule can be agreed upon through a standard protocol or can be information stored in the decoder.

[0162] In some embodiments, the S230 may include:

[0163] If the first element to be decoded is not the first element, then use the parameter prediction model to predict the first decoding parameter.

[0164] Exemplarily, if the first element to be decoded is not the first element, the decoder predicts the first decoding parameter based on the context information using the parameter prediction model; otherwise, the decoder performs entropy decoding on the first element to be decoded in other ways.

[0165] In some embodiments, the method 200 may further include:

[0166] If the first element to be decoded is the first element, entropy decoding is performed on the first element to be decoded based on default decoding parameters to obtain the decoded data of the first element to be decoded.

[0167] Exemplarily, if the first element to be decoded is not the first element, the decoder uses the parameter prediction model to predict the first decoding parameter; otherwise, the decoder performs entropy decoding on the first element to be decoded based on default decoding parameters to obtain the decoded data of the first element to be decoded. The default decoding parameters can be agreed upon through a standard protocol or can be information stored in the decoder. The default decoding parameters may include probability distribution parameters, or may include the probability values of each candidate value (i.e., possible value) of the first element to be decoded.

[0168] In some embodiments, S230 may include:

[0169] Using the context information as input, perform probability modeling on the first element to be decoded to obtain the first decoding parameter.

[0170] Exemplarily, using the context information as input, any probability distribution model that does not require prior information can be used to perform probability modeling on the first element to be decoded to obtain the first decoding parameter. For example, using the context information as input, a decomposed entropy model can be used to perform probability modeling on the first element to be decoded to obtain the first decoding parameter.

[0171] Exemplarily, if the first element to be decoded is not the first element, the decoder uses the context information as input to perform probability modeling on the first element to be decoded to obtain the first decoding parameter; otherwise, the decoder performs entropy decoding on the first element to be decoded based on default decoding parameters to obtain the decoded data of the first element to be decoded. The default decoding parameters can be agreed upon through a standard protocol or can be information stored in the decoder. The default decoding parameters may include probability distribution parameters, or may include the probability values of each candidate value (i.e., possible value) of the first element to be decoded.

[0172] Figure 6 This is an example of the encoding and decoding framework provided by the embodiments of the present application.

[0173] Such as Figure 6As shown, the audio codec framework may include an encoder, a decoder, a prediction model, and a context model. At least one of the encoder, the decoder, the context model, and the prediction model may be implemented as a neural network.

[0174] During the encoding process, for the input audio signal x, first send x (here x may be normalized, for example, divide the signal value by the loudness for normalization, and the loudness may refer to the overall intensity of the signal) into the encoder and output the feature y. Then, perform a quantization operation (quantify, Q) on y to obtain the quantization result At the same time, introduce context information. Specifically, assume when encoding the element at the i-th position (i≥0), it is necessary to use the decoded elements at the adjacent positions of the i-th position (for boundary elements, the adjacent position elements may not exist. If the adjacent position elements do not exist, ignore the calculation regarding the context) (such as ) as the input of the context model to obtain the corresponding feature φ i , and then send φ i into a neural network for prediction (i.e., the prediction model) to predict the probability distribution parameters (assuming follows a Gaussian distribution, then it is necessary to predict N(μ,θ), where μ represents the mean and θ represents the standard deviation); then, perform entropy coding on according to the predicted μ and θ and write it into the code stream. In this embodiment, arithmetic coding performed by AE on can be used.

[0175] During the decoding process, when decoding the element at the i-th position (i≥0), it is necessary to use the decoded elements at the adjacent positions of the i-th position (for boundary elements, the adjacent position elements may not exist. If the adjacent position elements do not exist, ignore the calculation regarding the context) (such as ) as the input of the context model to obtain the corresponding feature φ i , and then send φ i into a neural network for prediction (i.e., the prediction model) to predict the probability distribution parameters (assuming follows a Gaussian distribution, then it is necessary to predict N(μ,θ), where μ represents the mean and θ represents the standard deviation); then, perform entropy decoding on (that is, perform entropy decoding on the bits in the code stream and obtain and ). After repeating the above decoding operations, all The reconstructed value, and finally is input into a neural network (i.e., a decoder) and outputs a reconstructed signal (After that, inverse normalization processing can also be performed on . Similar to the encoding process, in this embodiment, AD can be used to perform arithmetic decoding on the bits in the bitstream and obtain

[0176] In some embodiments, the S230 may include:

[0177] Based on the second bitstream, obtain a priori features for determining the first decoding parameter; the second bitstream is a bitstream obtained after hyperprior encoding; using the context information and the a priori features as inputs, perform probability modeling on the first element to be decoded to obtain the first decoding parameter.

[0178] Exemplarily, if the first element to be decoded is not the first element, the decoder uses the context information and the a priori features as inputs, performs probability modeling on the first element to be decoded to obtain the first decoding parameter; otherwise, the decoder performs entropy decoding on the first element to be decoded based on default decoding parameters to obtain the decoded data of the first element to be decoded. The default decoding parameters can be agreed upon through a standard protocol or can be information stored in the decoder. The default decoding parameters may include probability distribution parameters, or may include probability values of each candidate value (i.e., possible value) of the first element to be decoded.

[0179] Exemplarily, using the context information and the a priori features as inputs, any probability distribution model that requires a priori information can be used to perform probability modeling on the first element to be decoded to obtain the first decoding parameter. For example, using the context information and the a priori features as inputs, an entropy parameter model can be used to perform probability modeling on the first element to be decoded to obtain the first decoding parameter.

[0180] Exemplarily, the first bitstream is a bitstream obtained by quantizing and entropy encoding the encoded features output by the encoder, and the second bitstream is a bitstream obtained by quantizing and entropy encoding the features obtained after hyperprior encoding the encoded features output by the encoder.

[0181] In some embodiments, obtaining a priori features for determining the first decoding parameter based on the second bitstream can be implemented as:

[0182] Based on the second bitstream, obtain features to be decoded; perform entropy decoding on the features to be decoded to obtain decoded features of the features to be decoded; perform feature processing on the decoded features to obtain a priori features.

[0183] Exemplarily, the decoder may extract the feature to be decoded from the second bitstream, and then perform entropy decoding on the feature to be decoded by using the factorized entropy model to obtain the decoded feature of the feature to be decoded; then, the hyper prior decoder is used to process the decoded feature to obtain the prior feature.

[0184] Figure 7 is an example of the audio coding framework provided by the embodiments of the present application.

[0185] Such as Figure 7 shown, the audio coding and decoding framework may include an encoder, a decoder, a hyper encoder, a hyper decoder, a factorized entropy model, a prediction model, and a context model. Among them, at least one of the encoder, decoder, hyper encoder, hyper decoder, factorized entropy model, prediction model, and context model may be implemented as a neural network.

[0186] During the encoding process, for the input audio signal x, first send x (here, x may be normalized, for example, the signal value is divided by the loudness for normalization, and the loudness may refer to the overall intensity of the signal) into the first-layer neural network (i.e., the encoder) and output the feature y of the first layer, and then further input y into the second-layer neural network (i.e., the hyper encoder) and output the feature z of the second layer. Then, a quantization operation (quantify, Q) is performed on z to obtain the quantization result After that, through the trained neural network-based model (i.e., the factorized entropy model) for probability modeling is performed, and according to the probability distribution parameters obtained by the modeling, entropy coding is performed and written into the bitstream. At the same time, is input into a neural network for decoding (i.e., the hyper decoder) and the transformed feature ψ is output. At the same time, the context information is introduced. Specifically, assume that when encoding and decoding the element at the i-th position (i≥0), in addition to the existing feature ψ, the decoded elements at the adjacent positions of the i-th position (for boundary elements, the adjacent position elements may not exist. If the adjacent position elements do not exist, the calculation regarding the context is ignored) (such as ) are used as the input of the context model to obtain the corresponding feature φ i , and then ψ and φ i are input into a neural network for prediction (i.e., the prediction model) to predict the probability distribution parameters (assuming If it follows a Gaussian distribution, then N(μ, θ) needs to be predicted, where μ represents the mean and θ represents the standard deviation); then, based on the predicted μ and θ, perform entropy coding and write it into the bitstream. In this embodiment, arithmetic coding performed on and can be adopted by AE.

[0187] During the decoding process, first decode the bitstream on the right side to obtain the information of , and then input into a neural network for decoding (i.e., the hyperprior decoder) and output the transformed feature ψ; when encoding and decoding the element at the i-th position (i≥0), in addition to the existing feature ψ, the decoded elements at the adjacent positions of the i-th position (for boundary elements, the adjacent position elements may not exist. If the adjacent position elements do not exist, ignore the calculation regarding the context), such as , are used as the input of the context model to obtain the corresponding feature φ i , and then ψ and φ i are sent into a neural network for prediction (i.e., the prediction model) to predict the probability distribution parameters of (assuming follows a Gaussian distribution, then N(μ, θ) needs to be predicted, where μ represents the mean and θ represents the standard deviation); then, based on the predicted μ and θ, perform entropy decoding (i.e., perform entropy decoding on the bits in the bitstream to obtain and ). After repeating the above decoding operations, the reconstructed values of all are obtained. Finally, is input into a neural network (i.e., the decoder) and the reconstructed signal is output (after that, inverse normalization processing can also be performed on ). Similar to the encoding process, in this embodiment, arithmetic decoding on the bits in the bitstream can be adopted by AD to obtain and

[0188] It should be understood that Figure 6 and Figure 7 are only examples of this application and should not be construed as limitations on this application. For example, in Figure 6 and Figure 7 , when the prediction model predicts the probability distribution parameters of , it uses Taking the case of following a Gaussian distribution (i.e., it is necessary to predict N(μ, θ)) as an example, it is equivalent to that the prediction model is a Gaussian model, but the embodiments of the present application are not limited thereto. For example, in other alternative embodiments, the prediction model can also be directly used to predict probability values, that is, the prediction model can also be replaced by a decomposition entropy model or other models for predicting probability values or parameters of probability distributions.

[0189] The following will be combined with Figure 8 Describe the encoding method according to the embodiments of the present application from the perspective of the encoder.

[0190] Figure 8 It is a schematic flowchart of the encoding method 300 provided by the present application. It should be understood that the encoding method 300 can be executed by any device or apparatus with encoding capabilities. For example, the encoding method 300 can be executed by Figure 1 The encoding device 110 or the audio encoder 112 shown. Again, for example, the encoding method 300 can be executed by Figure 2 The AE shown. As Figure 8 shown, the encoding method 300 may include:

[0191] S310, obtain the first element to be encoded;

[0192] S320, based on at least one encoded element encoded before the first element to be encoded, obtain the context information of the first element to be encoded;

[0193] S330, based on the context information, use a parameter prediction model based on a neural network to predict a first encoding parameter for entropy encoding the first element to be encoded;

[0194] S340, perform entropy encoding on the first element to be encoded based on the first encoding parameter to obtain the encoded data of the first element to be encoded.

[0195] In some embodiments, the S320 may include:

[0196] Use a convolutional layer with a convolution kernel of 1×k to perform convolution on K elements to obtain the context information; where the K elements include the at least one encoded element, and 1≤k≤K.

[0197] In some embodiments, when the K elements include unencoded elements, the weight configuration corresponding to the unencoded elements is 0, or the value configuration of the unencoded elements is 0; or

[0198] When the K elements include unencoded elements, the first element to be encoded is the element located in the middle position among the K elements.

[0199] In some embodiments, the S320 may include:

[0200] Determine the encoding parameter for entropy encoding the at least one encoded element as the context information.

[0201] In some embodiments, the at least one encoded element includes at least one of the following:

[0202] An encoded element adjacent to the position of the first element to be encoded, an encoded element whose distance from the first element to be encoded is less than or equal to a preset threshold, and an encoded element at a preset position.

[0203] In some embodiments, the S330 may include:

[0204] If the first element to be encoded is not the first element, use the parameter prediction model to predict the first encoding parameter;

[0205] The method 300 may further include:

[0206] If the first element to be encoded is the first element, perform entropy encoding on the first element to be encoded based on a default encoding parameter to obtain the encoded data of the first element to be encoded.

[0207] In some embodiments, the S330 may include:

[0208] Use the context information as an input to perform probability modeling on the first element to be encoded to obtain the first encoding parameter.

[0209] In some embodiments, the S330 may include:

[0210] Obtain a feature to be encoded;

[0211] Perform entropy encoding on the feature to be encoded to obtain an encoded feature of the feature to be encoded;

[0212] Perform entropy decoding and feature processing on the encoded feature to obtain a prior feature; and

[0213] Use the context information and the prior feature as inputs, and use the parameter prediction model to perform probability modeling on the first element to be encoded to obtain the first encoding parameter.

[0214] In some embodiments, the first element to be encoded is any one of the following: an audio frame to be encoded, a channel to be encoded, and a feature value to be encoded.

[0215] In some embodiments, the first encoding parameter includes a probability distribution parameter, or the first encoding parameter includes probability values corresponding to respective candidate values of the first element to be encoded.

[0216] It should be understood that the encoding method can be regarded as the inverse process of the decoding method. Therefore, for the specific solution of the encoding method 300, reference can be made to the relevant content of the decoding method 200. For the sake of simplicity of description, this application will not elaborate on it further.

[0217] The preferred embodiments of the present application have been described in detail above in conjunction with the accompanying drawings. However, the present application is not limited to the specific details in the above-mentioned embodiments. Within the scope of the technical concept of the present application, various simple modifications can be made to the technical solutions of the present application, and these simple modifications all fall within the protection scope of the present application. For example, the various specific technical features described in the above-mentioned specific embodiments can be combined in any suitable manner without contradiction. To avoid unnecessary repetition, this application will not separately describe various possible combination methods. For another example, any combination can be made between various different embodiments of the present application, as long as it does not violate the idea of the present application, it should also be regarded as the content disclosed in the present application.

[0218] It should also be understood that in various method embodiments of the present application, the magnitudes of the sequence numbers of the above-mentioned processes do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0219] The following Figures 9 to 11 describes in detail the device embodiments of the present application.

[0220] Figure 9 is a schematic block diagram of a decoder 400 provided by the present application.

[0221] As Figure 9 shown, the decoder 400 may include:

[0222] A first acquisition unit 410, configured to acquire a first element to be decoded based on a first bitstream;

[0223] A second acquisition unit 420, configured to acquire context information of the first element to be decoded based on at least one decoded element that has been decoded before the first element to be decoded;

[0224] A determination unit 430, configured to predict a first decoding parameter for entropy decoding the first element to be decoded by using the parameter prediction model;

[0225] An entropy decoding unit 440, configured to perform entropy decoding on the first element to be decoded based on the first decoding parameter to obtain decoded data of the first element to be decoded.

[0226] In some embodiments, the second acquisition unit 420 is specifically configured to:

[0227] Use a convolutional layer with a convolutional kernel of 1×k to convolve K elements to obtain the context information; where the K elements include the at least one decoded element, and 1 ≤ k ≤ K.

[0228] In some embodiments, when the K elements include undecoded elements, the weight configuration corresponding to the undecoded elements is 0, or the value configuration of the undecoded elements is 0; or

[0229] When the K elements include undecoded elements, the first element to be decoded is the element located in the middle position among the K elements.

[0230] In some embodiments, the second obtaining unit 420 is specifically configured to:

[0231] Determine the decoding parameter used for entropy decoding of the at least one decoded element as the context information.

[0232] In some embodiments, the at least one decoded element includes at least one of the following:

[0233] The decoded element adjacent to the position of the first element to be decoded, the decoded element whose distance from the first element to be decoded is less than or equal to a preset threshold, and the decoded element at a preset position.

[0234] In some embodiments, the determining unit 430 is specifically configured to:

[0235] If the first element to be decoded is not the first element, use the parameter prediction model to predict the first decoding parameter;

[0236] The determining unit 430 is specifically further configured to:

[0237] If the first element to be decoded is the first element, perform entropy decoding on the first element to be decoded based on the default decoding parameter to obtain the decoded data of the first element to be decoded.

[0238] In some embodiments, the determining unit 430 is specifically configured to:

[0239] Use the context information as an input to perform probability modeling on the first element to be decoded to obtain the first decoding parameter.

[0240] In some embodiments, the determining unit 430 is specifically configured to:

[0241] Obtain the prior feature for determining the first decoding parameter based on the second bitstream; the second bitstream is the bitstream obtained through hyperprior encoding;

[0242] Use the context information and the prior feature as inputs to perform probability modeling on the first element to be decoded to obtain the first decoding parameter.

[0243] In some embodiments, the determining unit 430 is specifically configured to:

[0244] Obtain a to-be-decoded feature based on the second bitstream;

[0245] Perform entropy decoding on the to-be-decoded feature to obtain a decoded feature of the to-be-decoded feature;

[0246] Perform feature processing on the decoded feature to obtain the prior feature.

[0247] In some embodiments, the first to-be-decoded element is any one of the following: a to-be-decoded audio frame, a to-be-decoded channel, and a to-be-decoded feature value.

[0248] In some embodiments, the first decoding parameter includes a probability distribution parameter, or the first decoding parameter includes probability values corresponding to respective candidate values of the first to-be-decoded element.

[0249] It should be understood that the apparatus embodiments of the decoder and the method embodiments of the decoding method can correspond to each other, and similar descriptions can refer to the method embodiments. Specifically, the decoder 400 can correspond to the corresponding main body in the decoding method 200 of the embodiments of the present application, and the foregoing and other operations and / or functions of each unit in the decoder 400 respectively implement the corresponding processes in the decoding method 200. To avoid repetition, details are not described herein again.

[0250] Figure 10 is a schematic block diagram of an encoder 500 provided by the present application.

[0251] As Figure 10 shown, the encoder 500 may include:

[0252] A first obtaining unit 510, configured to obtain a first to-be-encoded element;

[0253] A second obtaining unit 520, configured to obtain context information of the first to-be-encoded element based on at least one encoded element encoded before the first to-be-encoded element;

[0254] A determining unit 530, configured to predict a first encoding parameter for performing entropy encoding on the first to-be-encoded element based on the context information by using a parameter prediction model based on a neural network;

[0255] An entropy encoding unit 540, configured to perform entropy encoding on the first to-be-encoded element based on the first encoding parameter to obtain encoded data of the first to-be-encoded element.

[0256] In some embodiments, the second obtaining unit 520 is specifically configured to:

[0257] In some embodiments, the S320 may include:

[0258] Use a convolutional layer with a convolution kernel of 1×k to convolve K elements to obtain the context information; wherein, the K elements include the at least one encoded element, and 1 ≤ k ≤ K.

[0259] In some embodiments, when the K elements include unencoded elements, the weight configuration corresponding to the unencoded elements is 0, or the value configuration of the unencoded elements is 0; or

[0260] When the K elements include unencoded elements, the first element to be encoded is the element located in the middle position among the K elements.

[0261] In some embodiments, the S320 may include:

[0262] Determine the encoding parameter for entropy encoding the at least one encoded element as the context information.

[0263] In some embodiments, the at least one encoded element includes at least one of the following:

[0264] An encoded element adjacent to the position of the first element to be encoded, an encoded element whose distance from the first element to be encoded is less than or equal to a preset threshold, and an encoded element at a preset position.

[0265] In some embodiments, the S330 may include:

[0266] If the first element to be encoded is not the first element, use the parameter prediction model to predict the first encoding parameter;

[0267] The method 300 may further include:

[0268] If the first element to be encoded is the first element, perform entropy encoding on the first element to be encoded based on a default encoding parameter to obtain the encoded data of the first element to be encoded.

[0269] In some embodiments, the S330 may include:

[0270] Use the context information as an input to perform probability modeling on the first element to be encoded to obtain the first encoding parameter.

[0271] In some embodiments, the S330 may include:

[0272] Obtain the feature to be encoded;

[0273] Perform entropy encoding on the feature to be encoded to obtain the encoded feature of the feature to be encoded;

[0274] Perform entropy decoding and feature processing on the encoded feature to obtain a prior feature; and

[0275] Taking the context information and the prior feature as inputs, a probability model of the first element to be encoded is established by using the parameter prediction model to obtain the first encoding parameter.

[0276] In some embodiments, the first element to be encoded is any one of the following: an audio frame to be encoded, a channel to be encoded, a feature value to be encoded.

[0277] In some embodiments, the first encoding parameter includes a probability distribution parameter, or the first encoding parameter includes probability values corresponding to respective candidate values of the first element to be encoded.

[0278] It should be understood that the apparatus embodiments of the encoder and the method embodiments of the encoding method can correspond to each other, and similar descriptions can refer to the method embodiments. Specifically, the encoder 500 can correspond to the corresponding main body in the encoding method 300 of the embodiments of the present application, and the foregoing and other operations and / or functions of each unit in the encoder 500 respectively implement the corresponding processes in each method such as the encoding method 300. To avoid repetition, details are not described herein again.

[0279] It should also be understood that the terms "module" or "unit" involved in the embodiments of the present application are divided based on logical functions. In practical applications, the function of a module or unit can also be implemented by multiple modules or units, or the functions of multiple modules or units are implemented by one module or unit. Even, these functions can also be assisted by one or more other modules or units to be implemented. For example, part or all are combined into one or several other modules or units. Again, a certain module or unit can be further split into multiple smaller modules or units in terms of function, which can achieve the same operation without affecting the implementation of the technical effects of the embodiments of the present application. Again, other modules or units can also be included. In practical applications, these functions can also be assisted by other modules or units to be implemented, and can be implemented by the cooperation of multiple modules or units.

[0280] It should also be understood that the terms "module" or "unit" involved in the embodiments of the present application refer to a computer program with a predetermined function or a part of a computer program, and work together with other relevant parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit including the function of the module or unit.

[0281] For example, the decoder 400 or encoder 500 according to the embodiments of the present application can be constructed, and the encoding method or decoding method according to the embodiments of the present application can be implemented by running a computer program (including program code) capable of executing the steps involved in the corresponding method on a general computing device of a general computer including processing elements and storage elements such as a central processing unit (CPU), a random access storage medium (RAM), a read-only storage medium (ROM), etc. The computer program can be recorded on, for example, a computer-readable storage medium, and the computer-readable storage medium is loaded into an electronic device. Thus, when the computer program runs in the electronic device, it can execute the corresponding method according to the embodiments of the present application. Specifically, the steps of the method embodiments in the present application can be completed by the integrated logic circuit in hardware in the processor and / or instructions in software form. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed by a hardware decoding processor, or executed by a combination of hardware and software in the decoding processor. The software can be located in mature storage media in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The processor reads the information in the storage medium and combines its hardware to complete the steps in the method embodiments described above.

[0282] Figure 11 is a schematic structural diagram of the electronic device 600 provided by the present application.

[0283] As Figure 11 shown, the electronic device 600 at least includes a processor 610 and a computer-readable storage medium 620. Among them, the processor 610 and the computer-readable storage medium 620 can be connected through a bus or other means. The computer-readable storage medium 620 is used to store a computer program 621, and the computer program 621 includes computer instructions. The processor 610 is used to execute the computer instructions stored in the computer-readable storage medium 620. The processor 610 is the computing core and control core of the electronic device 600, and is adapted to implement one or more computer instructions, and is specifically adapted to load and execute one or more computer instructions to implement the corresponding method flow or corresponding function.

[0284] Exemplarily, the processor 610 may also be referred to as a Central Processing Unit (CPU). The processor 610 may include, but is not limited to: a general-purpose processor, a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or other programmable logic devices, transistor logic devices, discrete hardware components, and the like.

[0285] Exemplarily, the computer-readable storage medium 620 may be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory; alternatively, it may also be at least one computer-readable storage medium located far from the aforementioned processor 610. Specifically, the computer-readable storage medium 620 includes, but is not limited to: volatile memory and / or non-volatile memory. Among them, the non-volatile memory may be a Read-Only Memory (ROM), a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically EPROM (EEPROM), or a flash memory. The volatile memory may be a Random Access Memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as Static RAM (SRAM), Dynamic RAM (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced SDRAM (ESDRAM), synch link DRAM (SLDRAM), and Direct Rambus RAM (DR RAM).

[0286] Exemplarily, the electronic device 600 may be a decoder or a decoding framework involved in the embodiments of the present application; a first computer instruction is stored in the computer-readable storage medium 620; the processor 610 loads and executes the first computer instruction stored in the computer-readable storage medium 620 to implement the corresponding steps in the decoding method provided by the present application; in other words, the first computer instruction in the computer-readable storage medium 620 is loaded and executed by the processor 610 for the corresponding steps. To avoid repetition, details are not described herein again.

[0287] Exemplarily, the electronic device 600 may be an encoder or an encoding framework involved in the embodiments of the present application; a second computer instruction is stored in the computer-readable storage medium 620; the processor 610 loads and executes the second computer instruction stored in the computer-readable storage medium 620 to implement the corresponding steps in the encoding method provided by the present application; in other words, the second computer instruction in the computer-readable storage medium 620 is loaded and executed by the processor 610 for the corresponding steps. To avoid repetition, details are not described herein again.

[0288] According to another aspect of the present application, the present application further provides an encoding and decoding system, including the encoder and decoder described above.

[0289] According to another aspect of the present application, the present application further provides a computer-readable storage medium (Memory), which stores computer instructions. When the computer instructions are read and executed by the processor of the computer device, the computer device is caused to execute the encoding method or the decoding method described above.

[0290] Wherein, the computer-readable storage medium is a memory device in the decoder or the encoder, and is used to store programs and data. It can be understood that the computer-readable storage medium here may include both the built-in storage medium in the electronic device and, of course, the extended storage medium supported by the electronic device. The computer-readable storage medium can be used to provide a storage space, and the storage space can store the operating system of the electronic device. In addition, one or more computer instructions suitable for being loaded and executed by the processor are stored in this storage space. For example, one or more computer instructions for executing the encoding method or the decoding method described above are stored, and these computer instructions may be one or more computer programs (including program codes).

[0291] According to another aspect of the present application, the present application further provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer executes the encoding method or the decoding method provided in the various optional manners described above.

[0292] It should be understood that the computer device involved in the present application can be any device or apparatus capable of data processing. For example, it includes but is not limited to: general-purpose computers, special-purpose computers, computer networks, or other programmable devices. In addition, the computer instructions involved in the present application can be stored in a computer-readable storage medium, or can be transmitted between one computer-readable storage medium and another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.).

[0293] According to another aspect of the present application, the present application further provides a bitstream, which can be a bitstream decoded by using the decoding method provided in the present application or a bitstream generated by using the encoding method provided in the present application.

[0294] Those of ordinary skill in the art can realize that the units and process steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0295] Finally, it should be noted that the above content is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present application, and all of them should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A decoding method based on a neural network, characterized in that, it includes: Obtain a first element to be decoded based on a first bitstream; Obtain context information of the first element to be decoded based on at least one decoded element that has been decoded before the first element to be decoded; Based on the context information, use a parameter prediction model based on a neural network to predict a first decoding parameter for entropy decoding the first element to be decoded; Entropy decode the first element to be decoded based on the first decoding parameter to obtain decoded data of the first element to be decoded.

2. The method according to claim 1, characterized in that, the obtaining of the context information of the first element to be decoded includes: Use a convolutional layer with a convolution kernel of 1×k to perform convolution on K elements to obtain the context information; where the K elements include the at least one decoded element, and 1≤k≤K.

3. The method according to claim 2, characterized in that, when the K elements include undecoded elements, the weight configuration corresponding to the undecoded elements is 0, or the value configuration of the undecoded elements is 0; or when the K elements include undecoded elements, the first element to be decoded is the element located in the middle position among the K elements.

4. The method according to claim 1, characterized in that, the obtaining of the context information of the first element to be decoded includes: Determine the decoding parameter used for entropy decoding the at least one decoded element as the context information.

5. The method according to any one of claims 1 to 4, characterized in that, the at least one decoded element includes at least one of the following: A decoded element adjacent in position to the first element to be decoded, a decoded element whose distance from the first element to be decoded is less than or equal to a preset threshold, a decoded element at a preset position.

6. The method according to any one of claims 1 to 4, characterized in that, the using of the parameter prediction model based on a neural network to predict a first decoding parameter for entropy decoding the first element to be decoded includes: If the first element to be decoded is not the first element, then based on the using of the parameter prediction model, predict the first decoding parameter; The method further includes: If the first element to be decoded is the first element, then entropy decode the first element to be decoded based on a default decoding parameter to obtain decoded data of the first element to be decoded.

7. The method according to any one of claims 1 to 4, characterized in that, the using of the parameter prediction model based on a neural network to predict a first decoding parameter for entropy decoding the first element to be decoded includes: Using the context information as an input, use the parameter prediction model to perform probability modeling on the first element to be decoded to obtain the first decoding parameter.

8. The method according to any one of claims 1 to 4, characterized in that, the using of the parameter prediction model based on a neural network to predict a first decoding parameter for entropy decoding the first element to be decoded includes: Based on the second bitstream, obtain the prior features for determining the first decoding parameter; the second bitstream is the bitstream obtained through hyperprior encoding. Using the context information and the prior features as inputs, utilize the parameter prediction model to perform probability modeling on the first element to be decoded, and obtain the first decoding parameter.

9. The method according to claim 8, wherein, the obtaining the prior features for determining the first decoding parameter based on the second bitstream includes: Based on the second bitstream, obtain the features to be decoded; Perform entropy decoding on the features to be decoded to obtain the decoded features of the features to be decoded; Perform feature processing on the decoded features to obtain the prior features.

10. The method according to any one of claims 1 to 4, wherein, the first element to be decoded is any one of the following: an audio frame to be decoded, a channel to be decoded, a feature value to be decoded; the first decoding parameter includes probability distribution parameters, or the first decoding parameter includes the probability values corresponding to the respective candidate values of the first element to be decoded.

11. A neural network-based encoding method, wherein, it includes: Obtain a first element to be encoded; Based on at least one encoded element that has been encoded before the first element to be encoded, obtain the context information of the first element to be encoded; Based on the context information, utilize a neural network-based parameter prediction model to predict a first encoding parameter for performing entropy encoding on the first element to be encoded; Perform entropy encoding on the first element to be encoded based on the first encoding parameter to obtain the encoded data of the first element to be encoded.

12. The method according to claim 11, wherein, the obtaining the context information of the first element to be encoded includes: Use a convolutional layer with a convolution kernel of 1×k to perform convolution on K elements to obtain the context information; wherein, the K elements include the at least one encoded element, and 1 ≤ k ≤ K.

13. The method according to claim 12, wherein, when the K elements include unencoded elements, the weight configuration corresponding to the unencoded elements is 0, or the value configuration of the unencoded elements is 0; or when the K elements include unencoded elements, the first element to be encoded is the element located in the middle position among the K elements.

14. The method according to claim 11, wherein, the obtaining the context information of the first element to be encoded includes: Determine the encoding parameters used for performing entropy encoding on the at least one encoded element as the context information.

15. The method according to claim 11, wherein, the at least one encoded element includes at least one of the following: an encoded element adjacent in position to the first element to be encoded, an encoded element whose distance from the first element to be encoded is less than or equal to a preset threshold, an encoded element at a preset position.

16. The method according to claim 11, wherein, Predicting a first coding parameter for entropy coding the first element to be coded by using a parameter prediction model based on a neural network includes: If the first element to be coded is not the first element, use the parameter prediction model to predict the first coding parameter; The method further includes: If the first element to be coded is the first element, perform entropy coding on the first element to be coded based on default coding parameters to obtain coded data of the first element to be coded.

17. The method according to any one of claims 11 to 16, wherein, Predicting a first coding parameter for entropy coding the first element to be coded by using a parameter prediction model based on a neural network includes: Obtain features to be coded; Perform entropy coding on the features to be coded to obtain coded features of the features to be coded; Perform entropy decoding and feature processing on the coded features to obtain prior features; and Using the context information and the prior features as inputs, use the parameter prediction model to perform probability modeling on the first element to be coded to obtain the first coding parameter.

18. The method according to any one of claims 11 to 16, wherein, The first element to be coded is any one of the following: an audio frame to be coded, a channel to be coded, a feature value to be coded; The first coding parameter includes probability distribution parameters, or the first coding parameter includes probability values corresponding to respective candidate values of the first element to be coded.

19. An electronic device, wherein, includes: a processor adapted to execute a computer program; a computer-readable storage medium storing a computer program, which when executed by the processor implements: the neural network-based decoding method according to any one of claims 1 to 10, or the neural network-based coding method according to any one of claims 11 to 18.

20. A computer-readable storage medium, wherein, for storing a computer program, which causes a computer to execute: the neural network-based decoding method according to any one of claims 1 to 10, or the neural network-based coding method according to any one of claims 11 to 18.