Audio encoding and decoding method and device, computer equipment, program product and storage medium

By introducing configurable decoding and encoding sampling magnifications into the audio encoding and decoding device, the problem of structural delay fixation in the prior art is solved, and a wider range of applicable scenarios are achieved.

CN119943066APending Publication Date: 2025-05-06TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311468995.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-06
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing audio codec devices have fixed structural delays and cannot adapt to the needs of different scenarios, resulting in limited applicable scenarios.

Method used

Add configurable decoded sampling magnification to the upsampling module, and add configurable encoded sampling magnification to the downsampling module to achieve configurability of structural delays.

Benefits of technology

By configuring different sampling magnifications, the encoding and codec device can adapt to the requirements of different structural delays and expand applicable scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943066A_ABST
    Figure CN119943066A_ABST
Patent Text Reader

Abstract

The invention discloses an audio encoding and decoding method, an encoding and decoding device, computer equipment, a computer program product and a nonvolatile computer readable storage medium. The method comprises the steps that a current configuration parameter is obtained, the current configuration parameter is any one of a plurality of preset configuration parameters, the configuration parameters comprise at least one of a coding configuration parameter and a decoding configuration parameter, the coding configuration parameter at least comprises a coding sampling multiplying power, and the decoding configuration parameter at least comprises a decoding sampling multiplying power; each configuration parameter at least corresponds to different structural delays; encoding the input data according to the encoding configuration parameters to generate encoded data; and decoding the coded data according to the decoding configuration parameters to generate decoded data, and when coding and decoding are carried out based on each configuration parameter, other network parameters except the configuration parameters in the network parameters of the coding and decoding device are not changed. Therefore, the coding and decoding device can meet different structural delay requirements, so that the coding and decoding device can be suitable for more scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of coding and decoding technology, and more specifically, to an audio coding and decoding method, a coding and decoding device, a computer device, a computer program product, and a non-volatile computer-readable storage medium. Background Art

[0002] In recent years, deep learning solutions have been widely used in signal processing technologies of different dimensions (such as audio, image and video). In the audio encoding and decoding method based on deep learning, the encoder converts the original audio signal into encoded data, and the decoder decodes the encoded data to obtain the final reconstructed audio signal.

[0003] During the encoding and decoding process, structural delay will be generated. The parameters of existing encoding and decoding devices are fixed, resulting in fixed structural delay of the same encoding and decoding device. When the demand for structural delay changes, the encoding and decoding device cannot meet the current demand, resulting in limited application scenarios. Summary of the invention

[0004] The embodiments of the present application provide an audio coding method, a coding device, a computer device, a computer program product, and a non-volatile computer-readable storage medium. A configurable decoding sampling rate is added to an upsampling module, and a configurable encoding sampling rate is added to a downsampling module, so that the coding device can realize configurable structural delay, thereby making the coding device applicable to more scenarios.

[0005] The audio encoding and decoding method of the embodiment of the present application includes obtaining current configuration parameters, wherein the current configuration parameters are any one of a plurality of preset configuration parameters, wherein the configuration parameters include at least one of encoding configuration parameters and decoding configuration parameters, and when used for encoding, the configuration parameters include the encoding configuration parameters, and when used for decoding, the configuration parameters include the decoding configuration parameters, wherein the encoding configuration parameters at least include an encoding sampling rate, and the decoding configuration parameters at least include a decoding sampling rate, and each of the configuration parameters at least corresponds to a different structural delay; encoding input data according to the encoding configuration parameters to generate encoded data; decoding the encoded data according to the decoding configuration parameters to generate decoded data, and when encoding and decoding are performed based on each of the configuration parameters, among the network parameters of the encoding and decoding device, other network parameters other than the configuration parameters remain unchanged.

[0006] In some embodiments, the method further includes: acquiring input configuration parameters to determine the current configuration parameters, wherein the input configuration parameters include at least one of a codec sampling ratio, a structural delay, and a signal sampling rate.

[0007] In some embodiments, obtaining the input configuration parameters to determine the current configuration parameters includes: determining one or more target configuration parameters based on the input configuration parameters, the target configuration parameters being any of the configuration parameters; obtaining the input parameter index number to determine the current configuration parameter corresponding to the parameter index number in one or more of the target configuration parameters.

[0008] In certain embodiments, the number of parameters corresponding to the encoding sampling rate is determined according to the number of downsampling layers preset by the encoder, and the number of parameters corresponding to the decoding sampling rate is determined according to the number of upsampling layers preset by the decoder, and the encoding product of each encoding sampling rate is the same as the decoding product of each decoding sampling rate.

[0009] In some embodiments, the number of downsampling layers used for encoding is the same as the number of upsampling layers used for decoding, and the encoding sampling ratio and the decoding sampling ratio correspond one to one.

[0010] In some embodiments, the number of downsampling layers used for encoding and the number of upsampling layers used for decoding are different.

[0011] In some embodiments, the configuration parameters also include encoding quantization parameters and decoding quantization parameters, and the encoding of input data according to the encoding configuration parameters to generate encoded data includes: downsampling the input data according to the encoding sampling rate to generate encoded intermediate data; quantizing the encoded intermediate data according to the encoding quantization parameters to generate the encoded data; decoding the encoded data according to the decoding configuration parameters to generate decoded data includes: inverse quantizing the encoded data according to the decoding quantization parameters to generate decoded intermediate data; upsampling the decoded intermediate data according to the decoding sampling rate to generate the decoded data.

[0012] In some embodiments, the configuration parameters also include convolution layer parameters in the downsampling layer and convolution layer parameters in the upsampling layer, the convolution layer parameters include at least weight parameters and bias parameters, and encoding the input data according to the encoding configuration parameters to generate encoded data includes: encoding the input data according to the encoding sampling ratio and the convolution layer parameters in the downsampling layer to generate encoded data; decoding the encoded data according to the decoding configuration parameters to generate decoded data includes: decoding the encoded data according to the decoding sampling ratio and the convolution layer parameters in the upsampling layer to generate the decoded data.

[0013] In some embodiments, the configuration parameters include convolution layer parameters in the downsampling layer of a first target layer and convolution layer parameters in the upsampling layer of a second target layer, the first target layer being the downsampling layer in which the encoding sampling rate changes in different configuration parameters, and the second target layer being the upsampling layer in which the decoding sampling rate changes in different configuration parameters.

[0014] In some embodiments, the method also includes: obtaining a current input sample, and randomly selecting any of the configuration parameters to configure the encoder and the decoder, the current input sample being any sample in a preset sample set; encoding and decoding the current input sample by the configured encoder and the decoder to obtain a current output sample; determining loss values ​​of the encoder and the decoder based on the current input sample and the current output sample; adjusting the encoder and the decoder based on the loss value to update the network parameters of the encoder and the decoder until the encoder and the decoder are trained to convergence.

[0015] In some embodiments, encoding the input data according to the encoding configuration parameters to generate encoded data includes: encoding the input data according to the encoding sampling rate and the updated network parameters of the encoder to generate encoded data; decoding the encoded data according to the decoding configuration parameters to generate decoded data includes: decoding the encoded data according to the decoding sampling rate and the updated network parameters of the decoder to generate decoded data.

[0016] In some implementations, the encoder obtains the encoding configuration parameters according to the set structural delay information, and the decoder obtains the decoding configuration parameters according to the structural delay information in the bitstream formed by the encoded data.

[0017] In some embodiments, the structural delay of the structural delay information is determined according to at least one of the content complexity and real-time parameters of the input data.

[0018] In some embodiments, the structural delay is positively correlated with the content complexity, the structural delay is negatively correlated with the real-time parameter, and the content complexity is determined based on at least one of the sampling rate, the number of channels, and the number of sound source types of the input data.

[0019] The encoding and decoding device of the embodiment of the present application includes an encoder and a decoder. The encoder is used to encode input data according to the encoding configuration parameters of the current configuration parameters to generate encoded data. The decoder is used to decode the encoded data according to the decoding configuration parameters of the current configuration parameters to generate decoded data, the current configuration parameters are any one of a plurality of preset configuration parameters, the encoding configuration parameters at least include the encoding sampling rate, the decoding configuration parameters at least include the decoding sampling rate, and the structural delays corresponding to the various configuration parameters are different.

[0020] In some embodiments, the configuration parameters also include encoding quantization parameters and decoding quantization parameters, the encoder includes a downsampling module and a quantization module, the downsampling module is used to downsample the input data according to the encoding sampling ratio to generate encoded intermediate data; the quantization module is used to quantize the encoded intermediate data according to the encoding quantization parameters to generate the encoded data; the decoder includes an upsampling module and a dequantization module; the dequantization module is used to dequantize the encoded data according to the decoding quantization parameters to generate decoded intermediate data; the upsampling module is used to upsample the decoded intermediate data according to the decoding sampling ratio to generate the decoded data.

[0021] In some embodiments, the encoder includes multiple downsampling modules, the downsampling modules include a first convolution unit and a first residual unit, the decoder includes multiple upsampling modules, the upsampling modules include a deconvolution unit and a second residual unit, and the configuration parameters also include convolution layer parameters of the first convolution unit and convolution layer parameters of the deconvolution unit; the first residual unit is used to perform a first convolution operation on the input feature vector corresponding to the input data to generate a first intermediate feature vector; the first convolution unit is used to perform a second convolution operation on the intermediate feature vector according to the encoding sampling rate and the convolution layer parameters of the first convolution unit to generate the encoded data, the deconvolution unit is used to perform a third convolution operation on the encoded data according to the decoding sampling rate and the convolution layer parameters of the deconvolution unit to generate a second intermediate feature vector, and the second residual unit is used to perform a fourth convolution operation on the second intermediate feature vector to generate the decoded data.

[0022] The computer device of the embodiment of the present application includes a processor, a memory and a computer program, wherein the computer program is stored in the memory and executed by the processor, and the computer program includes instructions for executing the audio encoding and decoding method described in any of the above embodiments.

[0023] The computer program product of the embodiments of the present application includes a computer program, wherein when the computer program is executed by the processor, the audio encoding and decoding method described in any of the above embodiments is implemented.

[0024] The non-volatile computer-readable storage medium of the embodiment of the present application includes a computer program. When the computer program is executed by a processor, the processor executes the audio encoding and decoding method described in any of the above embodiments.

[0025] The audio coding method, coding device, computer equipment, computer program product and non-volatile computer-readable storage medium of the embodiment of the present application add a configurable decoding sampling rate in the upsampling module and a configurable encoding sampling rate in the downsampling module to provide a function of changing the structural delay. The current configuration parameters can be adjusted according to the current structural delay requirements, and at least the structural delays corresponding to each configuration parameter are different. The coding and decoding device can obtain the current configuration parameters corresponding to the current structural delay requirements, and encode the input data according to the encoding configuration parameters, and encode the encoded data according to the decoding configuration parameters to generate decoded data, and when encoding and decoding based on each configuration parameter, the network parameters of the coding and decoding device, other network parameters other than the configuration parameters remain unchanged. In this way, since the coding and decoding device can realize the configurability of the structural delay, the encoding configuration parameters and the decoding configuration parameters can be set according to the structural delay requirements, so that the coding and decoding device can meet different structural delay requirements, so that the coding and decoding device can be applied to more scenarios.

[0026] Additional aspects and advantages of the embodiments of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the embodiments of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the description of the embodiments in conjunction with the following drawings, in which:

[0028] Figure 1 It is a schematic diagram of an application scenario of an audio encoding and decoding method of certain implementation modes of the present application;

[0029] Figure 2 This is an example diagram of the basic process of the audio encoding and decoding method of some implementation methods of the present application;

[0030] Figure 3 It is a flowchart of an audio encoding and decoding method of certain implementation modes of the present application;

[0031] Figure 4 It is a scene schematic diagram of the audio encoding and decoding method of certain implementation modes of the present application;

[0032] Figure 5 It is a flowchart of an audio encoding and decoding method of certain implementation modes of the present application;

[0033] Figure 6 It is a flowchart of an audio encoding and decoding method of certain implementation modes of the present application;

[0034] Figure 7 It is a flowchart of an audio encoding and decoding method of certain implementation modes of the present application;

[0035] Figure 8 It is a flowchart of an audio encoding and decoding method of certain implementation modes of the present application;

[0036] Fig. 9 It is a flowchart of an audio encoding and decoding method of certain implementation modes of the present application;

[0037] Fig.10 It is a flowchart of an audio encoding and decoding method of certain implementation modes of the present application;

[0038] Fig.11 is a schematic diagram of the structure of a computer device of certain embodiments of the present application;

[0039] Fig.12 It is a schematic diagram of the connection status of a non-volatile computer-readable storage medium and a processor in certain embodiments of the present application. DETAILED DESCRIPTION

[0040] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals represent the same or similar elements or elements having the same or similar functions from beginning to end. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the embodiments of the present application, and cannot be understood as limiting the embodiments of the present application.

[0041] To facilitate understanding of this application, the terms appearing in this application are explained below:

[0042] Encoding and decoding: The audio encoding process is to compress the audio into smaller data, and the decoding process is to restore the smaller data to audio. The encoded smaller data is used for network transmission and occupies less bandwidth.

[0043] Sampling rate: The sampling rate describes the number of data contained in a unit of time (1 second). For example, a 16k sampling rate contains 16,000 sampling points, and each sampling point corresponds to a short integer.

[0044] Codebook: A collection of multiple vectors. The encoding device and the decoding device both store the same codebook.

[0045] Quantization: Find the closest vector in the codebook for the input vector, return it as a replacement for the input vector, and return the corresponding codebook index position.

[0046] Quantizer: The quantizer is responsible for quantization and updating the vectors in the codebook.

[0047] Figure 1 A schematic diagram schematically shows an exemplary system architecture to which the technical solution of the embodiments of the present application can be applied.

[0048] like Figure 1 As shown, the system architecture includes a plurality of terminal devices, which can communicate with each other through, for example, a network. For example, the system architecture may include a first terminal device 1001 and a second terminal device 1002 interconnected through a network. Figure 1 In the embodiment of the present invention, the first terminal device 1001 and the second terminal device 1002 perform unidirectional data transmission.

[0049] For example, the first terminal device 1001 can encode audio and video data (such as an audio and video data stream collected by the terminal device) for transmission to the second terminal device 1002 via a network. The encoded audio and video data is transmitted in the form of one or more encoded audio and video streams. The second terminal device 1002 can receive the encoded audio and video data from the network, decode the encoded audio and video data to restore the audio and video data, and play or display content based on the restored audio and video data.

[0050] Figure 2 The following is a basic flow chart of an end-to-end audio encoding and decoding method. During encoding, the encoder 100 at the data transmitting end first encodes the input audio signal to generate a binary code stream, and then the data transmitting end sends the binary code stream to the data receiving end. After receiving the binary code stream, the data receiving end decodes the binary code stream through the decoder 200 to obtain a reconstructed audio signal.

[0051] In one embodiment of the present application, the system architecture may include a third terminal device 1003 and a fourth terminal device 1004 that perform bidirectional transmission of encoded audio and video data, which may occur, for example, during an audio and video conference. For bidirectional data transmission, each of the third terminal device 1003 and the fourth terminal device 1004 may encode audio and video data (e.g., an audio and video data stream collected by the terminal device) for transmission to the other terminal device of the third terminal device 1003 and the fourth terminal device 1004 through a network. Each of the third terminal device 1003 and the fourth terminal device 1004 may also receive the encoded audio and video data transmitted by the other terminal device of the third terminal device 1003 and the fourth terminal device 1004, and may decode the encoded audio and video data to restore the audio and video data, and play or display content according to the restored audio and video data.

[0052] exist Figure 1 In the embodiment of the present invention, the first terminal device 1001, the second terminal device 1002, the third terminal device 1003 and the fourth terminal device 1004 may be servers, personal computers and smart phones, but the principles disclosed in the present application may not be limited thereto. The embodiments disclosed in the present application are applicable to laptop computers, tablet computers, media players and / or dedicated audio and video conferencing equipment. The network represents any number of networks that transmit encoded audio and video data between the first terminal device 1001, the second terminal device 1002, the third terminal device 1003 and the fourth terminal device 1004, including, for example, wired and / or wireless communication networks. The communication network may exchange data in circuit switching and / or packet switching channels. The network may include a telecommunications network, a local area network, a wide area network and / or the Internet. For the purpose of the present application, unless explained below, the architecture and topology of the network may be irrelevant to the operation disclosed in the present application.

[0053] The audio encoding and decoding method of this application will be described in detail below:

[0054] See also Figure 3 and Figure 4 The present application embodiment provides an audio coding and decoding method, the audio coding and decoding method comprising:

[0055] Step 011: obtaining a current configuration parameter, which is any one of a plurality of preset configuration parameters, and the configuration parameter includes at least one of an encoding configuration parameter and a decoding configuration parameter. When used for encoding, the configuration parameter includes an encoding configuration parameter, and when used for decoding, the configuration parameter includes a decoding configuration parameter. The encoding configuration parameter at least includes an encoding sampling magnification, and the decoding configuration parameter at least includes a decoding sampling magnification, and each configuration parameter at least corresponds to a different structural delay;

[0056] Specifically, the encoding and decoding device 1000 includes an encoder 100 and a decoder 200. The configuration parameters are parameters that can be configured in the network parameters that the encoder 100 and the decoder 200 need to use during the encoding and decoding process. When encoding and decoding are performed based on each configuration parameter, other network parameters other than the configuration parameters in the network parameters of the encoding and decoding device 1000 remain unchanged.

[0057] The encoder 100 may include a downsampling module 10, a first convolution module 20, a second convolution module 30, and a quantization module 40. Optionally, a fifth convolution module 60 may be added between any two downsampling modules 10. The encoding configuration parameters include at least network parameters of at least some modules in the encoder 100, such as the encoding sampling rate of the downsampling module 10, the convolution parameters of the first convolution module 20, the convolution parameters of the second convolution module 30, the encoding quantization parameters of the quantization module 40, and the convolution parameters of the fifth convolution module 60. The parameters of at least some modules in the various modules of the encoder 100 are not shared, and these parameters can be configured. These unshared and configurable parameters are encoding configuration parameters.

[0058] Similarly, the decoder 200 may include an upsampling module 210, a dequantization module 220, a third convolution module 230, and a fourth convolution module 240. Optionally, a sixth convolution module 260 may be added between any two upsampling modules 210. The decoding configuration parameters include at least network parameters of at least some modules in the decoder 200, such as the decoding sampling rate of the upsampling module 210, the decoding quantization parameters of the dequantization module 220, the convolution parameters of the third convolution module 230, the convolution parameters of the fourth convolution module 240, and the convolution parameters of the sixth convolution module 260. The parameters of at least some modules of the various modules of the decoder 200 are not shared, and these parameters can be configured. These unshared and configurable parameters are the decoding configuration parameters.

[0059] The configuration parameters may include at least one of the encoding configuration parameters and the decoding configuration parameters. For example, when encoding, the user may configure the encoding configuration parameters, and the configuration parameters are encoding configuration parameters. When decoding, the user may configure the decoding configuration parameters, and the configuration parameters are decoding configuration parameters. For another example, when encoding, the user may configure the encoding configuration parameters and the decoding configuration parameters at the same time, and the configuration parameters are encoding configuration parameters and decoding configuration parameters.

[0060] The downsampling module 10 is used to downsample the input data according to the encoding sampling ratio. The encoder 100 may include one or more downsampling modules 10, and one downsampling module 10 corresponds to one encoding sampling ratio. Wherein, when there are multiple downsampling modules 10, the multiple downsampling modules 10 are connected in sequence, and the output feature vector of the previous downsampling module 10 is used as the input feature vector of the next downsampling module 10. The first convolution module 20 and the second convolution module 30 can extract features from the data. The quantization module 40 is used to quantize the data to generate encoded data.

[0061] The up-sampling module 210 can up-sample the data according to the decoding sampling ratio to generate decoded data. The decoder 200 may include one or more up-sampling modules 210, and one up-sampling module 210 corresponds to one decoding sampling ratio. Wherein, when there are multiple up-sampling modules 210, the multiple up-sampling modules 210 are connected in sequence, and the output feature vector of the previous up-sampling module 210 is used as the input feature vector of the next up-sampling module 210. The dequantization module 220 can perform a dequantization operation on the encoded data. The third convolution module 230 and the fourth convolution module 240 can extract features from the data.

[0062] The encoder 100 may encode the input data according to the encoding configuration parameters, and the decoder 200 may decode the encoded data according to the decoding configuration parameters. For example, the encoder 100 implements dimensionality reduction according to the encoding sampling rate, thereby performing downsampling, and the decoder 200 implements dimensionality increase according to the decoding sampling rate, thereby performing upsampling. The structural delay is at least related to the encoding sampling rate and the decoding sampling rate. For example, assuming that the product of the encoding sampling rates of multiple downsampling modules 10 is N, and the sampling rate of the input signal is FS, then the structural delay of the encoder 100 is N*1000 / Fs. Therefore, the encoding configuration parameters at least include the encoding sampling rate, and the decoding configuration parameters at least include the decoding sampling rate.

[0063] At the same sampling rate, different configuration parameters correspond to at least different structural delays. For example, there are 4 downsampling modules 10 at the encoding end and 4 upsampling modules 210 at the decoding end. The structural delay corresponding to configuration parameter A is 8ms. According to the order in which data is transmitted in the downsampling module 10, the encoding sampling ratios of the 4 downsampling modules 10 are 2, 4, 4 and 6, respectively. According to the order in which data is transmitted in the upsampling module 210, the decoding sampling ratios of the 4 upsampling modules 210 are 6, 4, 4 and 2, respectively. The structural delay corresponding to configuration parameter B is 10ms. According to the order in which data is transmitted in the downsampling module 10, the encoding sampling ratios of the 4 downsampling modules 10 are 2, 4, 5 and 6, respectively. According to the order in which data is transmitted in the upsampling module 210, the decoding sampling ratios of the 4 upsampling modules 210 are 6, 5, 4 and 2, respectively. The structural delay corresponding to the configuration parameter C is 16ms. According to the order in which the data is transmitted in the downsampling module 10, the encoding sampling ratios of the four downsampling modules 10 are 2, 4, 6 and 8 respectively. According to the order in which the data is transmitted in the upsampling module 210, the decoding sampling ratios of the four upsampling modules 210 are 8, 6, 4 and 2 respectively.

[0064] Of course, at the same sampling rate, in addition to the structural delay, other parameters corresponding to different configuration parameters may also be different. For example, different configuration parameters may correspond to different encoding quantization parameters, and different configuration parameters may correspond to different decoding quantization parameters.

[0065] Therefore, the coding and decoding device 1000 can determine the current configuration parameters according to the current structural delay requirements to at least determine the encoding sampling rate and decoding sampling rate that match the current structural delay requirements. As for other network parameters in the coding and decoding process except the configuration parameters, they can be preset according to the parameters that can simultaneously meet the structural delay requirements and the coding and decoding effect requirements under different structural delay requirements. For example, the coding quantization parameter and the decoding quantization parameter are not within the configuration parameter range. Under different structural delay requirements, the coding quantization parameter A and the decoding quantization parameter B can meet the structural delay requirements and have a good coding and decoding effect. Then the coding quantization parameter A can be set to the preset coding quantization parameter, and the decoding quantization parameter B can be set to the preset decoding quantization parameter. After that, the coding quantization parameter and the decoding quantization parameter of the coding and decoding device 1000 can be determined according to the preset coding quantization parameter and the preset decoding quantization parameter.

[0066] When used for encoding, the configuration parameters include encoding configuration parameters, and the encoder 100 obtains the encoding sampling rate so that the encoder 100 encodes according to the structural delay requirement. When used for decoding, the configuration parameters include decoding configuration parameters. The decoder 200 obtains the decoding sampling rate so that the decoder 200 decodes according to the structural delay requirement. In this way, the encoding and decoding device 1000 can meet the current structural delay requirement.

[0067] Optionally, the structural delay information can be set according to the structural delay requirements, and the encoder 100 can determine the structural delay of the structural delay information from the structural delay information, thereby obtaining the encoding configuration parameters. When encoding, the encoder 100 will encode according to the encoding configuration parameters, and then specify the corresponding structural delay number and write it into the bitstream, thereby generating encoded data including the structural delay information. The decoding end can determine the encoding of the structural delay according to the structural delay information in the bitstream formed by the encoded data to obtain the decoding configuration parameters, and then complete the decoding according to the decoding configuration parameters to obtain the reconstructed decoded data.

[0068] Step 012: Encode the input data according to the encoding configuration parameters to generate encoded data;

[0069] Specifically, the audio data can be encoded and compressed by the encoder 100 at the data transmitting end. After receiving the input data, the encoder 100 can perform data sampling on the input data to be encoded to obtain the original feature vector. The encoder 100 inputs the original feature vector into the first convolution module 20 for convolution, and after the convolution process, the first feature vector corresponding to the input data can be obtained. Then, the encoder 100 can input the first feature vector into the downsampling module 10, and the downsampling module 10 will perform downsampling according to the encoding configuration parameters to obtain the second feature vector. Then, the second convolution module 30 performs convolution on the second feature vector to obtain the encoding feature vector. After that, the encoding feature vector is input into the quantization module 40 for quantization (such as using a residual-based vector quantizer) operation, and the quantization module 40 can perform quantization according to the encoding quantization parameter, and then a binary code stream can be generated according to the quantization result, thereby obtaining the encoded data. For example, the encoding feature vector can be input into the quantizer, so that the code stream corresponding to each encoding feature vector can be queried in the code book.

[0070] Optionally, a first linear transformation module 50 can be added between the second convolution module 30 and the quantization module 40. The first linear transformation module 50 can perform a linear transformation operation on the encoded feature vector output by the second convolution module 30, and then input the linearly transformed encoded feature vector into the quantization module 40 to improve the encoding effect.

[0071] Step 013: Decode the encoded data according to the decoding configuration parameters to generate decoded data. When encoding and decoding are performed based on the various configuration parameters, among the network parameters of the encoding and decoding device 1000, other network parameters except the configuration parameters remain unchanged.

[0072] Specifically, after receiving the encoded data, the data receiving end may input the encoded data into the dequantization module 220 for dequantization. For example, after receiving the code stream transmitted by the network, the code book feature vector corresponding to the code stream may be queried in the code book through the quantizer. After the dequantization module 220 performs dequantization according to the decoding quantization parameter, the decoder 200 may input the dequantized code book feature vector into the third convolution module 230 for feature extraction to obtain the third feature vector corresponding to the encoded data. Next, the decoding device inputs the third feature vector into the upsampling module 210. At this time, the upsampling module 210 performs an upsampling operation on the third feature vector according to the decoding configuration parameters to obtain a fourth feature vector. Then, the decoding device inputs the fourth feature vector into the fourth convolution module 240 for feature extraction, thereby determining the decoded data based on the fourth feature vector after feature extraction.

[0073] Of course, if there is a first linear transformation module 50 in the encoder 100, then the decoder 200 can also be provided with a second linear transformation module 250 in the inverse quantization module 220 and the third convolution module 230, so as to utilize the second linear transformation module 250 to perform a linear transformation operation on the feature vector after the inverse quantization operation, so as to improve the decoding effect.

[0074] It can be understood that since the encoding configuration parameters and the decoding configuration parameters are set according to the structural delay requirements, the structural delay of the codec device 1000 can be adjusted, and the codec device 1000 can meet different structural delay requirements during the encoding and decoding process, so that the codec device 1000 with configurable structural delay can be realized. That is, the codec device 1000 of the present application can provide a configurable structural delay of the codec device 1000 on the basis of the traditional fixed structural delay, so that the codec device 1000 of the present application has wide applicability and can be applied to the existing end-to-end audio codec solution based on deep learning.

[0075] For example, there are 4 down-sampling modules 10 and 4 up-sampling modules 210. According to the order in which data is transmitted in the down-sampling modules 10, the encoding sampling rates of the 4 down-sampling modules 10 are 2, 4, 6 and 8, respectively, and according to the order in which data is transmitted in the up-sampling modules 210, the decoding sampling rates of the 4 up-sampling modules 210 are 8, 6, 4 and 2, respectively.

[0076] In the input stage of the encoder 100, the input data to be encoded is sampled to obtain an original feature vector with a channel number of 1 and a dimension of 19200; the original feature vector is input to the first convolution module 20, and after convolution processing, a first feature vector with a channel number of 32 and a dimension of 19200 can be obtained. In some optional embodiments, in order to improve the encoding efficiency, the encoder 100 can encode a batch of feature vectors with a number B at the same time.

[0077] In the downsampling stage of the encoder 100, according to the corresponding encoding sampling ratio, the first downsampling module 10 reduces the vector dimension to 1 / 2, and obtains the fifth eigenvector with 64 channels and 9600 dimensions; the second downsampling module 10 reduces the vector dimension to 1 / 4, and obtains the sixth eigenvector with 128 channels and 2400 dimensions; the third downsampling module 10 reduces the vector dimension to 1 / 6, and obtains the seventh eigenvector with 256 channels and 400 dimensions; the fourth downsampling module 10 reduces the vector dimension to 1 / 8, and obtains the second eigenvector with 512 channels and 50 dimensions. Assuming that the sampling rate of the input data is 24 kHz, the structural delay at this time is 2*4*6*8*1000 / 24000=16ms.

[0078] At the output stage of the encoder 100 , the second convolution module 30 performs convolution processing on the second feature vector to obtain a coding feature vector with K channels and 50 dimensions. K is a preset vector quantization dimension, which may be 32, for example.

[0079] After the encoding feature vector is input to the quantization module 40 , the quantization module 40 quantizes the encoding feature vector according to the encoding quantization parameter to obtain encoding data.

[0080] After receiving the encoded data transmitted by the network, the decoder 200 of the data receiving end will input the encoded data into the dequantization module 220, and the dequantization module 220 will dequantize the encoded data according to the decoding quantization parameter to obtain the codebook feature vector corresponding to the encoded data, which can be, for example, a vector with a channel number of K and a dimension of 50. K is a preset vector quantization dimension, for example, which can be 32. In some optional embodiments, in order to improve the decoding efficiency, the data receiving end can decode a batch of codebook feature vectors with a number of B at the same time.

[0081] In the input stage of the decoder 200, the codebook feature vector to be decoded is input into the third convolution module 230, and after convolution processing, a third feature vector with 512 channels and 50 dimensions can be obtained.

[0082] In the upsampling stage of the decoder 200, according to the decoding sampling ratio, the first upsampling module 210 increases the vector dimension to 8 times, and obtains the eighth eigenvector with 256 channels and 400 dimensions; the second upsampling module 210 increases the vector dimension to 6 times, and obtains the ninth eigenvector with 128 channels and 2400 dimensions; the third upsampling module 210 increases the vector dimension to 4 times, and obtains the tenth eigenvector with 64 channels and 9600 dimensions; the fourth upsampling module 210 increases the vector dimension to 2 times, and obtains the fourth eigenvector with 32 channels and 19200 dimensions.

[0083] In the output stage of the decoder 200, the fourth convolution module 240 performs convolution processing on the fourth eigenvector to restore the decoded data with a channel number of 1 and a dimension of 16000.

[0084] Optionally, the encoding sampling rate and the decoding sampling rate may be the same or different, and it is only necessary to ensure that the encoding product of each encoding sampling rate and the decoding product of each decoding sampling rate are the same. The encoding product is determined by the number of downsampling layers used for encoding and the encoding sampling rate, and the decoding product is determined by the number of upsampling layers used for decoding and the decoding sampling rate.

[0085] The downsampling layer corresponds to the downsampling module 10 one-to-one, and the upsampling layer corresponds to the upsampling module 210 one-to-one. Each downsampling module 10 corresponds to an encoding sampling rate, and each upsampling module 210 corresponds to a decoding sampling rate. It can be understood that the number of parameters corresponding to the encoding sampling rate is determined according to the number of downsampling layers preset by the encoder 100, that is, the number of downsampling modules 10, and the number of parameters corresponding to the decoding sampling rate is determined according to the number of upsampling layers preset by the decoder 200, that is, the number of upsampling modules 210.

[0086] The number of downsampling layers used for encoding and the number of upsampling layers used for decoding may be the same or different. In one embodiment, the number of downsampling layers used for encoding is the same as the number of upsampling layers used for decoding, and the encoding sampling rate and the decoding sampling rate correspond one to one. For example, the number of downsampling layers and the number of upsampling layers are both 4, and the encoding sampling rates are 2, 4, 6 and 8, respectively, and the decoding sampling rates are 8, 6, 4 and 2, respectively. In another embodiment, the number of downsampling layers used for encoding is the same as the number of upsampling layers used for decoding, but the encoding sampling rate and the decoding sampling rate do not correspond one to one. For example, the number of downsampling layers and the number of upsampling layers are both 4, and the encoding sampling rates are 2, 4, 6 and 8, respectively, and the decoding sampling rates are 8, 8, 3 and 2, respectively. In another embodiment, the number of downsampling layers used for encoding is different from the number of upsampling layers used for decoding, for example, the number of downsampling layers is 3, the encoding sampling rates are 2, 6 and 8, respectively, and the number of upsampling layers is 4, which are 2, 2, 4 and 6, respectively.

[0087] In this way, the codec device 1000 can flexibly adjust the downsampling module 10 and the corresponding encoding sampling rate, the upsampling module 210 and the corresponding decoding sampling rate according to the requirements of the structural delay and the specific usage of the encoder 100 and the decoder 200, so that the structural delay of the codec device 1000 can change accordingly, thereby allowing the codec device 1000 to adapt to more user needs.

[0088] The audio codec method of the embodiment of the present application adds a configurable decoding sampling rate in the upsampling module 210 and a configurable encoding sampling rate in the downsampling module 10 to provide a function of changing the structural delay. The current configuration parameters can be adjusted according to the current structural delay requirements, and each configuration parameter at least corresponds to a different structural delay. The codec device 1000 can obtain the current configuration parameters corresponding to the current structural delay requirements, and encode the input data according to the encoding configuration parameters, and encode the encoded data according to the decoding configuration parameters to generate decoded data, and when encoding and decoding based on each configuration parameter, the network parameters of the codec device, other network parameters other than the configuration parameters remain unchanged. In this way, since the codec device 1000 can realize the configurability of the structural delay, the encoding configuration parameters and the decoding configuration parameters can be set according to the structural delay requirements, so that the codec device 1000 can meet different structural delay requirements, so that the codec device 1000 can be applied to more scenarios.

[0089] In addition, the upsampling ratio and downsampling ratio in the existing codec device 1000 are fixed, and the codec device 1000 can only meet one structural delay requirement, resulting in that all parameters in the codec 200 can only be retrained to obtain a codec device 1000 that meets different structural delay requirements. However, the configuration parameters of the present application can be adjusted according to the structural delay requirements, and in addition to the configuration parameters, other parameters involved in the encoding and decoding process can be set using corresponding parameters that can be adapted to more structural delay requirements. Therefore, compared with the solution in the prior art that requires retraining all parameters to obtain a codec device 1000 that meets the structural delay requirements, the present application only adjusts the configuration parameters to obtain a codec device 1000 that meets the structural delay requirements. Obviously, the present application can obtain a codec device 1000 that meets the structural delay requirements more quickly.

[0090] See also Figure 4 and Figure 5 In some implementations, the audio encoding and decoding method further includes:

[0091] Step 014: Obtain input configuration parameters to determine current configuration parameters, where the input configuration parameters include at least one of a codec sampling ratio, a structural delay, and a signal sampling rate.

[0092] Specifically, each codec sampling ratio has a corresponding configuration parameter, each structural delay has a corresponding configuration parameter, and each signal sampling rate also has a corresponding configuration parameter. Therefore, the user can input requirements as needed, such as the user can input at least one of the codec sampling ratio, structural delay and signal sampling rate to generate input configuration parameters. At this time, the codec device 1000 can obtain at least one of the codec sampling ratio, structural delay and signal sampling rate, and determine the configuration parameter that matches the input information from multiple configuration parameters, thereby determining the current configuration parameter. For example, the codec device 1000 can determine the current configuration parameter based on the codec sampling ratio, structural delay or signal sampling rate. The codec device 1000 can also determine the current configuration parameter based on the codec sampling ratio and structural delay, or structural delay and signal sampling rate. Alternatively, the codec device 1000 can determine the current configuration parameter based on the codec sampling ratio, structural delay and signal sampling rate. In this way, the codec device 1000 can determine the current configuration parameters according to the user's requirements, and the codec device 1000 is guaranteed to meet the user's needs.

[0093] See also Figure 4 and Figure 6 In some implementations, step 014: obtaining input configuration parameters to determine current configuration parameters, the input configuration parameters including at least one of a codec sampling rate, a structural delay, and a signal sampling rate, including:

[0094] Step 0141: determining one or more target configuration parameters according to the input configuration parameters, where the target configuration parameter is any configuration parameter;

[0095] Step 0142: Obtain the input parameter index number to determine the current configuration parameter corresponding to the parameter index number in one or more target configuration parameters.

[0096] Specifically, there may be multiple groups of configuration parameters corresponding to the same codec sampling ratio. For example, in addition to the coding sampling ratio and the decoding sampling ratio, the configuration parameters also include coding quantization parameters and decoding quantization parameters. If the structural delays of configuration parameters A and B are the same, but the coding quantization parameters and the decoding quantization parameters are different, then it can be considered that configuration parameters A and configuration parameters B are different. Similarly, there may be multiple groups of configuration parameters corresponding to the same structural delay, and there may be multiple groups of configuration parameters corresponding to the same signal sampling rate. At the same time, in order to facilitate management and use, each group of configuration parameters has a corresponding parameter index number, and the coding and decoding device 1000 can determine the current configuration parameters according to the parameter index number.

[0097] The codec device 1000 determines one or more configuration parameters (i.e., target configuration parameters) that meet the input requirements from multiple configuration parameters according to the input configuration parameters, wherein the target configuration parameter is any configuration parameter. For example, when the user only inputs the structural delay, the target configuration parameters may be multiple; when the user simultaneously inputs the codec sampling magnification, the structural delay, and the signal sampling rate, the target configuration parameter may be one. In this way, the preliminary determination of the configuration parameters can be completed. Then, the codec device 1000 will display the target configuration parameters, and the user can select the corresponding configuration parameters based on the displayed content to input the parameter index number. Next, the codec device 1000 can obtain the parameter index number of the input and obtain the parameter index number of the target configuration parameter to determine the current configuration parameter corresponding to the parameter index number among multiple target configuration parameters. Of course, if there is only one target configuration parameter, the target configuration parameter can be directly confirmed as the current configuration parameter. In this way, the codec device 1000 can accurately determine the current configuration parameters according to the input parameters, the structural delay, the signal sampling rate, and the parameter index number, thereby ensuring that the codec device 1000 can meet the needs of the user.

[0098] Of course, if the user knows the index number of the configuration parameter to be used, the user can directly enter the parameter index number to determine the current configuration parameter to ensure the accuracy of the current configuration parameter and also speed up the determination of the current configuration parameter.

[0099] Optionally, the input configuration parameters may also include other network parameters, such as encoding quantization parameters and decoding quantization parameters, to ensure that the codec device 1000 can accurately determine the current configuration parameters based on the input configuration parameters, thereby facilitating the codec device 1000 to meet user needs in all aspects.

[0100] See also Figure 4 and Figure 7 In some implementations, the configuration parameters further include encoding quantization parameters and decoding quantization parameters. Step 012: encoding the input data according to the encoding configuration parameters to generate encoded data, including:

[0101] Step 0121: down-sampling the input data according to the encoding sampling ratio to generate encoding intermediate data;

[0102] Step 0122: quantize the encoded intermediate data according to the encoding quantization parameter to generate encoded data.

[0103] Step 013: Decoding the encoded data according to the decoding configuration parameters to generate decoded data, including:

[0104] Step 0131: Dequantize the encoded data according to the decoded quantization parameter to generate decoded intermediate data;

[0105] Step 0132: Upsample the decoded intermediate data according to the decoding sampling ratio to generate decoded data.

[0106] Specifically, the configuration parameters may also include a coding quantization parameter and a decoding quantization parameter. The coding quantization parameter may be a parameter required for the quantization module 40 to use when performing quantization, and the decoding quantization parameter may be a parameter required for the inverse quantization module 220 to use when performing inverse quantization. For example, the coding quantization parameter and the decoding quantization parameter are the size of the vector quantization codebook and the length of the codeword. The coding quantization parameter and the decoding quantization parameter may also be configured according to different structural delay requirements. Among them, the quantization module 40 and the inverse quantization module 220 may also use heterogeneous quantizers (e.g., some quantizers use scalar quantization, and some quantizers use vector quantization).

[0107] The encoder 100 can obtain the encoding quantization parameter and the encoding sampling rate. Then the downsampling module 10 downsamples the input data according to the encoding sampling rate to generate encoding intermediate data. After obtaining the encoding intermediate data, the encoder 100 inputs the encoding intermediate data into the quantization module 40, and the quantization module 40 quantizes the encoding intermediate data according to the encoding quantization parameter to generate encoding data.

[0108] The decoder 200 may obtain the decoding quantization parameter and the decoding sampling rate. After obtaining the encoded data, the decoder 200 inputs the encoded data into the dequantization module 220, and the dequantization module 220 dequantizes the encoded data according to the decoding quantization parameter to generate decoded intermediate data. After obtaining the decoded intermediate data, the decoder 200 may input the decoded intermediate data into the upsampling module 210, and the upsampling module 210 upsamples the decoded intermediate data according to the decoding sampling rate to generate decoded data.

[0109] In this way, a configurable encoding quantization parameter can be added to the quantization module 40, and a configurable decoding quantization parameter can be added to the inverse quantization module 220 to further deepen the function of changing the structural delay, so that after the codec device 1000 changes the encoding quantization parameter, decoding quantization parameter, encoding sampling rate and decoding quantization parameter, the structural delay of the codec device 1000 can better meet the structural delay requirements.

[0110] See also Figure 4 and Figure 8 In some embodiments, the configuration parameters further include convolution layer parameters in the downsampling layer and convolution layer parameters in the upsampling layer, and the convolution layer parameters include at least weight parameters and bias parameters. Step 012: Encoding the input data according to the encoding configuration parameters to generate encoded data, including:

[0111] Step 0123: Encode the input data according to the encoding sampling ratio and the convolution layer parameters in the downsampling layer to generate encoded data;

[0112] Step 013: decoding the encoded data according to the decoding sampling rate to generate decoded data, including:

[0113] Step 0133: Decode the encoded data according to the decoding sampling rate and the convolution layer parameters in the upsampling layer to generate decoded data.

[0114] Specifically, the downsampling module 10 includes a first convolution unit 11 and a first residual unit 12. The first convolution unit 11 can be used to perform convolution on the input data, and the first residual unit 12 can also perform feature extraction on the input data. The upsampling module 210 includes a deconvolution unit 211 and a second residual unit 212. The deconvolution unit 211 can be used to perform deconvolution on the input data, and the second residual unit 212 can also perform feature extraction on the input data.

[0115] The configuration parameters also include the convolution layer parameters in the downsampling module 10 and the convolution layer parameters in the upsampling module 210, and the convolution layer parameters include at least weight parameters and bias parameters, that is, the convolution layer parameters in the downsampling module 10 and the convolution layer parameters in the upsampling module 210, such as the weight parameters and bias parameters in the downsampling module 10, and the convolution layer parameters in the upsampling module 210, which can also be configured according to different structural delay requirements. The convolution layer parameters in the downsampling module 10 may be other parameters of the downsampling module 10 in addition to the encoding sampling rate when downsampling, such as the size of the convolution kernel. The convolution layer parameters in the upsampling module 210 may be other parameters of the upsampling module 210 in addition to the encoding sampling rate when upsampling, such as the size of the convolution kernel.

[0116] The encoder 100 may input the input feature vector corresponding to the input data into the first residual unit 12 for a first convolution operation to convolve the input feature vector, thereby generating a first intermediate feature vector. Then, the encoder 100 inputs the first intermediate feature vector into the first convolution unit 11, and the first convolution unit 11 performs a second convolution operation on the intermediate feature vector according to the encoding sampling rate and the convolution layer parameters of the first convolution unit 11 to generate encoded data.

[0117] After obtaining the decoded data, the decoder 200 inputs the decoded data into the deconvolution unit 211, and the deconvolution unit 211 performs a third convolution operation on the encoded data according to the decoding sampling rate and the convolution layer parameters of the deconvolution unit 211 to deconvolve the encoded data and generate a second intermediate feature vector. Then, the decoder 200 inputs the second intermediate feature vector into the second residual unit 212, and the second residual unit 212 performs a fourth convolution operation on the second intermediate feature vector to generate decoded data.

[0118] In this way, configurable convolution layer parameters can be added to the downsampling module 10, and configurable convolution layer parameters can be added to the upsampling module 210 to further deepen the function of changing the structural delay of the system, so that after the codec device 1000 changes the convolution layer parameters of the first convolution unit 11, the convolution layer parameters of the deconvolution unit 211, the encoding sampling rate and the decoding sampling rate, the structural delay of the codec device 1000 can better meet the structural delay requirements.

[0119] At the same time, the residual unit is a component widely used in the field of deep learning, and is often used to enhance the expressiveness of neural networks and improve performance. Its function is to extract useful features from the input data and then output the results. The residual unit can be used to enhance the expressiveness of deep neural networks, reduce the impact of problems such as gradient disappearance and gradient explosion, and effectively alleviate the overfitting problem. By using residual units, the codec 1000 can build a deeper neural network and achieve a higher degree of abstraction, while reducing network training time and computing resource consumption.

[0120] See also Figure 4 In some embodiments, the configuration parameters include convolutional layer parameters in the downsampling module 10 of the first target layer and convolutional layer parameters in the upsampling module 210 of the second target layer, the first target layer is the downsampling module 10 in which the encoding sampling rate changes in different configuration parameters, and the second target layer is the upsampling module 210 in which the decoding sampling rate changes in different configuration parameters.

[0121] Specifically, when there are multiple downsampling modules 10, the configuration parameters may be the convolution layer parameters of some of the downsampling modules 10 in the downsampling modules 10, that is, the convolution layer parameters of some of the downsampling modules 10 are configurable, and the convolution layer parameters of the other downsampling modules 10 are fixed. For example, in the downsampling module 10 corresponding to structural delay A, the convolution layer parameters of the downsampling module 10 are 2, 4, 4, and 6, respectively, and in the downsampling module 10 corresponding to structural delay B, the convolution layer parameters of the downsampling module 10 are 2, 4, 4, and 8, respectively. Then, in the two downsampling modules 10, the convolution layer parameters of the downsampling modules 10 with convolution layer parameters of 2, 4, and 4 are fixed, the convolution layer parameters of the downsampling module 10 corresponding to the convolution layer parameter of 6 are configurable, and the convolution layer parameters of the downsampling module 10 corresponding to the convolution layer parameter of 8 are configurable. It can be understood that the downsampling modules 10 whose encoding sampling ratios change in different configuration parameters are the first target layer, so that the configuration parameters include the convolution layer parameters in the downsampling module 10 of the first target layer.

[0122] Similarly, when there are multiple upsampling modules 210, the configuration parameters may be the convolution layer parameters of some of the upsampling modules 210 in the upsampling modules 210, that is, the convolution layer parameters of some of the upsampling modules 210 are configurable, and the convolution layer parameters of the other upsampling modules 210 are fixed. Among them, the upsampling modules 210 whose decoding sampling ratios change in different configuration parameters are the second target layer, so that the configuration parameters include the convolution layer parameters in the upsampling modules 210 of the second target layer.

[0123] In this way, the encoding device can determine the first target layer according to the specific convolution layer parameters of different downsampling modules 10, and the decoding device can determine the second target layer according to the specific convolution layer parameters of different upsampling modules 210, so as to more flexibly adjust the configuration parameters, so that the structural delay of the encoding and decoding device 1000 can meet different structural delay requirements.

[0124] See also Figure 4 and Fig. 9 In some implementations, the audio encoding and decoding method further includes:

[0125] Step 015: Obtain a current input sample, and randomly select any configuration parameter to configure the encoder 100 and the decoder 200, where the current input sample is any sample in a preset sample set;

[0126] Step 016: The configured encoder 100 and decoder 200 encode and decode the current input sample to obtain the current output sample;

[0127] Step 017: Determine the loss values ​​of the encoder 100 and the decoder 200 according to the current input sample and the current output sample;

[0128] Step 018: Adjust the encoder 100 and the decoder 200 according to the loss value to update the network parameters of the encoder 100 and the decoder 200 until the encoder 100 and the decoder 200 are trained to convergence.

[0129] Specifically, the coding device involves many network parameters in the coding and decoding process, and all parameters except the configuration parameters need to be preset. And because these parameters (i.e., parameters except the configuration parameters) are fixed, it is also necessary to preset the parameters according to different structural delay requirements, so as to ensure that the structural delay requirements are met while ensuring the parameters of the coding and decoding effect.

[0130] Then, the encoding and decoding device 1000 can select any sample in the sample set as the current input sample. After the encoding and decoding device 1000 randomly selects any configuration parameter to configure the encoder 100 and the decoder 200, the current input sample can be input into the encoder 100 to train the encoder 100 and the decoder 200 using the configuration parameters corresponding to different structural delays. The encoder 100 and the decoder 200 that have completed the parameter configuration will perform encoding and decoding processing according to the current data sample to obtain the current output sample to complete the reconstruction of the data. It can be understood that the smaller the gap between the current output sample and the current input sample, the better the encoding and decoding effect. Then, the loss value of the encoder 100 and the decoder 200 can be determined according to the current input sample and the reconstructed current output sample to determine the encoding and decoding effect according to the encoder 100 and the decoder 200, and the encoder 100 and the decoder 200 can be adjusted according to the loss value to update the network parameters of the encoder 100 and the decoder 200. It should be noted that these updated network parameters are network parameters other than the configuration parameters in the network parameters corresponding to the encoder 100 and the decoder 200. Then, the encoder 100 and decoder 200 with updated parameters are trained again. When the loss value of a certain training is less than the preset loss value threshold, it can be considered that the encoding and decoding effect of the encoder 100 and decoder 200 is better, and the network parameters at this time can be considered to be adapted to different structural delay requirements. Therefore, it can be considered that the encoder 100 and decoder 200 are trained to convergence, and the network parameters in the converged encoder 100 and decoder 200 are the final network parameters.

[0131] In this way, the codec device 1000 can use the configuration parameters corresponding to different structural delays to train the encoder 100 and the decoder 200, so that the network parameters corresponding to the encoder 100 and the decoder 200 after convergence except for the configuration parameters can adapt to different structural delay requirements. Therefore, when facing different structural delay requirements, the codec device 1000 only needs to adjust the configuration parameters to ensure the coding and decoding effect while meeting different structural delay requirements, without adjusting other network parameters except the configuration parameters, thereby quickly adjusting the structural delay of the codec device 1000.

[0132] See also Figure 4 and Fig.10 In some embodiments, step 012: encoding the input data according to the encoding configuration parameters to generate encoded data, includes:

[0133] Step 0124: Encode the input data according to the encoding sampling magnification and the updated network parameters of the encoder 100 to generate encoded data;

[0134] Step 013: decoding the encoded data according to the decoding sampling rate to generate decoded data, including:

[0135] Step 0134: Decode the encoded data according to the decoding sampling rate and the updated network parameters of the decoder 200 to generate decoded data.

[0136] Specifically, after the encoder 100 and the decoder 200 are trained to convergence, it can be confirmed that the updated network parameters are adapted to different structural delay requirements. Therefore, when the codec device 1000 receives input data and obtains the current configuration parameters, the encoder 100 can encode the input data according to the encoding sampling rate in the current configuration parameters and according to the updated network parameters of the encoder 100 to generate encoded data. The decoder 200 can decode the encoded data according to the decoding sampling rate in the current configuration parameters and according to the updated network parameters of the decoder 200 to generate decoded data. In this way, the codec device 1000 can perform encoding and decoding according to the updated network parameters and the current configuration parameters to meet the current structural delay requirements on the one hand, and ensure that the encoding and decoding effect is always better on the other hand.

[0137] See also Figure 4 In some embodiments, the structural delay of the structural delay information is determined based on at least one of the content complexity and real-time parameters of the input data.

[0138] Specifically, content complexity and structural delay are positively correlated. The higher the content complexity, the more information is transmitted, the longer the encoding and decoding time is, and the longer the structural delay is. Among them, the content complexity can be determined according to at least one of the sampling rate, the number of channels, and the number of sound source types of the encoded data. It can be understood that the higher the sampling rate, the more channels, the more sound source types, and the higher the content complexity. The content complexity can be determined according to one of the sampling rate, the number of channels, and the number of sound source types of the encoded data, for example, the content complexity is determined according to the sampling rate. Alternatively, the content complexity can be determined according to any two of the sampling rate, the number of channels, and the number of sound source types of the encoded data, for example, according to the sampling rate and the number of channels, or according to the number of channels and the number of sound source types. Alternatively, the content complexity can be determined according to the sampling rate, the number of channels, and the number of sound source types of the encoded data.

[0139] There is a negative correlation between structural delay and real-time parameters. The higher the real-time performance, the larger the real-time parameter, and the shorter the structural delay. For example, when users are in a video conference, they need a higher real-time performance. At this time, the structural delay needs to be shorter to ensure the quality of communication. When users are listening to music, the real-time performance requirements are not high. At this time, the structural delay can be set longer to ensure the encoding and decoding effect, thereby ensuring the quality of the user's listening to the music.

[0140] In this way, the structural delay can be determined according to the content complexity or the real-time parameter, or according to both the content complexity and the real-time parameter, to ensure that the structural delay of the encoding and decoding device 1000 can meet the needs of the user.

[0141] See also Figure 4 In some embodiments, an activation layer may be added after each convolution module.

[0142] Specifically, the encoder 100 and the decoder 200 may also include an activation module, which corresponds to the activation layer one by one. An activation function, such as relu, leaky-relu or attention module, is a function added to an artificial neural network to help the network learn complex patterns in data. An activation module containing an activation function can be used to add a nonlinear operation (i.e., a nonlinear mapping operation) after the convolution operation, making the output of the neural network more complex and more expressive.

[0143] Therefore, after the convolution module performs the convolution operation, the activation module can be used to perform a nonlinear mapping operation on the output data of the convolution module. For example, an activation module can be added between the first convolution module 20 and the upsampling module 10 to perform a nonlinear mapping operation on the first feature vector, and the first feature vector subjected to the nonlinear mapping operation is input into the upsampling module 10. In this way, the encoder 100 and the decoder 200 can use the activation function to increase the nonlinear factors in the nonlinear operation, so that the encoder 100 and the decoder 200 can express more complex features.

[0144] See also Figure 4 In order to better implement the audio encoding and decoding method of the embodiment of the present application, the embodiment of the present application also provides a coding and decoding device 1000. The coding and decoding device 1000 may include an encoder 100 and a decoder 200. The encoder 100 is used to encode input data according to the encoding configuration parameters of the current configuration parameters to generate encoded data. The decoder 200 is used to decode the encoded data according to the decoding configuration parameters of the current configuration parameters to generate decoded data. The current configuration parameters are any one of the preset multiple configuration parameters. The encoding configuration parameters at least include the encoding sampling rate, and the decoding configuration parameters at least include the decoding sampling rate. The structural delays corresponding to the various configuration parameters are different.

[0145] See also Fig.11 The computer device 300 of the embodiment of the present application includes a processor 310, a memory 320 and a computer program, wherein the computer program is stored in the memory 320 and executed by the processor 310, and the computer program includes instructions for executing the audio coding method of any of the above embodiments.

[0146] Optionally, the computer device 300 can be any device with image processing capabilities, such as a server or a terminal device (such as a mobile phone, a tablet computer, a display device, a laptop computer, a smart watch, a head-mounted display device, a game console, etc.).

[0147] like Fig.11 As shown, the processor 310 included in the computer device 300 is a central processing unit (CPU), and the memory 320 includes a read-only memory 321 (ROM) and a random access memory 322 (RAM). The central processor can perform various appropriate actions and processes according to the program stored in the read-only memory 321 or the program loaded from the storage part 380 to the random access memory 322. In the random access memory 322, various programs and data required for system operation are also stored. The central processor, the read-only memory 321 and the random access memory 322 are connected to each other through a bus 330. The input / output interface 340 (Input / Output interface, i.e., I / O interface) is also connected to the bus 330.

[0148] The following components are connected to the input / output interface 340: an input section 350 including a keyboard, a mouse, etc.; an output section 360 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker; a storage section 370 including a hard disk, etc.; and a communication section 380 including a network interface card such as a LAN card, a modem, etc. The communication section 380 performs communication processing via a network such as the Internet. A drive 390 is also connected to the input / output interface 340 as needed. A removable medium 391, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 390 as needed so that a computer program read therefrom is installed into the storage section 370 as needed.

[0149] In particular, according to an embodiment of the present application, the process described in each method flow chart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer readable medium, and the computer program contains a program code for executing the method shown in the flow chart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 380, and / or installed from the removable medium 391. When the computer program is executed by the central processing unit, various functions defined in the system of the present application are executed.

[0150] In particular, according to an embodiment of the present application, the process described in each method flow chart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains a program code for executing the method shown in the flow chart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication part, and / or installed from a removable medium. When the computer program is executed by a processor, various functions defined in the system of the present application are executed.

[0151] The computer program product of the embodiments of the present application includes a computer program. When the computer program is executed by the processor 310, the audio encoding and decoding method of any of the above embodiments is executed.

[0152] See also Fig.12 The embodiment of the present application also provides a computer-readable storage medium 400 on which a computer program 410 is stored. When the computer program 410 is executed by a processor 420 (e.g., the processor 310 of the computer device 300), the steps of the audio encoding and decoding method of any of the above-mentioned embodiments are implemented. For the sake of brevity, they are not repeated here.

[0153] In the description of this specification, the descriptions with reference to the terms "certain embodiments", "in an example", "exemplarily", etc., mean that the specific features, structures, materials or characteristics described in conjunction with the embodiments or examples are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, unless they are contradictory.

[0154] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code that includes one or more executable instructions for implementing the steps of a specific logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by technicians in the technical field to which the embodiments of the present application belong.

[0155] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limitations on the present application. Ordinary technicians in this field can change, modify, replace and modify the above embodiments within the scope of the present application.

Claims

1. An audio coding and decoding method, characterized in that: The method comprises: Acquire a current configuration parameter, where the current configuration parameter is any one of a plurality of preset configuration parameters, where the configuration parameter includes at least one of an encoding configuration parameter and a decoding configuration parameter, where when used for encoding, the configuration parameter includes the encoding configuration parameter, and when used for decoding, the configuration parameter includes the decoding configuration parameter, where the encoding configuration parameter at least includes an encoding sampling magnification, and the decoding configuration parameter at least includes a decoding sampling magnification, and where at least corresponding structural delays of the configuration parameters are different; Encoding the input data according to the encoding configuration parameters to generate encoded data; The encoded data is decoded according to the decoding configuration parameters to generate decoded data. When encoding and decoding are performed based on the respective configuration parameters, among the network parameters of the encoding and decoding device, other network parameters except the configuration parameters remain unchanged.

2. The audio encoding and decoding method according to claim 1, characterized in that: The method further comprises: An input configuration parameter is obtained to determine the current configuration parameter, wherein the input configuration parameter includes at least one of a codec sampling ratio, a structural delay, and a signal sampling rate.

3. The audio encoding and decoding method according to claim 2, characterized in that: The obtaining of input configuration parameters to determine the current configuration parameters includes: Determine one or more target configuration parameters according to the input configuration parameters, the target configuration parameter being any of the configuration parameters; An input parameter index number is obtained to determine the current configuration parameter corresponding to the parameter index number in one or more target configuration parameters.

4. The audio encoding and decoding method according to claim 1, characterized in that: The number of parameters corresponding to the encoding sampling rate is determined according to the number of downsampling layers preset by the encoder, and the number of parameters corresponding to the decoding sampling rate is determined according to the number of upsampling layers preset by the decoder, and the encoding product of each encoding sampling rate is the same as the decoding product of each decoding sampling rate.

5. The audio encoding and decoding method according to claim 4, characterized in that: The number of downsampling layers used for encoding is the same as the number of upsampling layers used for decoding, and the encoding sampling ratio and the decoding sampling ratio correspond one to one.

6. The audio encoding and decoding method according to claim 4, characterized in that: The number of downsampling layers used for encoding and the number of upsampling layers used for decoding are not the same.

7. The audio encoding and decoding method according to claim 1, characterized in that: The configuration parameters also include encoding quantization parameters and decoding quantization parameters, and encoding the input data according to the encoding configuration parameters to generate encoded data includes: Downsampling the input data according to the encoding sampling magnification to generate encoding intermediate data; quantizing the encoded intermediate data according to the encoding quantization parameter to generate the encoded data; The step of decoding the encoded data according to the decoding configuration parameters to generate decoded data comprises: De-quantizing the encoded data according to the decoded quantization parameter to generate decoded intermediate data; The decoded intermediate data is up-sampled according to the decoded sampling magnification to generate the decoded data.

8. The audio encoding and decoding method according to claim 1, characterized in that: The configuration parameters also include convolution layer parameters in the downsampling layer and convolution layer parameters in the upsampling layer, the convolution layer parameters at least include weight parameters and bias parameters, and encoding the input data according to the encoding configuration parameters to generate encoded data includes: Encoding the input data according to the encoding sampling magnification and the convolution layer parameters in the downsampling layer to generate encoded data; The step of decoding the encoded data according to the decoding configuration parameters to generate decoded data comprises: The encoded data is decoded according to the decoding sampling rate and the convolution layer parameters in the upsampling layer to generate the decoded data.

9. The audio encoding and decoding method according to claim 8, characterized in that: The configuration parameters include convolution layer parameters in the downsampling layer of the first target layer and convolution layer parameters in the upsampling layer of the second target layer, the first target layer being the downsampling layer in which the encoding sampling ratio changes in different configuration parameters, and the second target layer being the upsampling layer in which the decoding sampling ratio changes in different configuration parameters.

10. The audio encoding and decoding method according to claim 1, characterized in that: The method further comprises: Obtaining a current input sample, and randomly selecting any of the configuration parameters to configure the encoder and the decoder, the current input sample being any sample in a preset sample set; Performing encoding and decoding processing on the current input sample by the configured encoder and decoder to obtain a current output sample; Determine loss values ​​of the encoder and the decoder according to the current input sample and the current output sample; The encoder and the decoder are adjusted according to the loss value to update network parameters of the encoder and the decoder until the encoder and the decoder are trained to converge.

11. The audio encoding and decoding method according to claim 10, characterized in that: The step of encoding the input data according to the encoding configuration parameters to generate encoded data comprises: Encoding input data according to the encoding sampling magnification and the updated network parameters of the encoder to generate encoded data; The step of decoding the encoded data according to the decoding configuration parameters to generate decoded data comprises: The encoded data is decoded according to the decoding sampling magnification and the updated network parameters of the decoder to generate decoded data.

12. The audio encoding and decoding method according to claim 1, characterized in that: The encoder obtains the encoding configuration parameters according to the set structural delay information, and the decoder obtains the decoding configuration parameters according to the structural delay information in the code stream formed by the encoded data.

13. The audio encoding and decoding method according to claim 12, characterized in that: The structural delay of the structural delay information is determined according to at least one of the content complexity and real-time parameters of the input data.

14. The audio encoding and decoding method according to claim 13, characterized in that: The structural delay is positively correlated with the content complexity, and negatively correlated with the real-time parameter. The content complexity is determined according to at least one of the sampling rate, the number of channels, and the number of sound source types of the input data.

15. A coding and decoding device, characterized in that: include: An encoder, configured to encode input data according to encoding configuration parameters of current configuration parameters to generate encoded data; A decoder, used to decode the encoded data according to the decoding configuration parameters of the current configuration parameters to generate decoded data, wherein the current configuration parameters are any one of a plurality of preset configuration parameters, the encoding configuration parameters at least include the encoding sampling rate, the decoding configuration parameters at least include the decoding sampling rate, and the structural delays corresponding to the various configuration parameters are different.

16. The encoding and decoding device according to claim 15, characterized in that: The configuration parameters also include encoding quantization parameters and decoding quantization parameters. The encoder includes a downsampling module and a quantization module. The downsampling module is used to downsample the input data according to the encoding sampling ratio to generate encoding intermediate data. The quantization module is used to quantize the encoded intermediate data according to the encoding quantization parameter to generate the encoded data; The decoder includes an upsampling module and a dequantization module; The dequantization module is used to dequantize the encoded data according to the decoded quantization parameter to generate decoded intermediate data; the upsampling module is used to upsample the decoded intermediate data according to the decoded sampling ratio to generate the decoded data.

17. The encoding and decoding device according to claim 15, characterized in that: The encoder includes a plurality of downsampling modules, each of which includes a first convolution unit and a first residual unit; the decoder includes a plurality of upsampling modules, each of which includes a deconvolution unit and a second residual unit; and the configuration parameters further include a convolution layer parameter of the first convolution unit and a convolution layer parameter of the deconvolution unit; The first residual unit is used to perform a first convolution operation on the input feature vector corresponding to the input data to generate a first intermediate feature vector; the first convolution unit is used to perform a second convolution operation on the intermediate feature vector according to the encoding sampling magnification and the convolution layer parameters of the first convolution unit to generate the encoded data, The deconvolution unit is used to perform a third convolution operation on the encoded data according to the decoding sampling rate and the convolution layer parameters of the deconvolution unit to generate a second intermediate feature vector, and the second residual unit is used to perform a fourth convolution operation on the second intermediate feature vector to generate the decoded data.

18. A computer device, characterized in that: include: Processor, memory; and A computer program, wherein the computer program is stored in the memory and executed by the processor, and the computer program includes instructions for executing the audio coding and decoding method according to any one of claims 1 to 14.

19. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the audio encoding and decoding method according to any one of claims 1 to 14 is implemented.

20. A non-volatile computer-readable storage medium containing a computer program, wherein when the computer program is executed by a processor, the processor executes the audio encoding and decoding method according to any one of claims 1 to 14.