Decoding method, coding method, device, equipment, product and storage medium
By using packet convolution technology in deep learning codec solutions, the computational complexity in the decoding process is reduced, and the problem that the existing technology is difficult to apply on mobile devices is solved, and efficient encoding and decoding effects are achieved.
Patent Information
- Application Number
- CN202311473488.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-06
- Publication Date
- 2025-05-06
AI Technical Summary
The existing deep learning codec solutions have high computational complexity and are difficult to effectively support and apply on mobile devices.
By using packet convolution during the decoding process, at least part of the convolution operation is realized, the computational complexity is reduced and it is suitable for mobile devices.
Significantly reduce the complexity of decoding calculation, improve coding efficiency, and realize practical use on mobile devices.
Smart Images

Figure CN119940418A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of coding and decoding technology, and more specifically, to a decoding method, an encoding method, an encoding device, a decoding device, a computer device, a computer program product and a non-volatile computer-readable storage medium. Background Art
[0002] In recent years, deep learning solutions have been widely used in signal processing technologies of different dimensions (such as audio, image, and video). In the deep learning-based encoding and decoding solution, the encoding device can convert the original signal into encoded data. The decoding device obtains the transmitted information according to the encoded data, and then decodes the encoded data to obtain the final reconstructed signal.
[0003] However, the current decoding scheme has high computational complexity and cannot be well supported by mobile devices and put into practical use. Summary of the invention
[0004] The embodiments of the present application provide a decoding method, an encoding method, an encoding device, a decoding device, a computer device, a computer program product and a non-volatile computer-readable storage medium, which implement at least part of the convolution operation in the decoding process by means of group convolution, thereby reducing the computational complexity of the decoding, so that it can be better supported by mobile devices and put into practical use.
[0005] The decoding method of the embodiment of the present application includes performing a nonlinear transformation operation on the encoded data to obtain decoded data; the nonlinear transformation operation includes a convolution operation, and the convolution operation of at least part of the convolution layer is implemented by grouped convolution, and the number of groups of the convolution layer performing grouped convolution is determined according to the number of channels of the convolution layer.
[0006] In some embodiments, the number of groups of the convolutional layer performing grouped convolution decreases along the direction from input to output.
[0007] In some embodiments, the nonlinear transformation operation includes a first convolution operation, an upsampling operation, and a second convolution operation, and the nonlinear transformation operation is performed on the encoded data to obtain decoded data, including: performing the first convolution operation on the encoded data to obtain a first feature vector; performing the upsampling operation on the first feature vector to obtain a second feature vector; performing the second convolution operation on the second feature vector to obtain the decoded data.
[0008] In some embodiments, performing the upsampling operation on the first feature vector to obtain the second feature vector includes: performing the upsampling operation on the first feature vector multiple times to obtain the second feature vector, the number of upsampling operations is determined according to the number of upsampling layers at the decoding end, the first feature vector is the input feature vector of the first upsampling operation, the second feature vector is the output feature vector of the last upsampling operation, and the input feature vector of each upsampling operation is the output feature vector of the last upsampling operation.
[0009] In some embodiments, the upsampling operation includes a third convolution operation, and the number of groups when the first convolution operation, the third convolution operation, and the second convolution operation perform group convolution increases. The number of upsampling operations is determined according to the number of upsampling layers at the decoding end, and the number of groups when the third convolution operation performs group convolution increases along the direction from input to output.
[0010] In some embodiments, along the direction from input to output, the ratio between the number of groups of two adjacent convolution layers performing grouped convolution is 2.
[0011] In some embodiments, the decoding method further includes: determining the number of groups of the convolutional layer for performing group convolution based on at least one of the number of channels of the convolutional layer, a preset number of codec channels, and a codec bit rate, wherein the number of groups is positively correlated with the number of channels of the convolutional layer, the number of codec channels, and the codec bit rate.
[0012] In some embodiments, the decoding method further includes determining the number of groups of the convolutional layer for performing group convolution based on the number of channels of the convolutional layer and the content complexity of the encoded data, wherein the number of groups is positively correlated with the number of channels of the convolutional layer and the content complexity.
[0013] In some embodiments, the content complexity is determined according to at least one of a sampling rate, a number of channels, and a number of sound source types of the encoded data.
[0014] The encoding method of the embodiment of the present application includes performing a nonlinear transformation operation on input data to obtain encoded data; the nonlinear transformation operation includes a convolution operation, and the convolution operation of at least part of the convolution layer is implemented by grouped convolution, and the number of groups of the convolution layer performing grouped convolution is determined according to the number of channels of the convolution layer.
[0015] The encoding device of the embodiment of the present application includes a first nonlinear transformation module. The first nonlinear transformation module is used to perform a nonlinear transformation operation on input data to obtain encoded data; the first nonlinear transformation module includes a convolution layer, and the convolution operation of at least part of the convolution layer is implemented by grouped convolution, and the number of groups of the convolution layer performing grouped convolution is determined according to the number of channels of the convolution layer.
[0016] The decoding device of the embodiment of the present application includes a second nonlinear transformation module, which is used to perform a nonlinear transformation operation on the encoded data to obtain decoded data; the second nonlinear transformation module includes a convolutional layer, and the convolution operation of at least part of the convolutional layer is implemented by grouped convolution, and the number of groups of the convolutional layer performing grouped convolution is determined according to the number of channels of the convolutional layer.
[0017] The computer device of an embodiment of the present application includes a processor, a memory and a computer program, wherein the computer program is stored in the memory and executed by the processor, and the computer program includes instructions for executing the decoding method or encoding method described in any of the above embodiments.
[0018] The computer program product of the embodiments of the present application includes a computer program, and when the computer program is executed by a processor, the decoding method or the encoding method described in any of the above embodiments is performed.
[0019] The non-volatile computer-readable storage medium of the embodiment of the present application includes a computer program. When the computer program is executed by a processor, the processor executes the decoding method or encoding method described in any of the above embodiments.
[0020] The encoding method, decoding method, encoding device, decoding device, computer equipment, computer program product and non-volatile computer-readable storage medium of the embodiments of the present application implement at least part of the convolution operations in the nonlinear transformation operation in the decoding process through grouped convolution. Compared with ordinary convolution, the number of channels after convolution is large. The present application can significantly reduce the number of channels through grouped convolution, thereby reducing the computational complexity.
[0021] The number of groups of the convolution layer performing grouped convolution is determined according to the number of channels of the convolution layer. Compared with all convolution layers using a larger number of groups for grouped convolution, the number of groups is set more reasonably, thereby reducing the computational complexity of the convolution layer and extracting more useful features to ensure the decoding effect.
[0022] Additional aspects and advantages of the embodiments of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the embodiments of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the description of the embodiments in conjunction with the following drawings, in which:
[0024] Figure 1 It is a schematic diagram of an application scenario of a decoding method in some embodiments of the present application;
[0025] Figure 2 This is an example diagram of the basic process of the encoding and decoding scheme of certain implementation methods of the present application;
[0026] Figure 3 It is a structural schematic diagram of an encoding device and a decoding device of a decoding method in certain embodiments of the present application;
[0027] Figure 4 is a flowchart of a decoding method in some embodiments of the present application;
[0028] Figure 5 is a schematic diagram of a decoding method according to some embodiments of the present application;
[0029] Figure 6 is a flowchart of a decoding method in some embodiments of the present application;
[0030] Figure 7 is a flowchart of a decoding method in some embodiments of the present application;
[0031] Figure 8 is a flowchart of a decoding method in some embodiments of the present application;
[0032] Fig. 9 is a flowchart of a decoding method in some embodiments of the present application;
[0033] Fig.10 It is a flowchart of the encoding method of certain embodiments of the present application;
[0034] Fig.11 It is a flowchart of the encoding method of certain embodiments of the present application;
[0035] Fig.12 It is a flowchart of the encoding method of certain embodiments of the present application;
[0036] Fig.13 is a schematic diagram of the structure of a computer device of certain embodiments of the present application;
[0037] Fig.14 It is a schematic diagram of the connection status of a non-volatile computer-readable storage medium and a processor in certain embodiments of the present application. DETAILED DESCRIPTION
[0038] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals represent the same or similar elements or elements having the same or similar functions from beginning to end. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the embodiments of the present application, and cannot be understood as limiting the embodiments of the present application.
[0039] To facilitate understanding of this application, the terms appearing in this application are explained below:
[0040] Encoding and decoding: The encoding process is to compress the input data into smaller data, and the decoding process is to restore the smaller data to the original size. The encoded smaller data is used for network transmission and occupies less bandwidth.
[0041] Sampling rate: The sampling rate describes the number of data contained in a unit of time (1 second). For example, a 16k sampling rate contains 16,000 sampling points, and each sampling point corresponds to a short integer.
[0042] Codebook: A collection of multiple vectors. The encoding device and the decoding device both store the same codebook.
[0043] Quantization: Find the closest vector in the codebook for the input vector, return it as a replacement for the input vector, and return the corresponding codebook index position.
[0044] Quantizer: The quantizer is responsible for quantization and updating the vectors in the codebook.
[0045] Figure 1 A schematic diagram schematically shows an exemplary system architecture to which the technical solution of the embodiments of the present application can be applied.
[0046] like Figure 1 As shown, the system architecture includes multiple terminal devices, and the terminal devices can communicate with each other through, for example, a network. For example, the system architecture may include a first terminal device 101 and a second terminal device 102 interconnected through a network. Figure 1 In the embodiment of the present invention, the first terminal device 101 and the second terminal device 102 perform unidirectional data transmission.
[0047] For example, the first terminal device 101 can encode audio and video data (such as an audio and video data stream collected by the terminal device) for transmission to the second terminal device 102 via the network. The encoded audio and video data is transmitted in the form of one or more encoded audio and video streams. The second terminal device 102 can receive the encoded audio and video data from the network, decode the encoded audio and video data to restore the audio and video data, and play or display content based on the restored audio and video data.
[0048] In one embodiment of the present application, the system architecture may include a third terminal device 103 and a fourth terminal device 104 that perform bidirectional transmission of encoded audio and video data, and the bidirectional transmission may occur, for example, during an audio and video conference. For bidirectional data transmission, each of the third terminal device 103 and the fourth terminal device 104 may encode audio and video data (e.g., an audio and video data stream collected by the terminal device) for transmission to the other terminal device of the third terminal device 103 and the fourth terminal device 104 through a network. Each of the third terminal device 103 and the fourth terminal device 104 may also receive the encoded audio and video data transmitted by the other terminal device of the third terminal device 103 and the fourth terminal device 104, and may decode the encoded audio and video data to restore the audio and video data, and play or display content according to the restored audio and video data.
[0049] exist Figure 1 In the embodiment of the present invention, the first terminal device 101, the second terminal device 102, the third terminal device 103 and the fourth terminal device 104 may be servers, mobile terminals, personal computers, etc., but the principles disclosed in the present application may not be limited thereto. The embodiments disclosed in the present application are applicable to laptop computers, tablet computers, media players and / or dedicated audio and video conferencing equipment.
[0050] The network represents any number of networks that transmit the encoded audio and video data between the first terminal device 101, the second terminal device 102, the third terminal device 103 and the fourth terminal device 104, including, for example, wired and / or wireless communication networks. The communication network can exchange data in a circuit switching and / or packet switching channel. The network may include a telecommunications network, a local area network, a wide area network and / or the Internet. For the purpose of this application, unless explained below, the architecture and topology of the network may be irrelevant to the operation disclosed in this application.
[0051] Figure 2 The following diagram shows an example of the basic process of an end-to-end audio encoding and decoding solution. During encoding, the encoder at the data transmitter first encodes the input audio signal to generate a binary code stream, and then the data transmitter sends the binary code stream to the data receiver. After receiving the binary code stream, the data receiver decodes the binary code stream through the decoder to obtain a reconstructed audio signal.
[0052] See also Figure 3 For ease of understanding, the encoding device 100 (encoding end) and the decoding device 200 (decoding end) of the present application are respectively introduced below.
[0053] The encoding device 100 includes a first non-linear transformation module, a first linear transformation module and a quantization module. The first non-linear transformation module 110, the first linear transformation module 120 and the quantization module 130 are sequentially connected in a direction from input to output.
[0054] The first nonlinear transformation module 110 includes a first convolution module 111, a downsampling module 112 and a second convolution module 113. The first convolution module 111 is used to perform a fourth convolution operation on the original audio data (i.e., the input data of the encoding device 100) to obtain a third eigenvector; then the third eigenvector is used as an input eigenvector of the downsampling module 112, and the downsampling module 112 performs a downsampling operation on the third eigenvector to obtain a fourth eigenvector. Finally, the second convolution module 113 performs a fifth convolution operation on the fourth eigenvector to obtain an output eigenvector of the second convolution module 113.
[0055] The output feature vector of the second convolution module 113 is linearly transformed by the first linear transformation module 120, and then input into the quantization module 130 (such as using a residual-based vector quantizer), which performs a quantization operation and generates a binary code stream based on the quantization result to obtain encoded data.
[0056] For example, the feature vector obtained after the linear transformation operation can be input into the quantizer, so that the vector index corresponding to each feature vector can be queried in the codebook. The vector index can then be transmitted to the data receiving end (i.e., the decoding end), and the data receiving end decodes the vector index through the decoding device 100 to obtain the restored data.
[0057] Optionally, the downsampling module 112 may include multiple downsampling modules 112, which are sequentially connected along the direction from input to output. The multiple downsampling modules 112 may perform multiple downsampling operations on the third feature vector to obtain a fourth feature vector, and the number of downsampling operations is determined according to the number of downsampling layers of the encoding end (i.e., the number of downsampling modules 112). The third feature vector is the input feature vector of the first downsampling module, the fourth feature vector is the output feature vector of the last downsampling module 112, and the output feature vector of the previous downsampling module 112 is used as the input feature vector of the next downsampling module 112.
[0058] In this way, multiple downsampling operations are performed through multiple downsampling modules 112, thereby successively reducing the dimension of the third eigenvector and increasing the number of channels of the fourth eigenvector, thereby obtaining a fourth eigenvector whose dimension and number of channels meet preset requirements.
[0059] Optionally, the downsampling module 112 includes one or more residual units and convolution units. One or more residual units and convolution units are sequentially connected in the direction from input to output. It can be understood that the connection order of the residual unit and the convolution unit can be other orders, such as the convolution unit and one or more residual units are sequentially connected in the direction from input to output, which is not limited here.
[0060] The residual unit of the down-sampling module 112 is used to perform a convolution operation on the input feature vector of the down-sampling module 112 to obtain an input feature vector of the convolution unit.
[0061] Among them, residual units are a widely used component in the field of deep learning, and are often used to enhance the expressiveness of neural networks and improve performance. Its function is to extract useful features from input data and then output the results. Residual units can be used to enhance the expressiveness of deep neural networks, reduce the impact of problems such as gradient vanishing and gradient exploding, and effectively alleviate overfitting problems.
[0062] The convolution unit of the down-sampling module 112 is used to perform a sixth convolution operation on the input feature vector to obtain an input feature vector of the next down-sampling module 112, or to obtain a fourth feature vector.
[0063] The decoding device 200 includes an inverse quantization module 210, a second linear transformation module 220, and a second nonlinear transformation module 230. The inverse quantization module 210, the second linear transformation module 220, and the second nonlinear transformation module 230 are sequentially connected in a direction from input to output.
[0064] The inverse quantization module 210 can perform an inverse quantization operation on the encoded data (ie, the input data of the decoding device 200). For example, after receiving the vector index transmitted by the network, the codebook feature vector corresponding to the vector index can be first queried in the codebook through the quantizer.
[0065] The second linear transformation module 220 may perform a linear transformation operation on the data after the inverse quantization operation to obtain a linearly changed feature vector.
[0066] The second non-linear transformation module 230 can perform a non-linear transformation operation on the linearly transformed feature vector to obtain decoded data. The second non-linear transformation module 230 includes a third convolution module 231 , an up-sampling module 232 and a fourth convolution module 233 .
[0067] The third convolution module 231 is used to perform a first convolution operation on the linearly changed feature vector to obtain a first feature vector; then the first feature vector is used as an input feature vector of the upsampling module 232, and the upsampling module 232 performs an upsampling operation on the first feature vector to obtain a second feature vector. Finally, the fourth convolution module 2333 performs a second convolution operation on the second feature vector to obtain decoded data.
[0068] Optionally, the upsampling module 232 may include multiple upsampling modules 232, which are sequentially connected along the direction from input to output. The multiple upsampling modules 232 may perform multiple upsampling operations on the first feature vector to obtain the second feature vector, and the number of upsampling operations is determined according to the number of upsampling layers at the decoding end (i.e., the number of upsampling modules 232). The first feature vector is the input feature vector of the first upsampling module 232, the second feature vector is the output feature vector of the last upsampling module 232, and the output feature vector of the last upsampling module 232 is used as the input feature vector of the next upsampling module 232.
[0069] In this way, multiple upsampling operations are performed through multiple upsampling modules 232, thereby successively increasing the dimension of the first feature vector and reducing the number of channels of the second feature vector, thereby restoring decoded data that is substantially the same as the input data of the encoding device 200.
[0070] Optionally, the upsampling module 232 includes one or more convolution units and one or more residual units. The convolution unit and the one or more residual units are sequentially connected in the direction from input to output. It can be understood that the connection order of the residual unit and the convolution unit can be other orders, such as one or more residual units and the convolution unit are sequentially connected in the direction from input to output, which is not limited here.
[0071] The convolution unit of the up-sampling module 232 is used to perform a third convolution operation on the input feature vector to obtain an input feature vector of the residual unit of the up-sampling module 232 .
[0072] The residual unit of the up-sampling module 232 is used to perform a convolution operation on the output feature vector of the convolution unit of the up-sampling module 232 to obtain an input feature vector or a second feature vector of the next up-sampling module 232 .
[0073] In one example, in the input stage of the encoding device 100, after receiving the original data of 24 kHz (khz), data sampling can be performed on the original data to be encoded, and an original feature vector with a channel number of 1 and a dimension of 16000 can be obtained; the original feature vector is input into the first convolution module 110, and after a fourth convolution operation, a third feature vector with a channel number of C (such as 32) and a dimension of 16000 can be obtained.
[0074] Optionally, in order to improve encoding efficiency, the encoding device 100 may simultaneously perform encoding processing on a batch of original feature vectors, the number of which is B (B is a positive integer).
[0075] In the downsampling stage of the encoding device 100, multiple downsampling modules 112 can be used to perform multiple downsampling operations on the third eigenvector, and the downsampling ratios S are 2, 4, 5 and 8 respectively, so as to obtain a fourth eigenvector with 512 channels and 50 dimensions.
[0076] In the output stage of the encoding device 100 , the second convolution module 113 performs a fifth convolution operation on the fourth eigenvector to obtain an encoded eigenvector with K channels and 50 dimensions.
[0077] Among them, K is a preset vector quantization dimension, for example, it can be set to 32. Next, the encoded feature vector is linearly transformed and then input into the quantizer, and a binary code stream can be generated according to the quantization result to obtain 75hz encoded data, thereby completing the encoding of the original signal.
[0078] After the decoding device 200 obtains the encoded data, in the input stage of the decoding device 100, the codebook in the inverse quantization module 210 is used to obtain the codebook feature vector corresponding to the encoded data, and then the codebook feature vector is input to the third convolution module 231 to perform a first convolution operation on the encoded data, thereby obtaining a first feature vector with a channel number of 512 and a dimension of 50.
[0079] In the upsampling stage of the decoding device 100 , multiple upsampling modules 11 may be used to perform multiple upsampling operations on the first feature vector, thereby obtaining a second feature vector with 32 channels and 16,000 dimensions.
[0080] In the output stage of the decoding device 100, the fourth convolution module 233 performs a second convolution operation on the second eigenvector to restore the decoded data with 24 kHz, 1 channel and 16,000 dimensions.
[0081] In this way, data encoding and decoding can be achieved.
[0082] In the above text, the encoding device 100 and the decoding device 200 are described from the perspective of functional modules in conjunction with the accompanying drawings. The functional modules can be implemented in hardware form, can be implemented in software form, or can be implemented in combination with hardware and software modules. Specifically, the steps of the method embodiment in the embodiment of the present application can be completed by the hardware integrated logic circuit and / or software form instructions in the processor, and the steps of the method disclosed in the embodiment of the present application can be directly embodied as a hardware encoding processor to execute, or a combination of hardware and software modules in the encoding processor to execute. Optionally, the software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, etc. The storage medium is located in a memory, and the processor reads the information in the memory, and completes the steps in the above method embodiment in conjunction with its hardware.
[0083] The decoding method of this application will be described in detail below:
[0084] See also Figure 4 , the decoding method of the present application includes:
[0085] Step 011: performing a nonlinear transformation operation on the encoded data to obtain decoded data;
[0086] The nonlinear transformation operation includes a convolution operation, and the convolution operation of at least part of the convolution layers is implemented by grouped convolution, and the number of groups of the convolution layers performing grouped convolution is determined according to the number of channels of the convolution layer.
[0087] Specifically, after the encoding end completes data encoding, the encoded data may be sent to the decoding end. The decoding end may dequantize the encoded data and perform a nonlinear transformation operation on the dequantized characteristic vector, thereby decoding to obtain decoded data.
[0088] At least part of the convolution operation in the nonlinear transformation operation is implemented by grouped convolution. Figure 3 Taking the decoding device 200 shown as an example, the convolution operation included in the nonlinear transformation operation at the decoding end may be the first convolution operation of the third convolution module 231, the third convolution operation of the convolution unit in the upsampling module 232, and the second convolution operation of the fourth convolution module 233.
[0089] Among them, the third convolution operation of the convolution unit in the upsampling module can be a deconvolution operation, and the deconvolution operation is also a kind of convolution operation, and both can perform grouped convolution.
[0090] For example, the principle of conventional convolution is as follows Figure 5 As shown in (a) Regular Convolution in Figure 1, the principle of group convolution is as follows: Figure 5(b) Group Convolution
[0091] For conventional convolution: the number of channels of each convolution kernel is equal to the number of input channels, and the number of convolution kernels is equal to the number of output channels. The computation of the convolution kernel consists of two parts: multiplication and addition.
[0092] For grouped convolution: divide the input features into g groups (g is a positive integer greater than 1) by channel, and then perform regular convolution on each group. Since the number of channels of each group of input features after grouping is the number of input channels / g, the number of channels of each convolution kernel is also reduced to the number of input channels / g, that is, the number of channels becomes 1 / g of regular convolution.
[0093] Since the number of channels of the convolution kernel of the grouped convolution is reduced to 1 / g of the number of channels of the conventional convolution, the amount of calculation and the amount of parameters of the grouped convolution are reduced, that is, the amount of calculation and the amount of parameters of the grouped convolution can be reduced to 1 / g of the amount of calculation and the amount of parameters of the conventional convolution.
[0094] Therefore, by implementing at least part of the convolution operations in the nonlinear transformation through grouped convolution, the computational complexity of the nonlinear transformation can be effectively reduced, thereby improving the coding efficiency.
[0095] The number of groups in the convolution layer for grouped convolution is determined according to the number of channels in the convolution layer. For example, the number of groups in the convolution layer for grouped convolution is positively correlated with the number of channels in the convolution layer. For the decoder, along the direction from input to output, the number of channels decreases, and the number of groups in the convolution layer for grouped convolution decreases. For example, see Figure 3 , along the direction from input to output, the number of groups of the third convolution module 231, the convolution units in the multiple up-sampling modules 232 and the fourth convolution module 233 decreases when performing convolution operations.
[0096] For the decoder, setting the convolutional layer for grouped convolution to decrease the number of groups has the following effects:
[0097] From input to output, with upsampling, the number of channels gradually decreases and the number of groups decreases, thereby reducing the computational complexity.
[0098] Optionally, the changing trends of the number of groups of the convolution layers performing group convolution at the encoding end and the decoding end are opposite.
[0099] Specifically, when the convolution layer for performing group convolution on the encoding end is set to increase the number of groups, the convolution layer for performing group convolution on the decoding end is set to decrease the number of groups.
[0100] In this way, the effects achieved by the encoding end and the decoding end correspond. For example, when the convolution layer for group convolution at the encoding end is set to increase the number of groups, the convolution layer for group convolution at the decoding end is set to decrease the number of groups, which can reduce the computational complexity.
[0101] Optionally, for the increasing number of groups at the encoding end, the ratio between the number of groups of two adjacent convolution layers performing grouped convolution along the direction from input to output is 1 / 2. Figure 3 The numbers of groups of the first convolution module 111, the four down-sampling modules 112 and the second convolution module 113 are 2, 4, 8, 16, 32 and 64 respectively.
[0102] For the decreasing number of groups at the decoding end, the ratio between the number of groups of two adjacent convolutional layers performing grouped convolution along the direction from input to output is 2. Figure 3 The numbers of groups of the third convolution module 231, the four up-sampling modules 232 and the fourth convolution module 233 are 64, 32, 16, 8, 4 and 2 respectively.
[0103] It can be understood that both the convolution operation of the convolution unit of the downsampling module at the encoding end and the deconvolution operation of the convolution unit of the upsampling module at the decoding end can be used as a convolution layer to implement grouped convolution.
[0104] The decoding method of the implementation mode of the present application realizes at least part of the convolution operations in the nonlinear transformation operation in the decoding process through grouped convolution. Compared with ordinary convolution, the number of channels after convolution is relatively large. The present application can significantly reduce the number of channels through grouped convolution, thereby reducing the computational complexity.
[0105] The number of groups of the convolution layer performing grouped convolution is determined according to the number of channels of the convolution layer. Compared with all convolution layers using a larger number of groups for grouped convolution, the number of groups is set more reasonably, thereby reducing the computational complexity of the convolution layer and extracting more useful features to ensure the decoding effect.
[0106] See also Figure 6 In some implementations, step 011: performing a nonlinear transformation operation on the encoded data to obtain decoded data, includes:
[0107] Step 0111: performing a first convolution operation on the encoded data to obtain a first eigenvector;
[0108] Step 0112: performing a downsampling operation on the first eigenvector to obtain a second eigenvector;
[0109] Step 0113: Perform a second convolution operation on the second eigenvector to obtain decoded data.
[0110] Specifically, the decoding device can perform inverse quantization and linear transformation operations on the encoded data, and input the linearly transformed feature vector into the third convolution module, and the first feature vector can be obtained after the first convolution operation.
[0111] Then, the first feature vector is input into the upsampling module for upsampling processing. Multiple upsampling modules can perform multiple upsampling processes in sequence to increase the dimension of the first feature vector and reduce the number of channels of the output second feature vector, thereby obtaining the second feature vector.
[0112] Next, the fourth convolution module performs a second convolution operation on the second feature vector to obtain decoded data, thereby completing the decoding of the encoded data.
[0113] In some embodiments, an activation layer may be added after any convolutional layer. The activation layer may perform activation operations on the features after convolution of the convolutional layer based on a preset activation function, such as a relu activation function, a leaky-relu activation function, or an attention module.
[0114] An activation function is a function added to an artificial neural network to help the network learn complex patterns in data. Activation functions can be used to perform nonlinear mapping operations, making the output of the neural network more complex and more expressive.
[0115] For example, adding an activation layer after the first convolution module, adding an activation layer after the second convolution module, adding an activation layer after the upsampling module or the convolution unit of the upsampling module, etc., is not limited here.
[0116] See also Figure 7 In some embodiments, step 0112: performing an upsampling operation on the first feature vector to obtain a second feature vector includes:
[0117] Step 01121: Perform multiple upsampling operations on the first feature vector to obtain a second feature vector. The number of upsampling operations is determined according to the number of upsampling layers at the decoding end. The first feature vector is the input feature vector of the first upsampling operation, and the second feature vector is the output feature vector of the last upsampling operation. The input feature vector of each upsampling operation is the output feature vector of the last upsampling operation.
[0118] Specifically, the decoding device may have multiple upsampling modules, which are connected in sequence to upsample the feature vector multiple times, thereby increasing the dimension of the first feature vector multiple times and reducing the number of channels of the second feature vector, thereby obtaining a second feature vector whose dimension and number of channels meet preset requirements.
[0119] See also Figure 3 , taking the case where there are four upsampling modules 232 as an example, the first feature vector output after convolution by the third convolution module 231 is input to the first upsampling module 232, and then the first upsampling module 232 performs upsampling processing on the first feature vector, and inputs the output feature vector after the first upsampling processing to the second upsampling module 232, and the second upsampling module 232 performs upsampling processing on the input feature vector, and inputs the output feature vector after the second upsampling processing to the third upsampling module 232, and the third upsampling module 232 performs upsampling processing on the input feature vector, and inputs the output feature vector after the third upsampling processing to the fourth upsampling module 232, and the fourth upsampling module 232 performs upsampling processing on the input feature vector, and outputs the second feature vector after the fourth upsampling processing.
[0120] For example, Figure 3 As shown, after the third convolution module 231 performs the first convolution operation, a first feature vector with a channel number of 16C (such as 512) and a dimension of 50 is obtained. The first upsampling module 232 reduces the dimension of the first feature vector to 8 times, and obtains an output feature vector with a channel number of 256 and a dimension of 400; the second upsampling module 232 increases the dimension of the input feature vector to 6 times, and obtains an output feature vector with a channel number of 128 and a dimension of 2400; the third upsampling module 232 increases the dimension of the input feature vector to 4 times, and obtains an output feature vector with a channel number of 64 and a dimension of 9600; the fourth upsampling module 232 increases the dimension of the input feature vector to 2 times, and obtains a second feature vector with a channel number of 32 and a dimension of 19200. Finally, the fourth convolution module 233 performs a second convolution operation on the second feature vector, and obtains a feature vector with a channel number of 1 and a dimension of 19200, thereby generating decoded data according to the feature vector.
[0121] See also Figure 8 In some embodiments, the decoding method further comprises:
[0122] Step 012: Determine the number of groups of the convolution layer for performing grouped convolution according to at least one of the number of channels of the convolution layer and the codec bit rate, where the number of groups is positively correlated with the number of channels of the convolution layer and the codec bit rate.
[0123] Specifically, the more channels a convolutional layer has after convolution, the more computation is required and the larger the number of groups needs to be set. Therefore, the number of groups is positively correlated with the number of channels of the convolutional layer.
[0124] Similarly, the higher the codec bit rate, the higher the amount of calculation required during encoding and decoding, and the more packets are needed. Therefore, the number of packets is positively correlated with the codec bit rate.
[0125] In this way, the number of groups of each convolutional layer is reasonably set through at least one of the number of channels of the convolutional layer and the encoding and decoding bit rate, so that the number of groups is set more reasonably, which can effectively reduce the calculation complexity and ensure the encoding and decoding effect.
[0126] See also Fig. 9 In some embodiments, the decoding method further comprises:
[0127] Step 014: According to the number of channels of the convolutional layer and the content complexity of the encoded data, the number of groups of the convolutional layer for performing grouped convolution is determined, and the number of groups is positively correlated with the number of channels and the content complexity of the convolutional layer.
[0128] Specifically, the higher the content complexity of the input data, the higher the content complexity of the encoded data after encoding, and the required amount of calculation is generally higher when encoding and decoding, and the number of groups should be set larger to reduce the amount of calculation. Therefore, the number of groups is positively correlated with the content complexity.
[0129] In this way, the number of groups can be reasonably set according to at least one of the content complexity and the number of channels of the convolutional layer, or according to at least one of the content complexity, the number of channels of the convolutional layer and the encoding and decoding bit rate, so that the number of groups is set more reasonably, which can effectively reduce the calculation complexity and ensure the encoding and decoding effect.
[0130] Optionally, the content complexity may be determined according to at least one of a sampling rate of the encoded data, a number of channels, and a number of sound source types.
[0131] It can be understood that the higher the sampling rate, the more channels, the more sound source types, the higher the content complexity. The content complexity can be determined according to one of the sampling rate, the number of channels, and the number of sound source types of the encoded data, for example, the content complexity is determined according to the sampling rate. Alternatively, the content complexity can be determined according to any two of the sampling rate, the number of channels, and the number of sound source types of the encoded data, for example, according to the sampling rate and the number of channels, or according to the number of channels and the number of sound source types. Alternatively, the content complexity can be determined according to the sampling rate, the number of channels, and the number of sound source types of the encoded data.
[0132] In some embodiments, the number of upsampling operations and the convolution kernel size of the convolution unit of the upsampling module are determined according to the content complexity of the input data.
[0133] Specifically, convolution kernels are a common tool in machine learning and computer vision for performing convolution operations on data such as images, audio, and video. A convolution kernel is a two-dimensional matrix that performs element-by-element product operations with the original data and adds the results to obtain a new value. The size of the convolution kernel can be various, such as 1x1, 3x3, 5x5, or 7x7. The size of the convolution kernel can be adjusted as needed to better capture the features in the data. In general, the larger the convolution kernel, the larger the receptive field, the more image information is seen, and the better the global features obtained. However, large convolution kernels will lead to a surge in the amount of calculation and reduced computing performance.
[0134] Therefore, the convolution kernel size of the convolution unit of the upsampling module can be determined according to the content complexity of the encoded data, and the content complexity and the size of the convolution kernel are positively correlated. The higher the content complexity, the more information is required for convolution while ensuring the convolution effect, so the convolution kernel is larger. The lower the content complexity, the smaller the convolution kernel, so that the amount of calculation can be ensured while ensuring the convolution effect.
[0135] Similarly, the more times the upsampling operation is performed, the more information is obtained, and the better the global features are. However, too many upsampling operations will result in excessive computation. The number of upsampling operations can also be determined according to the content complexity of the input data, and the content complexity is positively correlated with the number of upsampling operations. The higher the content complexity, the more information is required for convolution while ensuring the convolution effect, and the more upsampling operations are performed. The lower the content complexity, the smaller the convolution kernel, so that the convolution effect can be ensured while ensuring that the computation is not too large.
[0136] In this way, the appropriate number of upsampling operations and the appropriate size of the convolution kernel of the upsampling module can be determined according to the content complexity of the encoded data, so as to ensure the convolution effect on the one hand and ensure that the amount of calculation is not too large on the other hand, thereby ensuring the computing performance during the decoding process.
[0137] The encoding method of this application will be described in detail below:
[0138] See also Fig.10 , the embodiment of the present application provides a coding method, the coding method comprising:
[0139] Step 021: Perform a nonlinear transformation operation on the input data to obtain encoded data;
[0140] The nonlinear transformation operation includes a convolution operation. The convolution operation of at least some convolution layers is implemented by grouped convolution. The number of groups of the convolution layer performing the grouped convolution is determined according to the number of channels of the convolution layer.
[0141] Specifically, the data sending end (ie, the encoding end) may encode and compress one-dimensional data, such as audio data, or two-dimensional data, such as image or video data, through the encoding device 200 .
[0142] The first nonlinear transformation module 10 can perform a nonlinear transformation operation on the input data. The nonlinear transformation operation includes one or more convolution operations. Figure 3 For example, along the direction from input to output, the first convolution module 111 first performs the fourth convolution operation, and then the multiple downsampling modules 112 perform the sixth convolution operation multiple times in sequence, and then the second convolution module 113 performs the fifth convolution operation, and finally after the quantization operation, the encoded data is obtained.
[0143] The number of groups in the convolution layer for grouped convolution is determined according to the number of channels in the convolution layer. For example, the number of groups in the convolution layer for grouped convolution is positively correlated with the number of channels in the convolution layer. For the encoder, along the direction from input to output, the number of channels in the convolution layer for grouped convolution gradually increases. Therefore, the number of groups in the convolution layer for grouped convolution increases. For example, see Figure 3 , along the direction from input to output, the first convolution module 111, the convolution units in the multiple downsampling modules 112 and the second convolution module 113 increase the number of groups when performing convolution operations.
[0144] For the encoder, setting the convolution layer for grouped convolution to increase the number of groups has the following effects:
[0145] From input to output, with downsampling, the number of channels increases and the number of groups increases, thereby reducing the computational complexity.
[0146] The encoding method of the implementation mode of the present application realizes at least part of the convolution operations in the nonlinear transformation operation in the encoding process through grouped convolution. Compared with ordinary convolution, the number of channels after convolution is relatively large. The present application can significantly reduce the number of channels through grouped convolution, thereby reducing the computational complexity.
[0147] The number of groups of the convolution layer performing grouped convolution is determined according to the number of channels of the convolution layer. Compared with all convolution layers using a larger number of groups for grouped convolution, the number of groups is set more reasonably, thereby reducing the computational complexity of the convolution layer and extracting more useful features to ensure the encoding effect.
[0148] See also Fig.11 In some embodiments, step 021: performing a nonlinear transformation operation on the input data to obtain encoded data includes:
[0149] Step 0211: performing a fourth convolution operation on the input data to obtain a third eigenvector;
[0150] Step 0212: performing a downsampling operation on the third eigenvector to obtain a fourth eigenvector;
[0151] Step 0213: Perform a fifth convolution operation on the fourth eigenvector to obtain encoded data.
[0152] Specifically, the encoding device can perform data sampling on the input data to obtain an original feature vector, and the original feature vector is input into the first convolution module, and the third feature vector can be obtained after the first convolution operation.
[0153] Then, the third eigenvector is input into the downsampling module for downsampling processing. Multiple downsampling modules can perform multiple downsampling processes in sequence to reduce the dimension of the third eigenvector and increase the number of channels of the output fourth eigenvector, thereby obtaining the fourth eigenvector.
[0154] Next, the second convolution module performs a fifth convolution operation on the fourth eigenvector to obtain a coded eigenvector, thereby obtaining coded data, thereby completing the encoding of the input data.
[0155] See also Fig.12 In some embodiments, step 0212: performing a downsampling operation on the third feature vector to obtain a fourth feature vector includes:
[0156] Step 02121: Perform multiple downsampling operations on the third eigenvector to obtain a fourth eigenvector. The number of downsampling operations is determined according to the number of downsampling layers at the decoding end. The third eigenvector is the input eigenvector of the first downsampling operation, and the fourth eigenvector is the output eigenvector of the last downsampling operation. The input eigenvector of each downsampling operation is the output eigenvector of the previous downsampling operation.
[0157] Specifically, the encoding device may have multiple downsampling modules, which are connected in sequence to downsample the feature vector multiple times, thereby reducing the dimension of the third feature vector multiple times and increasing the number of channels of the fourth feature vector, thereby obtaining a fourth feature vector whose dimension and number of channels meet preset requirements.
[0158] See also Figure 3, taking the number of downsampling modules 112 as 4 as an example, the third feature vector output by the first convolution module 111 after convolution is input to the first downsampling module 112, and then the first downsampling module 112 downsamples the third feature vector and inputs the output feature vector after the first downsampling process to the second downsampling module 112, the second downsampling module 112 downsamples the input feature vector and inputs the output feature vector after the second downsampling process to the third downsampling module 112, the third downsampling module 112 downsamples the input feature vector and inputs the output feature vector after the third downsampling process to the fourth downsampling module 112, and the fourth downsampling module 112 downsamples the input feature vector and outputs the fourth feature vector after the fourth downsampling process.
[0159] For example, Figure 3 As shown, after the first convolution module 111 performs the fourth convolution operation, a third eigenvector with a channel number C (such as 32) and a dimension of 19200 is obtained. The first downsampling module 112 reduces the dimension of the third eigenvector to 1 / 2, and obtains an output eigenvector with a channel number of 64 and a dimension of 9600; the second downsampling module 112 reduces the dimension of the input eigenvector to 1 / 4, and obtains an output eigenvector with a channel number of 128 and a dimension of 2400; the third downsampling module 112 reduces the dimension of the input eigenvector to 1 / 6, and obtains an output eigenvector with a channel number of 256 and a dimension of 400; the fourth downsampling module 112 reduces the dimension of the input eigenvector to 1 / 8, and obtains a fourth eigenvector with a channel number of 512 and a dimension of 50. Finally, the encoding device 100 can generate encoding data according to the fourth eigenvector.
[0160] In some embodiments, the number of downsampling operations and the convolution kernel size of the convolution unit of the downsampling module are determined according to the content complexity of the input data.
[0161] Specifically, convolution kernels are a common tool in machine learning and computer vision for performing convolution operations on data such as images, audio, and video. A convolution kernel is a two-dimensional matrix that performs element-by-element product operations with the original data and adds the results to obtain a new value. The size of the convolution kernel can be various, such as 1x1, 3x3, 5x5, or 7x7. The size of the convolution kernel can be adjusted as needed to better capture the features in the data. In general, the larger the convolution kernel, the larger the receptive field, the more image information is seen, and the better the global features obtained. However, large convolution kernels will lead to a surge in the amount of calculation and reduced computing performance.
[0162] Therefore, the convolution kernel size of the convolution unit of the downsampling module can be determined according to the content complexity of the encoded data, and the content complexity and the size of the convolution kernel are positively correlated. The higher the content complexity, the more information is required for convolution while ensuring the convolution effect, so the convolution kernel is larger. The lower the content complexity, the smaller the convolution kernel, so that the amount of calculation can be ensured while ensuring the convolution effect.
[0163] Similarly, the more times the downsampling operation is performed, the more information is obtained, and the better the global features are. However, too many downsampling operations will result in excessive computation. The number of downsampling operations can also be determined according to the content complexity of the input data, and the content complexity is positively correlated with the number of downsampling operations. The higher the content complexity, the more information is required for convolution while ensuring the convolution effect, and the more downsampling operations are performed. The lower the content complexity, the smaller the convolution kernel, so that the convolution effect can be ensured while ensuring that the computation is not too large.
[0164] In this way, the appropriate number of downsampling operations and the appropriate size of the convolution kernel of the downsampling module can be determined according to the content complexity of the encoded data, so as to ensure the convolution effect on the one hand and the amount of calculation on the other hand, thereby ensuring the computing performance during the decoding process.
[0165] See also Fig.13 The computer device 300 of the embodiment of the present application includes a processor 310, a memory 320 and a computer program, wherein the computer program is stored in the memory 320 and executed by the processor 310, and the computer program includes instructions for executing the encoding method or decoding method of any of the above-mentioned embodiments.
[0166] Optionally, the computer device 300 can be any device with image processing capabilities, such as a server or a terminal device (such as a mobile phone, a tablet computer, a display device, a laptop computer, a smart watch, a head-mounted display device, a game console, etc.).
[0167] like Fig.13As shown, the processor 310 included in the computer device 300 is a central processing unit (CPU), and the memory 320 includes a read-only memory 321 (ROM) and a random access memory 322 (RAM). The central processor can perform various appropriate actions and processes according to the program stored in the read-only memory 321 or the program loaded from the storage part 380 to the random access memory 322. In the random access memory 322, various programs and data required for system operation are also stored. The central processor, the read-only memory 321 and the random access memory 322 are connected to each other through a bus 330. The input / output interface 340 (Input / Output interface, i.e., I / O interface) is also connected to the bus 330.
[0168] The following components are connected to the input / output interface 340: an input section 350 including a keyboard, a mouse, etc.; an output section 360 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker; a storage section 370 including a hard disk, etc.; and a communication section 380 including a network interface card such as a LAN card, a modem, etc. The communication section 380 performs communication processing via a network such as the Internet. A drive 390 is also connected to the input / output interface 340 as needed. A removable medium 391, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 390 as needed so that a computer program read therefrom is installed into the storage section 370 as needed.
[0169] In particular, according to an embodiment of the present application, the process described in each method flow chart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer readable medium, and the computer program contains a program code for executing the method shown in the flow chart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 380, and / or installed from the removable medium 391. When the computer program is executed by the central processing unit, various functions defined in the system of the present application are executed.
[0170] In particular, according to an embodiment of the present application, the process described in each method flow chart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains a program code for executing the method shown in the flow chart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication part, and / or installed from a removable medium. When the computer program is executed by a processor, various functions defined in the system of the present application are executed.
[0171] The computer program product of the embodiment of the present application includes a computer program. When the computer program is executed by the processor 310, the encoding method or decoding method of any of the above-mentioned embodiments is executed. For the sake of brevity, it is not repeated here.
[0172] See also Fig.14 The embodiment of the present application also provides a computer-readable storage medium 400 on which a computer program 410 is stored. When the computer program 410 is executed by a processor 420 (for example, the processor 310 of the above-mentioned computer device 300), the steps of the encoding method or decoding method of any of the above-mentioned embodiments are implemented. For the sake of brevity, they are not repeated here.
[0173] In the description of this specification, the descriptions with reference to the terms "certain embodiments", "in an example", "exemplarily", etc., mean that the specific features, structures, materials or characteristics described in conjunction with the embodiments or examples are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, unless they are contradictory.
[0174] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code that includes one or more executable instructions for implementing the steps of a specific logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by technicians in the technical field to which the embodiments of the present application belong.
[0175] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limitations on the present application. Ordinary technicians in this field can change, modify, replace and modify the above embodiments within the scope of the present application.
Claims
1. A decoding method, characterized in that: include: Performing a nonlinear transformation operation on the encoded data to obtain decoded data; The nonlinear transformation operation includes a convolution operation, and the convolution operation of at least part of the convolution layers is implemented by grouped convolution, and the number of groups of the convolution layers performing grouped convolution is determined according to the number of channels of the convolution layers.
2. The decoding method according to claim 1, characterized in that: Along the direction from input to output, the number of groups in the convolution layer performing grouped convolution decreases.
3. The decoding method according to claim 1, characterized in that: The nonlinear transformation operation includes a first convolution operation, an upsampling operation, and a second convolution operation. The nonlinear transformation operation is performed on the encoded data to obtain decoded data, including: Performing the first convolution operation on the encoded data to obtain a first eigenvector; Performing the upsampling operation on the first feature vector to obtain a second feature vector; The second convolution operation is performed on the second feature vector to obtain the decoded data.
4. The decoding method according to claim 3, characterized in that: The performing the upsampling operation on the first feature vector to obtain the second feature vector includes: The upsampling operation is performed multiple times on the first feature vector to obtain the second feature vector, the number of upsampling operations is determined according to the number of upsampling layers at the decoding end, the first feature vector is the input feature vector of the first upsampling operation, the second feature vector is the output feature vector of the last upsampling operation, and the input feature vector of each upsampling operation is the output feature vector of the last upsampling operation.
5. The decoding method according to claim 3, characterized in that: The upsampling operation includes a third convolution operation. The number of groups when the first convolution operation, the third convolution operation and the second convolution operation perform group convolution increases. The number of upsampling operations is determined according to the number of upsampling layers at the decoding end. Along the direction from input to output, the number of groups when the third convolution operation performs group convolution increases.
6. The decoding method according to claim 1, characterized in that: Along the direction from input to output, the ratio between the number of groups of two adjacent convolution layers performing grouped convolution is 2.
7. The decoding method according to any one of claims 1 to 6, characterized in that: Also includes: According to at least one of the number of channels and the codec bit rate of the convolution layer, the number of groups of the convolution layer for performing group convolution is determined, and the number of groups is positively correlated with the number of channels of the convolution layer and the codec bit rate.
8. The decoding method according to any one of claims 1 to 6, characterized in that: Also includes: According to the number of channels of the convolution layer and the content complexity of the encoded data, the number of groups of the convolution layer for performing grouped convolution is determined, and the number of groups is positively correlated with the number of channels of the convolution layer and the content complexity.
9. The decoding method according to claim 8, characterized in that: The content complexity is determined according to at least one of a sampling rate, a number of channels, and a number of sound source types of the encoded data.
10. A coding method, characterized in that: The method comprises: Perform a nonlinear transformation operation on the input data to obtain encoded data; The nonlinear transformation operation includes a convolution operation, and the convolution operation of at least part of the convolution layers is implemented by grouped convolution, and the number of groups of the convolution layers performing grouped convolution is determined according to the number of channels of the convolution layers.
11. A coding device, characterized in that: include: A first non-linear transformation module, used for performing a non-linear transformation operation on input data to obtain encoded data; The first nonlinear transformation module includes a convolution layer, and the convolution operation of at least part of the convolution layer is implemented by grouped convolution. The number of groups of the convolution layer performing grouped convolution is determined according to the number of channels of the convolution layer.
12. A decoding device, characterized in that: include: A second non-linear transformation module, used for performing a non-linear transformation operation on the encoded data to obtain decoded data; The second nonlinear transformation module includes a convolution layer, and the convolution operation of at least part of the convolution layer is implemented by grouped convolution. The number of groups of the convolution layer performing grouped convolution is determined according to the number of channels of the convolution layer.
13. A computer device, characterized in that: include: Processor, memory; and A computer program, wherein the computer program is stored in the memory and executed by the processor, and the computer program includes instructions for executing the decoding method according to any one of claims 1 to 9 or the encoding method according to claim 10.
14. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the decoding method according to any one of claims 1 to 9 or the encoding method according to claim 10 is implemented.
15. A non-volatile computer-readable storage medium containing a computer program, wherein when the computer program is executed by a processor, the processor executes the decoding method according to any one of claims 1 to 9 or the encoding method according to claim 10.