Method and device for grouping neural network for video encoding and decoding

By introducing packet technology into the video encoding and decoding system, the outputs processed by neural networks are divided into multiple groups and recombined, which solves the problem of high computational complexity of neural network processing and achieves more efficient video encoding performance.

CN115002473BActive Publication Date: 2025-05-13MEDIATEK INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210509362.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-01-26
Filing Date
2019-01-22
Publication Date
2025-05-13
Estimated Expiration
2039-01-22

AI Technical Summary

Technical Problem

In the existing video encoding and decoding technologies, neural network processing has high computational complexity and is difficult to meet the needs of efficient video encoding.

Method used

By introducing packet technology, the output of one layer processed by the neural network is divided into multiple groups, and after processing one layer, the output of all layers is mixed and regrouped, thereby reducing the computational complexity.

Benefits of technology

It realizes the performance of neural network processing while reducing the computational complexity, and is suitable for image repair processing in video encoding and codec systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115002473B_ABST
    Figure CN115002473B_ABST
Patent Text Reader

Abstract

A method and apparatus for signal processing using grouped neural network (NN) processing is disclosed. A plurality of input signals for a current layer of NN processing are grouped into a plurality of input groups, the plurality of input groups comprising a first input group and a second input group. The neural network processing for the current layer is divided into a plurality of NN processings, the plurality of NN processings comprising a first NN processing and a second NN processing. The first NN processing and the second NN processing are applied to the first input group and the second input group, respectively, to generate a first output group and a second output group for the current layer of NN processing. In another method, different code types are used to encode parameter sets related to layers of NN processing.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related references

[0002] This invention is a divisional application of the invention patent application with application number 201980009758.2 and invention name as method and device for grouping neural network for video encoding and decoding. This invention claims priority to U.S. Provisional Patent Application No. 62 / 622,224 filed on January 26, 2018 and U.S. Provisional Patent Application No. 62 / 622,226 filed on January 26, 2018. The U.S. Provisional Patent Application is hereby incorporated by reference. Technical Field

[0003] The present invention generally relates to neural networks. In particular, the present invention relates to reducing the complexity of neural network (NN) processing by dividing multiple inputs to a given layer of the neural network into multiple input groups. Background Art

[0004] A neural network (NN), also known as an "artificial" neural network (ANN), is an information processing system that has some of the same performance characteristics as a biological neural network. A neural network system consists of many simple and highly interconnected processing components to process information through their dynamic response to external inputs. The processing components can be thought of as neurons in the human brain, where each perceptron accepts multiple inputs and calculates the weighted sum of the inputs. In the field of neural networks, perceptrons are considered to be mathematical models of biological neurons. In addition, these interconnected processing components are usually organized in layers. For recognition applications, external inputs can correspond to patterns presented to the network, which communicates with one or more middle layers, also called "hidden layers", where the actual processing is done via a system of weighted "connections".

[0005] Artificial neural networks can use different architectures to specify the variables involved in the neural network and their topological relationships. For example, the variables involved in a neural network can be the weights of the connections between neurons, as well as the activity of neurons. A feed-forward network is a neural network topology in which the nodes in each layer are fed to the next stage and there are connections between nodes in the same layer. Most ANNs include some form of "learning rules" that modify the weights of the connections based on the input patterns presented. In a sense, ANNs, like their biological counterparts, learn from examples. Backward propagation neural networks are more advanced neural networks that allow backward error propagation of weight adjustments. Therefore, backpropagation neural networks can improve performance by minimizing the error fed back to the neural network.

[0006] NN can be a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), or other NN variations. Deep multilayer neural networks or deep neural networks (DNNs) correspond to neural networks with multiple levels of interconnected nodes that allow them to tightly represent highly nonlinear and highly varying functions. However, the computational complexity of DNNs grows rapidly with the number of nodes associated with a large number of layers.

[0007] CNN is a type of feedforward artificial neural network, most commonly used to analyze visual imagery. Recurrent neural networks (RNNs) are a type of artificial neural network in which the connections between nodes form a directed graph along a sequence. Unlike feedforward neural networks, RNNs can use their internal state (memory) to process input sequences. RNNs can have loops in them to allow information to persist. RNNs allow operations on sequences of vectors (e.g., sequences in input, output, or both).

[0008] The High Efficiency Video Codec (HEVC) standard was developed under a joint video project of the ITU-T Video Coding Experts Group (VCEG) and ISO / IEC Moving Picture Experts Group (MEPG) standardization bodies, specifically in collaboration with a group called the Joint Collaborative Team on Video Codecs (JCT-VC).

[0009] In HEVC, a slice is split into multiple coding tree units (CTUs). CTUs are further split into multiple coding units (CUs) to adapt to various local characteristics. HEVC supports multiple intra prediction modes, and for intra-coded CUs, the selected intra prediction mode is signaled. In addition to the concept of coding units, HEVC also introduces the concept of prediction units (PUs). Once the CU hierarchical tree is split, each leaf CU is further split into one or more prediction units (PUs) according to the prediction type and PU partition. After prediction, the residual associated with the CU is divided into multiple transform blocks for transform processing, called transform units.

[0010] Figure 1AAn exemplary adaptive intra / inter video encoder based on HEVC is shown. When the inter mode is used, the intra / inter prediction unit 110 generates an inter prediction based on motion estimation (ME) / motion compensation (MC). When the intra mode is used, the intra / inter prediction unit 110 generates an intra prediction. The intra / inter prediction data (i.e., the intra / inter prediction signal) is provided to the subtractor 116, and a prediction error, also referred to as a residual, is formed by subtracting the intra / inter prediction signal from a signal related to the input image. The process of generating intra / inter prediction data in the present invention is also referred to as a prediction process. The prediction error (i.e., the residual) is then processed by a transform (T) and then by a quantization (Q) process (T+Q, 120). The transformed and quantized residual is then encoded by the entropy coding unit 122 to be included in the video bitstream corresponding to the compressed video data. The bitstream associated with the transform coefficients is then packed together with side information such as motion, coding mode, and other information associated with the image region. The side information is also compressed by entropy coding to reduce the required bandwidth. Since the reconstructed image can also be used as a reference image for inter-frame prediction, one or more reference images need to be reconstructed at the encoder end. Therefore, the transformed and quantized residuals are processed by inverse quantization (IQ) and inverse transformation (IT) (IQ+IT, 124) to recover the residuals. The reconstructed residuals are then added back to the intra / inter prediction data in the reconstruction unit (REC) 128 to reconstruct the video data. In the present invention, the process of adding the reconstructed residuals to the intra / inter prediction signal is also referred to as reconstruction processing. The output image from the reconstruction processing is also referred to as a reconstructed image. In order to reduce artifacts in the reconstructed image, a loop filter including a deblocking filter (DF) 130 and a sample adaptive offset (SAO) 132 is used. In the present invention, the filtered reconstructed image at the output of all filtering processes is also referred to as a decoded image. The decoded image is stored in the frame buffer 140 and used for prediction of other frames.

[0011] Figure 1BAn exemplary adaptive intra / inter video decoder based on HEVC is shown. Because the encoder also includes a local decoder for reconstructing video data, in addition to the entropy decoder, some decoder components have been used in the encoder. On the decoder side, the entropy decoding unit 160 is used to recover the encoded symbols or syntax from the bitstream. The process of generating a reconstructed residual from the input bitstream in the present invention is called a residual decoding process. The prediction process for generating intra / inter prediction is also applied to the decoder side, however, because the inter prediction only needs to perform motion compensation using the motion information derived from the bitstream, the intra / inter prediction unit 150 is different from the intra / inter prediction unit on the encoder side. In addition, the adder 114 is used to add the reconstructed residual to the intra / inter prediction material.

[0012] During the development of the HEVC standard, another in-loop filter was disclosed, called Adaptive Loop Filter (ALF), but it was not adopted into the main standard. ALF can be used to further improve video quality. For example, Figure 2A The encoder side is Figure 2B As shown in the decoder side of FIG. 1 , the ALF 210 can be used after the SAO 132 and the output from the ALF 210 can be stored in the frame buffer 140. For the decoder side, the output from the ALF 210 can also be used as a decoder output for display or other processing. In the present invention, the deblocking filter, SAO and ALF are also referred to as filtering processes.

[0013] Among different image restoration and processing methods, neural network based methods are promising methods in recent years, such as deep neural network (DNN) and convolutional neural network (CNN) methods. It has been applied to various image processing applications, such as image de-nosing, image super-resolution, etc., and it has been demonstrated that DNN or CNN can achieve better performance than traditional image processing methods. Therefore, in the following, we propose to use CNN as an image restoration processing method in a video codec system to improve subjective quality or coding efficiency. It is expected to use NN as an image restoration method in a video codec system to improve the subjective quality or coding efficiency of new video codec standards such as High Efficiency Video Coding (HEVC). In addition, NN requires considerable computational complexity, and it is also expected to reduce the computational complexity of NN. Summary of the invention

[0014] A method and apparatus for signaling a parameter set associated with a neural network (NN) signal processing is disclosed. According to the method, the parameter set associated with a current layer of the neural network processing is mapped using at least two code types by mapping a first portion of the parameter set associated with the current layer of the neural network processing using a first code and mapping a second portion of the parameter set associated with the current layer of the neural network processing using a second code. The current layer of the neural network processing is applied to a plurality of input signals of the current layer of the neural network processing using the parameter set associated with the current layer of the neural network processing, the parameter set including the first portion of the parameter set associated with the current layer of the neural network processing and the second portion of the parameter set associated with the current layer of the neural network processing.

[0015] A system using this method may correspond to a video encoder or a video decoder. Thus, an initial input signal provided to an initial layer of the neural network processing may correspond to a target video signal in a path of a video signal processing flow in the video encoder or the video decoder. When the initial input signal corresponds to a loop filtered signal, the parameter set is signaled at a sequence level, a picture level, or a slice level. When the initial input signal corresponds to a post-filter signal, the parameter set is signaled as a supplemental enhancement information (SEI) message. The target video signal may correspond to a processed signal output from a reconstruction (REC), a deblocking filter (DF), a sample adaptive filtering (SAO), or an adaptive loop filtering (ALF).

[0016] When the system corresponds to a video encoder, mapping the parameter set associated with the current layer of the neural network processing may correspond to encoding the parameter set associated with the current layer of the neural network processing into the encoded data using the first code and the second code. When the system corresponds to a video decoder, mapping the parameter set associated with the current layer of the neural network processing corresponds to decoding the parameter set associated with the current layer of the neural network processing from the encoded data using the first code and the second code.

[0017] The first partial set of parameters related to the current layer of the neural network processing corresponds to a plurality of weights related to the current layer of the neural network processing, and the second partial set of parameters related to the current layer of the neural network processing corresponds to a plurality of offsets related to the current layer of the neural network processing. In this case, the first code may correspond to a variable length code. In addition, the variable length code may correspond to a Huffman code or an n-order exponential Golomb code (EGn) and n is an integer greater than or equal to 0. Different n is used for different layers of the neural network processing. The second code may correspond to a fixed length code. In another embodiment, the first code may correspond to a DPCM (Differential Pulse Code Modulation) code, and wherein the difference between the plurality of weights and the minimum value of the plurality of weights is encoded.

[0018] In another embodiment, different codes may be used in different layers. For example, the first code, the second code, or both may be selected from a group including a plurality of codes. The target code selected from the group including the first code or the second code is indicated by a flag.

[0019] The present invention divides the output of a layer processed by a neural network into multiple groups by applying a grouping technique, and mixes the outputs of all layers and regroups them after processing one layer, thereby achieving better performance while reducing computational complexity. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1A An exemplary adaptive intra / inter encoder based on the High Efficiency Video Codec (HEVC) standard is shown.

[0021] Figure 1B An exemplary adaptive intra / inter decoder based on the High Efficiency Video Codec (HEVC) standard is shown.

[0022] Figure 2A Shows something like Figure 1A An adaptive intra / inter video encoder in with additional ALF processing.

[0023] Figure 2B Shows something like Figure 1B An adaptive intra / inter video decoder in with additional ALF processing.

[0024] Figure 3 An example of applying a neural network to a reconstructed signal is shown, where the input to the NN is the reconstructed pixels from a reconstruction module (REC) and the output of the NN is the NN-filtered reconstructed pixels.

[0025] Figure 4An example of traditional neural network processing is shown, where the outputs of all channels in the previous layer are used as inputs to all filters in the current layer without grouping.

[0026] Figure 5 An example of grouped neural network processing according to an embodiment of the present invention is shown, in which the output of the previous layer before L1 is divided into two groups and the current layer of the neural network processing is also divided into two groups. In this embodiment, the outputs of L1 group A and L1 group B are used as inputs of L2 group A and L2 group B respectively without mixing.

[0027] Figure 6 An example of grouped neural network processing according to another embodiment of the present invention is shown, in which the output of the previous layer before L1 is divided into two groups and the current layer of the neural network processing is also divided into two groups. In this embodiment, the outputs of L1 group A and L1 group B can be mixed and the mixed output can be used as inputs of L2 group A and L2 group B.

[0028] Figure 7 An exemplary flow chart of grouping neural network (NN) processing for the system according to an embodiment of the present invention is shown.

[0029] Figure 8 An exemplary flow chart of neural network (NN) processing in a system having different code types for NN processing-related parameter sets according to another embodiment of the present invention is shown. DETAILED DESCRIPTION

[0030] The following description is the best way to implement the present invention. The description is to illustrate the basic principles of the present invention and should not be construed as limiting. The scope of the present invention is best determined by reference to the appended claims.

[0031] When NN is applied to video codec systems, NN can be applied to various signals along the signal processing path. Figure 3 An example of applying NN 310 to reconstruct a signal is shown. Figure 3 , the input of NN 310 is the reconstructed pixels from REC 128. The output of NN is the NN-filtered reconstructed pixels, which may be further processed by a deblocking filter (ie, DF 130). Figure 3 is an example of applying NN 310 in a video encoder, however, NN 310 can be applied in a corresponding video decoder in a similar manner. CNN can be replaced by other NN variations, such as DNN (deep fully connected feedforward neural network), RNN (recurrent neural network) or GAN (generative adversarial network).

[0032] In the present invention, a method for using CNN as an image restoration method in a video coding system is disclosed. For example, CNN can be applied to Figure 2A as well as Figure 2B The video encoder and the ALF output image in the decoder are used to generate the final decoded image. Alternatively, the video encoder and the ALF output image in the decoder are used to generate the final decoded image. Figure 1A -B and Figure 2A -B shows another repair method in the video coding system, in which CNN is directly applied after SAO, DF or REC. In another embodiment, CNN can be used to directly recover the quantization error or only improve the quality of the predictor. In the former method, CNN is applied after inverse quantization and transformation to recover the reconstructed residual. In the latter method, CNN is applied to the predictor generated by inter-frame or intra-frame prediction. In another embodiment, CNN is applied to the ALF output image as a post-loop filter. NN processing generally includes one or more layers. The layer that first processes the signal in the one or more layers of NN processing is called the initial layer. The initial input signal input to the initial layer corresponds to the signal in the video signal processing path that has not been processed by NN. When NN is applied after SAO, DF, or REC, the initial input signal is a signal processed by SAO, DF or REC. The same principle also applies to other processing of the present invention, such as applying NN after ALF, then the initial input signal is a signal processed by ALF.

[0033] In order to reduce the computational complexity of CNN, which is especially effective in video encoding and decoding systems, the present invention discloses a grouping technique. Traditionally, the network design of CNN is similar to a fully connected network. Figure 4 As shown in , the outputs of all channels of the previous layer are used as inputs to all filters in the current layer. Figure 4 , the input of L1 410 and the input of L2 430 are equal to the output of the previous layer before L1 420 and L2 440, respectively. Therefore, if the number of filters of the previous layer before L1 420 and L2 440 is equal to M and N, respectively, then for each filter in L1 and L2, the number of input channels in L1 and L2 is M and N. If the number of outputs in the previous layer (i.e., the number of inputs to the current layer) is M, the number of outputs in the current layer is N, and the filter tap lengths in the horizontal and vertical directions are h and w, respectively, then the computational complexity of the current layer is proportional to h×w×M×N.

[0034] In order to reduce the complexity, grouping technology is introduced in the CNN network design. Figure 5An example of a network design for a grouped CNN according to an embodiment of the present invention is shown. In this example, the output of the previous layer before L1 is divided into or treated as two groups, L1 channel group A 510 and L1 channel group B 512. The convolution process is divided into or treated as two independent processes, i.e., group A 520 convolved with an L1 filter and group B 522 convolved with an L1 filter. The next layer (i.e., L2) is also divided into or treated as two corresponding groups (530 / 532 and 540 / 542). However, in this design, there is no exchange between the two groups. This may cause performance loss. In one example, M inputs are split into two groups consisting of (M / 2) and (M / 2) inputs, and N outputs are also split into two groups consisting of (N / 2) and (N / 2) outputs. In this case, the computational complexity of the current layer is proportional to 1 / 2×(h×w×M×N).

[0035] In order to reduce performance loss, another network design of the present invention is disclosed, such as Figure 6 As shown, the processing of CNN groups can be mixed. The output of the previous layer before L1 can be divided or treated as two groups, L1 channel group A 610 and L1 channel group B 612. The convolution processing is divided or treated as independent processing, namely group A 620 convolved with L1 filters and group B 622 convolved with L1 filters. The next layer (i.e., L2) is also divided or treated as two corresponding groups (630 / 632 and 640 / 642). In this example, as Figure 6 As shown, the outputs of L1 group A and L1 group B can be mixed, and the mixed output can be used as the input of L2 group A and L2 group B.

[0036] In one example, the M inputs are divided into two groups consisting of (M / 2) and (M / 2) inputs, and the N outputs are also divided into two groups consisting of (N / 2) and (N / 2). For example, the mixing can be achieved by taking a portion of the (N / 2) output of L1 group A 620 and a portion of the (N / 2) output of L1 group B 622 to form the (N / 2) input of L2 group A (i.e., a combination of 630a and 632a), and taking the remaining portion of the (N / 2) output of L1 group A and the remaining portion of the (N / 2) output of L1 group B to form the (N / 2) input of L2 group B (i.e., a combination of 630b and 632b). Therefore, at least a portion of the output of at least L1 group A is crossed into L2 group B (as shown in the direction of 630b). In addition, at least a portion of the output of L1 group B is crossed into the input of L2 group A (as shown in the direction of 632a). In this case, the computational complexity of the current layer is proportional to 1 / 2×(h×w×M×N), which is the same as the case where the outputs of L1 group A and L1 group B are not mixed. However, because there are some interactions between group A and group B, the performance loss can be reduced.

[0037] The above disclosed grouping method or the grouping method with hybrid can be combined with the traditional design. For example, the grouping technique can be applied to even layers and the traditional design (i.e., no grouping) can be applied to odd layers. In another example, the grouping with hybrid technique can be applied to those layers whose layer index is equal to 1 and 2 after modulo 3, and the traditional design can be applied to those layers whose layer index is equal to 0 after modulo 3. For example, for those layers whose layer index is equal to 1 and 2 after modulo 3, the grouping technique is applied to divide the multiple input signals of the group into multiple groups. For those layers whose layer index is equal to 0 after modulo 3, the grouping technique is not applied to group the input signals of the group into multiple groups, that is, the multiple input signals are not grouped into non-divided networks, and the reorganization is not processed as multiple NNs.

[0038] When CNN is applied to video encoding and decoding, the parameter set of CNN can be signaled to the decoder so that the decoder can apply the corresponding CNN to achieve better performance. As is well known in the art, the parameter set may include weights and offsets of the connected network and filter information. If CNN is used as a loop filter, the parameter set can be signaled at the sequence level, image level or slice level. If CNN is used as a post-loop filter, the parameter set can be signaled as a supplemental enhancement information (SEI) message. The sequence level, image level or slice level mentioned above corresponds to different video data structures.

[0039] The parameters in the CNN parameter set can be divided into two groups, such as weights and offsets. For different groups, different encoding and decoding methods can be used to encode and decode the values. In one embodiment, a variable-length code (VLC) can be applied to the weights and a fixed-length code (FLC) can be used to encode the offsets. In another embodiment, the number of bits in the variable-length code table and the fixed-length code can vary for different layers. For example, for the first layer, the number of bits of the fixed-length code can be 8 bits; and in the following layers, the number of bits of the fixed-length code is only 6 bits. In another example, for the first layer, the EG-0 (i.e., 0-order exponential Golomb) code can be used as a variable-length code and the EG-5 (i.e., 5-order exponential Golomb) code can be used as a variable-length code for other layers. Although specific 0-order and 5-order exponential Golomb codes are given as examples, any n-order exponential Golomb can also be used, where n is an integer greater than or equal to 0.

[0040] In another embodiment, in addition to variable length codes and fixed length codes, DPCM (differential pulse coded modulation) can be used to further reduce the coding information. In this method, the minimum and maximum values ​​of the coefficients to be coded are first determined. Based on the difference between the minimum and maximum values, the number of bits used to encode the difference between the coefficient to be coded and the minimum value is determined. The minimum value and the number of bits used to encode the difference are first signaled, followed by the difference between the coefficient to be coded and the minimum value for each coefficient to be coded. For example, the coefficients to be coded are {20, 21, 18, 19, 20, 21}. When using fixed length codes, these parameters will require a 5-bit fixed length code for each coefficient. When using DPCM, the minimum value (18) and the maximum value (21) of the 6 coefficients are first determined. Because the difference ranges between 0 and 3, the number of bits required to encode the difference between the minimum value (18) and the maximum value (21) is only 2. Therefore, the minimum value (18) can be signaled by using a 5-bit fixed length code. The number of bits required to encode the difference between the minimum value (18) and the maximum value (21) can be signaled by using a 3-bit fixed length code. The difference between the coefficient to be encoded and the minimum value {2,3,0,1,2,3} can be signaled using 2 bits. Therefore, the total bits can be reduced from 30 bits = 6 (i.e., the number of coefficients to be encoded) × 5 bits to 20 bits = (5 bits + 3 bits + 6 × 2 bits). Fixed length codes can be changed to truncated binary codes, variable length codes, Huffman codes, etc.

[0041] Different coding methods can be selected and used together. For example, DPCM and fixed length codes can be supported at the same time, and a flag is encoded to indicate which method is used in the subsequent encoded bits.

[0042] CNN can be applied in various image applications such as image classification, face detection, object detection, etc. The above method can be applied when CNN parameter compression is required to reduce storage requirements. In this case, these compressed CNN parameters will be stored in some memory or device such as solid state disk (SSD), hard disk drive disk (HDD), memory stick, etc. These compressed parameters will be decoded and fed to the CNN network to perform CNN processing only when CNN processing is performed.

[0043] Figure 7 An exemplary flowchart of grouped neural network (NN) processing of a system according to an embodiment of the present invention is shown. The steps shown in the flowchart can be implemented as executable program code on one or more processors (such as one or more CPUs), encoder side, decoder side, or any other hardware or software component capable of executing program code. The steps shown in the flowchart can also be implemented as hardware, such as one or more electronic devices or processors for performing the steps in the flowchart. In step 710, the method uses multiple input signals of the current layer for NN processing as multiple input groups, and the multiple input groups include a first input group and a second input group for the current layer for NN processing. In step 720, the neural network processing of the current layer for NN processing is used as multiple NN processing, and the multiple NN processing includes a first NN processing and a second NN processing for the current layer for NN processing. In step 730, the first NN processing is applied to the first input group to generate a first output group for the current layer for NN processing. In step 740, the second NN processing is applied to the second input group to generate a second output group for the current layer for NN processing. In step 750, an output group is provided as a current output of the current layer processed by the NN, and the output group includes the first output group and the second output group for the current layer processed by the NN.

[0044] Figure 8An exemplary flow chart of a neural network processing in a system according to another embodiment of the present invention, the system having parameter sets related to NN processing with different code types. According to this method, in step 810, a parameter set related to the current layer of the neural network processing is mapped using at least two code types by mapping a first portion of a parameter set related to the current layer of the neural network processing using a first code and mapping a second portion of a parameter set related to the current layer of the neural network processing using a second code. In step 820, the current layer of the neural network processing is applied to an input signal of the current layer of the neural network processing using the parameter set related to the current layer of the neural network processing, the parameter set including the first portion of the parameter set related to the current layer of the neural network processing and the second portion of the parameter set related to the current layer of the neural network processing.

[0045] The flowchart shown is intended to illustrate an example of video encoding according to the present invention. Those skilled in the art can modify each step, rearrange steps, split steps or combine steps to implement the present invention without departing from the spirit of the present invention. In the present invention, specific syntax and semantics are used to illustrate examples of implementing embodiments of the present invention. Those skilled in the art can implement the present invention by replacing the syntax and semantics with equivalent syntax and semantics without departing from the spirit of the present invention.

[0046] The description given above enables those skilled in the art to implement the present invention as provided in the context of a specific application and its requirements. Various modifications of the described embodiments will be apparent to those skilled in the art, and the general principles defined herein may be applied to other embodiments. Therefore, the present invention is not intended to be limited to the specific embodiments shown and described, but should conform to the widest range consistent with the principles and novel features described herein. In the above-mentioned detailed description, various specific details are shown to provide a thorough understanding of the present invention, however, those skilled in the art will be able to understand that the present invention can be implemented.

[0047] The embodiments of the present invention described above can be implemented with various hardware, software codes and combinations thereof. For example, embodiments of the present invention can be one or more electronic circuits integrated into a video compression chip or a program code integrated into video compression software to perform the processing described herein. Embodiments of the present invention can also be program codes executed on a digital signal processor (DSP) to perform the processing described in the present invention. The present invention can also be related to multiple functions performed by a computer processor, a digital signal processor, a microprocessor or a field programmable gate array (FPGA). These processors can be configured to perform specific tasks according to the present invention by executing machine-readable software codes or firmware codes that define the specific methods implemented by the present invention. The software code or firmware code can be developed in different programming languages ​​and different formats or styles. Software codes can also be compiled for different target platforms. However, the different code formats, styles and languages ​​of the software code and other methods of configuring the code to perform specific tasks consistent with the present invention will not deviate from the spirit and scope of the present invention.

[0048] The present invention may be implemented in other specific forms without departing from its spirit or essential features. The described examples are considered to be exemplary and non-restrictive in all respects. Therefore, the scope of the present invention is indicated by the appended claims rather than the foregoing description. All changes within the equivalent meaning and scope of the claims are included within their scope.

Claims

1. A method for signal processing in a system using neural network processing, wherein the neural network processing comprises one or more layers of neural network processing, characterized in that The method comprises: mapping the set of parameters associated with the current layer of the neural network processing using at least two code types by mapping a first portion of the set of parameters associated with the current layer of the neural network processing using a first code and mapping a second portion of the set of parameters associated with the current layer of the neural network processing using a second code; and applying the current layer of the neural network processing to an input signal of the current layer of the neural network processing using the parameter set associated with the current layer of the neural network processing, the parameter set comprising the first partial parameter set associated with the current layer of the neural network processing and the second partial parameter set associated with the current layer of the neural network processing; wherein an initial input signal provided to an initial layer of the neural network processing corresponds to a target video signal in a path of a video signal processing flow in a video encoder or a video decoder, wherein the target video signal is a signal provided to the one or more layers of the neural network without being processed by the neural network.

2. The signal processing method using neural network processing in the system of claim 1, characterized in that: Wherein when the initial input signal corresponds to a loop filtered signal, the parameter set is signaled in a sequence level, a picture level or a slice level.

3. The signal processing method using neural network processing in the system of claim 1, characterized in that: Wherein when the initial input signal corresponds to a post-loop filtered signal, the parameter set is signaled as a supplemental enhancement information message.

4. The signal processing method using neural network processing in the system of claim 1, characterized in that: The target video signal is a signal provided to the one or more layers of the neural network that has not been processed by the neural network, and the target video signal corresponds to a processed signal output from reconstruction, deblocking filter, sample adaptive offset or adaptive loop filtering.

5. The signal processing method using neural network processing in the system of claim 1, characterized in that: Wherein, when the system corresponds to a video encoder, mapping the parameter set associated with the current layer processed by the neural network corresponds to encoding the parameter set associated with the current layer processed by the neural network using the first code or the second code.

6. The signal processing method using neural network processing in the system of claim 1, characterized in that: Wherein, when the system corresponds to a video decoder, mapping the parameter set associated with the current layer processed by the neural network corresponds to decoding the parameter set associated with the current layer processed by the neural network using the first code and the second code.

7. The signal processing method using neural network processing in the system of claim 1, characterized in that: Wherein the first portion of the parameter set related to the current layer processed by the neural network corresponds to a plurality of weights related to the current layer processed by the neural network, and the second portion of the parameter set related to the current layer processed by the neural network corresponds to a plurality of offsets related to the current layer processed by the neural network.

8. The signal processing method using neural network processing in the system of claim 7, characterized in that: The first code corresponds to a variable length code.

9. The signal processing method using neural network processing in the system of claim 8, characterized in that: The variable length code corresponds to a Huffman code or an n-th order Exponential Golomb code and n is an integer greater than or equal to 0.

10. The signal processing method using neural network processing in the system of claim 9, characterized in that: Where different n are used for different layers of the neural network processing.

11. The signal processing method using neural network processing in the system of claim 7, characterized in that: The second code corresponds to a fixed length code.

12. The signal processing method using neural network processing in the system of claim 7, characterized in that: wherein the first code corresponds to a differential pulse code modulation encoding, and wherein a difference between the plurality of weights and a minimum value of the plurality of weights is encoded.

13. The signal processing method using neural network processing in the system of claim 1, characterized in that: Wherein the first code, the second code, or both are selected from a group consisting of a plurality of codes.

14. The signal processing method using neural network processing in the system of claim 13, characterized in that: A target code selected from the group including a plurality of codes using the first code and the second code is indicated by a flag.

15. An apparatus for signal processing using a neural network, the neural network comprising one or more layers of neural network processing, characterized in that The apparatus comprises one or more electronic devices or processors for: mapping the set of parameters associated with the current layer of the neural network processing using at least two code types by mapping a first portion of the set of parameters associated with the current layer of the neural network processing using a first code and mapping a second portion of the set of parameters associated with the current layer of the neural network processing using a second code; as well as applying the current layer of the neural network processing to a plurality of input signals of the current layer of the neural network processing using the parameter set associated with the current layer of the neural network processing, the parameter set comprising the first partial parameter set associated with the current layer of the neural network processing and the second partial parameter set associated with the current layer of the neural network processing; wherein an initial input signal provided to an initial layer of the neural network processing corresponds to a target video signal in a path of a video signal processing flow in a video encoder or a video decoder, wherein the target video signal is a signal provided to the one or more layers of the neural network without being processed by the neural network.