Image reconstruction, encoding and decoding method, reconstruction model training method, and related devices
By selecting the corresponding image reconstruction model according to the component type of the input image, and each component is processed and combined separately, the problem of poor image reconstruction effect in the prior art is solved, and a more efficient image reconstruction effect is achieved.
Patent Information
- Application Number
- CN202111166009.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-30
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2041-09-30
AI Technical Summary
Existing super-resolution reconstruction techniques are less effective in image reconstruction, especially when processing different components.
By selecting the corresponding image reconstruction model according to the component type of the input image, each component is processed separately, and the reconstructed components are combined to generate a reconstructed image.
Improve the effect of image reconstruction, especially when processing different components, high-resolution images can be restored more accurately.
Smart Images

Figure CN114004743B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of video coding and decoding technology, and in particular to an image reconstruction, coding and decoding method, a reconstruction model training method, and related devices. Background Art
[0002] The amount of video image data is relatively large, and it is usually necessary to compress the video pixel data (RGB, YUV, etc.). The compressed data is called a video stream, which is transmitted to the user end through a wired or wireless network and then decoded for viewing. The entire video encoding process includes block division, prediction, transformation, quantization, encoding and other processes. In order to more effectively compress video data, high-resolution images are downsampled to low-resolution images during encoding and decoding. When high-resolution images are needed, they are enlarged through upsampling or reconstructed through super-resolution technology.
[0003] Super-resolution reconstruction technology not only needs to enlarge the low-resolution image, but also reconstruct the missing information through the model to restore the high-resolution image. The model of super-resolution reconstruction technology usually includes priors, neural networks, etc.
[0004] In the prior art, the reconstruction model of the super-resolution reconstruction technology has a poor effect when performing image reconstruction. Summary of the invention
[0005] The present invention provides an image reconstruction, encoding and decoding method, a reconstruction model training method, and related devices, which can improve the image reconstruction effect.
[0006] To solve the above technical problems, the first technical solution provided by the present invention is: to provide an image reconstruction method, comprising: obtaining an input image, the input image including at least one type of component; determining an image reconstruction model corresponding to the component based on the type of the component; using the selected image reconstruction model to process the corresponding component to obtain a reconstructed component, and combining all the reconstructed components to obtain a reconstructed image.
[0007] To solve the above technical problems, the second technical solution provided by the present invention is: to provide an image reconstruction device, comprising: an acquisition module, used to acquire an input image, the input image comprising at least one type of component; a determination module, used to determine the image reconstruction model corresponding to the component based on the type of the component; a reconstruction module, used to process the corresponding component using the selected image reconstruction model to obtain a reconstructed component, and a combination module, used to combine all the reconstructed components to obtain a reconstructed image.
[0008] To solve the above technical problems, the third technical solution provided by the present invention is: providing an image decoding method, comprising: acquiring a code stream; decoding the code stream to obtain an input image, wherein the input image includes at least one type of component; determining an image reconstruction model corresponding to the component based on the type of the component; processing the corresponding component using the selected image reconstruction model to obtain a reconstructed component; and combining all the reconstructed components to obtain a reconstructed image.
[0009] To solve the above technical problems, the fourth technical solution provided by the present invention is: to provide an image decoding device, comprising: an acquisition module, used to acquire a code stream; a decoding module, used to decode the code stream to obtain an input image, wherein the input image includes at least one type of component; a determination module, used to determine the image reconstruction model corresponding to the component based on the type of the component; a reconstruction module, used to process the corresponding component using the selected image reconstruction model to obtain a reconstructed component; and a combination module, used to combine all the reconstructed components to obtain a reconstructed image.
[0010] To solve the above technical problems, the fifth technical solution provided by the present invention is: to provide a training method for an image reconstruction model, comprising: obtaining multiple sample images, each of the sample images including at least one type of component; training the initial network model using different types of components respectively to obtain an image reconstruction model corresponding to each component, and the image reconstruction model is used to implement the above-mentioned image reconstruction method.
[0011] To solve the above technical problems, the sixth technical solution provided by the present invention is: to provide a training device for an image reconstruction model, comprising: an acquisition module, used to acquire multiple sample images, each of the sample images includes at least one type of component; a training module, used to train the initial network model using different types of components respectively to obtain an image reconstruction model corresponding to each component, and the image reconstruction model is used to implement the image reconstruction method described above.
[0012] To solve the above technical problems, the seventh technical solution provided by the present invention is: providing an image reconstruction method, comprising: obtaining an input image, the input image including at least one type of component; determining an image reconstruction model corresponding to the component from a model set based on the type of the component; the model set includes a first image reconstruction model and a second image reconstruction model, the first image reconstruction model is trained by any of the methods described above; using the selected image reconstruction model to process the corresponding component to obtain a reconstructed component; and combining all the reconstructed components to obtain a reconstructed image.
[0013] To solve the above technical problems, the eighth technical solution provided by the present invention is: to provide an image reconstruction device, comprising: an acquisition module, used to acquire an input image, the input image comprising at least one type of component; a selection module, used to determine the image reconstruction model corresponding to the component from a model set based on the type of the component; the model set comprises a first image reconstruction model and a second image reconstruction model, the first image reconstruction model is trained by any of the methods described above; a reconstruction module, used to process the corresponding component using the selected image reconstruction model to obtain a reconstructed component; and a combination module, used to combine all the reconstructed components to obtain a reconstructed image.
[0014] To solve the above technical problems, the ninth technical solution provided by the present invention is: providing an image decoding method, comprising: obtaining a code stream, the code stream including a filter tag, the filter tag representing the type of an image reconstruction model; decoding the code stream to obtain an input image, the input image including at least one type of component; determining the image reconstruction model corresponding to the component from a model set based on the type of the component and the filter tag; the model set including a first image reconstruction model and a second image reconstruction model, the first image reconstruction model being trained by the method described above; using the selected image reconstruction model to process the corresponding component to obtain a reconstructed component; and combining all the reconstructed components to obtain a reconstructed image.
[0015] To solve the above technical problems, the tenth technical solution provided by the present invention is: to provide an image decoding device, comprising: an acquisition module, used to acquire a code stream, the code stream includes a filter tag, and the filter tag represents the type of an image reconstruction model; a decoding module, used to decode the code stream to obtain an input image, the input image includes at least one type of component; a selection module, used to determine the image reconstruction model corresponding to the component from a model set based on the type of the component and the filter tag; the model set includes a first image reconstruction model and a second image reconstruction model, and the first image reconstruction model is obtained by training the method described above; a reconstruction module, used to process the corresponding component using the selected image reconstruction model to obtain a reconstructed component; and a combination module, used to combine all the reconstructed components to obtain a reconstructed image.
[0016] To solve the above technical problems, the eleventh technical solution provided by the present invention is: to provide an image encoding method, comprising: obtaining an image to be encoded, the image to be encoded including at least one type of component; determining an image reconstruction model corresponding to the component based on the type of the component; using the selected image reconstruction model to process the component corresponding to the reference image of the image to be encoded to obtain a reconstructed component; encoding the image to be encoded based on the reconstructed component to obtain a code stream.
[0017] To solve the above technical problems, the twelfth technical solution provided by the present invention is: to provide an image encoding device, including: an acquisition module, used to acquire an image to be encoded, the image to be encoded includes at least one type of component; a determination module, used to determine the image reconstruction model corresponding to the component based on the type of the component; a reconstruction module, used to use the selected image reconstruction model to process the corresponding component of the reference image of the image to be encoded to obtain a reconstructed component; an encoding module, used to encode the image to be encoded based on the reconstructed component to obtain a code stream.
[0018] To solve the above technical problems, the thirteenth technical solution provided by the present invention is: to provide an electronic device, comprising a processor and a memory coupled to each other, wherein the memory is used to store program instructions for implementing any of the above methods; and the processor is used to execute the program instructions stored in the memory.
[0019] In order to solve the above technical problems, the fourteenth technical solution provided by the present invention is: providing a computer-readable storage medium storing a program file, and the program file can be executed to implement any of the above methods.
[0020] The beneficial effect of the present invention is different from the prior art. The present invention determines the image reconstruction model corresponding to the component based on the type of the component; uses the selected image reconstruction model to process the corresponding component to obtain the reconstructed component. The method takes the type of component into consideration during the image reconstruction process, and uses different types of image reconstruction models for different components to perform image reconstruction, which can improve the image reconstruction effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work, among which:
[0022] Figure 1 is a schematic flow chart of a first embodiment of an image reconstruction method of the present invention;
[0023] Figure 2 for Figure 1 A schematic flow chart of an embodiment of step S13;
[0024] Figure 3 for Figure 1 A schematic flow chart of another embodiment of step S13;
[0025] Figure 4for Figure 3 A schematic diagram of an embodiment of a first convolution module in FIG.
[0026] Figure 5 for Figure 3 A schematic diagram of a first embodiment of a residual module;
[0027] Figure 6a-6b for Figure 3 A schematic diagram of a second embodiment of a residual module in FIG.
[0028] Figure 7 for Figure 3 A schematic diagram of a third embodiment of a residual module;
[0029] Figure 8 for Figure 7 A schematic diagram of an embodiment of an attention processing module;
[0030] Fig. 9 for Figure 3 A schematic diagram of a fourth embodiment of a residual module in FIG.
[0031] Fig.10 for Figure 3 A schematic diagram of an embodiment of an up-sampling module;
[0032] Figure 11a-Figure 11b for Figure 3 A schematic diagram of an embodiment of the third convolution module;
[0033] Fig.12 A schematic structural diagram of an embodiment of an image reconstruction method of the present invention;
[0034] Fig.13 A schematic diagram of a flow chart of an embodiment of a method for training an image reconstruction model of the present invention;
[0035] Fig.14 A schematic structural diagram of an embodiment of a training device for an image reconstruction model of the present invention;
[0036] Fig.15 A schematic diagram of another embodiment of the image reconstruction method of the present invention is shown in FIG.
[0037] Fig.16 is a schematic structural diagram of another embodiment of the image reconstruction device of the present invention;
[0038] Fig.17 A schematic diagram of a flow chart of an embodiment of an image decoding method of the present invention;
[0039] Fig.18 A schematic diagram of the structure of an image decoding device according to an embodiment of the present invention;
[0040] Fig.19A schematic diagram of a flow chart of another embodiment of an image decoding method of the present invention;
[0041] Fig. 20 is a structural schematic diagram of another embodiment of an image decoding device of the present invention;
[0042] Fig.21 A schematic diagram of a flow chart of an embodiment of an image encoding method of the present invention;
[0043] Fig. 22 A schematic diagram of the structure of an embodiment of an image encoding device of the present invention;
[0044] Fig.23 A schematic structural diagram of an embodiment of an electronic device of the present invention;
[0045] Fig.24 It is a schematic diagram of the structure of the computer-readable storage medium of the present invention. Specific implementation methods
[0046] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0047] The present invention is described in detail below with reference to the accompanying drawings and embodiments.
[0048] See also Figure 1 , is a flowchart of a first embodiment of an image reconstruction method of the present invention, which specifically includes:
[0049] Step S11: Acquire an input image, where the input image includes at least one type of component.
[0050] In the prior art, the YUV components of the input image are trained together to obtain an image reconstruction model, and the YUV components are processed using the image reconstruction model during image reconstruction. However, the weights of the YUV components are different during training, which easily leads to poor training effects of the UV components, and poor image reconstruction effects corresponding to the UV components during image reconstruction.
[0051] In the present application, different components are trained separately to improve the image reconstruction effect. Specifically, an image reconstruction model is trained using the Y component, and the input of the image reconstruction model is the Y component, and the output is the Y component. An image reconstruction model is trained using the U component, and the input of the image reconstruction model is the U component, and the output is the U component. An image reconstruction model is trained using the V component, and the input of the image reconstruction model is the V component, and the output is the V component.
[0052] Alternatively, in another embodiment, an image reconstruction model is trained using the Y component, the input of the image reconstruction model is the Y component, and the output is the Y component. An image reconstruction model is trained using the UV component, the input of the image reconstruction model is the UV component, and the output is the UV component.
[0053] In the prior art, when the image reconstruction model is trained, the input lacks auxiliary information, that is, the network input and the network output have the same components, which is not conducive to improving the reconstruction result. The present application adds additional information when training the image reconstruction model to improve the reconstruction effect of the image reconstruction model. The additional information includes at least one of an additional component corresponding to the component, a quantization parameter map corresponding to the component, an intra-frame and inter-frame prediction value corresponding to the component, and prior information of the component.
[0054] Specifically, in one embodiment, an image reconstruction model is trained using the Y component and the UV component, the input of the image reconstruction model is the YUV component, and the output is the Y component, wherein the UV component is additional information, specifically, the UV component is an additional component corresponding to the Y component.
[0055] Alternatively, in one embodiment, an image reconstruction model is trained using the Y component and QPmap, the input of the image reconstruction model is the Y component and QPmap, and the output is the Y component. Wherein, QPmap is a quantization parameter map corresponding to the Y component. Alternatively, in one embodiment, an image reconstruction model is trained using the U component and QPmap, the input of the image reconstruction model is the U component and QPmap, and the output is the U component. Wherein, QPmap is a quantization parameter map corresponding to the U component. Alternatively, in one embodiment, an image reconstruction model is trained using the V component and QPmap, the input of the image reconstruction model is the V component and QPmap, and the output is the V component. Wherein, QPmap is a quantization parameter map corresponding to the V component. Alternatively, in one embodiment, an image reconstruction model is trained using the UV component and QPmap, the input of the image reconstruction model is the UV component and QPmap, and the output is the UV component. Wherein, QPmap is a quantization parameter map corresponding to the UV component.
[0056] After the image reconstruction model is pre-trained, during the image reconstruction process, an input image is obtained, where the input image includes at least one type of component, for example, the input image includes a Y component, a U component, and a V component.
[0057] Step S12: determining the image reconstruction model corresponding to the component based on the type of the component.
[0058] The image reconstruction model corresponding to the component is determined based on the type of the component. In the prior art, the YUV components are reconstructed using the image reconstruction model trained by the YUV components together. In the present application, the YUV components are reconstructed using their own trained image reconstruction models. It can be understood that the number of image reconstruction models in the present application is greater than 1.
[0059] Specifically, the image reconstruction model corresponding to the Y component is determined from the trained image reconstruction model, for example, the image reconstruction model output as the Y component is determined to be the image reconstruction model corresponding to the Y component. The image reconstruction model corresponding to the U component is determined from the trained image reconstruction model, for example, the image reconstruction model output as the U component is determined to be the image reconstruction model corresponding to the U component. The image reconstruction model corresponding to the V component is determined from the trained image reconstruction model, for example, the image reconstruction model output as the V component is determined to be the image reconstruction model corresponding to the V component. Another example. The image reconstruction model corresponding to the UV component is determined from the trained image reconstruction model, for example, the image reconstruction model output as the UV component is determined to be the image reconstruction model corresponding to the UV component.
[0060] Different types of components are determined based on quantization parameters within different ranges; the initial network model is trained using the determined different types of components to obtain an image reconstruction model for each component. Specifically, different quantization parameters characterize the degree of compression and distortion of the component. For example, components with quantization parameters within the range of 22 to 32 can be determined, and the initial network model is trained based on different types of components with quantization parameters within the range of 22 to 32, thereby obtaining an image reconstruction model corresponding to each component.
[0061] In one embodiment, for example, a model may be trained for different types of components corresponding to each QP (quantization parameter). For example, the Y component, U component, and V component are determined when the QP is 22, and a model is trained for each of the determined Y component, U component, and V component. The Y component, U component, and V component are determined when the QP is 27, and a model is trained for each of the determined Y component, U component, and V component. The Y component, U component, and V component are determined when the QP is 32, and a model is trained for each of the determined Y component, U component, and V component. The Y component, U component, and V component are determined when the QP is 37, and a model is trained for each of the determined Y component, U component, and V component. The Y component, U component, and V component are determined when the QP is 42, and a model is trained for each of the determined Y component, U component, and V component.
[0062] In another embodiment, a model may be trained for each of a plurality of different types of components whose quantization parameters QP are within a range. For example, a model a is trained for each of the Y component, U component, and V component whose QP is within the range of 22-32; wherein the QP within the range of 22-32 may be 22, 27, and 32. A model b is trained for each of the Y component, U component, and V component whose QP is within the range of 32-42; wherein the QP within the range of 32-42 may be 32, 37, and 42. During image reconstruction and encoding and decoding, the quantization parameter is compared with a preset value, and the image reconstruction model corresponding to the component is determined based on the type of component and the comparison result. For example, model a is selected when the QP is not greater than the preset value 32, and model b is selected when the QP is greater than the preset value 32. For another example, the QP combinations are {17, 22, 27}, {22, 27, 32}, {27, 32, 37}, {32, 37, 42}, {37, 42, 47}, and models 1, 2, 3, 4, and 5 are trained respectively. Model 1 is selected when QP is closest to 22, model 2 is selected when QP is closest to the preset value 27, model 3 is selected when QP is closest to the preset value 32, model 4 is selected when QP is closest to the preset value 37, and model 5 is selected when QP is closest to the preset value 42. The degree of quantization distortion is determined by the quantization parameter QP. Generally, the larger the QP, the greater the distortion caused by quantization, and vice versa. When encoding a video sequence, the QP of each image frame is based on the QP of the sequence and varies within a certain range according to the encoding configuration. That is, the quantization parameter characterizes the degree of distortion of the input image.
[0063] When determining an image reconstruction model corresponding to a component based on the type of the component, the image reconstruction model corresponding to the component is determined based on the quantization parameter and the type of the component in combination with the quantization parameter.
[0064] Step S13: using the selected image reconstruction model to process the corresponding components to obtain reconstructed components.
[0065] Specifically, the Y component is processed by the image reconstruction model whose output is the Y component to obtain the reconstructed component of the Y component; the U component is processed by the image reconstruction model whose output is the U component to obtain the reconstructed component of the U component; the V component is processed by the image reconstruction model whose output is the V component to obtain the reconstructed component of the V component; for another example, the UV component is processed by the image reconstruction model whose output is the UV component to obtain the reconstructed component of the UV component.
[0066] In one embodiment, when the image reconstruction model corresponding to the component is determined, the input information of the image reconstruction model is further determined, for example, whether the input information of the image reconstruction model includes additional information is determined. If the input information of the image reconstruction model includes additional information, the component is combined with the additional information, and the combined component and the additional information are processed using the selected image reconstruction model. For example, if the input information of the image reconstruction model corresponding to the Y component includes additional information, the Y component is combined with the additional information, and the combined Y component and the additional information are processed using the selected image reconstruction model.
[0067] Step S14: combining all the reconstruction components to obtain a reconstructed image.
[0068] Specifically, the reconstruction components of the Y component, the reconstruction components of the U component, and the reconstruction components of the V component are combined to obtain a reconstructed image. Alternatively, the reconstruction components of the Y component and the reconstruction components of the UV component are combined to obtain a reconstructed image.
[0069] In this embodiment, each component is used to train the image reconstruction model corresponding to each component, so that during the image reconstruction process, the corresponding component can be reconstructed using the corresponding image reconstruction model, thereby improving the image reconstruction effect. Furthermore, the component and the additional information corresponding to the component are used to train the image reconstruction model corresponding to each component, so that during the image reconstruction process, the corresponding component and the additional information can be reconstructed using the corresponding image reconstruction model, thereby improving the image reconstruction effect.
[0070] See also Figure 2 , Figure 2 for Figure 1 The flowchart of an embodiment of step S13 in the embodiment specifically includes:
[0071] Step S21: Process the components using the first convolution module.
[0072] For details, please combine Figure 3 , using the first convolution module to process the components. In the prior art, low-level feature extraction generally uses a [3×3×3×M] convolution layer to process the input image. There is a lack of reuse of low-level features, that is, the output extracted by the current convolution layer is only used as the input of the next level, and is not used as input across levels, which reduces the fluidity of the features.
[0073] This application improves the low-level feature extraction part and uses the first convolution module to process the components. Figure 4 , Figure 4Schematic diagram of the first convolution module. The components are processed sequentially using n first processing modules 41 and first convolution layers; wherein n is an integer greater than or equal to 0, and the first processing module 41 includes convolution layers and activation layers cascaded sequentially. Specifically, the components can be processed sequentially using one first processing module 41 and a first convolution layer; the components can also be processed sequentially using two first processing modules 41 and a first convolution layer, without specific limitation. In this way, the convolution depth of these parts can be increased or a more complex convolution layer can be used to effectively extract features.
[0074] Step S22: using a residual module to process the output of the first convolution module.
[0075] In one embodiment, if Figure 5 As shown, using the residual module to process the output of the first convolution module includes: sequentially using m second processing modules 42 (including Conv1, ReLU1) and a second convolution layer (including Conv2) to process the output of the first convolution module. In this embodiment, the number of second processing modules 42 is 1. In another embodiment, m is greater than 1. Among them, the second processing module 42 includes convolution layers and activation layers cascaded in sequence. Specifically, the input of the residual module is a 128×128×64 feature map, which is processed by a [64×3×3×64] convolution layer and an activation layer ReLU1 of a second processing module 42, and then processed by a [64×3×3×64] second convolution layer Conv2, and finally outputs a 128×128×64 feature map.
[0076] In this embodiment, the residual module increases the number of convolution layers to improve feature extraction performance.
[0077] In another embodiment, the residual module further includes feature reuse. Specifically, the residual module sequentially processes the output of the first convolution module using m second processing modules 42 and a second convolution layer. The input of the mth second processing module 42 is the output of the math second processing module 42 and the output of the ma-1th second processing module 42, and / or the input of the second convolution layer is the output of the mth second processing module 42 and the output of the math second processing module 42; the output of the second convolution layer is superimposed with the output of the first convolution module as the output of the residual module; wherein, m is greater than or equal to 1, a is less than m, and a and m are positive integers, and the second processing module 42 includes sequentially cascaded convolution layers and activation layers.
[0078] See also Figure 6a, the input of the residual module is a 128×128×64 feature map, and the network structure consists of three second processing modules 42 (including Conv1, Relu1, conv2, Relu2, Conv3, Relu3) and a second convolutional layer (Conv4), as shown in detail Figure 6a As shown. The input of the third second processing module 42 is the output of the second second processing module 42 and the output of the first second processing module 42; and the input of the second convolution layer is the output of the third second processing module 42 and the output of the second second processing module 42. Finally, the output of the second convolution layer is superimposed with the input of the residual module (that is, the output of the first convolution module) as the output of the residual module, and finally a 128×128×64 feature map is obtained. Figure 6a In the embodiment, a=1, m=3.
[0079] See also Figure 6b , the input of the residual module is a 128×128×64 feature map, and the network structure consists of three second processing modules 42 (including Conv1, Relu1, conv2, Relu2, Conv3, Relu3) and a second convolutional layer (Conv4), as shown in detail Figure 6b As shown. The input of the second convolution layer is the output of the first second processing module 42 and the output of the third second processing module 42. Finally, the output of the second convolution layer is superimposed with the input of the residual module (that is, the output of the first convolution module) as the output of the residual module, and finally a 128×128×64 feature map is obtained. Figure 6b In the embodiment, a=2, m=3.
[0080] In this embodiment, feature reuse is used to increase the flow performance of features, improve the utilization rate of features, and accelerate the convergence speed. Feature reuse methods include but are not limited to addition and stacking. Addition means adding two feature maps of the same dimension at each feature point, and the feature dimension does not change. For example, two 128×128×64 feature maps, the i×j×kth points of the two maps are added, and the feature dimension after addition is still 128×128×64; stacking means splicing on the feature channel, for example, two 128×128×64 feature maps are spliced into a 128×128×128 feature map.
[0081] In another embodiment, if Figure 7 As shown, the residual module further includes an attention processing module AB. In this embodiment, the output of the first convolution module is processed using m second processing modules 42, a second convolution layer, and an attention processing module AB; wherein m is greater than or equal to 1, and m is a positive integer, and the second processing module 42 includes a convolution layer and an activation layer cascaded in sequence.
[0082] Specifically, Figure 7 As shown, m=1, and the output of the first convolution module is processed by a second processing module 42 and a second convolution layer in sequence. That is, the output of the first convolution module is processed by Conv1, Relu1, and Conv4 (i.e., the second convolution layer) in sequence. The output of the second convolution layer is processed by the attention processing module AB. The output of the second convolution layer is multiplied by the output of the attention processing module, and the stacked result is superimposed with the output of the first convolution module, and the superimposed result is used as the output of the residual module.
[0083] The attention processing module AB can be used to distinguish the importance of different features and adaptively assign weights to feature maps. The structure of the attention processing module AB is as follows: Figure 8 As shown, it includes a pooling layer, t third processing modules 43, a third convolution layer, and an output activation layer. The step of using the attention processing module AB to process the output of the first convolution module and the output of the second convolution layer includes: using the pooling layer, t third processing modules 43, the third convolution layer, and the output activation layer to process the output of the first convolution module and the output of the second convolution layer in sequence. Wherein, t is greater than or equal to 1, and t is a positive integer, and the output activation layer includes at least one of Sigmoid and Softmax.
[0084] The attention processing module AB first compresses the input feature image through the pooling layer, then extracts and activates the feature through t third processing modules 43, then extracts the feature through the third convolution layer, and finally outputs it through the output activation layer. In addition, the input of the module includes but is not limited to the previous level output, network input, low-level feature extraction, feature reuse, etc., the pooling layer includes but is not limited to maximum pooling, average pooling, etc., and the output activation layer includes but is not limited to Sigmoid and Softmax.
[0085] In another embodiment, if Fig. 9 As shown, in this embodiment, both the attention processing module and the feature reuse are introduced. Specifically, in this embodiment, the input of the second convolutional layer Conv4 is the output of the mth second processing module 42 and the output of the math second processing module 42. Assuming m=2, a=1, in this embodiment, the input of the second convolutional layer is the output of the second second processing module 42 and the output of the first second processing module 42. In this embodiment, the second processing module 42 includes Conv1, Relu1, conv2, Relu2.
[0086] This embodiment introduces an attention processing module AB, which can be used to distinguish the importance of different features and adaptively assign weights to feature maps.
[0087] Step S23: using a second convolution module to process the output of the residual module.
[0088] In this embodiment, the structure of the second convolution module is the same as that of the first convolution module, as shown in the following example: Figure 4 The output of the residual module is processed by using h fourth processing modules and fourth convolutional layers in sequence; wherein h is greater than or equal to 0 and is a positive integer, and the fourth processing module includes convolutional layers and activation layers that are cascaded in sequence.
[0089] Step S24: using an upsampling module to process the output of the first convolution module and the output of the second convolution module.
[0090] like Fig.10 As shown, the output of the first convolution module and the output of the second convolution module are processed in sequence using g convolution layers and shuffle layers, where g is greater than or equal to 1 and g is a positive integer.
[0091] See also Figure 3 , the upsampling module processes the first superposition result, and the first superposition result is the output of the first convolution module and the superposition result of the second convolution module. That is, in this embodiment, the first superposition result is processed by g convolution layers and shuffle layers in sequence.
[0092] Step S25: using a third convolution module to process the output of the up-sampling module to obtain the reconstructed component.
[0093] As shown in FIG. 11 , the output of the upsampling module is processed in sequence using the fifth convolutional layer, d residual units, the sixth convolutional layer, the seventh convolutional layer, the shuffle layer, e residual units or e eighth convolutional layers to obtain the reconstructed component.
[0094] Specifically, Fig.11a As shown, the output of the upsampling module is processed in sequence using the fifth convolutional layer, d residual units, the sixth convolutional layer, the seventh convolutional layer, the shuffle layer, and e residual units to obtain the reconstructed component.
[0095] like Fig.11b As shown, the output of the upsampling module is processed in sequence using the fifth convolutional layer, d residual units, the sixth convolutional layer, the seventh convolutional layer, the shuffle layer, and e eighth convolutional layers to obtain the reconstructed components.
[0096] The input of the seventh convolutional layer is a second superposition result, and the second superposition result is a superposition result of the output of the fifth convolutional layer and the output of the sixth convolutional layer.
[0097] In a specific embodiment, d is 16 and e is 1.
[0098] In a specific embodiment, among the d residual units, the structure of each residual unit is as follows: Fig. 9 As shown in Figure 2. Among the e residual units, the structure of each residual unit is as follows: Figure 5 shown.
[0099] The input of the image reconstruction model of the present application can be single-input single-output, multiple-input single-output, or multiple-input multiple-output, that is, the input may or may not contain additional information. For the form of single-input single-output, the model effect of a single component can be effectively improved; for the form of multiple-input single-output, the model effect of a single component can be further improved; for multiple-input multiple-output, the effect of the multi-component output model can be improved. In addition, the multiple-input case can also be used to reduce the number of models, that is, using variables as additional information to improve the network's fitting performance for variables, and avoiding training a model with a single or multiple variable values.
[0100] The image reconstruction model of the present application improves the feature extraction performance by adding convolutional layers and activation layers.
[0101] The advanced residual block designed in the image reconstruction model of the present application may include increasing the number of convolutional layers, feature reuse, and attention processing modules, which increases the flow of features and distinguishes the importance of feature maps, which can effectively improve the performance of the residual block.
[0102] The image reconstruction model of the present application can be used to replace or serve as a candidate upsampling module in the encoding and decoding process, or can be used independently for image super-resolution reconstruction.
[0103] See also Fig.12 , is a schematic diagram of the structure of an embodiment of the image reconstruction device of the present invention, which specifically includes: an acquisition module 121, a determination module 122, a reconstruction module 123 and a combination module 124.
[0104] The acquisition module 121 is used to acquire an input image, where the input image includes at least one type of component.
[0105] The determination module 122 is used to determine the image reconstruction model corresponding to the component based on the type of the component, and the total number of the image reconstruction models is greater than 1.
[0106] The reconstruction module 123 is used to process the corresponding components using the selected image reconstruction model to obtain reconstructed components.
[0107] The combining module 124 is used to combine all the reconstruction components to obtain a reconstructed image.
[0108] See also Fig.13 , is a flow chart of an embodiment of a training method for an image reconstruction model of the present invention, comprising:
[0109] Step S51: Acquire multiple sample images, each of which includes at least one type of component.
[0110] A plurality of sample images are obtained, each of which includes at least one type of component, and specifically, the component includes a YUV component.
[0111] Step S52: using different types of components to train the initial network model respectively, so as to obtain an image reconstruction model corresponding to each component.
[0112] In the present application, different components are trained separately to improve the image reconstruction effect. Specifically, an image reconstruction model is trained using the Y component, and the input of the image reconstruction model is the Y component, and the output is the Y component. An image reconstruction model is trained using the U component, and the input of the image reconstruction model is the U component, and the output is the U component. An image reconstruction model is trained using the V component, and the input of the image reconstruction model is the V component, and the output is the V component.
[0113] Alternatively, in another embodiment, an image reconstruction model is trained using the Y component, the input of the image reconstruction model is the Y component, and the output is the Y component. An image reconstruction model is trained using the UV component, the input of the image reconstruction model is the UV component, and the output is the UV component.
[0114] In one embodiment, different types of components and quantization parameters are used to train the initial network model respectively to obtain an image reconstruction model corresponding to each component; the quantization parameter represents the degree of coding distortion.
[0115] In one embodiment of the present application, different types of components are determined based on quantization parameters within different ranges; the initial network model is trained using the determined different types of components to obtain an image reconstruction model for each component. Specifically, different quantization parameters characterize the degree of compression and distortion of the component. For example, components with quantization parameters within the range of 22 to 32 can be determined, and the initial network model is trained based on different types of components with quantization parameters within the range of 22 to 32, thereby obtaining an image reconstruction model corresponding to each component.
[0116] In one embodiment, for example, a model may be trained for different types of components corresponding to each QP (quantization parameter). For example, the Y component, U component, and V component are determined when the QP is 22, and a model is trained for each of the determined Y component, U component, and V component. The Y component, U component, and V component are determined when the QP is 27, and a model is trained for each of the determined Y component, U component, and V component. The Y component, U component, and V component are determined when the QP is 32, and a model is trained for each of the determined Y component, U component, and V component. The Y component, U component, and V component are determined when the QP is 37, and a model is trained for each of the determined Y component, U component, and V component. The Y component, U component, and V component are determined when the QP is 42, and a model is trained for each of the determined Y component, U component, and V component.
[0117] In another embodiment, a plurality of different types of components with quantization parameters QP within a range may be trained with a model respectively. For example, a Y component, a U component, and a V component with QP within a range of 22-32 may each train a model a, wherein the QP within a range of 22-32 may be 22, 27, and 32. A Y component, a U component, and a V component with QP within a range of 32-42 may each train a model b, wherein the QP within a range of 32-42 may be 32, 37, and 42. During image reconstruction and encoding and decoding, the quantization parameter is compared with a preset value, and the image reconstruction model corresponding to the component is determined based on the type of component and the comparison result. For example, when the QP is not greater than the preset value 32, model a is selected, and when the QP is greater than the preset value 32, model b is selected.
[0118] For another example, the QP combinations are {17, 22, 27}, {22, 27, 32}, {27, 32, 37}, {32, 37, 42}, {37, 42, 47}, and models 1, 2, 3, 4, and 5 are trained respectively. Model 1 is selected when QP is closest to 22, model 2 is selected when QP is closest to the preset value 27, model 3 is selected when QP is closest to the preset value 32, model 4 is selected when QP is closest to the preset value 37, and model 5 is selected when QP is closest to the preset value 42. The degree of quantization distortion is determined by the quantization parameter QP. Generally, the larger the QP, the greater the distortion caused by quantization, and vice versa. When encoding a video sequence, the QP of each image frame is based on the QP of the sequence and varies within a certain range according to the encoding configuration. That is, the quantization parameter characterizes the degree of distortion of the input image.
[0119] In another embodiment, the initial network model can also be trained using different types of components and additional information to obtain an image reconstruction model corresponding to each component; the additional information includes at least one of the additional component corresponding to the component, the quantization parameter map corresponding to the component, the intra-frame and inter-frame prediction values corresponding to the component, and the prior information of the component. For example, an image reconstruction model is trained using the Y component and QPmap, and the input of the image reconstruction model is the Y component and QPmap, and the output is the Y component. Wherein, QPmap is the quantization parameter map corresponding to the Y component. Alternatively, in one embodiment, an image reconstruction model is trained using the U component and QPmap, and the input of the image reconstruction model is the U component and QPmap, and the output is the U component. Wherein, QPmap is the quantization parameter map corresponding to the U component. Alternatively, in one embodiment, an image reconstruction model is trained using the V component and QPmap, and the input of the image reconstruction model is the V component and QPmap, and the output is the V component. Wherein, QPmap is the quantization parameter map corresponding to the V component. Alternatively, in one embodiment, an image reconstruction model is trained using UV components and QPmap, the input of the image reconstruction model is UV components and QPmap, and the output is UV components, wherein QPmap is a quantization parameter map corresponding to the UV components.
[0120] In the prior art, when the image reconstruction model is trained, the input lacks auxiliary information, that is, the network input and the network output have the same components, which is not conducive to improving the reconstruction result. The present application adds additional information when training the image reconstruction model to improve the reconstruction effect of the image reconstruction model. The additional information includes at least one of an additional component corresponding to the component, a quantization parameter map corresponding to the component, an intra-frame and inter-frame prediction value corresponding to the component, and prior information of the component.
[0121] See also Fig.14 , is a structural diagram of an embodiment of a training device for an image reconstruction model of the present invention, which specifically includes an acquisition module 141 and a training module 142.
[0122] The acquisition module 141 is used to acquire multiple sample images, each of which includes at least one type of component. The training module 142 is used to train the initial network model using different types of components to obtain an image reconstruction model corresponding to each component.
[0123] See also Fig.15 , is a flow chart of another embodiment of the image reconstruction method of the present invention, comprising:
[0124] Step S61: Acquire an input image, wherein the input image includes at least one type of component.
[0125] Step S62: Determine the image reconstruction model corresponding to the component from a model set based on the type of the component, the total number of image reconstruction models in the model set is greater than 1; the model set includes a first image reconstruction model and a second image reconstruction model, and the first image reconstruction model is trained by any one of the methods described.
[0126] Specifically, the first image reconstruction model can be Figure 3 to Figure 1 The method shown in 1 performs image reconstruction on the input image. The second image reconstruction model can be an existing image reconstruction model.
[0127] Step S63: using the selected image reconstruction model to process the corresponding components to obtain reconstructed components.
[0128] Step S64: combine all the reconstructed components to obtain a reconstructed image.
[0129] See also Fig.16 , Fig.16 It is a structural schematic diagram of another embodiment of the image reconstruction device of the present invention, comprising: an acquisition module 161 , a selection module 162 , a reconstruction module 163 and a combination module 164 .
[0130] The acquisition module 161 is used to acquire an input image, where the input image includes at least one type of component.
[0131] The selection module 162 is used to determine the image reconstruction model corresponding to the component from a model set based on the type of the component, and the total number of image reconstruction models in the model set is greater than 1; the model set includes a first image reconstruction model and a second image reconstruction model, and the first image reconstruction model is trained by any of the methods described above.
[0132] The reconstruction module 163 is used to process the corresponding components using the selected image reconstruction model to obtain reconstructed components.
[0133] The combining module 164 is used to combine all the reconstruction components to obtain a reconstructed image.
[0134] See also Fig.17 , is a flow chart of an embodiment of an image decoding method of the present invention, which specifically includes:
[0135] Step S171: Obtain code stream.
[0136] Step S172: Decode the code stream to obtain an input image, where the input image includes at least one type of component.
[0137] The bitstream obtained after encoding is obtained, and the bitstream is decoded to obtain the input image. Specifically, if the bitstream is a bitstream obtained by image encoding, the input image is obtained after decoding, and if the bitstream is a bitstream obtained by video encoding, the video is obtained after decoding, and the input image is a frame in the video.
[0138] Step S173: Determine the image reconstruction model corresponding to the component based on the type of the component.
[0139] Step S174: using the selected image reconstruction model to process the corresponding components to obtain reconstructed components.
[0140] Step S175: combine all the reconstructed components to obtain a reconstructed image.
[0141] See also Fig.18 , a schematic structural diagram of an embodiment of an image decoding device of the present invention, comprising an acquisition module 181, a decoding module 182, a determination module 183, a reconstruction module 184 and a combination module 185.
[0142] The acquisition module 181 is used to acquire a code stream.
[0143] The decoding module 182 is used to decode the code stream to obtain an input image, where the input image includes at least one type of component.
[0144] The acquisition module 181 acquires the bit stream obtained after encoding, and the decoding module 182 decodes the bit stream to obtain the input image. Specifically, if the bit stream is a bit stream obtained by image encoding, the input image is obtained after decoding, and if the bit stream is a bit stream obtained by video encoding, the video is obtained after decoding, and the input image is a frame in the video.
[0145] The determination module 183 is used to determine the image reconstruction model corresponding to the component based on the type of the component.
[0146] The reconstruction module 184 is used to process the corresponding components using the selected image reconstruction model to obtain reconstructed components.
[0147] The combination module 185 is used to combine all the reconstruction components to obtain a reconstructed image.
[0148] See also Fig.19 , is a flow chart of another embodiment of the image decoding method of the present invention, which specifically includes:
[0149] Step S191: Acquire a code stream, wherein the code stream includes a filter tag, and the filter tag represents the type of the image reconstruction model.
[0150] Step S192: Decode the code stream to obtain an input image, where the input image includes at least one type of component.
[0151] Step S193: determining the image reconstruction model corresponding to the component from a model set based on the type of the component and the filter tag; the model set includes a first image reconstruction model and a second image reconstruction model.
[0152] The first image reconstruction model is obtained by training through the method described above.
[0153] Specifically, the code stream includes a filter flag, which represents the type of the image reconstruction model. For example, the filter flag (i.e., the syntactic element) is defined as ARN_ENABLE_FLAG, which takes a value of 0 or 1. When the value is 0, it means that the image reconstruction model provided by this application is disabled; when the value is 1, it means that the image reconstruction model provided by this application is enabled.
[0154] In the decoding method of the present application, the corresponding image reconstruction model is determined based on the component type and the filter flag.
[0155] Step S194: using the selected image reconstruction model to process the corresponding components to obtain reconstructed components.
[0156] Step S195: combine all the reconstructed components to obtain a reconstructed image.
[0157] See also Fig. 20 , is a structural diagram of another embodiment of the image decoding device of the present invention, which specifically includes: an acquisition module 201, a decoding module 202, a selection module 203, a reconstruction module 204 and a combination module 205.
[0158] The acquisition module 201 is used to acquire a code stream, wherein the code stream includes a filter tag, and the filter tag represents the type of the image reconstruction model.
[0159] The decoding module 202 is used to decode the code stream to obtain an input image, where the input image includes at least one type of component.
[0160] The selection module 203 is used to determine the image reconstruction model corresponding to the component from a model set based on the type of the component and the filter label; the model set includes a first image reconstruction model and a second image reconstruction model.
[0161] The first image reconstruction model is obtained by training through the method described above.
[0162] Specifically, the code stream includes a filter flag, and the filter flag represents the type of the image reconstruction model. For example, the filter flag (that is, the syntactic element) is defined as ARN_ENABLE_FLAG, which takes a value of 0 or 1. When the value is 0, it means that the image reconstruction model provided by this application is disabled; when the value is 1, it represents the image reconstruction model provided by this application. It can be understood that if the value is 0, it means that the image reconstruction model of the prior art is currently used, and if the value is 1, it means that the image reconstruction model provided by this application is currently used.
[0163] In the decoding device of the present application, the corresponding image reconstruction model is determined based on the component type and the filter flag.
[0164] The reconstruction module 204 is used to process the corresponding components using the selected image reconstruction model to obtain reconstructed components.
[0165] The combination module 205 is used to combine all the reconstruction components to obtain a reconstructed image.
[0166] See also Fig.21 , is a flow chart of an embodiment of the image encoding method of the present invention, which specifically includes:
[0167] Step S211: Acquire a to-be-encoded image, where the to-be-encoded image includes at least one type of component.
[0168] Step S212: determining an image reconstruction model corresponding to the component based on the type of the component.
[0169] Step S213: using the selected image reconstruction model to process the corresponding components of the reference image of the image to be encoded to obtain reconstructed components.
[0170] Step S214: Encode the image to be encoded based on the reconstructed component to obtain a code stream.
[0171] Specifically, when encoding, if it is necessary to upsample the reference image of the image to be encoded (the reference image may be the previous frame of the image to be encoded), the image reconstruction model of the present application may be used to process the components of the reference image of the image to be encoded to obtain the reconstructed components. The image to be encoded is then encoded based on the reconstructed components to obtain a bitstream.
[0172] In another embodiment of the present application, an image reconstruction model corresponding to the component can be determined from a model set based on the type of the component; the model set includes a first image reconstruction model and a second image reconstruction model. The first image reconstruction model is trained by the image reconstruction model training method described in the present application. That is, in this embodiment, when it is necessary to upsample the reference image of the image to be encoded, a model of the type of the corresponding component is selected from the image reconstruction model of the present application and the image reconstruction model in the prior art to process the component to obtain a reconstructed component.
[0173] It can be understood that after the corresponding image reconstruction model is selected and during the encoding process, the filter flag is encoded. That is, the code stream includes a filter flag, and the filter flag represents the type of the image reconstruction model selected when encoding the image to be encoded. For example, the filter flag (that is, the syntactic element) is defined as ARN_ENABLE_FLAG, which takes a value of 0 or 1. When the value is 0, it means that the image reconstruction model provided by the present application is disabled; when the value is 1, it represents the image reconstruction model provided by the present application. It can be understood that if the value is 0, it means that the image reconstruction model of the prior art is currently used, and if the value is 1, it means that the image reconstruction model provided by the present application is currently used.
[0174] See also Fig. 22 , is a structural diagram of an embodiment of an image encoding device of the present invention, which specifically includes: an acquisition module 221, a determination module 222, a reconstruction module 223, and an encoding module 224.
[0175] The acquisition module 221 is used to acquire an image to be encoded, where the image to be encoded includes at least one type of component;
[0176] The determination module 222 is used to determine the image reconstruction model corresponding to the component based on the type of the component;
[0177] The reconstruction module 223 is used for processing the corresponding components of the reference image of the image to be encoded by using the selected image reconstruction model to obtain the reconstructed components;
[0178] The encoding module 224 is used to encode the image to be encoded based on the reconstructed component to obtain a code stream.
[0179] Specifically, when encoding, if it is necessary to upsample the reference image of the image to be encoded, the image reconstruction model of the present application can be used to process the components of the reference image of the image to be encoded to obtain the reconstructed components. Then, the image to be encoded is encoded based on the reconstructed components to obtain a bitstream.
[0180] In another embodiment of the present application, an image reconstruction model corresponding to the component can be determined from a model set based on the type of the component; the model set includes a first image reconstruction model and a second image reconstruction model. The first image reconstruction model is trained by the image reconstruction model training method described in the present application. That is, in this embodiment, when it is necessary to upsample the reference image of the image to be encoded, a model of the type of the corresponding component is selected from the image reconstruction model of the present application and the image reconstruction model in the prior art to process the component to obtain a reconstructed component.
[0181] It can be understood that after the corresponding image reconstruction model is selected and during the encoding process, the filter flag is encoded. That is, the code stream includes a filter flag, and the filter flag represents the type of the image reconstruction model selected when encoding the image to be encoded. For example, the filter flag (that is, the syntactic element) is defined as ARN_ENABLE_FLAG, which takes a value of 0 or 1. When the value is 0, it means that the image reconstruction model provided by the present application is disabled; when the value is 1, it represents the image reconstruction model provided by the present application. It can be understood that if the value is 0, it means that the image reconstruction model of the prior art is currently used, and if the value is 1, it means that the image reconstruction model provided by the present application is currently used.
[0182] See also Fig.23 , is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention. The electronic device includes a memory 82 and a processor 81 connected to each other.
[0183] The memory 82 is used to store program instructions for implementing any one of the above methods.
[0184] The processor 81 is used to execute program instructions stored in the memory 82 .
[0185] The processor 81 may also be referred to as a CPU (Central Processing Unit). The processor 81 may be an integrated circuit chip having signal processing capabilities. The processor 81 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0186] The memory 82 can be a memory stick, a TF card, etc., which can store all the information in the electronic device, including the input raw data, computer programs, intermediate operation results and final operation results are all stored in the memory. It stores and retrieves information according to the location specified by the controller. Only with the memory can the electronic device have a memory function and ensure normal operation. The memory of the electronic device can be divided into main memory (internal memory) and auxiliary memory (external memory) according to its use. There is also a classification method of dividing it into external memory and internal memory. External memory is usually a magnetic medium or an optical disk, etc., which can store information for a long time. Memory refers to the storage component on the motherboard, which is used to store the data and programs currently being executed, but it is only used to temporarily store programs and data. If the power is turned off or the power is cut off, the data will be lost.
[0187] In the several embodiments provided in the present application, it should be understood that the disclosed methods and devices can be implemented by other methods. For example, the device implementation method described above is only schematic. For example, the division of modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0188] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0189] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0190] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a system server, or a network device, etc.) or a processor to execute all or part of the steps of each implementation method of the present application.
[0191] See also Fig.24 , which is a schematic diagram of the structure of the computer-readable storage medium of the present invention. The storage medium of the present application stores a program file 91 that can implement all the above methods, wherein the program file 91 can be stored in the above storage medium in the form of a software product, including a number of instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) or a processor (processor) to execute all or part of the steps of each implementation method of the present application. The aforementioned storage device includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes, or terminal devices such as computers, servers, mobile phones, tablets, etc.
[0192] The above is only an implementation method of the present invention, and does not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. An image reconstruction method, characterized in that: include: Acquire an input image, wherein the input image includes at least one type of component; Determining an image reconstruction model corresponding to the component based on the type of the component and a quantization parameter, wherein the quantization parameter represents a degree of distortion of the input image; Processing the corresponding components using the selected image reconstruction model to obtain reconstructed components; All the reconstruction components are combined to obtain a reconstructed image.
2. The image reconstruction method according to claim 1, characterized in that: The step of using the selected image reconstruction model to process the corresponding component to obtain the reconstructed component comprises: Determining input information corresponding to the selected image reconstruction model; In response to the input information including additional information, the component is combined with the additional information, and the combined component and the additional information are processed using the selected image reconstruction model to obtain a reconstructed component; the additional information includes at least one of an additional component corresponding to the component, a quantization parameter map corresponding to the component, an intra-frame and inter-frame prediction value corresponding to the component, and prior information of the component.
3. The image reconstruction method according to claim 2, characterized in that: The step of determining the image reconstruction model corresponding to the component based on the type of the component and the quantization parameter comprises: Comparing the quantization parameter with a preset value; An image reconstruction model corresponding to the component is determined based on the type of the component and the comparison result.
4. The image reconstruction method according to claim 1, characterized in that: The step of using the selected image reconstruction model to process the corresponding component to obtain the reconstructed component comprises: Processing the components using a first convolution module; Processing the output of the first convolution module using a residual module; Processing the output of the residual module using a second convolution module; Using an upsampling module to process the output of the first convolution module and the output of the second convolution module; The output of the up-sampling module is processed by using a third convolution module to obtain the reconstructed component.
5. The image reconstruction method according to claim 4, characterized in that: The step of processing the components by using the first convolution module comprises: Processing the components using n first processing modules and first convolutional layers in sequence; Wherein, n is a positive integer greater than or equal to 0, and the first processing module includes a convolutional layer and an activation layer cascaded in sequence.
6. The image reconstruction method according to claim 4, characterized in that: The step of processing the output of the first convolution module by using the residual module includes: Processing the output of the first convolution module using m second processing modules and a second convolution layer in sequence; wherein m is greater than or equal to 1; The second processing module includes a convolutional layer and an activation layer cascaded in sequence.
7. The image reconstruction method according to claim 4, characterized in that: The step of processing the output of the first convolution module by using the residual module includes: Using m second processing modules and a second convolutional layer in sequence to process the output of the first convolutional module; The input of the mth second processing module is the output of the math second processing module and the output of the ma-1th second processing module, and / or the input of the second convolutional layer is the output of the mth second processing module and the output of the math second processing module; Superimposing the output of the second convolutional layer and the output of the first convolutional module as the output of the residual module; Wherein, m and a are positive integers greater than or equal to 1, a is less than m, and the second processing module includes a convolutional layer and an activation layer cascaded in sequence.
8. The image reconstruction method according to claim 4, characterized in that: The step of processing the output of the first convolution module by using the residual module includes: Processing the output of the first convolution module using m second processing modules, a second convolution layer, and an attention processing module; Wherein, m is a positive integer greater than or equal to 1, and the second processing module includes a convolutional layer and an activation layer cascaded in sequence.
9. The image reconstruction method according to claim 8, characterized in that: The step of processing the output of the first convolution module by using m second processing modules, a second convolution layer and an attention processing module comprises: Using m second processing modules and a second convolutional layer in sequence to process the output of the first convolutional module; Processing the output of the second convolutional layer using an attention processing module; The output of the second convolutional layer is multiplied by the output of the attention processing module, and the stacked result is superimposed with the output of the first convolutional module, and the superimposed result is used as the output of the residual module.
10. The image reconstruction method according to claim 9, characterized in that: The step of sequentially using m second processing modules and a second convolutional layer to process the output of the first convolutional module includes: The input of the second convolutional layer is the output of the mth second processing module and the output of the math second processing module.
11. The image reconstruction method according to claim 9, characterized in that: The step of processing the output of the first convolution module and the output of the second convolution layer by using the attention processing module comprises: Processing the output of the first convolution module and the output of the second convolution layer using a pooling layer, t third processing modules, a third convolution layer, and an output activation layer in sequence; Wherein, t is a positive integer greater than or equal to 1, and the output activation layer includes at least one of Sigmoid and Softmax.
12. The image reconstruction method according to claim 4, characterized in that: The step of processing the output of the residual module by using the second convolution module comprises: Processing the output of the residual module using h fourth processing modules and fourth convolutional layers in sequence; Wherein, h is a positive integer greater than or equal to 0, and the fourth processing module includes a convolutional layer and an activation layer cascaded in sequence.
13. The image reconstruction method according to claim 4, characterized in that: The step of processing the output of the first convolution module and the output of the second convolution module by using the upsampling module comprises: The output of the first convolution module and the output of the second convolution module are processed in sequence using g convolution layers and shuffle layers, where g is a positive integer greater than or equal to 1.
14. The image reconstruction method according to claim 13, characterized in that: The step of sequentially using g convolutional layers and shuffle layers to process the output of the first convolutional module and the output of the second convolutional module includes: The first superposition result is processed in sequence using g convolution layers and shuffle layers, where the first superposition result is the output of the first convolution module and the superposition result of the second convolution module.
15. The image reconstruction method according to claim 4, characterized in that: The step of using a third convolution module to process the output of the up-sampling module to obtain the reconstructed component comprises: Processing the output of the upsampling module by using the fifth convolutional layer, d residual units, the sixth convolutional layer, the seventh convolutional layer, the shuffle layer, e residual units or e eighth convolutional layers in sequence to obtain the reconstructed component; The input of the seventh convolutional layer is a second superposition result, and the second superposition result is a superposition result of the output of the fifth convolutional layer and the output of the sixth convolutional layer.
16. An image reconstruction device, characterized in that: include: An acquisition module, configured to acquire an input image, wherein the input image includes at least one type of component; A determination module, configured to determine an image reconstruction model corresponding to the component based on the type of the component and a quantization parameter, wherein the quantization parameter represents a degree of distortion of the input image; A reconstruction module, used for processing the corresponding components using the selected image reconstruction model to obtain reconstructed components; The combination module is used to combine all the reconstruction components to obtain a reconstructed image.
17. An image decoding method, characterized in that: include: Get the code stream; Decoding the code stream to obtain an input image, wherein the input image includes at least one type of component; Determining an image reconstruction model corresponding to the component based on the type of the component and a quantization parameter, wherein the quantization parameter represents a degree of distortion of the input image; Processing the corresponding components using the selected image reconstruction model to obtain reconstructed components; All the reconstruction components are combined to obtain a reconstructed image.
18. An image decoding device, characterized in that: include: Acquisition module, used to obtain code stream; A decoding module, used for decoding the code stream to obtain an input image, wherein the input image includes at least one type of component; A determination module, configured to determine an image reconstruction model corresponding to the component based on the type of the component and a quantization parameter, wherein the quantization parameter represents a degree of distortion of the input image; A reconstruction module, used for processing the corresponding components using the selected image reconstruction model to obtain reconstructed components; The combination module is used to combine all the reconstruction components to obtain a reconstructed image.
19. A training method for an image reconstruction model, characterized in that: include: Acquire a plurality of sample images, each of the sample images comprising at least one type of component; The initial network model is trained using different types of components and quantization parameters to obtain an image reconstruction model corresponding to each component, wherein the quantization parameter represents the degree of coding distortion, and the image reconstruction model is used to implement the image reconstruction method described in any one of claims 1 to 15 above.
20. The training method according to claim 19, characterized in that: The step of training the initial network model using different types of components and quantization parameters to obtain an image reconstruction model corresponding to each component includes: Determining different types of components based on quantization parameters within different ranges; The initial network model is trained using the determined different types of components respectively to obtain an image reconstruction model for each component.
21. The training method according to claim 19, characterized in that: The step of respectively training the initial network model using different types of components to obtain an image reconstruction model corresponding to each component includes: The initial network model is trained using different types of components and additional information to obtain an image reconstruction model corresponding to each component; the additional information includes at least one of an additional component corresponding to the component, a quantization parameter map corresponding to the component, an intra-frame and inter-frame prediction value corresponding to the component, and prior information of the component.
22. A training device for an image reconstruction model, characterized in that: include: An acquisition module, used for acquiring a plurality of sample images, each of the sample images comprising at least one type of component; A training module is used to train the initial network model using different types of components and quantization parameters to obtain an image reconstruction model corresponding to each component, wherein the quantization parameter represents the degree of coding distortion, and the image reconstruction model is used to implement the image reconstruction method described in any one of claims 1 to 15 above.
23. An image reconstruction method, characterized in that: include: Acquire an input image, wherein the input image includes at least one type of component; Determining an image reconstruction model corresponding to the component from a model set based on the type of the component and a quantization parameter, wherein the quantization parameter represents the degree of distortion of the input image; the model set comprises a first image reconstruction model and a second image reconstruction model, wherein the first image reconstruction model is trained by the method described in any one of claims 19 to 21 above; Processing the corresponding components using the selected image reconstruction model to obtain reconstructed components; All the reconstruction components are combined to obtain a reconstructed image.
24. An image reconstruction device, characterized in that: include: An acquisition module, configured to acquire an input image, wherein the input image includes at least one type of component; A selection module, configured to determine an image reconstruction model corresponding to the component from a model set based on the type of the component and a quantization parameter, wherein the quantization parameter represents the degree of distortion of the input image; the model set comprises a first image reconstruction model and a second image reconstruction model, wherein the first image reconstruction model is trained by the method according to any one of claims 19 to 21; A reconstruction module, used for processing the corresponding components using the selected image reconstruction model to obtain reconstructed components; The combination module is used to combine all the reconstruction components to obtain a reconstructed image.
25. An image decoding method, characterized in that: include: Acquire a code stream, wherein the code stream includes a filter tag, and the filter tag represents a type of an image reconstruction model; Decoding the code stream to obtain an input image, wherein the input image includes at least one type of component; Determining an image reconstruction model corresponding to the component from a model set based on the type of the component and the quantization parameter and the filter flag, the quantization parameter characterizing the degree of distortion of the input image; the model set includes a first image reconstruction model and a second image reconstruction model, the first image reconstruction model being trained by the method described in any one of claims 19 to 21 above; Processing the corresponding components using the selected image reconstruction model to obtain reconstructed components; All the reconstruction components are combined to obtain a reconstructed image.
26. An image decoding device, characterized in that: include: An acquisition module, used for acquiring a code stream, wherein the code stream includes a filter mark, and the filter mark represents a type of an image reconstruction model; A decoding module, used for decoding the code stream to obtain an input image, wherein the input image includes at least one type of component; A selection module, configured to determine an image reconstruction model corresponding to the component from a model set based on the type of the component and a quantization parameter and the filter flag, wherein the quantization parameter represents the degree of distortion of the input image; the model set comprises a first image reconstruction model and a second image reconstruction model, wherein the first image reconstruction model is trained by the method according to any one of claims 19 to 21; A reconstruction module, used for processing the corresponding components using the selected image reconstruction model to obtain reconstructed components; The combination module is used to combine all the reconstruction components to obtain a reconstructed image.
27. An image encoding method, characterized in that: include: Acquire a to-be-encoded image, wherein the to-be-encoded image includes at least one type of component; Determining an image reconstruction model corresponding to the component from a model set based on the type of the component and a quantization parameter, wherein the quantization parameter represents the degree of distortion of the input image; the model set comprises a first image reconstruction model and a second image reconstruction model, wherein the first image reconstruction model is trained by the method according to any one of claims 19 to 21; Processing the corresponding components of the reference image of the image to be encoded using the selected image reconstruction model to obtain reconstructed components; The to-be-encoded image is encoded based on the reconstructed component to obtain a code stream.
28. The encoding method according to claim 27, characterized in that The code stream includes a filter flag, and the filter flag represents the type of the image reconstruction model selected when encoding the image to be encoded.
29. An image encoding device, characterized in that: include: An acquisition module, used for acquiring an image to be encoded, wherein the image to be encoded includes at least one type of component; A determination module, configured to determine an image reconstruction model corresponding to the component from a model set based on the type of the component and a quantization parameter, wherein the quantization parameter represents a degree of distortion of the input image; the model set includes a first image reconstruction model and a second image reconstruction model; A reconstruction module, used for processing the corresponding components of the reference image of the image to be encoded by using the selected image reconstruction model to obtain a reconstructed component; The encoding module is used to encode the image to be encoded based on the reconstructed component to obtain a code stream.
30. An electronic device, characterized in that: It includes a processor and a memory coupled to each other, wherein: The memory is used to store program instructions for implementing the method according to any one of claims 1-15, 17, 19-21, 23, 25, 27-29; The processor is configured to execute the program instructions stored in the memory.
31. A computer-readable storage medium, characterized in that: A program file is stored, and the program file can be executed to implement the method according to any one of claims 1-15, 17, 19-21, 23, 25, 27-28.
Citation Information
Patent Citations
Method and apparatus of neural network for video coding
CN111937392A