An image processing, model training method and device
By determining the target encoding method of the image, selecting the matching super-resolution model and training, the problem of poor reconstruction quality of low-resolution image in the prior art is solved, and the generation of high-quality super-resolution images and the improvement of model efficiency are achieved.
Patent Information
- Application Number
- CN202510314154.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-03-17
AI Technical Summary
In the existing super-resolution reconstruction technology, the image quality after low-resolution image reconstruction is poor.
By determining the target encoding method of the image, selecting the target super-resolution model that matches it, using the model to process the image to generate high-quality super-resolution images, and by training the super-resolution model to improve image quality, the model training process is optimized using encoding features and loss functions matching encoding type group.
Improves the quality of super-resolution images, reduces the noise introduced during the encoding process, enhances the resolution and details of the image, and reduces the number of models and storage costs.
Smart Images

Figure CN119831845B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technologies, and in particular, to an image processing, model training method and device. Background Art
[0002] Super-resolution reconstruction refers to an implementation technology for restoring a super-resolution image with a higher resolution from a low-resolution image. However, currently, based on the super-resolution reconstruction technology, the image quality of the super-resolution image reconstructed from the low-resolution image is poor. Summary of the Invention
[0003] On the one hand, the present application provides an image processing method, including:
[0004] Obtaining a decoded first image, and determining a target encoding method corresponding to the first image;
[0005] Determining a target super-resolution model corresponding to the target encoding method from multiple super-resolution models, where the multiple super-resolution models correspond to different encoding type groups, the target encoding method matches the encoding type group corresponding to the target super-resolution model, and the encoding methods matching the same encoding type group have at least one same encoding feature;
[0006] Processing the first image by using the target super-resolution model to generate a second image.
[0007] In a possible implementation manner, determining a target super-resolution model corresponding to the target encoding method from multiple super-resolution models includes:
[0008] Determining a target super-resolution model corresponding to the target encoding method based on a mapping list, where the mapping list records at least two encoding methods included in each of multiple encoding type groups, and the super-resolution model corresponding to each encoding type group.
[0009] In another possible implementation manner, the determining a target super-resolution model corresponding to the target encoding method based on the mapping list includes:
[0010] If the target encoding method is queried from the mapping list, determining a target super-resolution model corresponding to the target encoding method from the mapping list;
[0011] If the target encoding method is not queried from the mapping list, determining a target encoding type group whose corresponding encoding method has a similarity meeting the requirements with the target encoding method from the mapping list, and determining a target super-resolution model corresponding to the target encoding type group in the mapping list.
[0012] In yet another possible implementation, determining the target super-resolution model corresponding to the target encoding type group in the mapping list includes:
[0013] If multiple target encoding type groups are determined from the mapping list, determine multiple target super-resolution models corresponding to the multiple target encoding type groups in the mapping list;
[0014] Processing the first image using the target super-resolution model to generate a second image includes:
[0015] Process the first image using each target super-resolution model respectively, and fuse the candidate images obtained by processing the first image with each target super-resolution model based on the weights corresponding to each target super-resolution model to obtain a second image;
[0016] Or,
[0017] Based on the model parameters in each target super-resolution model, construct a comprehensive super-resolution model, and process the first image using the comprehensive super-resolution model to generate a second image.
[0018] In yet another possible implementation, the super-resolution model recorded in the mapping list is a super-resolution model corresponding to the encoding type group, which is trained with the first sample images encoded by each encoding method in the encoding type group as training data and the second sample images corresponding to the first sample images as training targets, and the quality of the second sample images is higher than that of the first sample images.
[0019] In yet another possible implementation, the second sample image corresponding to the first sample image is: the original image corresponding to the first sample image without encoding, or an optimized image obtained by optimizing the original image through image processing, where the quality of the original image corresponding to the first sample image is higher than that of the first sample image;
[0020] The optimized image is obtained by sharpening the detail areas of the original image.
[0021] In yet another possible implementation, in the loss function used to train the super-resolution model, different regional weights are assigned to different regions in the second sample image, and the detail complexity of different regions in the second sample image is different;
[0022] And / or, a noise suppression term is included in the loss function used for training the super-resolution model. In the noise suppression term, different suppression weights are assigned to different pixel regions in the second sample image. Among them, the greater the difference between the pixel region in the second sample image and the corresponding pixel region in the target image, the greater the suppression weight assigned to the pixel region in the second sample image. The pixel region includes at least one pixel, and the target image is the image generated by the super-resolution model through processing the first sample image;
[0023] And / or, the loss function used for training the super-resolution model corresponding to each coding type group matches the coding characteristics of the coding method corresponding to the coding type group.
[0024] On the other hand, the present application also provides a model training method, including:
[0025] Determine the super-resolution models to be trained corresponding to at least two coding type groups respectively. The coding methods matching the same coding type group have at least one same coding characteristic;
[0026] For each coding method corresponding to each coding type group, obtain the first sample image encoded by the coding method corresponding to the coding type group, and the second sample image corresponding to the first sample image. The quality of the second sample image is higher than that of the first sample image;
[0027] For each coding type group, use the first sample image corresponding to the coding type group as the training data, and use the second sample image corresponding to the first sample image as the training target to train the super-resolution model corresponding to the coding type group, so as to obtain the trained super-resolution model corresponding to the coding type group.
[0028] In a possible implementation manner, the obtaining the first sample image encoded by the coding method corresponding to the coding type group, and the second sample image corresponding to the first sample image includes:
[0029] Obtain the original image without encoding;
[0030] Encode the original image using the coding method corresponding to the coding type group to obtain the first sample image;
[0031] Determine the original image or the optimized image obtained by optimizing the original image as the second sample image corresponding to the first sample image.
[0032] In yet another possible implementation, using the first sample image corresponding to the encoding type group as training data and the second sample image corresponding to the first sample image as the training target to train the super-resolution model corresponding to the encoding type group, including:
[0033] Using the first sample image corresponding to the encoding type group as training data, using the second sample image corresponding to the first sample image as the training target, and training the super-resolution model corresponding to the encoding type group in combination with a loss function;
[0034] Wherein, in the loss function, different regional weights are assigned to different regions in the second sample image, and the detail complexities of different regions in the second sample image are different;
[0035] And / or, a noise suppression term is included in the loss function. In the noise suppression term, different suppression weights are assigned to different pixel point regions in the second sample image. Among them, the greater the difference between the pixel point region in the second sample image and the corresponding pixel point region in the target image, the greater the suppression weight assigned to the pixel point region in the second sample image. The target image is the image generated by the super-resolution model by processing the first sample image, and the pixel point region includes at least one pixel point;
[0036] And / or, the loss function used to train the super-resolution model corresponding to each encoding type group matches the encoding characteristics of the encoding method corresponding to the encoding type group.
[0037] In another aspect, the present application also provides an image processing apparatus, including:
[0038] A type determination unit, configured to obtain the decoded first image and determine the target encoding method corresponding to the first image;
[0039] A model determination unit, configured to determine, from multiple super-resolution models, the target super-resolution model corresponding to the target encoding method, where the multiple super-resolution models correspond to different encoding type groups, the target encoding method matches the encoding type group corresponding to the target super-resolution model, and the encoding methods matched by the same encoding type group have at least one same encoding characteristic;
[0040] An image processing unit, configured to process the first image using the target super-resolution model to generate a second image.
[0041] In another aspect, the present application also provides a model training apparatus, including:
[0042] A model determination unit, configured to determine a super-resolution model to be trained corresponding to each of at least two encoding type groups, and encoding methods matching the same encoding type group have at least one same encoding feature;
[0043] A sample acquisition unit, configured to, for each encoding method corresponding to each encoding type group, acquire a first sample image encoded by the encoding method corresponding to the encoding type group, and a second sample image corresponding to the first sample image, where the quality of the second sample image is higher than that of the first sample image;
[0044] A model training unit, configured to, for each encoding type group, use the first sample image corresponding to the encoding type group as training data, and use the second sample image corresponding to the first sample image as a training target, and train the super-resolution model corresponding to the encoding type group to obtain the trained super-resolution model corresponding to the encoding type group. Description of the Drawings
[0045] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages, and aspects of the embodiments of the present disclosure will become more obvious. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic, and the original elements and elements are not necessarily drawn to scale.
[0046] Figure 1 It is a schematic flowchart of an image processing method provided by the present application;
[0047] Figure 2 It is a schematic flowchart of a model training method provided by the present application;
[0048] Figure 3 It is another schematic flowchart of a model training method provided by the present application;
[0049] Figure 4 It is an example diagram of an implementation framework for training a super-resolution model corresponding to a single encoding method in the present application;
[0050] Figure 5 It is another schematic flowchart of an image processing method provided by the present application;
[0051] Figure 6 It is a schematic structural diagram of a composition of an image processing device provided by the present application;
[0052] Figure 7 It is a schematic structural diagram of a composition of a model training device provided by the present application;
[0053] Figure 8 It is a schematic structural diagram of a composition of an electronic device provided by the present application. Detailed Description of the Embodiments
[0054] The embodiments of the present application will be described below with reference to the accompanying drawings in the embodiments of the present application. The terms used in the embodiments of the present application are only used to explain the specific embodiments of the present application, rather than to limit the present application. As is known to those of ordinary skill in the art, with the development of technology and the emergence of new scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.
[0055] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, which is only a way of distinguishing objects with the same attributes when describing the embodiments of the present application. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product or device comprising a series of units does not have to be limited to those units, but may include other units not clearly listed or inherent to these processes, methods, products or devices.
[0056] As Figure 1 , a schematic flowchart of an image processing method provided by an embodiment of the present application is shown. The method of this embodiment is applied to an electronic device, and the electronic device can be a laptop computer, a desktop computer, or a tablet computer, etc., and can also be a device node in the cloud or a distributed cluster, without limitation thereto.
[0057] The method of this embodiment may include the following steps S101 to S103:
[0058] S101, obtain the decoded first image and determine the target encoding method corresponding to the first image.
[0059] Among them, the first image is an image that needs to be processed by an image reconstruction process using a super-resolution model. Since the first image is an image decoded from an encoded and compressed image, therefore, affected by the encoder, for example, noise is introduced during the compression encoding process of the encoder for the image, so that the image quality of the decoded first image is relatively low. For example, the image resolution of the first image is relatively low and there is more noise, etc.
[0060] Among them, the target encoding method is the encoding method used when encoding the first image. It can be understood that since the types of encoders used for encoding the first image are different, the encoding methods used for encoding the first image are also different. Therefore, the target encoding method is also the encoding method corresponding to the encoder for encoding the first image.
[0061] In this application, there can be multiple encoding methods. For example, the encoding method can include, but is not limited to, Joint Photographic Experts Group (JPEG) encoding, High Efficiency Video Coding (HEVC), or H.264 (Advanced Video Coding of Moving Picture Experts Group-4), etc. Among them, High Efficiency Video Coding is also known as H.265 encoding.
[0062] Among them, determining the target encoding method corresponding to the first image can be by analyzing the metadata of the first image to determine the target encoding method corresponding to the first image. Among them, the metadata of the first image can at least include the encoding type of the first image, and can also include other attribute information of the first image, without specific limitation. For example, since the first image is obtained by decoding a decoded image, the metadata of the encoded image can be obtained, and the target encoding method of the encoded image can be determined based on the metadata of the encoded image, so as to obtain the target encoding method corresponding to the first image. Among them, the metadata of the encoded image is actually the metadata of the first image.
[0063] Determining the target encoding method can also be determined by analyzing the extension name of the first image; or, detecting the target encoding method corresponding to the first image with the help of a specific encoding detection tool; it can also be by analyzing features related to image encoding such as the compression noise pattern or frequency domain characteristics corresponding to the first image to determine the target encoding method corresponding to the first image, without specific limitation.
[0064] It should be noted that in this application, the first image can be an independent single-frame image or an image frame in a video, without specific limitation.
[0065] S102. Determine the target super-resolution model corresponding to the target encoding method from multiple super-resolution models.
[0066] Among them, multiple super-resolution models correspond to different encoding type groups, and the target encoding method matches the encoding type group corresponding to the target super-resolution model. For the sake of distinction, the super-resolution model corresponding to the target encoding method is called the target super-resolution model. Therefore, the target super-resolution model belongs to the multiple super-resolution models.
[0067] In this application, the coding methods matched by the same coding type group have at least one same coding feature. For example, the coding methods matched by the same coding type group may include at least two coding methods, and the at least two coding methods may have one or more same coding features. Based on this, this application actually matches the coding methods with at least one same coding feature to the same coding type group, and the various coding methods matched to the same coding type group can share the same super-resolution model, so there is no need to construct a super-resolution model for each coding method separately, and naturally the number of super-resolution models that need to be constructed and stored can be reduced.
[0068] In a possible case, the possible coding type groups can be pre-divided. Therefore, the coding methods included in each coding type group are pre-determined. Of course, the various coding methods included in the same coding type group have at least one same feature. In this case, after determining the target coding method corresponding to the first image, the target coding type group including the target coding method can be determined, and the target super-resolution model corresponding to the target coding type group can be determined.
[0069] In another possible case, this application can only determine that each coding type group corresponds to at least one coding feature, but the coding methods included in the coding type group are not pre-divided. In this case, after determining the target coding method corresponding to the first image, this application can determine the target coding type group matched by the target coding method based on the coding feature of the target coding method, and determine the target super-resolution model corresponding to the target coding type group.
[0070] Among them, the coding method can have coding features in multiple dimensions. For example, the coding features of the coding method can include but are not limited to various coding features such as coding algorithm principle, coding complexity, and coding applicable scenario.
[0071] Taking the determination of different coding type groups based on the coding feature of the coding algorithm principle as an example:
[0072] The coding algorithm principles corresponding to different coding type groups are different. For example, the coding algorithm principles corresponding to different coding type groups can be: transform-based coding (also known as transform coding), block prediction-based coding, or entropy-based coding, etc.
[0073] For example, for the coding type group matched by the coding feature of transform-based coding, the coding methods matched to it can adopt transform-based coding. Therefore, the coding methods matched to this coding type group can include but are not limited to: JPEG coding, AV1 (Open Media Video Generation 1) coding, and intra-frame coding (VCC Intra) in Versatile Video Coding (VCC), etc.
[0074] For the coding type group that matches the coding feature of block-based prediction coding, the matching coding method can adopt block-based prediction coding. Therefore, the coding methods that match this coding type group can include but are not limited to: H.264 coding, HEVC coding, and inter-frame coding in VCC (VCC Inter).
[0075] Similarly, for the coding type groups that match each coding method of entropy-based coding, the coding methods can include but are not limited to: Huffman coding and Shannon coding, etc. In this application, the target super-resolution model is applicable to perform super-resolution reconstruction on images encoded by various coding methods that match the coding type group corresponding to the target super-resolution model. Among them, for the sake of easy distinction, the coding type group corresponding to the target super-resolution model can be called the target coding type group. For example, the target super-resolution model can perform targeted noise reduction on an image encoded by any coding method corresponding to the target coding type group, and reduce the corresponding type of noise introduced by the encoder of any coding method corresponding to the target coding type group in this image.
[0076] The target super-resolution model is a trained model, and there can be various possibilities for the model type of the target super-resolution model. For example, the target super-resolution model can be a convolutional neural network model, a deep recursive convolutional network, or a generative adversarial network model, etc., without limitation.
[0077] In this application, there is no limitation on the specific implementation of training the target super-resolution model. For example, in a possible implementation, the target super-resolution model is a super-resolution model corresponding to the target coding type group, which is trained with the first sample image encoded by the coding method corresponding to the target coding type group as the training data and the second sample image corresponding to the first sample image as the training target. Among them, the quality of the second sample image is higher than that of the first sample image.
[0078] Among them, the fact that the quality of the second sample image is higher than that of the first sample image can be reflected from multiple aspects. For example, the fact that the quality of the second sample image is higher than that of the first sample image can at least characterize that the resolution of the second sample image is higher than that of the first sample image; it can also characterize that: the noise in the second sample image is less than that in the first sample image, and the edge sharpness and detail richness of the second sample image are higher than those of the first sample image, etc.
[0079] S103. Use the target super-resolution model to process the first image to generate a second image.
[0080] Among them, the second image is a high-quality image generated by the target super-resolution model based on the first image. Therefore, the quality of the second image is higher than that of the first image.
[0081] For example, the fact that the quality of the second image is higher than that of the first image can at least indicate that the resolution of the second image is higher than that of the first image. Of course, the fact that the quality of the second image is higher than that of the first image can also indicate that the noise contained in the second image is less than that contained in the first image, etc.
[0082] It can be understood that after generating the second image, the present application can also output the second image so that the user can obtain the second image with relatively high image quality. Of course, in practical applications, after generating the second image, it can also be saved according to actual needs, which will not be elaborated here.
[0083] From the above content, it can be seen that after decoding the first image, the present application will determine the target encoding method corresponding to the first image, and use the target super-resolution model corresponding to the target encoding type group that matches the target encoding method to process the first image, so as to be able to select a target super-resolution model suitable for processing the first image encoded by this target encoding method, so that the target super-resolution model can more effectively remove the noise or other influences introduced during the encoding process of the first image, thereby improving the image quality of the generated second image.
[0084] Moreover, each encoding method in the present application that matches the same encoding type group has at least one same encoding feature, so that the encoding features of each encoding method in the same encoding type group are similar, so that each encoding method that matches the same encoding type group can share the same super-resolution model, which not only improves the generality of the encoding methods applicable to a single super-resolution model, but also reduces the number of super-resolution models that need to be constructed and stored.
[0085] In a possible implementation manner, in order to be able to efficiently determine the super-resolution model corresponding to the encoding type group that matches the encoding method and reduce the time consumed for determining the super-resolution model corresponding to the encoding method, the present application can also construct a mapping list, in which at least two encoding methods included in multiple encoding type groups and the super-resolution model corresponding to each encoding type group are recorded. Among them, different super-resolution models in the mapping list correspond to different encoding type groups respectively.
[0086] On this basis, the present application can determine the target super-resolution model corresponding to the target encoding method based on the mapping list.
[0087] For example, the mapping list records the group identifier of each coding type group, at least two coding methods corresponding to the group identifier of each coding type group, and the model identifier of the super-resolution model corresponding to the group identifier of each coding type group. Among them, the group identifier can be the coding feature common to each coding method in the coding type group, or the group serial number corresponding to the coding type group, etc., which is used to uniquely identify the coding type group, and there is no specific limitation. Similarly, the model identifier of the super-resolution model can be the model name, number, or call address of the super-resolution model, etc., and there is no specific limitation.
[0088] By querying the mapping list to determine the target super-resolution model corresponding to the target coding method, the time-consuming caused by calculating the target coding type group for the coding feature matching of the target coding method can be reduced. Naturally, the efficiency of determining the target super-resolution model can be improved, thereby reducing the time-consuming required to process the first image to be reconstructed, and further improving the efficiency of super-resolution image reconstruction for the first image.
[0089] In this application, there can be multiple possibilities for the model structure of the super-resolution model corresponding to each coding type group. For specific details, please refer to the relevant introduction of the possible model structures of the target super-resolution model above, and no further elaboration will be provided.
[0090] In this application, each super-resolution model can be pre-trained, and the specific training process is not limited. Among them, through the training of each super-resolution model, the processing effects of the same super-resolution model for images encoded by different coding methods corresponding to different coding type groups are different.
[0091] In a possible implementation manner, in order to enable the super-resolution model corresponding to each coding type group to pay more attention to the noise and other influences introduced by encoding the image using each coding method (or the encoders corresponding to each coding method) in the coding type group, for the super-resolution model corresponding to any coding type group, the super-resolution model can use the first sample images encoded by each coding method in the coding type group as training data, and use the second sample images corresponding to the first sample images as the training target to train the super-resolution model corresponding to the coding type group. As mentioned above, the quality of the second sample image is higher than that of the first sample image.
[0092] Among them, for any coding type group, each first sample image corresponds to one coding method in the coding type group, and there can be multiple first sample images for training the super-resolution model, and the coding methods corresponding to the multiple first sample images can include at least two coding methods that match the coding type group.
[0093] For the sake of easy understanding, the following combines Figure 2The flowchart shown illustrates an implementation method for training super-resolution models corresponding to various coding type groups. Figure 2 The figure shows an implementation flowchart of the model training method provided by an embodiment of the present application. The model training method may include the following steps S201 to S203:
[0094] S201, determine the super-resolution models to be trained corresponding to at least two coding type groups respectively.
[0095] Among them, the coding methods matching the same coding type group have at least one same coding feature. For example, before training the super-resolution model, for each coding type group, at least two coding methods matching this coding type group can be determined.
[0096] Among them, the model types of the super-resolution models to be trained corresponding to different coding type groups may be the same or different, and can be specifically selected according to actual needs, without limitation. For the model types that can be selected for the super-resolution model, reference can be made to the relevant introduction above, and details will not be elaborated here.
[0097] S202, for each coding method corresponding to each coding type group, obtain the first sample image encoded by the coding method corresponding to this coding type group, and the second sample image corresponding to the first sample image.
[0098] Among them, the first sample image may be an image obtained by encoding a raw image with a quality meeting the requirements through an encoder corresponding to one coding method in this coding type group. Among them, the quality of the raw image is higher than that of the first sample image.
[0099] In one implementation, each coding type group may correspond to multiple first sample images, and the multiple first sample images include at least one first sample image encoded by each coding method of this coding type group.
[0100] For example, taking the case where each coding method in the coding type group belongs to transform-based compression coding as an example, assuming that this coding type group includes two coding methods, JPEG and AV1, at least one raw image with a resolution exceeding a set threshold can be compressed and encoded using a JPEG encoder to obtain at least one first sample image corresponding to this JPEG coding method; it is also necessary to compress and encode at least one raw image with a resolution exceeding a set threshold using an AV1 encoder to obtain at least one first sample image corresponding to this AV1 coding method. Among them, each first sample image and the raw image corresponding to it contain the same object content, but the resolution of the first sample image is lower than that of the raw image.
[0101] Among them, the quality of the second sample image is higher than that of the first sample image.
[0102] Among them, the goal of training the super-resolution model corresponding to this encoding type group is to expect that the super-resolution model corresponding to this encoding type group can generate the second sample image based on the first sample image. Therefore, the second sample image is the training goal for training the super-resolution model corresponding to this encoding type group.
[0103] S203. For each encoding type group, using the first sample image corresponding to this encoding type group as training data and the second sample image corresponding to this first sample image as the training goal, train the super-resolution model corresponding to this encoding type group to obtain the trained super-resolution model corresponding to this encoding type group.
[0104] Among them, using the second sample image corresponding to the first sample image as the training goal essentially means hoping to minimize the gap between the target image generated by the super-resolution model based on the first sample image and this first sample image, so that the target image generated by the super-resolution model based on the first sample image is as similar as possible to the second sample image.
[0105] From the above content, it can be seen that for each encoding type group, this application will use each first sample image encoded by each encoding method in this encoding type group as training data, and use the relatively higher-quality second sample image corresponding to the first sample image as the training goal to train the super-resolution model corresponding to this encoding type group, so that the super-resolution model corresponding to each encoding type group can reconstruct the image encoded by any one of the encoding methods in this encoding type group into a high-quality image.
[0106] In addition, although noise will be introduced when the encoders corresponding to each encoding method in each encoding type group encode the image, the types of noise introduced by the encoders corresponding to each encoding method that match the same encoding type group are similar. Therefore, using the same super-resolution model for super-resolution image reconstruction of the images encoded by the encoders corresponding to each encoding method in each encoding type group can not only enable the super-resolution model to effectively remove the specific type of noise introduced by the encoder corresponding to this encoding type group and improve the quality of the reconstructed super-resolution image, but also effectively reduce the problems of a large number of models and high maintenance costs caused by training one super-resolution model for each encoding method in the encoding type group.
[0107] It can be understood that during the training process of the super-resolution model corresponding to each coding type group, it is necessary to combine the corresponding loss function to control the training of the super-resolution model. The loss function is used to quantify the difference between the target image generated by the super-resolution model based on the first sample image and the second sample image corresponding to the first sample image. For example, the function value of the combined loss function can be used to determine whether the super-resolution model reaches the training target and to adjust the model parameters of the super-resolution model, etc.
[0108] Based on this, in the present application, for each coding type group, the first sample image corresponding to the coding type group can be used as the training data, and the second sample image corresponding to the first sample image can be used as the training target, and the super-resolution model corresponding to the coding type can be trained in combination with the loss function.
[0109] Among them, the loss function can be selected according to actual needs, and there is no specific limitation.
[0110] In a possible case of the loss function, in the loss function used for training the super-resolution model, different regional weights are assigned to different regions in the second sample image. Among them, the detail complexity of different regions in the second sample image is different.
[0111] Among them, each region in the second sample image has a one-to-one correspondence with the corresponding region in the target image generated by the super-resolution model based on the first sample image. Based on this, the function value of the loss function is related to the gap between the target image and the second sample image in different regions and the regional weights corresponding to different regions.
[0112] Different from the fact that the function value of the conventional loss function is only related to the overall gap between the target image generated by the super-resolution model and the second sample image, in the present application, since different regional weights are assigned to different regions of the second sample image in the loss function, the learning ability of the super-resolution model for regions where the detail complexity meets the requirements can be enhanced during the training process of the super-resolution model.
[0113] Among them, there can be various bases for dividing different regions in the second sample image. The following describes several possible cases of dividing different regions in the second sample image:
[0114] For example, in a possible scenario, a detail area and a flat area can be divided in the second sample image. On this basis, different area weights are assigned to the detail area and the flat area in the second sample image in the loss function, and the area weight of the detail area is greater than that of the flat area, so that during the process of training the super-resolution model, the super-resolution model can learn more detail information in the second sample image, thereby enhancing the sharpness of the detail area of the generated image; and the flat area of the generated image can be kept smooth to reduce the new noise introduced into the generated image due to the sharpening process.
[0115] Among them, the detail area of an image refers to the area where the pixel values change violently and contains rich information. The detail area of an image usually includes edges, textures, patterns or complex structures. The flat area of an image refers to the area where the pixel values change gently and contains less information. For example, the area with a detail complexity not lower than the set threshold in the image can be the detail area, while the area with a detail complexity lower than the set threshold can be the flat area. An image can include one or more detail areas, and correspondingly, an image can also include one or more flat areas.
[0116] In this application, there is no limitation on the specific implementation of determining the detail area and the flat area in the image. For the sake of understanding, two implementation methods are described.
[0117] For example, in one implementation method, an edge detection algorithm can be used to detect the edge area and the non-edge area in the image. The edge area usually belongs to the detail area, and the non-edge area belongs to the flat area. The specific implementation process can be: perform edge detection on the image to obtain a binary edge map, and determine the area where the white pixels in the edge map are located as the detail area, and the area where the black pixels are located as the flat area.
[0118] In another implementation method, a texture analysis method can be used to determine the detail area and the flat area of the image based on the texture complexity of the image. Among them, the texture complexity reflects the texture richness of the local area of the image. The area with a texture complexity not lower than the set complexity threshold is the detail area, and the area with a texture complexity lower than this complexity threshold is the flat area. For example, texture feature extraction methods such as local binary pattern and gray-level co-occurrence matrix can be used to extract the texture feature values of different pixel areas in the image. The texture feature value can represent the texture complexity. Correspondingly, the detail area and the flat area of the image are determined in combination with the set texture feature threshold.
[0119] Of course, there can also be other implementation methods to determine the detail area and the flat area in the image, which will not be elaborated here.
[0120] In yet another possible scenario, based on the objects in the second sample image, the second sample image is divided into object regions with objects and background regions without objects, resulting in multiple different regions in the second sample image.
[0121] Among them, determining the object regions and background regions in the second sample image can be achieved using edge detection algorithms or object recognition models, etc., without specific limitations. For example, taking the determination of the object regions and background regions in the second sample image by an edge detection algorithm as an example, based on the edge detection algorithm, the edges of each object in the second sample image can be detected, obtaining the object regions of each object in the second sample image and the background regions outside the object regions.
[0122] In yet another possible scenario, the target of interest in the second sample image can be set as needed, and the regions of interest and non - regions of interest in the second image are determined, resulting in multiple different regions in the second sample image. For example, the regions of interest and non - regions of interest in the second sample image can be recognized by combining a recognition model, or, based on a semantic segmentation algorithm, the second sample image is semantically segmented to segment out the regions of interest and non - regions of interest in the second sample image. Of course, there can be other ways to determine the regions of interest and non - regions of interest in the second sample image, which are not restricted.
[0123] In this application, there can be multiple possible specific forms of the loss function. For example, the loss function can be a pixel - level loss function. In the pixel - level loss function, at least one of the absolute difference and the squared difference of the pixel values at different positions of the target image and the second sample image is calculated. Another example is that the loss function can be a perceptual loss function, in which the distance between the target image and the second sample image in the feature space can be calculated. Another example is that the loss function can be a gradient loss function. Of course, the loss function can also have other forms, which are not restricted.
[0124] To facilitate the understanding of the specific implementation of assigning different region weights to different regions of the second sample image in the loss function, the following takes the loss function as a pixel - level loss function and combines a specific function form of the loss function as an example for illustration. For example, the following formula is an expression of the loss function :
[0125]
[0126] Among them, is the total number of pixel points in the second sample image (or the target image).
[0127] and are the height and width of the second sample image, that is, the number of pixel points in each column and each row of the second sample image respectively;
[0128] is the weight corresponding to the pixel point at the th column and the th row in the second sample image.
[0129] is the pixel value of the pixel point at the th column and the th row in the target image generated by the super-resolution model.
[0130] is the pixel value of the pixel point at the th column and the th row in the second sample image.
[0131] In the present application, the regional weights of different regions in the second sample image or the target image are different. Therefore, in the above formula, the weight of the pixel point is the regional weight corresponding to the region of the pixel point in the second sample image (or the target image). It can be seen that the weights corresponding to the pixel points located in different regions of the second sample image are also different.
[0132] In another possible case, in order to reduce the noise that may be amplified by the super-resolution model during the process of reconstructing the super-resolution image, in the present application, the loss function may include a noise suppression term. In this noise suppression term, different suppression weights are assigned to different pixel point regions in the second sample image. Each pixel point region includes at least one pixel point.
[0133] For any pixel point region in the second sample image, the greater the difference between the pixel point region and the corresponding pixel point region in the target image, the greater the suppression weight assigned to the pixel point region. The target image is the image generated by the super-resolution model by processing the first sample image.
[0134] Among them, the corresponding pixel point region in the target image is the pixel point region with the same coordinate region as the pixel point region in the second sample image. Based on this, the suppression weight corresponding to the pixel point region in the second sample image is actually the suppression weight corresponding to the corresponding pixel point region in the target image.
[0135] Among them, the value of the noise suppression term is the weighted sum of the differences between the target image and the second sample image in each pixel point region determined based on the suppression weights of each pixel point region in the second sample image. And the greater the value of the noise suppression term, the greater the function value of the loss function.
[0136] It can be understood that the greater the gap between the target image and the second sample image in the same pixel region, the more noise there is in the pixel region of the target image. In order to reduce the noise in each pixel region of the target image during the training of the super-resolution model, the present application assigns a larger weight (i.e., suppression weight) to the pixel regions with a larger gap between the target image and the second sample image, so as to increase the learning of the super-resolution model for the pixel regions with a larger gap between the target image and the second sample image, and reduce the noise in the pixel regions of the target image.
[0137] Among them, as the super-resolution model is trained, the gap between the target image and the second sample image in different pixel regions may change. Therefore, as the gap between the target image and the second sample image in different pixel regions changes, the corresponding suppression weights of different pixel regions will be adaptively adjusted.
[0138] Taking the above formula as an example, as the super-resolution model is continuously trained, the gap between the target image and the second sample image in different pixel regions will change, resulting in a change in the suppression weight corresponding to the pixel region, and the weights of each pixel in the pixel region will also change accordingly.
[0139] It can be understood that in practical applications, the loss function for training the super-resolution model can include the above two possible situations, that is, while different regional weights are assigned to different regions in the second sample image in the loss function, the loss function also includes the above-mentioned loss suppression term. For example, in the loss function, the weight assigned to a certain pixel in the second sample image includes weight a and weight b. Among them, weight a is a fixed weight assigned to the region where the pixel is located in the second sample image. And weight b is a suppression weight determined based on the gap between the pixel in the second sample image and the corresponding pixel in the target image. Therefore, as the super-resolution model is continuously trained, weight b can be a changing weight value.
[0140] In another possible situation, considering that in the present application, the first sample image can be a static image containing static objects such as scenery or objects, or a dynamic image containing dynamic objects such as animals or people, or a dynamic image sourced from a video, etc. In order to train the super-resolution model corresponding to each coding type group more reasonably, different loss functions can be determined by combining the proportions of static images and dynamic images in the first sample image corresponding to the coding type group.
[0141] In yet another possible scenario, the loss function used to train the super-resolution model corresponding to each coding type group matches the coding features of the coding method corresponding to that coding type group. Based on this, for any coding type group, if the same coding features of the coding methods in the coding type group are different, the loss function selected for training the super-resolution model corresponding to that coding type group can also be different. Specifically, the appropriate loss function can be determined by combining the edge features shared by the coding methods in the coding type group. It can be understood that the loss function for training the super-resolution model can include one or more of the above-mentioned multiple possible scenarios, without specific limitations.
[0142] In the present application, there can be multiple possible ways to obtain the second sample image.
[0143] For example, in one possible scenario, the second sample image can be the original image corresponding to the first sample image mentioned above, that is, the unencoded original image corresponding to the first sample image.
[0144] In this scenario, for each coding type group, the unencoded original image can be obtained first. Then, the encoder of the coding method corresponding to the coding type group is used to encode the original image to obtain the first sample image. Correspondingly, the original image corresponding to the first sample image can be determined as the second sample image.
[0145] It can be understood that by compressing and encoding the original image, not only will the resolution of the original image be reduced, but also compression noise, etc. will be introduced into the original image, resulting in the quality of the generated first sample image being lower than its corresponding original image. And in the present application, the first sample image is used as training data to train the super-resolution model, with the aim of expecting that through training the super-resolution model, the super-resolution model can restore the first sample image to the original image.
[0146] In yet another possible scenario, the second sample image is an optimized image obtained by optimizing the original image corresponding to the first sample image.
[0147] It can be understood that by optimizing the original image, the image quality of the original image can be improved. Therefore, the quality of the optimized image is not only higher than the quality of the first sample image but also higher than the quality of the second sample image. Based on this, using the optimized image as the second sample image corresponding to the first sample image to train the super-resolution model is more conducive to improving the quality of the image generated by the super-resolution model.
[0148] Among them, the way to optimize the original image can be any processing method that can improve the quality of the original image, without limitation.
[0149] In an alternative manner, considering that the information in the detail regions of the image is more abundant, in this application, an optimized image can be obtained by sharpening the detail regions of the original image. Among them, during the process of sharpening the detail regions of the original image, the flat regions in the original image remain unchanged.
[0150] On this basis, since the optimized image is only obtained by sharpening the detail regions of the original image, therefore, using the optimized image as the second sample image corresponding to the first sample image to train the super-resolution model can enhance the ability of the super-resolution model to learn the detail information in the image, not only can improve the quality of the image generated by the super-resolution model, but also can, to a certain extent, avoid overprocessing the flat regions in the image and reduce the situation of introducing new noise due to the super-resolution model sharpening and enhancing the flat regions in the image during the image generation process.
[0151] For ease of understanding, hereinafter, taking the optimized image corresponding to the first sample image as the second sample image as an example, the model training method of this application will be described. As Figure 3 shown in, another schematic flowchart of the model training method provided by this application is shown, and this embodiment may include:
[0152] S301, determine the super-resolution models to be trained corresponding to at least two encoding type groups respectively.
[0153] S302, for each encoding type group, obtain the original image that has not been encoded.
[0154] Among them, the original image is a high-quality image that has not been encoded, such as the original image that has not been encoded by each encoding method in this encoding type group.
[0155] For example, for each encoding type group, at least one original image can be obtained.
[0156] S303, use the encoding method corresponding to this encoding type group to encode the original image to obtain the first sample image.
[0157] For example, use the encoders corresponding to each encoding method in the encoding type group to encode the original image respectively to obtain at least one first sample image corresponding to each encoding method in the encoding type group, and the original images corresponding to each first sample image are the same.
[0158] It can be understood that since after each encoder corresponding to each encoding method encodes the original image, it will not only reduce the resolution of the original image, but also introduce noise of the noise type corresponding to this encoder, therefore, the quality of the first sample image obtained by encoding the original image with any one encoding method is lower than the quality of this original image.
[0159] S304, perform sharpening processing on the flat area in the original image and keep the flat area in the original image unchanged, to obtain an optimized image obtained by processing the original image.
[0160] By performing sharpening processing on the flat area of the original image, the edges and details in the original image can be enhanced, thereby improving the image quality. Moreover, only performing sharpening processing on the flat area of the original image can introduce as little new noise as possible, so that the optimized image has higher image quality than the original image.
[0161] S305, use the first sample image corresponding to this coding type group as training data, use this optimized image as the training target, and train the super-resolution model corresponding to this coding type group in combination with the loss function, to obtain the trained super-resolution model corresponding to this coding type group.
[0162] Since the optimized image has higher image quality than the original image, therefore, using the optimized image as the training target for training the super-resolution model can further improve the quality of the image generated by the super-resolution model through continuous training.
[0163] Among them, the loss function can be any one or a combination of several of the above-mentioned possible situations, and will not be elaborated here.
[0164] In a possible implementation manner, the present application can also separately train the super-resolution model corresponding to some relatively common coding methods, so as to obtain the super-resolution models corresponding to at least one coding method respectively. On this basis, after the present application determines the target coding method corresponding to the first image, it can first query from the super-resolution models corresponding to at least one coding method respectively whether there is a super-resolution model corresponding only to the target coding method. If there is a super-resolution model corresponding only to the target coding method, determine this super-resolution model as the target super-resolution model.
[0165] If there is no super-resolution model corresponding only to the target coding method, the target super-resolution model corresponding to the target coding method can be determined based on the mapping list.
[0166] Among them, the specific implementation of training the super-resolution model corresponding to a single coding method is similar to the process of training the super-resolution model corresponding to the coding type group before. It's just that when training the super-resolution model corresponding to a single coding method, the first sample image used to train this super-resolution model only includes the sample images encoded by this coding method (or encoded by the encoder corresponding to this coding method). The following combines Figure 4The example diagram of the model training implementation framework shown below briefly explains the process of training a super-resolution model corresponding to a single encoding method.
[0167] It can be seen from Figure 4 that after collecting at least one high-quality original image, for each encoding method for which a super-resolution model needs to be trained, the encoder of this encoding method can be used to compress and encode each original image respectively to generate at least one low-quality image corresponding to this encoding method. Among them, some of the at least one low-quality images are used as training data to form a training dataset. The remaining low-quality images in the at least one low-quality images are used as validation data to construct a validation dataset.
[0168] In addition, before training the super-resolution model of this encoding method, the key areas (such as detailed areas) of the original image will be optimized (such as sharpened), and other areas (such as flat areas) will remain unchanged to obtain an optimized image ( Figure 4 the golden image in
[0169] Based on this, for each encoding method, the low-quality images in the training dataset can be used to train the super-resolution model of this encoding method. A noise suppression term can be added to the loss function used for training the super-resolution model. Moreover, in this loss function, different regional weights can be assigned to the key areas (such as detailed areas) and non-key areas (such as flat areas) of the optimized image respectively to enhance the learning ability and denoising ability of the super-resolution model for key areas such as detailed areas.
[0170] Among them, before training the super-resolution model by combining the function value of the loss function, the model parameters of the super-resolution model can also be initialized.
[0171] During the training process, the low-quality images used as training data are input into the super-resolution model to be trained, so that the super-resolution model outputs a target image (i.e., the forward propagation process). Based on this, the function value of the loss function is calculated based on the target image and the optimized image. Based on the function value of the loss function, the model parameters of the super-resolution model are updated through backpropagation. This process is repeated multiple times to achieve iterative training of the super-resolution model.
[0172] After training a super-resolution model corresponding only to this encoding method, the super-resolution model can be verified using the low-resolution images in the validation set and model optimization can be performed. Based on this, the finally trained super-resolution model can be model-quantized, and the model-quantized super-resolution model can be stored in the model library so that the model library can include super-resolution models applicable to the encoders of at least one encoding method.
[0173] It can be understood that after determining the target encoding method corresponding to the first image, if the target encoding method is found in the mapping list, the target super-resolution model corresponding to the target encoding method can be directly determined from the mapping list.
[0174] However, since there are many types of encoding methods, and the encoding methods that can be recorded in the mapping list are limited, after determining the target encoding method corresponding to the first image, it may also occur that the target encoding method does not exist in the mapping list. In this case, the present application can also determine a target encoding type group from the mapping list whose similarity between the corresponding encoding method and the target encoding method meets the requirements. Correspondingly, the target resolution model corresponding to the target encoding type group in the mapping list can be determined, so as to process the first image by using the target super-resolution model corresponding to the target encoding type group.
[0175] For example, the encoding type group with the highest similarity between the included encoding method and the target encoding method can be determined as the target encoding type group; or, the encoding type group with the similarity between the included encoding method and the target encoding method exceeding the set threshold can be determined as the target encoding type group.
[0176] Among them, there are various implementation manners for determining the similarity between the encoding method corresponding to the encoding type group and the target encoding method.
[0177] For example, in a possible implementation manner, for each encoding type group in the mapping list, the similarity between the target encoding method and each encoding method in the encoding type group can be calculated respectively. Correspondingly, the encoding type group where the encoding method with the similarity meeting the requirements with the target encoding method is located is determined as the target encoding type group.
[0178] Among them, calculating the similarity between the target encoding method and the encoding method in the encoding type group can be calculating the vector similarity between the vector corresponding to the target encoding method and the vector corresponding to the encoding method in the encoding type group. Among them, the vector corresponding to the target encoding method can be a vector converted from the name of the target encoding method, or the encoding feature of the target encoding method is obtained. For example, the encoding feature of the target encoding method can include encoding characteristics such as encoding principle or encoding applicable scenario, and the vector corresponding to the target encoding method is determined based on the encoding feature of the target encoding method. The same is true for determining the vector corresponding to the encoding method in the encoding type group, which will not be elaborated here.
[0179] In yet another possible implementation, for each coding type group in the mapping list, the similarity between the target coding method and each coding method in the coding type group is calculated separately. On this basis, for each coding type group, the average similarity or weighted average of the similarities corresponding to each coding method in the coding type group is determined, and the corresponding average similarity or weighted average is determined as the similarity between the coding method corresponding to the coding type group and the target coding method, so as to determine the target coding type group whose similarity between the corresponding coding method and the target coding method meets the requirements. For example, the coding type group with the largest average similarity is determined as the target coding type group.
[0180] It can be understood that there may be multiple coding type groups whose similarity between the corresponding coding method and the target coding method meets the requirements. In this case, the present application may randomly select one coding type group from the determined multiple coding type groups as the target coding type group.
[0181] In an alternative manner, when there are multiple coding type groups whose similarity between the corresponding coding method and the target coding method meets the requirements, the present application may also determine all of the multiple coding type groups as the target coding type groups. On this basis, the present application may combine the target super-resolution models corresponding to the multiple target coding type groups to comprehensively process the first image.
[0182] The following is combined with Figure 5 for illustration. As Figure 5 shows another flowchart of the image processing method provided by the present application. The method of this embodiment may include:
[0183] S501, obtain the decoded first image and determine the target coding method corresponding to the first image.
[0184] S502, query whether the target coding method exists in the mapping list. If so, execute step S503; if not, execute step S504.
[0185] Among them, the mapping list records at least two coding methods included in each of multiple coding type groups, and the super-resolution model corresponding to each coding type group.
[0186] S503, determine the target super-resolution model corresponding to the target coding method from the mapping list and execute step S506.
[0187] S504, determine the target coding type group whose similarity between the corresponding coding method and the target coding method meets the requirements from the mapping list.
[0188] In this embodiment, if the target coding method is not queried from the mapping list, a target coding type group whose corresponding coding method has a similarity to the target coding method meeting the requirements will be determined based on the coding methods included in each coding type group, so as to match a coding type whose included coding method is similar to the target coding method, that is, to determine the target coding type group that the coding features of the target coding method may match.
[0189] For example, it may be to determine from the mapping list the target coding type group whose included coding method has the highest similarity to the target coding method.
[0190] S505, if a target coding type group whose corresponding coding method has a similarity to the target coding method meeting the requirements is determined, determine the target super-resolution model corresponding to the target coding type group in the mapping list, and execute step S506.
[0191] S506, input the first image into the target super-resolution model to obtain the second image generated by the target super-resolution model.
[0192] In this embodiment, if the target coding method is not queried from the mapping list, it will be detected whether there is a target coding type group in the mapping list whose corresponding coding method has a similarity to the target coding method meeting the requirements. If a target coding type group is queried, directly use the target super-resolution model corresponding to the target coding type group to process the first image, so that in the case of a new coding method not existing in the mapping list, a suitable super-resolution model can also be adapted from the mapping list without having to retrain the corresponding super-resolution model.
[0193] In particular, when there is only one target coding type group whose corresponding coding method has a similarity to the target coding method meeting the requirements, it indicates that the coding features of the target coding method are similar to the coding features of each coding method in the target coding type. Therefore, the present application can also add the target coding method as the coding method included in the target coding type group in the mapping list.
[0194] S507, if multiple target coding type groups whose corresponding coding methods have a similarity to the target coding method meeting the requirements are determined, determine the multiple target super-resolution models corresponding to the multiple target coding type groups from the mapping list.
[0195] S508, based on the model parameters in each target super-resolution model, construct a comprehensive super-resolution model, and use the comprehensive super-resolution model to process the first image to generate the second image.
[0196] Among them, based on the model parameters of each target super-resolution model, constructing a comprehensive super-resolution model may be to fuse the model parameters such as the weights of multiple target super-resolution models to generate a new super-resolution model.
[0197] For the sake of easy distinction, this application refers to the super-resolution model constructed based on each target super-resolution model as a comprehensive super-resolution model.
[0198] Since this comprehensive super-resolution model synthesizes the image processing characteristics of each target super-resolution model and has a wider applicable coding method, therefore, using this comprehensive super-resolution model to process the first image can also generate a high-quality second image; moreover, there is no need to retrain a new super-resolution model.
[0199] It can be understood that step 508 is described by taking an implementation manner in which multiple target super-resolution models corresponding to multiple target coding type groups are in the first image as an example. In practical applications, after determining the multiple target super-resolution models corresponding to the multiple target coding type groups, this application can also process the first image by using each target super-resolution model respectively to obtain candidate images generated by each target super-resolution model. Based on the weights corresponding to each target super-resolution model, the candidate images obtained by each target super-resolution model processing the first image are fused to obtain the second image.
[0200] Among them, the weights corresponding to each target super-resolution can be set in advance; it can also be based on the similarity between the coding method corresponding to each target coding type group and this target coding method to determine the weights of the target super-resolution models corresponding to each target coding type group, and there is no specific limitation.
[0201] Corresponding to the image processing method provided in this application, this application also provides an image processing device.
[0202] Such as Figure 6 , which shows a schematic structural diagram of a composition of the image processing device provided in this application. The image processing device in this embodiment may include:
[0203] A type determination unit 601, configured to obtain the decoded first image and determine the target coding method corresponding to the first image;
[0204] A model determination unit 602, configured to determine, from multiple super-resolution models, the target super-resolution model corresponding to the target coding method, where the multiple super-resolution models correspond to different coding type groups, the target coding method matches the coding type group corresponding to the target super-resolution model, and the coding methods matched by the same coding type group have at least one same coding feature;
[0205] An image processing unit 603 for processing the first image using a target super-resolution model to generate a second image.
[0206] In a possible implementation, the model determination unit includes:
[0207] A model determination subunit for determining a target super-resolution model corresponding to the target encoding method based on a mapping list, where the mapping list records at least two encoding methods included in each of multiple encoding type groups, and the super-resolution model corresponding to each encoding type group.
[0208] In another possible implementation, the model determination subunit includes:
[0209] A first determination subunit for, if the target encoding method is queried from the mapping list, determining the target super-resolution model corresponding to the target encoding method from the mapping list;
[0210] A second determination subunit for, if the target encoding method is not queried from the mapping list, determining a target encoding type group whose corresponding encoding method has a similarity to the target encoding method that meets the requirements from the mapping list, and determining the target super-resolution model corresponding to the target encoding type group in the mapping list.
[0211] In another possible implementation, when the second determination subunit determines the target super-resolution model corresponding to the target encoding type group in the mapping list, specifically, if multiple target encoding type groups are determined from the mapping list, determining multiple target super-resolution models corresponding to the multiple target encoding type groups in the mapping list;
[0212] The image processing unit includes:
[0213] A first processing subunit for processing the first image using each target super-resolution model respectively, and fusing the candidate images obtained by processing the first image with each target super-resolution model based on the weights corresponding to each target super-resolution model to obtain a second image;
[0214] Or,
[0215] A second processing subunit for constructing a comprehensive super-resolution model based on the model parameters in each target super-resolution model, and using the comprehensive super-resolution model to process the first image to generate a second image.
[0216] In a possible implementation, the model determines that the super-resolution model recorded in the mapping list adopted in the subunit is a super-resolution model trained with the first sample images encoded by each encoding method in the encoding type group as training data and the second sample images corresponding to the first sample images as training targets, and the quality of the second sample images is higher than that of the first sample images.
[0217] In a possible implementation, the second sample image corresponding to the first sample image is: the original image corresponding to the first sample image without encoding, or an optimized image obtained by optimizing the original image through image optimization processing, where the quality of the original image corresponding to the first sample image is higher than that of the first sample image.
[0218] In yet another possible implementation, when the second sample image is an optimized image, the optimized image is obtained by sharpening the detail area of the original image.
[0219] In yet another possible implementation, in the loss function used to train the super-resolution model, different regional weights are assigned to different regions in the second sample image, and the detail complexity of different regions in the second sample image is different.
[0220] In yet another possible implementation, the loss function used to train the super-resolution model includes a noise suppression term. In the noise suppression term, different suppression weights are assigned to different pixel point regions in the second sample image. Among them, the greater the difference between the pixel point region in the second sample image and the corresponding pixel point region in the target image, the greater the suppression weight assigned to the pixel point region in the second sample image. The pixel point region includes at least one pixel point, and the target image is the image generated by the super-resolution model by processing the first sample image.
[0221] In yet another possible implementation, the loss function used to train the super-resolution model corresponding to each encoding type group matches the encoding characteristics of the encoding method corresponding to the encoding type group.
[0222] In another aspect, corresponding to a model training method of the present application, the present application further provides a model training device. As Figure 7 , a schematic diagram of a composition structure of the model training device provided by the present application is shown. The device in this embodiment may include:
[0223] A model determination unit 701, configured to determine super-resolution models to be trained corresponding to at least two encoding type groups respectively, and the encoding methods matching the same encoding type group have at least one same encoding feature;
[0224] A sample acquisition unit 702, configured to obtain, for each encoding method corresponding to each encoding type group, a first sample image encoded by the encoding method corresponding to the encoding type group, and a second sample image corresponding to the first sample image, where the quality of the second sample image is higher than that of the first sample image;
[0225] A model training unit 703, configured to, for each encoding type group, use the first sample image corresponding to the encoding type group as training data, and use the second sample image corresponding to the first sample image as a training target, to train the super-resolution model corresponding to the encoding type group, so as to obtain the trained super-resolution model corresponding to the encoding type group.
[0226] In a possible implementation manner, the sample acquisition unit includes:
[0227] An original image acquisition subunit, configured to obtain an original image that has not been encoded;
[0228] An encoding processing subunit, configured to encode the original image by using the encoding method corresponding to the encoding type group, to obtain a first sample image;
[0229] An image determination subunit, configured to determine the original image or an optimized image obtained by optimizing the original image as the second sample image corresponding to the first sample image.
[0230] In another possible implementation manner, the model training unit is specifically configured to use the first sample image corresponding to the encoding type group as training data, use the second sample image corresponding to the first sample image as a training target, and train the super-resolution model corresponding to the encoding type group in combination with a loss function;
[0231] Wherein, in the loss function, different regional weights are assigned to different regions in the second sample image, and the detail complexities of different regions in the second sample image are different;
[0232] And / or, the loss function includes a noise suppression term. In the noise suppression term, different suppression weights are assigned to different pixel point regions in the second sample image. Wherein, the greater the difference between the pixel point region in the second sample image and the corresponding pixel point region in the target image, the greater the suppression weight assigned to the pixel point region in the second sample image. The target image is an image generated by the super-resolution model by processing the first sample image, and the pixel point region includes at least one pixel point;
[0233] And / or, the loss function used to train the super-resolution model corresponding to each encoding type group matches the encoding characteristics of the encoding method corresponding to the encoding type group.
[0234] An embodiment of the present application further provides an electronic device. As Figure 8As shown, it shows a schematic diagram of a composition structure of the electronic device, and the electronic device at least includes a processor 801 and a memory 802;
[0235] The processor 801 is configured to execute the image processing method or the model training method described in any one of the above embodiments;
[0236] The memory 802 is used to store programs required for the processor to perform operations.
[0237] It can be understood that the electronic device may further include a display unit 803 and an input unit 804.
[0238] Of course, the electronic device may also have Figure 8 more or fewer components, and no limitation is imposed thereon.
[0239] An embodiment of the present application also provides a computer program product, including computer-readable instructions. When the computer-readable instructions run on the electronic device, the electronic device is enabled to implement any one of the image processing methods or model training methods provided by the embodiments of the present application.
[0240] An embodiment of the present application also provides a computer-readable storage medium. The storage medium carries one or more computer programs. When the one or more computer programs are executed by the electronic device, the electronic device can be enabled to implement any one of the image processing methods or model training methods provided by the embodiments of the present application.
[0241] In addition, it should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the drawings of the device embodiments provided in the present application, the connection relationship between the modules indicates that they have a communication connection, which can be specifically implemented as one or more communication buses or signal lines.
[0242] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general hardware. Of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be diverse, such as analog circuits, digital circuits, or dedicated circuits, etc. However, for this application, software program implementation is a better embodiment in more cases. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disc of a computer, and includes several instructions for causing a computer device (which can be a personal computer, training device, or network device, etc.) to execute the methods described in various embodiments of this application.
[0243] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.
[0244] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of this application are generated in whole or in part. The computer can be a general-purpose computer, a dedicated computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, training device, or data center to another website, computer, training device, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can store, or a data storage device such as a training device or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.
Claims
1. An image processing method, comprising: Obtaining a decoded first image and determining a target encoding method corresponding to the first image; Determining a target super-resolution model corresponding to the target encoding method from multiple super-resolution models, wherein the multiple super-resolution models correspond to different encoding type groups, the target encoding method matches the encoding type group corresponding to the target super-resolution model, and the same encoding type group includes at least two encoding methods, and the encoding methods matched by the same encoding type group have at least one same encoding feature; Processing the first image by using the target super-resolution model to generate a second image; Wherein, determining the target super-resolution model corresponding to the target encoding method includes: If the target encoding method exists among the encoding methods included in each encoding type group, determining the target super-resolution corresponding to the target encoding method; If the target encoding method does not exist among the encoding methods included in each encoding type group, determining a target encoding type group whose corresponding encoding method has a similarity to the target encoding method that meets the requirements based on the encoding methods included in each encoding type group, and determining the target super-resolution model corresponding to the target encoding type group.
2. The image processing method according to claim 1, wherein determining the target super-resolution model corresponding to the target encoding method from multiple super-resolution models includes: Determining the target super-resolution model corresponding to the target encoding method based on a mapping list, wherein the mapping list records at least two encoding methods included in each of multiple encoding type groups and the super-resolution model corresponding to each encoding type group.
3. The image processing method according to claim 2, wherein determining the target super-resolution model corresponding to the target encoding method based on the mapping list includes: If the target encoding method is queried from the mapping list, determining the target super-resolution model corresponding to the target encoding method from the mapping list; If the target encoding method is not queried from the mapping list, determining a target encoding type group whose corresponding encoding method has a similarity to the target encoding method that meets the requirements from the mapping list, and determining the target super-resolution model corresponding to the target encoding type group in the mapping list.
4. The image processing method according to claim 3, wherein determining the target super-resolution model corresponding to the target encoding type group in the mapping list includes: If multiple target encoding type groups are determined from the mapping list, determining multiple target super-resolution models corresponding to the multiple target encoding type groups in the mapping list; The processing the first image by using the target super-resolution model to generate a second image includes: Processing the first image by using each target super-resolution model respectively, and fusing the candidate images obtained by each target super-resolution model processing the first image based on the weights corresponding to each target super-resolution model to obtain a second image; Or, Based on the model parameters in each target super-resolution model, a comprehensive super-resolution model is constructed, and the first image is processed by the comprehensive super-resolution model to generate a second image.
5. The image processing method according to claim 2, wherein the super-resolution models recorded in the mapping list are super-resolution models corresponding to the encoding type group, which are trained with the first sample images encoded by each encoding method in the encoding type group as training data and the corresponding second sample images of the first sample images as training targets, and the quality of the second sample images is higher than that of the first sample images.
6. The second sample image corresponding to the first sample image in the image processing method according to claim 5 is: the original image corresponding to the first sample image that is not encoded, or an optimized image obtained by performing image optimization processing on the original image, where The quality of the original image corresponding to the first sample image is higher than that of the first sample image; The optimized image is obtained by sharpening the detail area of the original image.
7. The image processing method according to claim 5, in the loss function used to train the super-resolution model, different regional weights are assigned to different regions in the second sample image, and the detail complexities of different regions in the second sample image are different; And / or, a noise suppression term is included in the loss function used for training the super-resolution model. In the noise suppression term, different suppression weights are assigned to different pixel point regions in the second sample image, where The greater the difference between the pixel region in the second sample image and the corresponding pixel region in the target image, the greater the suppression weight assigned to the pixel region in the second sample image. The pixel region includes at least one pixel, and the target image is the image generated by the super-resolution model processing the first sample image; And / or, the loss function used to train the super-resolution model corresponding to each encoding type group matches the encoding characteristics of the encoding method corresponding to the encoding type group.
8. A model training method, comprising: Determine the super-resolution models to be trained corresponding to at least two encoding type groups respectively. The encoding methods matching the same encoding type group have at least one same encoding characteristic; each encoding type group includes at least two encoding methods; For each encoding method corresponding to each encoding type group, obtain the first sample image encoded by the encoding method corresponding to the encoding type group and the corresponding second sample image of the first sample image, and the quality of the second sample image is higher than that of the first sample image; For each encoding type group, use the first sample image corresponding to the encoding type group as training data and the corresponding second sample image of the first sample image as the training target to train the super-resolution model corresponding to the encoding type group to obtain the trained super-resolution model corresponding to the encoding type group; When performing image processing, to determine the target super-resolution model corresponding to the target encoding method of the first image, including: If the target encoding method exists among the encoding methods included in each encoding type group, determine the target super-resolution corresponding to the target encoding method; If the target encoding method does not exist among the encoding methods included in each encoding type group, based on the encoding methods included in each encoding type group, determine the target encoding type group whose similarity between the corresponding encoding method and the target encoding method meets the requirements, and determine the target super-resolution model corresponding to the target encoding type group.
9. The model training method according to claim 8, wherein the obtaining of the first sample image encoded by the encoding method corresponding to the encoding type group and the second sample image corresponding to the first sample image comprises: Obtaining an original image that has not been encoded; Encoding the original image by using the encoding method corresponding to the encoding type group to obtain a first sample image; Determining the original image or an optimized image obtained by optimizing the original image as the second sample image corresponding to the first sample image.
10. The model training method according to claim 8, wherein using the first sample image corresponding to the encoding type group as training data and using the second sample image corresponding to the first sample image as a training target to train the super-resolution model corresponding to the encoding type group comprises: Using the first sample image corresponding to the encoding type group as training data, using the second sample image corresponding to the first sample image as a training target, and training the super-resolution model corresponding to the encoding type group in combination with a loss function; Wherein, in the loss function, different regional weights are assigned to different regions in the second sample image, and the detail complexities of different regions in the second sample image are different; And / or, a noise suppression term is included in the loss function. In the noise suppression term, different suppression weights are assigned to different pixel point regions in the second sample image. Wherein, the greater the difference between the pixel point region in the second sample image and the corresponding pixel point region in the target image, the greater the suppression weight assigned to the pixel point region in the second sample image. The target image is an image generated by the super-resolution model by processing the first sample image, and the pixel point region includes at least one pixel point; And / or, the loss function used for training the super-resolution model corresponding to each encoding type group matches the encoding characteristics of the encoding method corresponding to the encoding type group.
11. An image processing apparatus, comprising: A type determination unit configured to obtain a decoded first image and determine a target encoding method corresponding to the first image; A model determination unit configured to determine, from a plurality of super-resolution models, a target super-resolution model corresponding to the target encoding method, wherein the plurality of super-resolution models correspond to different encoding type groups, the target encoding method matches the encoding type group corresponding to the target super-resolution model, the same encoding type group includes at least two encoding methods, and the encoding methods matched by the same encoding type group have at least one same encoding characteristic; An image processing unit configured to process the first image by using the target super-resolution model to generate a second image; Wherein, determining the target super-resolution model corresponding to the target encoding method comprises: If the target encoding method exists among the encoding methods included in each encoding type group, determining the target super-resolution corresponding to the target encoding method; If the target encoding method does not exist among the encoding methods included in each encoding type group, based on the encoding methods included in each encoding type group, determine a target encoding type group whose corresponding encoding method has a similarity to the target encoding method that meets the requirements, and determine a target super-resolution model corresponding to the target encoding type group.
12. A model training device, comprising: A model determination unit, configured to determine a super-resolution model to be trained corresponding to each of at least two encoding type groups, and the encoding methods matching the same encoding type group have at least one same encoding feature; each encoding type group includes at least two encoding methods; A sample acquisition unit, configured to, for each encoding method corresponding to each encoding type group, acquire a first sample image encoded by the encoding method corresponding to the encoding type group, and a second sample image corresponding to the first sample image, where the quality of the second sample image is higher than the quality of the first sample image; A model training unit, configured to, for each encoding type group, use the first sample image corresponding to the encoding type group as training data, and use the second sample image corresponding to the first sample image as a training target to train the super-resolution model corresponding to the encoding type group, so as to obtain the trained super-resolution model corresponding to the encoding type group; When performing image processing, determining a target super-resolution model corresponding to a target encoding method of a first image includes: If the target encoding method exists among the encoding methods included in each encoding type group, determine a target super-resolution corresponding to the target encoding method; If the target encoding method does not exist among the encoding methods included in each encoding type group, based on the encoding methods included in each encoding type group, determine a target encoding type group whose corresponding encoding method has a similarity to the target encoding method that meets the requirements, and determine a target super-resolution model corresponding to the target encoding type group.
Citation Information
Patent Citations
Video image super-division method and device, storage medium and electronic equipment
CN113055713A
Method for generating image super-division data set, image super-division model and training method
CN116503252A
Remote sensing image super-resolution learning method and device based on expert knowledge supervision
CN118247146A