Model training method and platform, image restoration method and device, equipment and medium

CN119948520AActive Publication Date: 2025-05-06BOE TECHNOLOGY GROUP CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202380010369.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-08-30
Publication Date
2025-05-06
Estimated Expiration
2043-08-30

AI Technical Summary

Technical Problem

The existing deep learning methods have limited training mechanisms in image repair, making it difficult to effectively improve the clarity and quality of images.

Method used

By acquiring multiple image sample pairs, updating model parameters using text data, predicted images and high-quality image samples, combined with human feedback on image repair quality, multi-stage training is carried out to optimize model effects.

Benefits of technology

The image quality of the image repair model is improved, the adaptability and effect of the model to image repair tasks is enhanced, and the model is closer to human visual evaluation through the text data feedback mechanism.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119948520A_ABST
    Figure CN119948520A_ABST
Patent Text Reader

Abstract

The invention discloses a model training method and platform, an image restoration method and device, equipment and a medium, and belongs to the technical field of image processing, and the model training method comprises the steps: obtaining a plurality of image sample pairs which comprise a first image sample and a second image sample of the same image; wherein the image quality of the second image sample is higher than that of the first image sample; a plurality of image sample pairs are used as training samples, a first preset model is trained, the first preset model is used for improving the image quality of the first image sample, and the training process comprises the steps of obtaining text data corresponding to a prediction image currently output by the first preset model; wherein the text data comprises data for evaluating the image quality of the predicted image; and updating parameters of the first preset model based on the text data, the prediction image and the second image sample.
Need to check novelty before this filing date? Find Prior Art

Description

Model training method and platform, image restoration method, device, equipment and medium Technical Field

[0001] The present disclosure relates to the field of image processing technology, and in particular to a model training method and platform, an image restoration method, device, equipment and medium. Background Art

[0002] With the development of image processing technology, image restoration has become increasingly necessary. For example, increasing image resolution from low to high to improve clarity, repairing scratches and noise to improve image quality, or restoring missing parts of an image have become necessary. Image restoration is often achieved through deep learning methods, but existing deep learning methods have limited training mechanisms.

[0003] Overview

[0004] The present disclosure provides a model training method, the method comprising:

[0005] Acquire a plurality of image sample pairs, the image sample pairs comprising a first image sample and a second image sample of the same image; wherein the image quality of the second image sample is higher than the image quality of the first image sample;

[0006] A first preset model is trained using the plurality of image sample pairs as training samples, wherein the first preset model is used to improve the image quality of the first image samples, and the training process includes:

[0007] Acquire text data corresponding to the predicted image currently output by the first preset model; wherein the text data includes data evaluating the image quality of the predicted image;

[0008] Based on the text data, the predicted image and the second image sample, the parameters of the first preset model are updated.

[0009] For example, the text data is text input by a user for the predicted image; and updating the parameters of the first preset model based on the text data, the predicted image, and the second image sample includes:

[0010] Determining the similarity between a target text and the predicted image; wherein the target text is the text data, or a text having semantics opposite to that of the text data;

[0011] determining a loss value based on the predicted image and the second image sample;

[0012] Based on the similarity and the loss value, the parameters of the first preset model are updated.

[0013] For example, determining the similarity between the target text and the predicted image includes:

[0014] Encoding the target text to obtain a text feature vector of the target text;

[0015] Encoding the predicted image to obtain an image feature vector of the predicted image; wherein the dimensions of the text feature vector and the image feature vector are consistent;

[0016] The similarity is determined based on the text feature vector and the image feature vector.

[0017] For example, the training of the first preset model using the plurality of image sample pairs as training samples includes:

[0018] Performing a first training on the first preset model using some of the image sample pairs as training samples; wherein, in the first training, updating parameters of the first preset model based on the predicted image output by the first preset model and the second image sample;

[0019] Performing a second training on the first preset model obtained by the first training using some of the image sample pairs as training samples;

[0020] In the second training, the parameters of the first preset model obtained by the first training are updated based on the text data, the predicted image and the second image sample.

[0021] For example, the method further includes:

[0022] Acquire a plurality of third image samples and text data samples corresponding to the third image samples; wherein the text data samples are used to describe the image quality of the third image samples;

[0023] Based on the plurality of the third image samples and the text data samples, a third training is performed on the second preset model; wherein the second preset model is used to determine the similarity between the third image samples and the text data samples;

[0024] The updating of the parameters of the first preset model based on the text data, the predicted image, and the second image sample includes:

[0025] Inputting the predicted image and the text data into the second preset model when the third training is completed;

[0026] Based on the similarity output by the second preset model, the predicted image and the second image sample, the parameters of the first preset model are updated.

[0027] For example, the text data sample carries a category label, and the category label is used to indicate whether the description of the text data sample conforms to the image quality of the third image sample; and the third training of the second preset model based on the plurality of the third image samples and the corresponding at least two text data samples includes:

[0028] Inputting the third image sample and the text data sample into the second preset model, and obtaining the predicted similarity between the third image sample and each of the text data samples output by the second preset model;

[0029] Based on the predicted similarity and the category label, the parameters of the second preset model are updated.

[0030] For example, the third image sample corresponds to two text data samples, the two text data samples including a first category of text data samples and a second category of text data samples, the first category characterizing that the description of the text data sample meets the image quality of the third image sample, and the second category characterizing that the description of the text data sample does not meet the image quality of the third image sample.

[0031] For example, the second preset model includes a text encoder and an image encoder, and similarity determination modules respectively connected to the text encoder and the image encoder;

[0032] The text encoder is used to perform text encoding on the text data sample to obtain a predicted text vector;

[0033] The image encoder is configured to perform image encoding on the third image sample to obtain a predicted image vector; wherein the predicted image vector and the predicted text vector have the same dimension;

[0034] The similarity determination module is used to determine the predicted similarity between the predicted image vector and the predicted text vector.

[0035] For example, during the process of training the first preset model, the step of performing the third training on the second preset model is performed; the multiple third image samples include at least one of the following: the first image sample, the second image sample, and the predicted image output by the first preset model before the current moment.

[0036] For example, the third training is performed in an interval between training the first preset model, the third image sample includes a predicted image output by the first preset model for the first image sample, and the third training of the second preset model based on the plurality of the third image samples and the text data sample includes:

[0037] Inputting a plurality of the first image samples into the first preset model;

[0038] Inputting the predicted image output by the first preset model and the text data sample corresponding to the predicted image into the second preset model to perform the third training on the second preset model;

[0039] In the third training, the parameters of the first preset model are fixed.

[0040] For example, after the third training, the training of the first preset model includes:

[0041] Inputting a plurality of the first image samples into the first preset model;

[0042] Inputting the predicted image output by the first preset model and the text data into the second preset model;

[0043] updating parameters of the first preset model based on the predicted image, the second image sample, and the predicted similarity output by the second preset model;

[0044] When the parameters of the first preset model are updated, the parameters of the second preset model are fixed.

[0045] For example, the acquiring of a plurality of third image samples and text data samples corresponding to the third image samples includes:

[0046] Acquire first text data corresponding to the third image sample;

[0047] determining a category corresponding to the first text data, the category being used to indicate whether a description of the first text data is consistent with the image quality of the third image sample;

[0048] generating second text data based on the first text data; wherein the category of the second text data is different from the category of the first text data;

[0049] The first text data and the second text data are used as text data samples corresponding to the third image sample.

[0050] For example, obtaining text data corresponding to the predicted image currently output by the first preset model includes:

[0051] Determining whether there is a target third image sample corresponding to the predicted image from the plurality of third image samples; wherein the target third image sample and the predicted image correspond to the same first image sample;

[0052] If yes, taking the text data sample corresponding to the target third image sample as the text data;

[0053] If not, then obtaining text data input for the predicted image.

[0054] For example, obtaining text data corresponding to the predicted image currently output by the first preset model includes:

[0055] displaying the predicted image;

[0056] Acquire text data input for the prediction image.

[0057] For example, after displaying the predicted image, the method further includes:

[0058] monitoring an input operation on an operation interface displaying the predicted image;

[0059] If the input operation is not detected, the preset text data is used as the text data corresponding to the predicted image;

[0060] The obtaining of text data input for the predicted image includes:

[0061] If the input operation is detected, the text data input for the predicted image is obtained.

[0062] For example, the text data includes at least one entry; the at least one entry is used to describe the image quality of the predicted image in different image regions and / or the image quality in different quality dimensions.

[0063] For example, the method further includes at least one of the following:

[0064] In response to the first training, displaying a predicted image output by the first preset model, and in response to performing a first preset operation on the predicted image, ending the first training;

[0065] In response to the second training, the predicted image output by the first preset model is displayed, and in response to a second preset operation performed on the predicted image, the second training is ended.

[0066] For example, during the process of performing the third training on the second preset model, the method further includes:

[0067] Displaying the loss value corresponding to at least one training of the second preset model before the current moment; wherein the gradient is determined by the predicted similarity and the category label;

[0068] In response to a third preset operation being performed on each of the displayed loss values, the third training is ended.

[0069] The present disclosure also provides an image restoration method, the method comprising:

[0070] The target image to be restored;

[0071] Inputting the target image into an image restoration model; wherein the image restoration model is a first preset model trained according to the model training method;

[0072] Obtain a repaired image output by the image repair model; wherein the image quality of the repaired image is higher than that of the target image.

[0073] The present disclosure also provides a model training platform, which includes:

[0074] A sample library, configured to store a plurality of image sample pairs, wherein the image sample pairs include a first image sample and a second image sample of the same image; wherein the image quality of the second image sample is higher than that of the first image sample;

[0075] A training module is configured to train a first preset model multiple times using the plurality of image sample pairs as training samples, wherein the first preset model is configured to improve the image quality of the first image samples; wherein the training process includes:

[0076] Acquire text data corresponding to the predicted image currently output by the first preset model; wherein the text data includes data evaluating the image quality of the predicted image;

[0077] Based on the text data, the predicted image and the second image sample, the parameters of the first preset model are updated.

[0078] The present disclosure also provides an image restoration device, comprising:

[0079] A first acquisition module is used for the target image to be restored;

[0080] An input module, configured to input the target image into an image restoration model; wherein the image restoration model is a first preset model trained according to the model training method;

[0081] The second acquisition module is used to acquire the repaired image output by the image repair model; wherein the image quality of the repaired image is higher than that of the target image.

[0082] The present disclosure provides a model training method that can obtain multiple image sample pairs and use the multiple image sample pairs as training samples to train a first preset model. The training process includes: obtaining text data corresponding to a predicted image currently output by the first preset model; and updating the parameters of the first preset model based on the text data, the predicted image, and a second image sample. Specifically, the image sample pairs include a first image sample of lower quality and a second image sample of higher quality of the same image. Accordingly, the first preset model is used to improve the image quality of the first image sample.

[0083] By adopting the training method provided by the present invention, in the process of training the first preset model using multiple image samples, the parameters of the first preset model can be updated using text data, predicted images and second image samples; since the text data includes data for evaluating the image quality of the predicted image, that is, the text data can be used to evaluate the image quality of the predicted image, it is possible not only to supervise the training based on the difference between the predicted image output by the model and the second image sample, but also to obtain text data for evaluating the predicted image to supervise the training based on the image quality evaluation provided by the text data; in this way, the model training can be supervised from the pixel difference dimension between the predicted image and the second image sample, as well as the quality evaluation dimension of the predicted image, so that the model can be combined with the provided quality evaluation, and the direction of model optimization is provided by the text data, thereby optimizing the model effect and improving the image restoration quality of the trained model.

[0084] An embodiment of the present disclosure also discloses an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the model training method or image restoration method as described above when executed.

[0085] An embodiment of the present disclosure also discloses a computer-readable storage medium, which stores a computer program that enables a processor to execute the model training method or image restoration method as described in the present disclosure.

[0086] The above description is only an overview of the technical solution of the present disclosure. In order to more clearly understand the technical means of the present disclosure, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present disclosure more obvious and easy to understand, the specific implementation methods of the present disclosure are listed below.

[0087] BRIEF DESCRIPTION OF THE DRAWINGS

[0088] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or related technologies, the following is a brief introduction to the drawings required for the description of the embodiments or related technologies. Obviously, the drawings described below are some embodiments of the present disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without inventive efforts. It should be noted that the scales in the drawings are for illustration only and do not represent the actual scale.

[0089] FIG1 is a schematic diagram showing the overall process of the model training method according to an embodiment of the present disclosure;

[0090] FIG2 is a schematic diagram showing the steps of the model training method according to an embodiment of the present disclosure;

[0091] FIG3 is a schematic diagram showing a process of training a first preset model in stages according to an embodiment of the present disclosure;

[0092] FIG4 is a schematic diagram showing a process of adding a training phase of a second preset model to a training process of a first preset model in an embodiment of the present disclosure;

[0093] FIG5 shows a schematic flow chart of the steps for training the second preset model in an embodiment of the present disclosure;

[0094] FIG6 shows a schematic diagram of the composition of a third image sample input to the second preset model in an embodiment of the present disclosure;

[0095] FIG7 is a schematic diagram showing a process of obtaining a text data sample corresponding to a third image sample in an embodiment of the present disclosure;

[0096] FIG8 shows a schematic diagram of the model structure of the second preset model in an embodiment of the present disclosure;

[0097] FIG9 is a schematic diagram showing the complete process of the first training, the second training, and the third training in an embodiment of the present disclosure;

[0098] FIG10 shows a schematic diagram of another process of training a second preset model in an embodiment of the present disclosure;

[0099] FIG11 shows a loss value change trend diagram of the second preset model in the third training process in an embodiment of the present disclosure;

[0100] FIG12 is a schematic diagram showing a process of a model training method according to an embodiment of the present disclosure;

[0101] FIG13 shows a schematic diagram of the framework structure of a model training platform in an embodiment of the present disclosure;

[0102] FIG14a shows a schematic diagram of a first operation interface of a model training platform during a model training process according to an embodiment of the present disclosure;

[0103] FIG14 b shows a schematic diagram of a second operation interface of the model training platform during the model training process according to an embodiment of the present disclosure;

[0104] FIG14c is a schematic diagram showing a third operation interface of the model training platform during the model training process according to an embodiment of the present disclosure;

[0105] FIG14d shows a schematic diagram of a fourth operation interface of the model training platform during the model training process according to an embodiment of the present disclosure;

[0106] FIG15 is a schematic diagram showing the steps of the image restoration method according to an embodiment of the present disclosure;

[0107] FIG16 shows a schematic diagram of the framework structure of the image restoration device in an embodiment of the present disclosure.

[0108] Detailed description

[0109] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present disclosure without making any creative efforts shall fall within the scope of protection of the present disclosure.

[0110] In related technologies, image restoration is usually achieved through deep learning methods, such as obtaining an image restoration model by training an autoencoder fully convolutional neural network or a generative adversarial network. However, the training effect of the model is often improved by adjusting the neural network structure design, loss function, training parameters, etc., making it difficult to significantly improve the quality of image restoration using the trained model.

[0111] In view of this, the present disclosure improves the training process of the model. Specifically, during the model training process, human quality feedback language for image restoration is incorporated, so that during the model training process, in addition to utilizing the pixel differences between the network-repaired image and the complete image to supervise the training, the model can also utilize human quality feedback language for image restoration, thereby combining the quality feedback language to supervise the training. In this way, during the model training process, with the help of the evaluation language for the network-repaired image, the model can perceive human visual evaluation and assist the model in updating parameters, thereby optimizing the model training effect and helping to improve the model's image restoration quality.

[0112] 1 and 2 , FIG1 shows a schematic diagram of the overall flow of the model training method of the present disclosure, and FIG2 shows a schematic diagram of the step flow of the model training method of the present disclosure. As shown in FIG1 and FIG2 , the model training method of the present disclosure can be applied to electronic devices, and specifically may include the following steps:

[0113] Step S201: Acquire multiple image sample pairs, where the image sample pairs include a first image sample and a second image sample of the same image;

[0114] The image quality of the second image sample is higher than that of the first image sample.

[0115] In this embodiment, since the image sample pair includes a first image sample and a second image sample of the same image with different image qualities, the first image sample and the second image sample have the same image content, but differ only in image quality. Image quality can include image quality descriptions such as resolution, noise, and scratches. For example, the second image sample having higher image quality than the first image sample may include: the second image sample having a higher resolution than the first image sample, or the second image sample having less noise than the first image sample; or the second image sample having no scratches while the first image sample has scratches.

[0116] For example, two images with different image qualities can be collected for the same scene at the same perspective, such as taking a low-resolution image and a high-resolution image of the same person, so that the low-resolution image is used as the first image sample and the high-resolution image is used as the second image sample; for another example, a high-definition image can be blurred, and the blurred image is used as the first image sample, and the high-definition image is used as the second image sample; for another example, a high-definition image can be processed with scratches, noise, etc., and the image with scratches and noise is used as the first image sample, and the high-definition image is used as the second image sample.

[0117] For example, when the application scenario is the restoration of old photos, the goal of model training is to train a model that can repair old photos. The model needs to improve the resolution of old photos and repair missing parts in old photos. When obtaining multiple image sample pairs, a low-resolution image and a high-resolution image can be taken for the same person, so that the low-resolution image is used as the first image sample and the high-resolution image is used as the second image sample; and a high-resolution image is taken for the same person, and the high-resolution image is blurred, and corresponding noise and scratches are added and some areas are randomly cut out to obtain the first image sample, and the high-resolution image is used as the second image sample.

[0118] The sizes of the first image sample and the second image sample can be processed into preset sizes to meet the model training requirements.

[0119] Step S202: training a first preset model using a plurality of image sample pairs as training samples, wherein the first preset model is used to improve the image quality of the first image sample.

[0120] After obtaining multiple image sample pairs, multiple first image samples in the multiple image sample pairs can be used as input of the first preset model, and the second image samples can be used as supervision of the model to train the first preset model; wherein, the first preset model can adopt the existing model structure, which will not be repeated here.

[0121] Among them, the first preset model is mainly used to perform image restoration on the input first image sample, so as to improve the image quality of the first image sample. Specifically, the image restoration can mainly include: clarity restoration, scratch repair and noise removal, among which clarity restoration can refer to improving the resolution of the first image sample; of course, in addition to the image restoration in the above examples, other types of restoration can also be included, such as restoration of missing parts of the image; in practice, the first preset model can be used to perform at least one of the above-mentioned image restorations, that is, perform at least one of the repair processes of clarity restoration, scratch repair, noise removal and restoration of missing parts.

[0122] Among them, after the first preset model performs image restoration on the first image sample, it obtains a restored predicted image, and the first preset model outputs the predicted image. The difference between the predicted image and the second image sample can be used to supervise the training of the first preset model, such as for updating the parameters of the first preset model.

[0123] In the process of training the first preset model, there is at least one training session in which text data of the user's image quality evaluation of the predicted image output by the first preset model is obtained. Then, based on the text data, the gap between the predicted image and the expected image quality level (the image quality level represented by the second image sample) is determined, and this gap is used together with the gap between the above-mentioned predicted image and the second image sample to supervise the training of the first preset model.

[0124] Accordingly, in the process of training the first preset model, at least in one training session, the following steps are performed:

[0125] Step S2021: Acquire text data corresponding to the predicted image currently output by the first preset model;

[0126] Step S2022: updating the parameters of the first preset model based on the text data, the predicted image, and the second image sample;

[0127] The text data includes data for evaluating the image quality of the predicted image.

[0128] In this embodiment, during a training session, a predicted image can be displayed, for example, the predicted image output by the first preset model can be displayed on the display interface of an electronic device, so that the user can observe with the naked eye the effect of the first preset model repairing the first image sample, and thus text for evaluating the image quality of the predicted image can be input for the predicted image to obtain text data corresponding to the predicted image.

[0129] Among them, the user can input text data through an input tool, such as typing text data on the input interface; or, the user's voice can be collected through a voice acquisition module, and then the voice is recognized to obtain text data. In this case, the user can speak the language of quality evaluation of the predicted image to the voice acquisition module without the need for manual entry by the user.

[0130] Among them, the text data may include sentences evaluating the predicted image. For example, the text data includes the sentence "the image is not clear enough". When the parameters of the first preset model are updated in combination with the text data, the first preset model needs to be optimized in the direction of "the image is clearer"; for another example, the text data includes the sentence "the hair is not clear". When the parameters of the first preset model are updated in combination with the text data, the first preset model needs to be optimized in the direction of "the hair is clearer"; for another example, the text data includes the sentence "the image noise is not eliminated cleanly". When the parameters of the first preset model are updated in combination with the text data, the first preset model needs to be optimized in the direction of "eliminating clean noise".

[0131] In practice, when the parameters of the first preset model are updated according to the text data, the predicted image and the second image sample, since the text data can characterize the image quality gap between the predicted image and the second image sample, the first gap between the predicted image and the second image sample at the user perspective level can be determined according to the text data and the predicted image, and the second gap between the two at the pixel level can be determined according to the predicted image and the second image sample; when updating the parameters of the first preset model, it can be done according to the first gap and the second gap. In this way, the first preset model can be optimized not only in the optimization direction provided by the user vision, but also in the optimization direction provided by the second image sample, so that the first preset model can be combined with the provided quality assessment to obtain the ability to perceive human vision, thereby providing multiple optimization directions for the first preset model, improving the model training method in the field of image restoration, and helping to improve the image restoration quality of the first preset model.

[0132] In some embodiments, the text data may include at least one term; the at least one term is used to describe the image quality of the predicted image in different image regions and / or the image quality in different quality dimensions.

[0133] In this embodiment, the text data may include at least one term. These terms may vary depending on the image restoration task to be performed by the first preset model. In other words, the at least one term corresponds to the image restoration task. Specifically, the term may include a term for describing the image quality of different image regions and / or image quality of different quality dimensions.

[0134] Among them, the image region refers to different regions in the predicted image. In image restoration, for the same image restoration task, the restoration effects of different image regions may be different. The text data can indicate the differences in the restoration of different image regions.

[0135] For example, in old photo restoration, clarity restoration requires performing clarity restoration on two people's faces simultaneously, resulting in one person's face being restored very clearly while the other's face being restored less clearly. For another example, clarity restoration requires performing clarity restoration on one person's face and their clothing and accessories simultaneously, resulting in the person's face being restored more clearly while their clothing and accessories are not clearly restored. Accordingly, the text data includes image quality evaluations of different image regions of the predicted image. For example, the terms related to image regions in the text data may include terms such as "clothing and accessories are not clear enough" and "face A is not clear enough."

[0136] Among them, the quality dimension is related to the image restoration task performed by the first preset model. One image restoration task can correspond to one quality dimension. For example, if the image restoration task is a clarity restoration task, the quality dimension includes the clarity dimension. If the image restoration task also includes a scratch restoration task, the quality dimension also includes the scratch dimension.

[0137] By way of example, the image restoration task of the first preset model includes clarity restoration and scratch restoration, then the entries related to the quality dimension in the text data may include: the image is not clear enough, scratches exist, at least one of them; by way of example, the image restoration task of the first preset model includes clarity restoration and noise elimination, then the entries included in the text data may include: the image is not clear enough, noise exists, at least one of them; by way of another example, the image restoration task of the first preset model includes noise elimination and missing part restoration, then the entries included in the text data may include: noise exists, missing parts are not completely restored, at least one of them; by way of another example, the image restoration task of the first preset model includes clarity restoration and missing part restoration, then the entries included in the text data may include: not clear enough, missing parts are not completely restored, at least one of them.

[0138] Of course, the above is only an exemplary description. In other examples, other image restoration tasks may be included, and the text data may include entries describing the predicted image in other quality dimensions.

[0139] When adopting the above method, the text data can indicate the weak links of the first preset model in the image restoration process. Therefore, when updating the first preset model, the text data can be used to determine the area to be optimized in the predicted image that is different from the expected image quality, so that the area to be optimized can be fed back to the first preset model, and the first preset model can be supervised to perform reinforcement learning on the restoration of the area to be optimized, thereby guiding the first preset model to continuously optimize the weak links in image restoration.

[0140] In some embodiments, since text data is used to evaluate the image quality of the predicted image, the text data may include text that reversely evaluates the predicted image. The reverse evaluation of the predicted image means that the text contained is an evaluation opposite to the actual image quality of the predicted image. For example, if the face in the predicted image is not clear, the reverse evaluation is: the face is clear; for example, if there are scratches in the predicted image, the reverse evaluation is: there are no scratches.

[0141] Alternatively, the text data may include text that positively evaluates the predicted image. A positive evaluation of a predicted image means that the text contained therein is an evaluation that is consistent with the actual image quality of the predicted image. For example, if the face in the predicted image is not clear, the positive evaluation is: the face is not clear; for another example, if there are scratches in the predicted image, the positive evaluation is: there are scratches.

[0142] Since the parameters of the first preset model can be updated in conjunction with text data, it is necessary to allow the first preset model to combine text data during training to identify weak links in the image restoration process. This can be understood as the need to combine text data to determine the optimization direction of the model during the image restoration process. Therefore, whether it is positive or negative evaluation, the text data can indicate the direction in which the model needs to be optimized. Specifically, based on the text data, it can be determined which aspects of the predicted image quality need to be improved.

[0143] In this embodiment, the gap between the predicted image and the expected image quality level (which can be understood as the image quality level represented by the second image sample) can be determined based on the text data. This gap reflects the evaluation of the predicted image from a human perspective, so that the first preset model can perceive the evaluation from a human perspective through the text data. In the process of updating the parameters of the first prediction model, the update can be based on this gap and the gap in pixels between the predicted image and the second image sample.

[0144] Among them, since the text data can be data generated by positive evaluation or data generated by negative evaluation, and since it is necessary to guide the first preset model to optimize towards the expected image quality level based on the text data, the gap between the predicted image and the text data on which it is negatively evaluated can be determined, and the gap can be represented by the similarity between the text data on which the negative evaluation is conducted and the predicted image.

[0145] In a specific implementation, the similarity between the target text and the predicted image can be determined, and a loss value can be determined based on the predicted image and the second image sample; then, the parameters of the first preset model can be updated based on the similarity and the loss value. The target text is text data, or text with semantics opposite to the text data.

[0146] Among them, the target text includes an expected description of the image quality that needs to be further optimized for the predicted image. For example, if the background of the predicted image is not clear enough, the target text needs to include an expected description of "clear background" to clarify the optimization direction of image restoration.

[0147] In one embodiment, if the text data is data generated by a positive evaluation, a text with the opposite semantics to the text data can be determined, and then the text data of the positive evaluation can be converted into text data of a negative evaluation, so that the text data of the negative evaluation obtained after the conversion can be used as the target text to determine the similarity between the target text and the predicted image; for example, if the hair of the portrait in the predicted image is not clear enough, and the text data contains a positive description of "the hair is not clear", then "the hair is not clear" needs to be converted into "the hair is clear" to clarify the optimization direction for optimizing the image restoration.

[0148] If the text data is data generated by reverse evaluation, the text data can be used as the target text to determine the similarity between the target text and the predicted image.

[0149] In this embodiment, when updating the parameters of the first preset model, it is also necessary to update the parameters based on the difference in pixels between the predicted image and the second image. Specifically, the loss value can be determined based on the predicted image and the second image samples. Specifically, a loss function can be constructed based on the predicted image and the second image samples to obtain the loss value. The loss function can adopt the loss function shown in the following formula (1) or formula (2):

[0150] In formulas (1) and (2), C represents the number of channels (for RGB images with three channels, C = 3; for grayscale images with a single channel, C = 1), H represents the image height, and W represents the image width. y is the true image, x is the input image, and f(x) is the output image.

[0151] Since similarity can reflect the difference between the predicted image and the expected image quality, and its value can be 0-1, and the loss value can reflect the difference between the predicted image and the second image sample, and its value can also be 0-1, in practice, similarity can also be regarded as a type of loss value. Then, the parameters of the first preset model can be updated based on the two loss values. Specifically, weights can be preset for the similarity and loss values ​​respectively, and the two can be weighted and summed according to the weights to obtain a total loss. Based on this total loss, the parameters of the first preset model are updated.

[0152] In one example, the weight corresponding to the similarity may be smaller than the weight corresponding to the loss value, that is, the importance corresponding to the first gap between the predicted image determined based on the text data at the user perspective level may be slightly smaller than the importance corresponding to the second gap between the predicted image and the second image sample at the pixel level.

[0153] Among them, since the text data is text type data and the predicted image is image type data, when determining the similarity between the two, the two can be converted into the same feature space for comparison, that is, the text data is converted into target type data, and the predicted image is also converted into target type data, so that the similarity between the two converted into target type data can be determined.

[0154] For example, the predicted image can be converted into text-type data, so that the predicted image can be compared with the text data in the same text space. Specifically, features can be extracted from the predicted image to obtain a feature vector, and then the feature vector can be converted into a text vector according to certain rules, so that it can be compared with the text data.

[0155] For example, in this embodiment, when determining the similarity between the target text and the predicted image, the two can be first encoded to obtain the vectors obtained by encoding each of the two, and then the similarity can be determined based on the distance between the vectors. Specifically, the target text can be encoded to obtain the text feature vector of the target text, and the predicted image can be encoded to obtain the image feature vector of the predicted image; then, the similarity is determined based on the text feature vector and the image feature vector.

[0156] Among them, the dimensions of the text feature vector and the image feature vector are consistent.

[0157] In this embodiment, when encoding the target text, the keywords in the target text can be encoded, and the text feature vector obtained by encoding can be a one-dimensional vector; when encoding the predicted image, the features of the predicted image can be extracted first, and then the extracted image features can be encoded, and a one-dimensional image feature vector can also be obtained after encoding; then, the cosine distance between the text feature vector and the image feature vector can be calculated to obtain the similarity between the two.

[0158] Next, the training process of the first preset model is introduced.

[0159] In some examples, the first preset model can be trained in stages. The process of training the first preset model in stages can include two stages. In the first stage, the first preset model can be trained using some image samples. Text data may not be added to this training process. After the training in this stage is completed, the second stage is entered. In the second stage, the first preset model can be continued to be trained using the remaining image samples. During the training process of this stage, text data can be added for parameter updating.

[0160] 3 , a schematic diagram of the process of training the first preset model in stages is shown. As shown in FIG3 , a first training is performed on the first preset model using some image sample pairs as training samples; and a second training is performed on the first preset model obtained by the first training using some image sample pairs as training samples;

[0161] In the first training, the parameters of the first preset model are updated based on the predicted image and the second image samples output by the first preset model; in the second training, the parameters of the first preset model obtained by the first training are updated based on the text data, the predicted image and the second image samples.

[0162] Among them, there may be no overlapping and repeated image sample pairs between the partial image sample pairs used in the first training and the partial image sample pairs used in the second training; for example, n first image sample pairs from multiple image sample pairs are used in the first training, and the m image sample pairs used in the second training may be the second image sample pairs from the multiple image sample pairs excluding the n first image samples.

[0163] In some other examples, the image sample pairs used in the first training and the image sample pairs used in the second training are repeated. For example, the first training uses n first image sample pairs from the plurality of image sample pairs, and the second training may use all of the plurality of image sample pairs, or a portion of m second image sample pairs, where the m second image sample pairs include at least one first image sample pair. In this case, since the image sample pairs used in the first and second training are not repeated, the diversity of the training samples can be increased, thereby improving the generalization performance of the first preset model.

[0164] Among them, the predicted images output by the first preset model in the second training and the first training can be displayed. In this way, when there are repeated image samples between the image sample pairs used in the first training and the image sample pairs used in the second training, the predicted image corresponding to the same first image sample in the first training and the predicted image corresponding to the second training can both be displayed, so as to facilitate the user to determine whether the same image is more optimally restored in the second training.

[0165] For example, when outputting the predicted image corresponding to the first image sample in the second training, it can be determined whether the first image sample is input into the first preset model in the first training. If so, the predicted image corresponding to the first image sample in the first training is also output. Thus, the predicted images corresponding to the first image sample in the first training and the second training can be displayed simultaneously on the display interface, so as to determine whether the first preset model is optimized in the second training by comparing the two predicted images.

[0166] In the first training, since no text data is required for supervision, it can be understood as pre-training of the first preset model. When the pre-training meets the termination condition, the first training can be stopped. The termination condition can be: the difference between the predicted image and the second image sample is small, such as the loss value is less than the preset loss threshold. Of course, in order to complete the first training as quickly as possible to achieve the goal of pre-training, the loss threshold can be set larger, such as achieving basic image restoration capabilities, and then the first training can be terminated.

[0167] Among them, the end condition can also be triggered by the user. As mentioned above, during the first training process, the first preset model can output a predicted image corresponding to the first image sample, and the user can determine the image restoration capability of the first preset model through the predicted image. In this way, the user can decide when to end the first training.

[0168] In some other embodiments, since it is necessary to determine the distance between the predicted image and the expected image quality level (the image quality level represented by the second image sample) based on text data, in practice, it is necessary to determine the similarity between the target text and the predicted image. In some examples, the similarity between the target text and the predicted image can be determined by a neural network, that is, using deep learning technology, the second preset model can be trained so that the second preset model can obtain the text feature vector of the target text and the image feature vector of the predicted image, and then determine the similarity between the target text and the predicted image.

[0169] Specifically, the process of training the first preset model in stages may also include a stage of training the second preset model. The second preset model may be a CLIP (Contrastive Language-Image Pre-Training) model structure, which may include an image encoder and a text encoder, wherein the image encoder is used to encode the predicted image and the text encoder is used to encode the text data.

[0170] 4 and 5 , FIG4 shows a schematic diagram of a process of adding a training phase of a second preset model to a training process of a first preset model, and FIG5 shows a schematic flow diagram of steps for training the second preset model. As shown in FIG4 and 5 , before the first training or the second training of the first preset model, the following steps may also be included:

[0171] Step S501: Acquire a plurality of third image samples and text data samples corresponding to the third image samples;

[0172] Step S502: performing a third training on the second preset model based on a plurality of third image samples and corresponding text data samples;

[0173] The second preset model after the third training is completed is used to determine the similarity between the predicted image and text data in the training of the first preset model.

[0174] Accordingly, as shown in Figure 4, when the parameters of the first preset model are updated based on the text data, the predicted image and the second image sample, the predicted image and the text data can be input into the second preset model when the third training is completed; and the parameters of the first preset model can be updated based on the similarity output by the second preset model, the predicted image and the second image sample.

[0175] In this embodiment, the third image sample may be different from the aforementioned image sample pair, or may include image samples from the aforementioned image sample pair, or may include the predicted image output during the first training. In practice, there is no requirement for the third image sample to be related to the aforementioned image sample pair. For example, if the image sample pair is a portrait photograph, the third image sample is not limited to a portrait photograph and may also be an image of an animal, etc., thereby reducing the difficulty of obtaining image samples.

[0176] Among them, each third image sample corresponds to a text data sample. As mentioned above, the text data may include text data generated by positive evaluation of the predicted image, and text data generated by negative evaluation of the predicted image. In some examples, in the third training, the text data sample corresponding to the third image sample may be data generated by positive evaluation of the third image sample, or data generated by negative evaluation of the third image sample, or the third image sample may also be two text data samples generated by positive and negative evaluations.

[0177] The text data sample may also include at least one entry; the at least one entry is used to describe the image quality of the third image sample in different image regions and / or the image quality in different quality dimensions.

[0178] In this case, when annotating the text data sample corresponding to the third image sample, the annotation can be performed according to the image restoration task of the first preset model. For example, if the image restoration task is a clarity restoration task, the entries of the annotated text data sample on the quality dimension can include entries on the clarity dimension, such as whether each image area is clear.

[0179] Taking positive evaluation as an example, the image restoration task of the first preset model includes clarity restoration and noise elimination. If the image quality of the third image sample is not high (not clear enough, noise exists, there are scratches or some areas are not restored), then the terms included in the text data sample may include: the image is not clear enough, there is noise, at least one of them; for another example, the image restoration task of the first preset model includes noise elimination and missing part restoration, then the terms included in the text data sample may include: noise exists, missing parts are not completely restored, at least one of them; for another example, the image restoration task of the first preset model includes clarity restoration and missing part restoration, then the terms included in the text data sample may include: not clear enough, missing parts are not completely restored, at least one of them.

[0180] Taking reverse evaluation as an example, the image restoration task of the first preset model includes clarity restoration and noise elimination. If the image quality of the third image sample is not high (not clear enough, there is noise, there are scratches or some areas are not restored), then the terms included in the text data sample may include: the image is clear enough and there is no noise; for another example, the image restoration task of the first preset model includes noise elimination and missing part restoration, then the terms included in the text data sample may include: there is no noise and the missing part is completely restored. At least one; for another example, the image restoration task of the first preset model includes clarity restoration and missing part restoration, then the terms included in the text data sample may include: clear enough and the missing part is completely restored.

[0181] The image region refers to different regions in the third image sample, and the entries related to the image region in the text data sample may also be related to the image restoration task to be performed by the first preset model.

[0182] Taking positive evaluation as an example, for the clarity restoration task, the third image sample may include multiple different image areas. If the third image sample is not clear enough in image area 1 and image area 2, the text data sample may include: image area 1 (such as clothes and accessories) is not clear enough, image area 2 (such as face A) is not clear enough, etc.

[0183] Taking reverse evaluation as an example, for the clarity restoration task, the third image sample may include multiple different image areas. If the third image sample is not clear enough in image area 1 and image area 2, the text data sample may include: image area 1 (such as clothes and accessories) is clear enough, image area 2 (such as face A) is clear enough, and other terms.

[0184] By adopting this method of labeling text data samples, the text data samples labeled with the third image samples are related to the image restoration task of the first preset model. In this way, the second preset model can be optimized in the third training towards the similarity determination method required for the image restoration task of the first preset model. The second preset model can extract the image feature vector of the third image sample and the text feature vector of the text data sample in the direction related to the image restoration task. For example, the second preset model is optimized in the direction of fully extracting the features reflecting the text and image in the image restoration task. In this way, the second preset model can match the needs of the first preset model, thereby providing the first preset model with a more accurate comparison result between the predicted image and text data.

[0185] In some embodiments, in the third training, it is necessary to perform a third training on the second preset model based on multiple third image samples and text data samples. Since the second preset model is used to determine the similarity between the third image sample and the text data sample, the parameters of the second preset model can be updated based on the similarity output by the second preset model and the actual similarity between the third image sample and the text data sample. As shown in Figure 4, specifically, the text data sample carries a category label, and the category label is used to indicate whether the description of the text data sample meets the image quality of the third image sample. In this way, in the third training, the third image sample and the text data sample can be input into the second preset model to obtain the predicted similarity between the third image sample output by the second preset model and each of the text data samples; then, the parameters of the second preset model can be updated based on the predicted similarity and the category label.

[0186] In this embodiment, the category label carried by the text data sample is used to indicate whether the description of the text data sample conforms to the image quality of the third image sample. Specifically, if the text data sample conforms to the image quality of the third image sample, it is positively evaluated data, and its category label can be 1; if the text data sample does not conform to the image quality of the third image sample, it is negatively evaluated data, and its category label can be 0.

[0187] For example, taking the first image sample as an example, assuming that the restoration task is clarity restoration, the clarity of the first image sample is poor. If the text data sample is "image clear", it is negatively evaluated data, and its category label can be 0. If the text data sample is "image unclear", it is positively evaluated data, and its category label can be 1. Taking the second image sample as an example, assuming that the restoration task is clarity restoration, the clarity of the second image sample is high. If the text data sample is "image clear", it is positively evaluated data, and its category label can be 1. If the text data sample is "image unclear", it is negatively evaluated data, and its category label can be 0.

[0188] When updating the parameters of the second preset model based on the predicted similarity and category label, the loss corresponding to the text data sample can be determined based on the predicted similarity and category label, and the parameters of the second preset model can be updated according to the loss. Specifically, the corresponding parameters of the text encoder and image encoder in the second preset model can be updated.

[0189] After multiple parameter updates, the second preset model can align the image and text so that the two are compared in the same feature space, thereby generating accurate comparison results.

[0190] Exemplarily, there is at least one third image sample corresponding to at least two text data samples, the at least two text data samples include text data samples of the first category and text data samples of the second category, the description of the text data samples of the first category is consistent with the image quality of the third image sample, and the description of the text data samples of the second category does not meet the image quality of the third image sample.

[0191] For example, referring to Figure 6, a schematic diagram of the composition of the third image sample input into the second preset model is shown. As shown in Figure 6, the text data sample of the first category can be understood as being generated by positively evaluating the third image sample, and the category label it carries can be 1, indicating that the two are matched; the text data sample of the second category can be understood as being generated by negatively evaluating the third image sample, and the category label it carries can be 0, indicating that the third image sample and the text data sample are not matched.

[0192] Specifically, each third image sample can correspond to two text data samples, or, some third image samples correspond to two text data samples, and the remaining third image samples correspond to one text data sample; for example, 500 third image samples are included, of which 200 third image samples correspond to two text data samples, text data samples of the first category and text data samples of the second category; the remaining 200 third image samples correspond to text data samples of the first category, and the remaining 100 third image samples correspond to text data samples of the second category.

[0193] In which, when the third image sample corresponds to at least two text data samples, one of the text data samples can be automatically generated based on the other text data sample, for example, the text data sample of the first category can be generated based on the text data sample of the second category, or the text data sample of the second category can be generated based on the text data sample of the first category.

[0194] Accordingly, referring to FIG7 , a schematic diagram of a process for obtaining a text data sample corresponding to a third image sample is shown. As shown in FIG7 , the process may specifically include the following steps:

[0195] Step S701: Acquire first text data corresponding to the third image sample;

[0196] Step S702: determining a category corresponding to the first text data, where the category is used to indicate whether the description of the first text data is consistent with the image quality of the third image sample;

[0197] Step S703: generating second text data based on the first text data; wherein the category of the second text data is different from the category of the first text data;

[0198] Step S704: taking the first text data and the second text data as text data samples corresponding to the third image sample.

[0199] Among them, the first text data is input by the user, which can be data of the first category or data of the second category; then, the category corresponding to the first text data can be determined, wherein the category to which it belongs can be determined by detecting whether the first text data contains target keywords. The target keywords can be "no", "noise exists", "scratches exist", etc., which are related to the image restoration task and have negative meanings.

[0200] After determining the category corresponding to the first text data, second text data opposite to the category can be generated. For example, the category corresponding to the first text data is the first category, and the representation is consistent with the image quality of the third image sample, then the category corresponding to the second text data is the second category, and the representation is inconsistent with the image quality of the third image sample; for another example, the category corresponding to the first text data is the second category, then the category corresponding to the second text data is the first category.

[0201] Among them, the method of generating the second text data can be: if the first text data is the first category, the words with negative meaning in the first text data can be converted into words with positive meaning, thereby obtaining the second text data, such as removing the "not" in the first text data, and changing the "noise exists" in the first text data to "no noise does not exist".

[0202] If the first text data belongs to the second category, the words with positive meaning in the first text data can be converted into words with negative meaning to obtain the second text data, such as adding "not" to the first text data to make "clear hair" become "unclear hair", or changing "noise does not exist" in the first text data to "noise exists".

[0203] In this way, for the first third image sample, although it is only labeled with a text data sample of one category, the labeled text data sample can generate a text data sample of another category, so that the third image sample can correspond to text data samples of two categories.

[0204] Accordingly, referring to FIG8 , a schematic diagram of a model structure of a second preset model is shown. The second preset model may include a text encoder and an image encoder, and similarity determination modules connected to the text encoder and the image encoder, respectively.

[0205] Among them, the text encoder is used to perform text encoding on the text data sample to obtain a predicted text vector;

[0206] The image encoder is used to perform image encoding on the third image sample to obtain a predicted image vector, where the predicted image vector and the predicted text vector have the same dimension;

[0207] The similarity determination module is used to determine the predicted similarity between the predicted image vector and the predicted text vector.

[0208] In this embodiment, in the third training, the third image sample can be input into the image encoder, and the text data sample corresponding to the third image sample can be input into the text encoder, wherein, if the third image sample corresponds to two text data samples, then both text data samples corresponding to the third image sample can be input into the text encoder; wherein, the image encoder can encode the third image sample to obtain a predicted image vector, and the text encoder can encode the text data sample to obtain a predicted text vector; the similarity determination module can calculate the cosine distance between the predicted text vector and the predicted image vector to obtain the predicted similarity between the two.

[0209] When the third image sample corresponds to two text data samples, the text encoder can output predicted text vectors corresponding to the two text data samples respectively, and the similarity determination module can obtain two predicted similarities, which respectively correspond to the two text data samples.

[0210] 9 , a schematic diagram of the complete process of the first training, the second training and the third training is shown. As shown in FIG9 , the third training can be performed before the first training or before the second training, that is, the second preset model can be trained before the first training starts, or the third training can be performed after the first training ends and before the second training starts.

[0211] Illustratively, the third training may be performed after the first training completes and before the second training begins. The third image samples used in the third training may overlap with the image sample pairs used in the first training process and / or overlap with the predicted images output by the first preset model in the first training process. Specifically, the plurality of third image samples may include at least one of the following: the first image samples, the second image samples, and the predicted images output by the first preset model before the current moment.

[0212] Exemplarily, the plurality of third image samples may include all or part of the first image samples in the plurality of image sample pairs, such as the first image samples used in the first training process; or the plurality of third image samples may include all or part of the second image samples in the plurality of image sample pairs, such as the second image samples used in the first training process. Alternatively, the plurality of third image samples may include the predicted images output by the first preset model for the first image samples in the first training process; or the plurality of third image samples may include the first image samples and the second image samples, so that the third training and the first training can share image samples; or the plurality of third image samples may include the first image samples and the predicted images; or the plurality of third image samples may include the second image samples and the predicted images; or the plurality of third image samples may include the first image samples, the second image samples, and the predicted images.

[0213] Among them, when the multiple third image samples include the predicted image output by the first preset model in the first training, the image sample pairs used in the second training can include the image sample pairs used in the first training. With this sample setting method, since the third image sample needs to correspond to the text data sample, it is necessary to annotate the text data sample for the predicted image. Since the annotated text data sample is used to evaluate the quality of the predicted image, in practice, when conducting the second training, it is also necessary to obtain text data for evaluating the predicted image. Then, the text data sample annotated for the third image sample in the third training can be directly used as the text data of the predicted image in the second training.

[0214] For example, the first image sample A1 corresponds to the predicted image A1 in the first training, and the predicted image A1 serves as the third image sample, then it corresponds to at least one text data sample T1. Then, the first image sample A1 participates in the second training, and corresponds to the predicted image A2 in the second training. Since the image quality difference between the predicted image A2 and the predicted image A1 is very small when the second training starts, the text data sample T1 corresponding to the predicted image A1 can be used as the text data corresponding to the predicted image A2, thereby participating in the parameter update of the first preset model.

[0215] When the third image sample includes the first image sample as a corresponding prediction image in the first training, full utilization of the labeled text data samples can be improved, thereby improving training efficiency.

[0216] Exemplarily, the first preset model can be connected to the second preset model to form a new model, and the first training, second training and third training can be performed based on the new model. Specifically, when the first training and the second training are performed, the parameters of the second preset model can be fixed while the parameters of the first preset model are kept updated; when the third training is performed, the parameters of the first preset model can be fixed while the parameters of the second preset model are kept updated.

[0217] In which, when the third image sample includes a predicted image, the labeling of the text data sample can be carried out as the first training proceeds. That is, during the first training, the predicted image output by the first preset model can be used both in the parameters of the first preset model and in the process of text data samples. For example, during the first training process, the output predicted images are all associated with the text data samples, so that during the first training process, the third image samples used for the third training can be collected.

[0218] Exemplarily, the third training can be performed after the first training is completed and before the second training is started. In this way, the third training can be performed in the gap between training the first preset model, wherein, as described above, when the third image sample includes the predicted image output by the first preset model for the first image sample, the text data sample corresponding to the predicted image can be used again in the second training of the first preset model.

[0219] During specific implementation, referring to FIG10 , another schematic diagram of the process of training the second preset model is shown. As shown in FIG10 , when the second preset model is subjected to the third training based on multiple third image samples and text data samples, multiple first image samples can be input into the first preset model; and the predicted image output by the first preset model and at least two text data samples corresponding to the predicted image are input into the second preset model to perform the third training on the second preset model; wherein, in the third training, the parameters of the first preset model are fixed.

[0220] In this embodiment, the third training can be performed during the first training. In this case, at the beginning of the first training, the predicted image output by the first preset model and at least two text data samples corresponding to the predicted image can be input into the second preset model. Then, the parameters of the first preset model are updated based on the predicted image and the second image samples; and the parameters of the second preset model are updated based on the category label corresponding to the text data sample and the similarity output by the second preset model.

[0221] In which, the process of updating the parameters of the first preset model and the parameters of the second preset model can be independent of each other, so that when the parameters of the first preset model are updated, the parameters of the second preset model are fixed unchanged; when the parameters of the second preset model are updated, the parameters of the first preset model are fixed unchanged.

[0222] In this way, the third training can be started in parallel during the first training, thereby improving the training progress.

[0223] Correspondingly, the third image sample may further include the first image sample. Then, in addition to inputting the predicted image output by the first preset model and the at least two text data samples corresponding to the predicted image into the second preset model, the first image sample and the text data sample corresponding to the first image sample may also be input into the second preset model when the first image sample is input into the first preset model.

[0224] In this way, in a first training, the first image sample and the predicted image corresponding to the first image sample can be input into the second preset model for training, so that when the first preset model is trained based on n first image samples, the second preset model can be trained based on 2n third image samples, thereby speeding up the training progress of the second preset model.

[0225] In this embodiment, if the first training is completed, the third training can be terminated or continued based on its optimization effect. When the third training is required to continue, the parameters of the first preset model are first fixed, and then the remaining third image samples and corresponding text data samples are input into the second preset model for retraining.

[0226] In some embodiments, the predicted image sent to the second preset model for training can be the predicted image output by the first preset model when the first training is about to end. In this way, the image quality difference between the predicted image for the third training and the predicted image output for the same first image sample in the second training is small, so that the text data sample corresponding to the predicted image can be reused.

[0227] For example, at the end of the first training, the predicted images corresponding to a preset number of first image samples are input into the second preset model for training, and the preset number of first image samples are input into the first preset model in the second training. The difference between the predicted images output in the second training and the predicted images input into the second preset model is small. Therefore, the text data samples can be reused as text data corresponding to the predicted images in the second training to update the parameters of the first preset model.

[0228] Thus, when the first preset model is first trained using a portion of the image samples, the predicted images output during the first training can be saved. Then, a third training can be started. During the third training, a plurality of third image samples targeted by the third training can be obtained, wherein the plurality of third image samples can include the plurality of predicted images output during the first training. A text data sample is annotated for each third image sample, so that the plurality of predicted images also have their own corresponding text data samples.

[0229] For example, after the third training, a second training session may be initiated, wherein the second preset model may be used to determine the similarity between the predicted image and the text data. In the second training session, multiple first image samples may be input into the first preset model; the predicted image and text data output by the first preset model may be input into the second preset model; and the parameters of the first preset model may be updated based on the predicted image, the second image samples, and the predicted similarity output by the second preset model.

[0230] When the parameters of the first preset model are updated, the parameters of the second preset model are fixed.

[0231] In this example, after the training of the second preset model is completed, the parameters of the second preset model can be fixed, so that in the second training, the predicted image output by the first preset model and the text data corresponding to the predicted image can be input into the second preset model, and the second preset model outputs the similarity between the predicted image and the text data. This similarity can be involved in the update of the first preset model together with the loss value determined by the predicted image and the second image sample.

[0232] Exemplarily, regardless of whether the third image sample includes the predicted image output by the first preset model during the first training process, in the second training, since it is necessary to obtain text data corresponding to the predicted image, when obtaining text data corresponding to the predicted image currently output by the first preset model, it can be determined from the plurality of third image samples whether there is a target third image sample corresponding to the predicted image; if so, the text data sample corresponding to the target third image sample is used as the text data; if not, the text data input for the predicted image is obtained;

[0233] The target third image sample and the predicted image correspond to the same first image sample.

[0234] In this example, in the second training, it is necessary to obtain the text data corresponding to the predicted image output by the first preset model. If the predicted image is used as the third image sample to train the second preset model, it has a corresponding text data sample. In practice, the labeled text data sample is directly used as the text data corresponding to the predicted image.

[0235] In practice, for the same first image sample, there may be differences between the predicted image output by the first preset model in the first training and the predicted image output by the first preset model in the second training. In this case, when searching for the target third image sample, the identifier of the first image sample associated with the predicted image in the third training and the identifier of the first image sample input to the first preset model in the second training can be used as the basis to search for the predicted image output by the first preset model this time, and whether the text data sample corresponding to the predicted image used in the third training can be used, that is, to search for the target third image sample. If found, the text data sample corresponding to the target third image sample is used as the text data corresponding to the predicted image output by the first preset model this time.

[0236] As shown in Figure 6, in the first training, the first image sample A is input into the first preset model to obtain the predicted image A1, and the predicted image A1 is annotated with text data sample 3 and text data sample 4. The predicted image A1, text data sample 3 and text data sample 4 are input into the second preset model as training samples for the third training; when the second training comes, the first image sample A is input into the first preset model again and output to the predicted image A2. Since the predicted image A1 and the predicted image A2 correspond to the same first image sample A, the text data sample 3 and text data sample 4 can be input into the second preset model as the text data corresponding to the predicted image A2 to obtain the similarity.

[0237] It should be noted that if the target third image sample is associated with multiple text data samples, the text data sample with negative evaluation can be used as the text data corresponding to the predicted image. Specifically, if the text data sample associated with the target third image sample is a text data sample with positive evaluation, it can be converted into a text data sample with negative evaluation, and the converted text data sample can be used as the text data corresponding to the predicted image.

[0238] Among them, if the currently output predicted image has not been trained as the third image sample of the second preset model, it is necessary to re-acquire the text data corresponding to the predicted image. Since the text data can be input by the user, the predicted image needs to be displayed for the user to observe. Specifically, the predicted image can be displayed on the display interface, and the text data entered by the user for the predicted image can be obtained.

[0239] Of course, for example, even if the predicted image is used as the third image sample in the third training and has a text data sample, for the same first image sample, the predicted image output by the first preset model in the first training and the predicted image output in the second training may have large differences. For example, at the beginning of the first training, the first image sample B is input to the first preset model, and its corresponding predicted image B1 is input to the second preset model. However, as the first training progresses, the image restoration quality of the first preset model will be optimized. In the second training, if the first image sample B is input to the first preset model after the first training, the predicted image B2 it outputs is large different from the predicted image B1, and it is no longer appropriate to use the text data sample corresponding to the predicted image B2.

[0240] In this case, the text data samples corresponding to the predicted image may not be used as the text data for updating the parameters of the first preset model, but the predicted image may be displayed so that the user re-enters text data for the predicted image.

[0241] As described above, since the first preset model can be subjected to the first training and the second training in stages, and the third training can be performed before the first training or between the first training and the second training, the third image sample used in the third training may include the predicted image output in the first training. For example, the predicted image output during the training process can be displayed in real time during the first training, the second training and the third training, and in the process of displaying the predicted image, the text data input by the user can be monitored, so that the entire training stage can be visualized by the user.

[0242] In one example, when obtaining text data corresponding to a predicted image currently output by a first preset model, the predicted image may be displayed, and text data input for the predicted image may be obtained.

[0243] In this example, in response to the training of the first preset model, the predicted image output by the first preset model during the training process can be displayed. Specifically, the predicted image output by the first preset model can be displayed in both the first training and the second training, so that the user can input text data for the predicted image. As described above, the text data can be input by using an input tool on the display interface displaying the predicted image, or by voice input.

[0244] In another example, after displaying the predicted image, input operations on the operation interface displaying the predicted image can also be monitored; if no input operation is monitored, the preset text data is used as the text data corresponding to the predicted image; accordingly, when obtaining the text data input for the predicted image, the text data input for the predicted image can be obtained if the input operation is monitored.

[0245] In this embodiment, the interface for displaying the predicted image can be called an operation interface. In this operation interface, the user can input text data through input operations. The input operation can be an operation of typing text in the operation interface, or it can be an operation of clicking the voice collection control in the operation interface. When the voice collection control is clicked, the device starts to collect voice data and starts to recognize the voice data to obtain text data.

[0246] When the predicted image is displayed, a countdown of a preset duration may be started. During the countdown, the system monitors whether an input operation is received. If no input operation is received, preset text data may be used as the text data corresponding to the predicted image. The preset text data may be text data that negatively evaluates the predicted image. Specifically, the preset text data may be consistent with the image quality level of the second image sample, thereby eliminating the need for the user to input text data.

[0247] In one example, during the first training and / or the second training, the output predicted image can be displayed so that the user can observe whether the predicted image meets expectations. If so, the user can specify whether the training is complete. Alternatively, during the first training, the predicted image output by the first preset model can be displayed, and the first training can be terminated in response to a first preset operation performed on the predicted image. Alternatively, during the second training, the predicted image output by the first preset model can be displayed, and the second training can be terminated in response to a second preset operation performed on the predicted image. Alternatively, during both the first and second training, the output predicted image can be displayed so that the user can terminate the first and second training sessions at the appropriate time based on visual observation of the predicted image, thereby improving model training efficiency.

[0248] Among them, the first preset operation can be the operation of clicking the "end control" of the first training, and the second preset operation can be the operation of clicking the "end control" of the second training; or, the first preset operation and the second preset operation can be the operation of double-clicking or right-clicking the mouse in the operation interface, which will not be repeated here.

[0249] When the first preset operation is detected, the first training of the first preset model can be stopped and the parameters of the first preset model can be fixed. When the second preset operation is detected, the first training of the first preset model can be stopped and the parameters of the first preset model can be fixed.

[0250] In another example, during the third training of the second preset model, the loss value corresponding to at least one training of the second preset model before the current moment can also be displayed; and in response to the third preset operation performed on each of the displayed loss values, the third training is ended; wherein the gradient is determined by the predicted similarity and the category label.

[0251] In this example, in response to the start of the third training, the loss value based on which each gradient update is performed in the third training can be displayed. The loss value can be determined based on the predicted similarity and the category label corresponding to the text data sample. In this way, the optimization process of the second preset model can be determined based on the displayed loss value. As shown in Figure 11, a trend graph of the loss value change of the second preset model during the third training process is output. As shown in Figure 11, as the third training deepens, the loss value becomes smaller and smaller. When the loss value change curve indicates that the loss value has converged, the third training can be ended.

[0252] The third preset operation is as described above for the first and second preset operations and is not further described here. In response to the third preset operation, the parameters of the second preset model can be fixed, and the input of the second preset model can be connected to the output of the first preset model. Specifically, the input of the image encoder in the second preset model can be connected to the output of the first preset model.

[0253] When this implementation is adopted, the training process of the first preset model and the second preset model can be visualized, so that when the first preset model and the second preset model are trained in stages, it is convenient for the user to actively control the start and end timing of each training stage, thereby optimizing the model training process.

[0254] As described above, in the case of including a first preset model and a second preset model, the output end of the first preset model can be connected to the input end of the second preset model, and the two models constitute a target model, which can be called a model to be trained. In practice, the target model can be first trained in stages using image sample pairs, multiple third image samples, and text data samples corresponding to the third image samples to obtain the trained first preset model and the second preset model, wherein the trained first preset model can be used as an image restoration model, and the trained second preset model can be used in the spatial alignment task of text and image.

[0255] Among them, the training of the target model includes the following processes:

[0256] The first stage of training (first training) process: the first image sample in the image sample pair is input into the first preset model, and the pixel difference between the predicted image output by the first preset model and the second image sample corresponding to the first image sample is calculated, that is, the loss value is obtained, and the parameters of the first preset model are updated according to the loss value. In this process, the second preset model is not processed, that is, the predicted image is not input into the second preset model for training;

[0257] The second stage of training (third training) process: fix the parameters of the first preset model, mainly including the following training methods:

[0258] Inputting the first image sample and the text data sample corresponding to the first image sample into the first preset model and the second preset model, inputting the second image sample and the text data sample corresponding to the second image sample into the second preset model, and at the same time, inputting the predicted image output by the first preset model and the text data sample corresponding to the predicted image into the second preset model, the second preset model outputs the predicted similarity between the image sample and the text data sample, then determining the loss value corresponding to the second preset model according to the predicted similarity and the category label corresponding to the text data sample, and then updating the parameters of the second preset model;

[0259] Inputting a first image sample into a first preset model, inputting a predicted image output by the first preset model and a text data sample corresponding to the predicted image into a second preset model, the second preset model outputting a predicted similarity between the image sample and the text data sample, then determining a loss value corresponding to the second preset model based on the predicted similarity and the category label corresponding to the text data sample, and then updating the parameters of the second preset model;

[0260] A new image sample that is different from the first image sample, the second image sample and the predicted image, and a text data sample corresponding to the new image sample are input into a second preset model. The second preset model outputs the predicted similarity between the image sample and the text data sample. Then, based on the predicted similarity and the category label corresponding to the text data sample, the loss value corresponding to the second preset model is determined, and then the parameters of the second preset model are updated.

[0261] The first image samples, the second image samples, the predicted image samples and the new image samples mentioned above are collectively referred to as third image samples in the third training.

[0262] After the second stage of training is completed, the third stage of training (second training) begins:

[0263] Fixing the parameters of the second preset model, inputting the first image sample into the first preset model, inputting the predicted image output by the first preset model and the text data generated by the user's reverse evaluation of the predicted image into the second preset model, and the second preset model outputting the similarity between the predicted image and the text data;

[0264] Next, the loss value between the predicted image and the second image sample is calculated, and the parameters of the first preset model are updated according to the similarity and the loss value.

[0265] After the third stage of training is completed, the first preset model can be used as the image restoration model.

[0266] 12 , a schematic diagram of a process of a model training method is exemplarily shown. As shown in FIG12 , taking portrait photo restoration as an example, it is necessary to train an image restoration model for portrait photo restoration. The model training method exemplarily includes the following process:

[0267] S11: Preparing a plurality of high-resolution portrait photos, photographing the plurality of portrait photos, and pre-processing the obtained images as second image samples. Then, blurring the plurality of portrait photos, such as by scratching or creating noise on the portrait photos, photographing the blurred portrait photos, and pre-processing the obtained images as first image samples. The first image sample and the second image sample belonging to the same portrait photo form an image sample pair.

[0268] In this example, the number of image sample pairs may be 1000 pairs.

[0269] S12: Constructing a model, including constructing a first preset model and a second preset model; wherein the first preset model may adopt a DDPM model (Diffusion Models Beat GANs on Image Synthesis) or a CNN model in related art; the first preset model is used to repair the input image, and the repair includes: improving the clarity of the image and removing scratches and noise; the second preset model may include an image encoder and a text encoder, and similarity determination modules connected to the image encoder and the text encoder respectively;

[0270] The second preset model can adopt the mature CLIP model structure, which includes two parts: a text encoder and an image encoder. The image encoder is used to encode the image feature information of the input image and output an image feature vector. The text encoder is used to encode the text feature information of the input text and output a text feature vector.

[0271] The second preset model is connected to the output end of the first preset model. In the subsequent training process, the model composed of the first preset model and the second preset model can be trained separately.

[0272] S13: Perform a first training. When performing the first training, disconnect the first preset model from the second preset model. In the first training, input a first image sample from a plurality of image sample pairs into the first preset model, display a predicted image output by the first preset model after repairing the first image sample, and construct a loss function based on the predicted image and the corresponding second image sample, thereby updating the parameters of the first preset model. It should be noted that the second image sample and the predicted image used to construct the loss function are for the same first image sample. The loss function may be a function described in the following formula (1) or formula (2):

[0273] In formulas (1) and (2), C represents the number of channels (for RGB images with three channels, C = 3; for grayscale images with a single channel, C = 1), H represents the image height, and W represents the image width. y is the true image, x is the input image, and f(x) is the output image.

[0274] S14: During the first training process, the user can observe the restoration optimization process of the first preset model on the first image sample through the displayed predicted image. For example, if the quality of the predicted image is getting higher and higher, it means that the first preset model has a preliminary image restoration capability. For another example, if the image quality of the predicted image is low, the first image sample is continuously input to optimize the first preset model. When it is monitored that the user performs a first preset operation on the displayed predicted image, such as clicking a corresponding control on the operation interface, the first training stops. At this time, the parameters of the first preset model are fixed and the third training begins.

[0275] It is assumed that when the first training stops, the number of first image samples input to the first preset model is 340;

[0276] In the first training, the predicted image corresponding to each first image sample can be saved, so that the predicted image can be used as a third image sample in the subsequent third training to label the text data sample for the third image sample.

[0277] S15: Constructing new third image samples. The third image samples may be obtained by collecting blurry old photos and clear new photos. A preset number of predicted images near the end of the first training may be used as the third image samples. For example, the 40 predicted images last input to the first preset model during the first training process are used as the third image samples.

[0278] The third image samples are assumed to be 500, which is less than the number of the first image samples in the plurality of image sample pairs;

[0279] Each third image sample is labeled with a text data sample. When the user clicks on a third image sample, an input area may pop up on the operation interface for inputting the text data sample. The input text data sample may include text that positively evaluates the third image sample. For example, for a third image sample taken from an old photo, the picture is relatively blurry, and the input text may include: blurry background; jagged edges; a little noise; unnatural eyes; unclear hair, etc.; for a third image sample taken from a new photo, the picture is relatively clear, and the input text may include: clear background; less noise; clear hair, etc.

[0280] Of course, for the same third image sample, text that conducts a reverse evaluation of the third image sample may also be included; for example, for the third image sample taken of an old photo, the picture is relatively blurry, so the input text may be: clear background; low noise; clear hair, etc.; for the third image sample taken of a new photo, the picture is relatively clear, so the input text may be: blurry background; jagged edges; a little noise; unnatural eyes; unclear hair, etc.

[0281] In this example, each third image sample corresponds to a text data sample of the first category and a text data sample of the second category;

[0282] Each text data sample corresponding to the third image sample carries a category label, which is used to characterize the true similarity between the image quality of the text data sample and the third image sample. For example, for the text data sample of the first category, its category label can be set to 1, indicating that the text data sample is a positive evaluation, and the similarity between the two is 1; for the text data sample of the second category, its category label can be set to 0, indicating that the text data sample is a negative evaluation, and the similarity between the two is 0;

[0283] S16: Starting the third training, as shown in the black part in FIG8 , a plurality of third image samples and text data samples corresponding to each of the third image samples are input into the second preset model, wherein the third image samples are input into the image encoder, and the text data samples are input into the text encoder, and the similarity determination module can determine the similarity between the image feature vector output by the image encoder and the text feature vector output by the text encoder;

[0284] Next, based on the similarity and the category label carried by the text data sample, the loss value for updating the parameters of the second preset model is determined; it should be noted that in the third training, the parameters of the first preset model are fixed unchanged;

[0285] In the third training, the loss value of the second preset model during the training process can be displayed. In this way, the user can observe whether the third training needs to be completed based on the displayed loss value. If the third training needs to be completed, the corresponding control can be triggered to end the third training and start the second training;

[0286] S17: Start the second training. During this training process, the output end of the first preset model needs to be connected to the input end of the second preset model;

[0287] S171: Inputting the first image sample of the 1000 image sample pairs into the first preset model, wherein, since the third image sample used when the second preset model is trained includes the 40 predicted images finally output by the first preset model during the first training, the 40 first image samples corresponding to the 40 predicted images can be first input into the first preset model to utilize the text data samples annotated therein during the third training;

[0288] Among them, the quality difference between the predicted image output by the first preset model and the predicted image output during the first training is not large, therefore, the text data samples corresponding to the 40 predicted images can be used as text data; specifically, when the 40 first image samples corresponding to the 40 predicted images are input into the first preset model, the text data samples corresponding to the 40 predicted images are input into the second preset model, and the predicted image output by the first preset model is also input into the second preset model, so that the second preset model can output the similarity, and the loss value between the predicted image and the second image sample is calculated by formula (1) or formula (2), and then, according to the loss value and the similarity, the parameters of the first preset model are updated;

[0289] After the 40 first image samples corresponding to the 40 predicted images are input into the first preset model for training, the remaining 960 image sample pairs in the 1000 image sample pairs can be used to continue training the first preset model;

[0290] S172: Continue training process:

[0291] The first image sample of the 960 image sample pairs is input into the first preset model, and the predicted image output by the first preset model can be input into the second preset model and displayed in the display area of ​​the front-end operation interface;

[0292] The user interface also features an input area alongside the display area, where the user enters text data corresponding to the displayed predicted image. In this second training, the text data entered is the reverse evaluation data. For example, if the person in the current image responds that the hair is not clear, the text data entered is the clear hair. This way, the image feature vector of the predicted image must be close to the semantic vector of the text "clear hair," thereby stimulating the first preset model to train in the direction of clear hair. If the user does not enter text in the input area, the default reverse evaluation text is used as the text data.

[0293] The text data inputted in the input area is inputted into the second preset model, and the second preset model outputs the similarity between the predicted image and the text data. The loss value between the predicted image and the second image sample is calculated by formula (1) or formula (2), and then the parameters of the first preset model are updated according to the loss value and the similarity.

[0294] It should be noted that, in the second training, the parameters of the second preset model are fixed.

[0295] Among them, during the second training process, the user can determine the image restoration ability of the first preset model through the displayed predicted image. When the image quality of the predicted image is always at a high level, the user can trigger the second preset operation to end the second training; of course, the second training can also be ended when the loss value and the similarity value both approach their respective preset thresholds.

[0296] Based on the same inventive concept, the present disclosure also provides a model training platform, as shown in FIG13 , which shows a schematic diagram of a framework structure of a model training platform. As shown in FIG13 , the model training platform may specifically include a sample library and a training module; wherein:

[0297] A sample library, configured to store a plurality of image sample pairs, each image sample pair comprising a first image sample and a second image sample of the same image; wherein the image quality of the second image sample is higher than that of the first image sample;

[0298] The training module is configured to, in response to a training operation, train a first preset model multiple times using a plurality of image sample pairs as training samples, wherein the first preset model is configured to improve image quality of the first image sample; wherein the training process includes:

[0299] Acquire text data corresponding to the predicted image currently output by the first preset model; wherein the text data is used to describe the image quality difference between the predicted image and the second image sample;

[0300] Based on the text data, the predicted image and the second image sample, the parameters of the first preset model are updated.

[0301] The model training platform provided in this embodiment may include a sample library and a training module connected to the sample library, wherein a first preset model may be deployed in the training module, and the training module may extract multiple image sample pairs from the sample library in response to a training operation, and input the first image sample in the image sample pair into the first preset model, thereby starting training the first preset model.

[0302] As described in the above embodiment, during the training of the first preset model, there may be at least one of the following training methods: obtaining text data corresponding to the predicted image currently output by the first preset model; and updating the parameters of the first preset model based on the text data, the predicted image and the second image sample; wherein the text data includes data for evaluating the image quality of the predicted image.

[0303] Accordingly, if the text data is data generated by positively evaluating the predicted image, and the text data is consistent with the image quality of the predicted image, then a target text having the opposite semantics to the text data can be determined, thereby determining the similarity between the target text and the predicted image. Then, the loss value between the predicted image and the second image sample is determined. When determining the loss, it can be performed according to the above formula (1) or formula (2), thereby updating the parameters of the first preset model according to the loss value and the similarity.

[0304] Among them, the multiple image sample pairs included in the sample library can be uploaded by the user in advance, and the process of obtaining the image sample pairs can refer to the description in the above-mentioned model training method embodiment, which will not be repeated here.

[0305] By adopting this training platform, the training module can train the first preset model with multiple image samples obtained from the sample library as training samples. During the training process, the text data can be used to determine the area to be optimized in the predicted image that is different from the expected image quality. Therefore, when updating the first preset model, the area to be optimized can be fed back to the first preset model to supervise the first preset model to perform reinforcement learning on the repair of the area to be optimized, thereby guiding the first preset model to continuously optimize the weak links in image repair and help improve the model's image repair quality.

[0306] As described in the above embodiment, the training of the first preset model can be carried out in stages, including a first training stage, a second training stage and a third training stage, wherein the first training stage uses part of the image samples to train the first preset model. During this training process, the parameters of the first preset model are updated based on the predicted image and the second image sample output by the first preset model; during the second training process, the parameters of the first preset model are updated based on the predicted image, the second image sample and the text data corresponding to the predicted image output by the first preset model; in the third training process, the second preset model is trained based on the third image sample and the text data sample.

[0307] For example, in order to enable intuitive visualization of the training stages and control switching between the training stages during the staged training of the first preset model, the training platform may further provide an operation interface, wherein the operation interface may include a control area, and the control area may include multiple controls; wherein different controls are used to start different training when triggered; wherein:

[0308] The training includes a first training and a second training, wherein the first training includes: updating parameters of the first preset model based on the predicted image and the second image sample output by the first preset model;

[0309] The second training includes updating parameters of the first preset model obtained by the first training based on the text data, the predicted image and the second image sample.

[0310] In this embodiment, the training module can start the corresponding training phase in response to the triggering of the corresponding control in the operation interface. For example, in response to the triggering of the first control in the operation interface, multiple image sample pairs can be obtained from the sample library, and the first training of the first preset model can be started using the multiple image sample pairs obtained as training samples. Initiating the first training can mean that the training module starts the first thread corresponding to the first training. When the first thread is running, the first image sample in the obtained image sample pair can be input into the first preset model, and the predicted image and the second image sample output by the first preset model can be used to calculate the loss function. Then, according to the value of the loss function, the parameters of the first preset model are updated.

[0311] For another example, in response to triggering a second control in an operation interface, a second thread corresponding to a second training can be started, and the first thread can be kept running; wherein, when the second thread is running, the first image sample in the acquired image sample pair can be input into the first preset model, and the predicted image output by the first preset model can be displayed, and the text data input for the predicted image can be monitored; then, it is determined whether the text data is data generated by positively evaluating the predicted image; if so, the text data is converted into target text (data for negatively evaluating the predicted image), and the similarity between the target text and the predicted image is calculated; if not, the similarity between the text data and the predicted image is calculated;

[0312] Since the first thread keeps running, the first thread will calculate the loss value between the predicted image and the second image sample. In this case, the second thread will feed back the similarity to the first thread. When the first thread receives the similarity, it will start to update the parameters of the first preset model according to the similarity and the loss value.

[0313] Of course, as described above, the control may also include a control for triggering a third training, which includes: inputting multiple third image samples and corresponding at least two text data samples into a second preset model to train the second preset model; wherein the second preset model when the third training is completed is used to determine the similarity between the predicted image and the text data in the training of the first preset model.

[0314] Among them, the control that triggers the third training can be called the third control. When the third control is triggered, the third thread corresponding to the third training can be started. When the third thread is started, the third image sample and the corresponding text data sample can be obtained from the sample library, and the third image sample and the text data sample can be input into the second preset model. The loss value is determined based on the predicted similarity output by the second preset model and the category label corresponding to the text data sample, and the parameters of the second preset model are updated based on the loss value.

[0315] As described above, the step of performing the third training on the second preset model can be performed during the process of training the first preset model; wherein the multiple third image samples include at least one of the following: the first image sample, the second image sample, and the predicted image output by the first preset model before the current moment.

[0316] Among them, since the third image sample can include the predicted image output during the first training process, in one implementation method, the first thread can save the output prediction images to the sample library and associate them with the first image sample; before the third training starts, in response to the annotation of the third image sample, the annotated fourth thread can obtain the predicted image from the sample library and display it, and then, the text data sample input for the displayed predicted image can be obtained, and the text data sample can be associated with the predicted image and saved in the sample library.

[0317] In another implementation, the training module can respond to the data labeling function turned on in the first training, instruct the first thread to display the predicted image of the first preset model on the display interface after obtaining the predicted image. Then, the fourth thread can be started in response to the input operation triggered on the display interface, and obtain the text data sample input for the displayed predicted image, and associate the text data sample with the predicted image and save them together in the sample library; and, at the same time, update the parameters of the first preset model according to the loss value between the predicted image and the second image sample.

[0318] Regardless of which of the above specific implementation methods is used, when the third thread is started, the predicted image and the corresponding text data sample can be input from the sample library into the second preset model for training.

[0319] As described above, the text data sample needs to be associated with the predicted image so that the text data sample can be found based on this association during the second training and used to update the parameters of the first preset model. Accordingly, the training platform may further include an association unit and an output unit. The association unit may be used to associate the text data sample input in the input area with the predicted image when displaying the predicted image;

[0320] The output unit can be used to output the text data sample associated with the predicted image currently output by the first preset model as text data when performing the second training on the first preset model.

[0321] Specifically, when the second thread is running, the output unit can monitor the predicted image output by the first preset model, search for text data samples associated with the predicted image from the sample library, and send the text data samples to the second thread, so that the second thread can input the text data samples output by the output unit into the second preset model.

[0322] For example, in order to enable intuitive visualization of the training stages during the phased training of the first preset model, the operation interface may further include a display area and an input area; wherein:

[0323] The display area is used to display at least one of: a predicted image output by the first preset model, a third image sample, and a loss value during the third training process; wherein the loss value is determined by the predicted similarity output by the second preset model and a category label; the category label is used to indicate whether the description of the text data sample meets the image quality of the third image sample;

[0324] The input area is used for users to input text data and text data samples.

[0325] Among them, in both the second training and the first training, the first thread can output the predicted image to the display interface for display. For example, the first thread will send the predicted image to the front-end interface rendering thread of the training platform, and the front-end interface rendering thread will render the predicted image to the display area; correspondingly, the first thread can execute the operation of outputting the predicted image to the front-end interface rendering thread regardless of the first training and the second training, thereby helping the user to intuitively observe the optimization degree of the first preset model in the first training and the second training.

[0326] Among them, in the third training, the third thread can send the calculated loss between the predicted similarity and the category label to the front-end interface rendering thread of the training platform, and the front-end interface rendering thread renders the predicted image into the display area.

[0327] In one example, the front-end interface rendering thread can record the received loss values, and in response to the triggering operation of the loss value change curve on the front-end interface (display interface), it can generate a gradient change curve based on multiple loss values ​​and the moments corresponding to the multiple loss values, and render the gradient change curve to the display area. In this way, the user can determine whether the second preset model converges through the gradient change curve.

[0328] In another example, the front-end interface rendering thread can display the received loss value in real time, for example, near the display area, so that the user can determine the degree of optimization of the second preset model in the third training through the real-time displayed loss value in the third training.

[0329] In some examples, the third image sample used for the third training can correspond to two of the text data samples, the two text data samples including a text data sample of a first category and a text data sample of a second category, the first category characterizing that the description of the text data sample is consistent with the image quality of the third image sample, and the second category characterizing that the description of the text data sample is inconsistent with the image quality of the third image sample.

[0330] Accordingly, in this embodiment, in response to the first text data input in the input area for the displayed third image sample, the first text data may be sent to the fifth thread, and the fifth thread may perform the following steps:

[0331] determining a category corresponding to the first text data, the category being used to indicate whether a description of the first text data is consistent with the image quality of the third image sample;

[0332] generating second text data based on the first text data; wherein the category of the second text data is different from the category of the first text data;

[0333] Using the first text data and the second text data as text data samples corresponding to the third image sample;

[0334] The fifth thread may bind the first text data and the second text data with the third image sample and store them in a sample library in an associated manner.

[0335] Accordingly, in the third training, the third thread can obtain a third image sample, as well as the first text data and the second text data corresponding to the third image sample from the sample library, and input the third image sample, the first text data and the second text data into the second preset model to update the parameters of the second preset model according to the predicted similarities and category labels corresponding to the two text data samples.

[0336] Accordingly, the second preset model may include an image encoder, a text encoder and a similarity determination module, wherein the third thread may input the third image sample into the image encoder of the second preset model, and input the text data sample corresponding to the third image sample into the text encoder of the second preset model, and the similarity determination module determines the predicted similarity between the text data sample and the third sample image. Then, the third thread may update the parameters of the second preset model based on this predicted similarity and category label.

[0337] The following describes the process of executing model training on the model training platform of the present disclosure with reference to a specific example. Referring to FIG. 14 a to FIG. 14 e , schematic diagrams of changes in the operation interface of the model training platform during the execution of model training are shown in sequence:

[0338] S21: As shown in FIG14a , first, in response to the user creating a model training task, a first operation interface is output. The first operation interface provides the settings required to create a training task, including but not limited to: task name, training dataset, network structure of the first preset model, network structure of the second preset model, and loss calculation options. Once the settings are complete, clicking the "Start Training 1" control (first control) in the first operation interface initiates the first training session.

[0339] S22: As shown in Figure 14b, when the "Start Training 1" control in the first operation interface is triggered, the operation platform outputs a second operation interface. In the second operation interface, multiple image sample pairs in the sample library for the first training can be selected. A list of first image samples in the selected image sample pairs can be displayed on the second operation interface, such as Image 1, Image 2-Image 5. The first image sample currently input to the first preset model for training can be highlighted in the list, such as bold black represents the first image sample currently input to the first preset model for training.

[0340] In response to the triggering of the first control, the first thread starts running, and the multiple first image samples displayed in the list can be automatically input into the first preset model in sequence. The first thread can calculate the loss value based on the predicted image and the second image sample output by the first preset model, and then update the parameters of the first preset model according to the loss value;

[0341] At the same time, during the first training, the second operation interface includes a display area and an input area. The display area can display the first image sample currently input to the first preset model for training, as well as the predicted image output after being repaired. In this example, the input area in the second operation interface can be locked to not support text annotation of the predicted image, that is, during the first training, the predicted image can be not annotated with text data samples.

[0342] When the user believes that the first preset model has been trained to a valid state based on the displayed prediction image, the user can click the Start Training 2 button (the third control) below to start the third training.

[0343] Among them, in the second operation interface, the predicted image and the first image sample can be allowed to be displayed in the display area; for example, clicking the first image sample in the display area will switch to displaying the predicted image, and clicking the predicted image will switch to displaying the first image sample.

[0344] In the first training process, the predicted image output by the first preset model can be saved in a sample library so that the predicted image can be labeled as a third image sample later.

[0345] S23: As shown in FIG14c, when the control (third control) of “Start Training 2” in the second operation interface is triggered, the third operation interface is output and the third training begins.

[0346] During the third training, a list of third image samples may be displayed on the third operation interface. Specifically, in response to the third training, third image samples accurate for the third training may be obtained from the sample library, as well as a preset number of predicted images at the end of the first training. These images may be displayed as third image samples in the list of the third operation interface, such as image 6, image 7-image 10. The third image sample currently to be labeled may be highlighted in the list, such as bold black representing the third image sample currently to be labeled. The third operation interface still includes a display area, in which the selected third image sample may be displayed. The third operation interface still includes an input area. At this time, during the third training stage, the input area is unlocked to allow text input.

[0347] The user can enter a textual expression in the input area, and if a category label is not set for the textual expression entered, the third image sample displayed in the display area is bound to the text entered in the input area. If a third image sample is selected and the user observer considers that there is no problem, and considers the Tisza image sample to be an image sample of higher image quality, then there is no need to enter text in the input area. The training platform can directly use the first preset text data as the first category text data sample of the third image sample, and / or use the second preset text data as the second category text data sample of the third image sample;

[0348] Then, you can directly click on the next third image sample to mark it. When all the third image samples are marked, click Finish Marking. Only then will the third training be truly started.

[0349] At this time, the third thread starts running, inputs the third image sample and the corresponding text data sample into the second preset model, and determines the loss value based on the predicted similarity output by the second preset model and the category label carried by the text data sample, and updates the parameters of the second preset model based on the loss value.

[0350] S24: As shown in FIG14c , in the third operation interface, the loss value obtained in the third training can also be displayed in real time, such as “Loss value dynamic display during training: 0.00012” in FIG14c ; wherein 0.00012 is the loss value in the third training;

[0351] Among them, a fourth control may also be provided in the third operation interface. The fourth control may be, for example, the "View Loss Curve" control in FIG14c. When the control is triggered, the respective loss values ​​corresponding to the second preset model during the third training process may be displayed in the third operation interface. As shown in FIG11 , the current existing loss value record of the third training may be dynamically displayed.

[0352] Among them, the user can judge whether it is necessary to end the third training based on the displayed loss value. If necessary, the user can click the "Start Training 3" control (the second control), so that the second training can be started in response to the triggering of the second control.

[0353] S4: As shown in FIG14d , when the “Start Training 3” control in the third operation interface is triggered, the fourth operation interface is output and the second training begins.

[0354] Among them, the second training is a mode of training while annotating text, and the fourth operation interface still includes a list, which includes multiple first image samples as training samples, such as image n1-image n5, among which the first image samples corresponding to the predicted image used for the third training can be arranged at the front, that is, the first image samples arranged at the front of the list are a preset number of first image samples, and the predicted images corresponding to these first image samples are bound to text data samples in the sample library.

[0355] For the first preset number of first image samples, in response to the selection operation of the image identifier displayed in the list, the selected first image sample can be input into the first preset model, and the predicted image output by the first preset model for the first image sample will be displayed in the display area of ​​the third operation interface, and the text data sample corresponding to the predicted image obtained from the sample library will be displayed in the input area. Then, the operation in the display area can be monitored. If the user clicks on the input area and updates the text data, the newly input text data of the user will be sent to the second thread, and the second thread will calculate the similarity based on the input text data and the predicted image, and feed it back to the first thread; if the user does not click on the input area and update the text data, the text data sample will be sent to the second thread, and the second thread will calculate the similarity based on the input text data sample and the predicted image, and feed it back to the first thread;

[0356] For the subsequent first image sample, in response to the selection operation of the image identifier displayed in the list, the selected first image sample can be input into the first preset model, and the predicted image output by the first preset model for the first image sample will be displayed in the display area of ​​the third operation interface, and the operation in the display area can be monitored to obtain the text data entered by the user in the input area. The text data is sent to the second thread, and the second thread calculates the similarity based on the input text data and the predicted image, and feeds it back to the first thread.

[0357] In the second training, the first thread updates the parameters of the first preset model based on the similarity sent by the second thread and the loss value between the predicted image and the second image sample.

[0358] In the first training, the first thread updates the parameters of the first preset model based on the loss value between the predicted image and the second image sample.

[0359] When the third thread is executing the third training, the first thread and the second thread are not running to fix the parameters of the first preset model; when the first thread and the second thread are running, the third thread is not running to fix the parameters of the second preset model.

[0360] In summary, the model training platform provided by the present invention can help users train the first preset model in stages through the operation interface, training module, and the display area, input area and multiple controls set on the operation interface. During the training, the optimization degree of the image restoration model in the training process can be observed through the predicted image displayed in the display area. By inputting the evaluation text of the predicted image in the input area, the first preset model can be guided to update the parameters in the optimization direction indicated by the evaluation text during the training of the first preset model, thereby realizing the supervision of model training from the human perspective in the image restoration scenario, making new attempts and explorations for model training in the field of image restoration, thereby optimizing the model training mechanism and contributing to improving the image quality of the first preset model.

[0361] Based on the same inventive concept, the present disclosure further provides an image restoration method. As shown in FIG15 , a schematic flow chart of the steps of the image restoration method is shown. As shown in FIG15 , the method may specifically include the following steps:

[0362] Step S1501: Acquire a target image to be repaired;

[0363] Step S1502: inputting the target image into an image restoration model; wherein the image restoration model is a first preset model trained according to the model training method;

[0364] Step S1503: Obtaining a repaired image output by the image repair model; wherein the image quality of the repaired image is higher than that of the target image.

[0365] In this embodiment, the target image may refer to an image with poor image quality, such as an old photo, scratched film data, etc.

[0366] The image restoration model is the first preset model described in the above embodiment. Its training process can refer to the description of the above embodiment and is not repeated here. The restored image output by the image restoration model can be obtained. Since the image restoration model is used to restore the image, the image quality of the restored image is higher than that of the target image.

[0367] For example, if the image restoration task of the image restoration model is clarity restoration, the clarity of the restored image is higher than that of the target image; for another example, if the image restoration task of the image restoration model is scratch restoration, the scratches in the restored image are fewer than those in the target image; for another example, if the image restoration task of the image restoration model is missing part restoration, the image details of the missing image area in the target image are restored; for another example, if the image restoration task of the image restoration model is noise restoration, the noise in the restored image is less than that in the target image.

[0368] The image restoration method of this embodiment uses an image restoration model trained according to the embodiment of the above-mentioned model training method. During training, the parameters of the first preset model can be updated using text data, predicted images, and second image samples. The text data also includes data for evaluating the image quality of the predicted image. Therefore, not only can supervised training be based on the pixel difference between the predicted image and the second image sample output by the model, but also supervised training can be based on the difference of visual evaluation of the predicted image by human vision. This enables the model to acquire the ability to perceive human vision, and thus can guide the optimization direction of the first preset model based on facial vision, which in turn helps to improve the image restoration quality of the trained model, thereby improving the image restoration quality of the target image.

[0369] Based on the same inventive concept, the present disclosure further provides an image restoration device, as shown in FIG16 , which shows a schematic structural diagram of an image restoration device. As shown in FIG16 , the image restoration device may specifically include the following modules:

[0370] A first acquisition module is used for the target image to be restored;

[0371] An input module, configured to input the target image into an image restoration model; wherein the image restoration model is a first preset model trained according to the model training method;

[0372] The second acquisition module is used to acquire the repaired image output by the image repair model; wherein the image quality of the repaired image is higher than that of the target image.

[0373] The description of the device embodiment can be found in the description of the above method embodiment, and will not be repeated here.

[0374] The embodiments of the present disclosure also provide a computer-readable storage medium, which stores a computer program that enables a processor to execute the model training method or image restoration method as described in the embodiments of the present disclosure.

[0375] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0376] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity, or device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, commodity, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, commodity, or device that includes the element.

[0377] The above is a detailed introduction to a model training method and platform, image restoration method, device, equipment and medium provided by the present disclosure. Specific examples are used in this article to illustrate the principles and implementation methods of the present disclosure. The description of the above embodiments is only used to help understand the method of the present disclosure and its core ideas. At the same time, for those skilled in the art, according to the ideas of the present disclosure, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as a limitation on the present disclosure.

[0378] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the following claims.

[0379] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

[0380] References herein to "one embodiment," "an embodiment," or "one or more embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present disclosure. Furthermore, please note that instances of the phrase "in one embodiment" do not necessarily all refer to the same embodiment.

[0381] In the description provided herein, numerous specific details are described. However, it is understood that embodiments of the present disclosure may be practiced without these specific details. In some instances, well-known methods, structures, and techniques are not shown in detail so as not to obscure the understanding of this description.

[0382] In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in the claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present disclosure may be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In a unit claim enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third etc. does not indicate any order. These words may be interpreted as names.

[0383] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present disclosure, rather than to limit them. Although the present disclosure has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present disclosure.

Claims

1. A model training method, wherein: The method comprises: Acquire a plurality of image sample pairs, wherein the image sample pairs include a first image sample and a second image sample of the same image; wherein the image quality of the second image sample is higher than the image quality of the first image sample; The first preset model is trained using the plurality of image sample pairs as training samples, wherein the first preset model is used to improve the image quality of the first image samples, and the training process includes: Acquire text data corresponding to the predicted image currently output by the first preset model; wherein the text data includes data for evaluating the image quality of the predicted image; Based on the text data, the predicted image and the second image sample, the parameters of the first preset model are updated.

2. The method according to claim 1, wherein: The updating of the parameters of the first preset model based on the text data, the predicted image and the second image sample includes: Determining the similarity between the target text and the predicted image; wherein the target text is the text data, or a text with opposite semantics to the text data; Determining a loss value based on the predicted image and the second image sample; Based on the similarity and the loss value, the parameters of the first preset model are updated.

3. The method according to claim 2, wherein: The determining the similarity between the target text and the predicted image comprises: Encoding the target text to obtain a text feature vector of the target text; Encoding the predicted image to obtain an image feature vector of the predicted image; wherein the dimensions of the text feature vector and the image feature vector are consistent; The similarity is determined based on the text feature vector and the image feature vector.

4. The method according to claim 1, wherein: The step of training the first preset model using the plurality of image sample pairs as training samples comprises: Using some of the image sample pairs as training samples, a first training is performed on the first preset model; wherein, in the first training, based on the predicted image output by the first preset model and the second image sample, the parameters of the first preset model are updated; Using part or all of the image sample pairs as training samples, performing second training on the first preset model obtained by the first training; Wherein, in the second training, based on the text data, the predicted image and the second image sample, the parameters of the first preset model obtained by the first training are updated.

5. The method according to any one of claims 1 to 4, wherein: The method further comprises: Acquire a plurality of third image samples and text data samples corresponding to the third image samples; wherein the text data samples are used to describe the image quality of the third image samples; Based on the plurality of the third image samples and the text data samples, performing a third training on the second preset model; wherein the second preset model is used to determine the similarity between the third image samples and the text data samples; The updating of the parameters of the first preset model based on the text data, the predicted image and the second image sample includes: Inputting the predicted image and the text data into the second preset model after the third training is completed; Based on the similarity output by the second preset model, the predicted image and the second image sample, the parameters of the first preset model are updated.

6. The method according to claim 5, wherein: The text data sample carries a category label, and the category label is used to indicate whether the description of the text data sample conforms to the image quality of the third image sample; and the third training of the second preset model based on the plurality of the third image samples and the corresponding at least two text data samples includes: Inputting the third image sample and the text data sample into the second preset model to obtain the predicted similarity between the third image sample and each of the text data samples; Based on the predicted similarity and the category label, the parameters of the second preset model are updated.

7. The method according to claim 5, wherein: The third image sample corresponds to a text data sample of a first category and a text data sample of a second category, the first category characterizing that the description of the text data sample meets the image quality of the third image sample, and the second category characterizing that the description of the text data sample does not meet the image quality of the third image sample.

8. The method according to claim 5, wherein: The second preset model includes a text encoder and an image encoder, and a similarity determination module connected to the text encoder and the image encoder respectively; wherein: The text encoder is used to perform text encoding on the text data sample to obtain a predicted text vector; The image encoder is used to perform image encoding on the third image sample to obtain a predicted image vector; wherein the predicted image vector and the predicted text vector have the same dimension; The similarity determination module is used to determine the predicted similarity between the predicted image vector and the predicted text vector.

9. The method according to claim 5, wherein: During the process of training the first preset model, the step of performing the third training on the second preset model is performed; the plurality of third image samples include at least one of the following: the first image sample, the second image sample, and the predicted image output by the first preset model before the current moment.

10. The method according to any one of claims 5 to 9, wherein: The third training is performed in the interval of training the first preset model, the third image sample includes the predicted image output by the first preset model for the first image sample, and the third training of the second preset model based on the plurality of the third image samples and the text data samples includes: Inputting a plurality of the first image samples into the first preset model; Inputting the predicted image output by the first preset model and the text data sample corresponding to the predicted image into the second preset model to perform the third training on the second preset model; Wherein, in the third training, the parameters of the first preset model are fixed.

11. The method according to any one of claims 5 to 9, wherein: After the third training, the training of the first preset model includes: Inputting a plurality of the first image samples into the first preset model; The predicted image output by the first preset model and the text data are input into the second Preset models; Based on the predicted image, the second image sample and the predicted similarity output by the second preset model, updating the parameters of the first preset model; When the parameters of the first preset model are updated, the parameters of the second preset model are fixed.

12. The method according to claim 5, wherein: The acquiring of a plurality of third image samples and text data samples corresponding to the third image samples comprises: Acquire first text data corresponding to the third image sample; determining a category corresponding to the first text data, the category being used to indicate whether a description of the first text data is consistent with an image quality of the third image sample; Based on the first text data, generating second text data; wherein the category of the second text data is different from the category of the first text data; The first text data and the second text data are used as text data samples corresponding to the third image sample.

13. The method according to claim 5, wherein: The acquiring text data corresponding to the predicted image currently output by the first preset model includes: Determine whether there is a target third image sample corresponding to the predicted image from the plurality of third image samples; wherein the target third image sample and the predicted image correspond to the same first image sample; If yes, taking the text data sample corresponding to the target third image sample as the text data; If not, the text data input for the predicted image is obtained.

14. The method according to claim 1, wherein: The acquiring text data corresponding to the predicted image currently output by the first preset model includes: displaying the predicted image; The text data input for the prediction image is obtained.

15. The method according to claim 14, wherein: After displaying the predicted image, the method further includes: monitoring an input operation on an operation interface displaying the predicted image; If the input operation is not monitored, using the preset text data as the text data corresponding to the predicted image; The step of obtaining text data input for the predicted image includes: If the input operation is monitored, the text data input for the predicted image is obtained.

16. The method according to claim 1, wherein: The text data includes at least one entry; the at least one entry is used to describe the image quality of the predicted image in different image regions and / or the image quality in different quality dimensions.

17. The method according to any one of claims 4 or 14-16, wherein: The method further comprises at least one of the following: In response to the first training, displaying a predicted image output by the first preset model, and in response to a first preset operation performed on the predicted image, ending the first training; In response to the second training, the predicted image output by the first preset model is displayed, and in response to a second preset operation performed on the predicted image, the second training is ended.

18. The method according to claim 6, wherein: In the process of performing the third training on the second preset model, the method further includes: Displaying the loss value corresponding to at least one training of the second preset model before the current moment; wherein the gradient is determined by the predicted similarity and the category label; In response to a third preset operation being performed on each of the displayed loss values, the third training is ended.

19. An image restoration method, wherein: The method comprises: The target image to be repaired; Inputting the target image into an image restoration model; wherein the image restoration model is a first preset model trained according to the model training method according to any one of claims 1 to 18; Obtain a repaired image output by the image repair model; wherein the image quality of the repaired image is higher than that of the target image.

20. A model training platform, wherein: The training platform comprises: A sample library is used to store multiple image sample pairs, wherein the image sample pairs include the first an image sample and a second image sample; wherein the image quality of the second image sample is higher than the image quality of the first image sample; A training module is used to train a first preset model multiple times using the plurality of image sample pairs as training samples, wherein the first preset model is used to improve the image quality of the first image samples; wherein the training process includes: Acquire text data corresponding to the predicted image currently output by the first preset model; wherein the text data includes data for evaluating the image quality of the predicted image; Based on the text data, the predicted image and the second image sample, the parameters of the first preset model are updated.

21. An image restoration device, wherein: The device comprises: A first acquisition module, used for the target image to be repaired; An input module, used for inputting the target image into an image restoration model; wherein the image restoration model is a first preset model trained by the model training method according to any one of claims 1 to 18; The second acquisition module is used to acquire the repaired image output by the image repair model; wherein the image quality of the repaired image is higher than that of the target image.

22. An electronic device, wherein: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor is executed, the model training method described in any one of claims 1 to 18 or the image restoration method described in claim 19 is implemented.

23. A computer-readable storage medium, wherein: The computer program stored therein enables the processor to execute the model training method as described in any one of claims 1 to 18, or the image restoration method as described in claim 19.

Citation Information

Patent Citations

  • Text-guided image restoration method and system

    CN111861945A

  • Training method of image enhancement model, image enhancement method and electronic equipment

    CN112801918A

  • Image text information generation method and deep learning model training method

    CN115359323A

  • Transitory salient attention capture to draw attention to digital document parts

    US20220284071A1