Fine-tuning training method, device, medium and electronic device for image processing model
By fine-tuning training of the model to be adjusted in the image processing model to adapt it to the output image of the previous model, the problem of inability to adapt to the single-task AI image processing model when deploying in series is solved, and the overall effect of image processing is improved.
Patent Information
- Application Number
- CN202210466553.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-29
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-04-29
AI Technical Summary
When the single-task AI image processing model is deployed in series, it cannot adapt to the output image of the previous model, resulting in defects in the image processing results.
By acquiring multiple pretrained models in the image processing model, fine-tuning training of the adjustment model based on the tandem order of the preset sample image and the pretrained model, a fine-tuning model with strong adaptability is generated to avoid image processing defects caused by mismatch in the tandem order.
It effectively avoids the defects in image processing results caused by mismatch in series order, ensures the output quality of the image processing model, and improves the overall effect of image processing.
Smart Images

Figure CN114781531B_ABST
Abstract
Description
Background Art
[0002] With the continuous development of artificial intelligence (AI) technology in the field of image processing, more and more AI image processing technologies are applied to the processing flow of image signal processors (ISPs) such as image denoising, image deblurring, and image super-resolution reconstruction to replace traditional processing technologies. In the academic field, attempts have even been made to replace the entire ISP processing flow with a single AI model to achieve true end-to-end, that is, from the raw data collected by the sensor to the final output image, which is realized by a single AI model. However, it is difficult to achieve in actual deployment.
[0003] In related technologies, a single-task AI model is usually used to replace a single processing node in the ISP processing flow, and a single processing task is completed through the single-task AI model. For example, using AI denoising to replace a traditional denoising node in the ISP process, using AI deblurring to replace a traditional deblurring node in the ISP process, and so on. Summary of the Invention
[0004] The purpose of the present disclosure is to provide a fine-tuning training method for an image processing model, a fine-tuning training device for an image processing model, a computer-readable medium, and an electronic device, thereby at least to a certain extent avoiding the problem that the single-task AI image processing model with a later serial order cannot adapt to the output image of the single-task AI image processing model with an earlier serial order, resulting in defective image processing results.
[0005] According to a first aspect of the present disclosure, there is provided a fine-tuning training method for an image processing model, including: obtaining an image processing model, the image processing model including a plurality of pre-trained models connected in series; based on a preset sample image, a pre-trained model, and the serial order of the pre-trained model in the image processing model, performing fine-tuning training on the model to be adjusted in the image processing model to obtain a fine-tuned model corresponding to the model to be adjusted; wherein, the model to be adjusted includes at least one pre-trained model, and the model to be adjusted at least includes the pre-trained model at the end of the serial order in the image processing model; generating a trained image processing model based on the fine-tuned model.
[0006] According to a second aspect of the present disclosure, there is provided a fine-tuning training device for an image processing model, including: a model acquisition module configured to acquire an image processing model, where the image processing model includes a plurality of pre-trained models connected in series; a fine-tuning training module configured to perform fine-tuning training on a model to be adjusted in the image processing model based on a preset sample image, the pre-trained models, and the connection order of the pre-trained models in the image processing model, so as to obtain a fine-tuned model corresponding to the model to be adjusted; wherein the model to be adjusted includes at least one pre-trained model, and the model to be adjusted at least includes the pre-trained model at the end of the connection order in the image processing model; and a model generation module configured to generate a trained image processing model based on the fine-tuned model.
[0007] According to a third aspect of the present disclosure, there is provided a computer-readable medium having a computer program stored thereon, and when the computer program is executed by a processor, the above method is implemented.
[0008] According to a fourth aspect of the present disclosure, there is provided an electronic device, characterized by including: a processor; and a memory configured to store one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the above method.
[0009] The image processing model method provided by an embodiment of the present disclosure takes into account the connection order of the pre-trained models in the image processing model, and based on this, performs fine-tuning training on at least the pre-trained model at the end of the connection order in the image processing model based on a preset sample image and the pre-trained models, so that the pre-trained model at the end of the connection order can be fine-tuned to a state adapted to the output image of the pre-trained model with a previous connection order, thereby avoiding the problem of defective image processing results caused by non-adaptation.
[0010] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts. In the drawings:
[0012] Figure 1 A schematic diagram showing an image processing process to which the embodiments of the present disclosure can be applied is shown;
[0013] Figure 2 A schematic diagram schematically showing an image processing model processing process in an exemplary embodiment of the present disclosure is shown;
[0014] Figure 3 Schematically shows a flowchart of a method for fine-tuning and training an image processing model in an exemplary embodiment of the present disclosure;
[0015] Figure 4 Schematically shows a flowchart of a method for obtaining an image processing model in an exemplary embodiment of the present disclosure;
[0016] Figure 5 Schematically shows a flowchart of a fine-tuning and training method in an exemplary embodiment of the present disclosure;
[0017] Figure 6 Schematically shows a flowchart of another fine-tuning and training method in an exemplary embodiment of the present disclosure;
[0018] Figure 7 Schematically shows a flowchart of a method for generating a trained image processing model in an exemplary embodiment of the present disclosure;
[0019] Figure 8 Schematically shows a schematic diagram of a method for generating a fine-tuning model L in an exemplary embodiment of the present disclosure;
[0020] Figure 9 Schematically shows a schematic diagram of another method for generating a fine-tuning model L in an exemplary embodiment of the present disclosure;
[0021] Figure 10 Schematically shows a schematic diagram of yet another method for generating a fine-tuning model L in an exemplary embodiment of the present disclosure;
[0022] Figure 11 Schematically shows a schematic diagram of still another method for generating a fine-tuning model L in an exemplary embodiment of the present disclosure;
[0023] Figure 12 Schematically shows a schematic diagram of a method for generating a fine-tuning sample image in an exemplary embodiment of the present disclosure;
[0024] Figure 13 Schematically shows a comparison diagram of the processing effects of an image processing model before and after fine-tuning training in an exemplary embodiment of the present disclosure;
[0025] Figure 14 Schematically shows a schematic diagram of the composition of a fine-tuning training device for an image processing model in an exemplary embodiment of the present disclosure;
[0026] Figure 15 Shows a schematic diagram of an electronic device to which the embodiments of the present disclosure can be applied. Detailed implementation manners
[0027] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the concept of example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0028] In addition, the accompanying drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0029] In an exemplary embodiment, the ISP processing flow is generally a cascaded flow. For example, an image needs to be denoised first, and then deblurred based on the denoised image, and then other processing is performed based on the deblurred image. Correspondingly, when replacing the processing nodes of the ISP processing flow with a single-task AI image processing model, it is very likely that multiple consecutive processing nodes before and after are replaced by the single-task AI image processing model, and correspondingly, there will be a situation where multiple single-task AI image processing models are deployed in series. For example, if two consecutive processing nodes of the ISP processing flow are deblurring processing and super-resolution reconstruction processing respectively, then there will be a situation where the AI image deblurring model and the AI image super-resolution reconstruction model are connected in series.
[0030] However, in practical applications, the single-task AI image processing models trained separately are usually cascaded directly to obtain the result of serial deployment. When the developer is training, it is already clear which processing node of the ISP process the single-task AI image processing model needs to replace. Therefore, when making the training data set, the principle that the training samples are consistent with the input images of the application scenario is generally followed as much as possible. Therefore, after cascading, after the single-task AI image processing model with a prior serial order processes the input image of the application scenario, it is very likely that the single-task AI image processing model with a posterior serial order cannot adapt to the output data of the single-task AI image processing model with a prior serial order, and thus it is very likely that the image processing result is defective.
[0031] For example, as Figure 1As shown in the figure, after inputting Image 1 into the trained Model A, the data distribution of the obtained Image 2 (where different data distributions of images are represented by different grayscales) is different from that of Image 1; after inputting Image 1 into the trained Model B, the data distribution of the obtained Image 3 is also different from that of Image 1. At this time, no matter how Model A and Model B are cascaded, it may lead to the situation where the data distribution corresponding to the input image of the model with a later cascading order is different from that of Image 1. Therefore, it is very likely that there are defects in the image processing results. For example, Figure 2 As shown in the figure, if the models are cascaded in the order of Model A and Model B, when actually performing image processing, inputting Image 1 into Model A to obtain Image 2. Since the data distribution of Image 2 is different from that of Image 1, when Model B processes the Image 2 with a different data distribution, the output Image 4 is very likely to have defects, that is, the domain gap problem caused by Model B's inability to adapt to the data distribution of the output image of Model A.
[0032] Based on the above one or more problems, the exemplary embodiment of the present example provides a fine-tuning training method for an image processing model. The fine-tuning training method of the image processing model can be applied to a terminal device with image processing functions, including but not limited to desktop computers, portable computers, smart phones, tablet computers, and so on. No special limitation is made in this exemplary embodiment. Refer to Figure 3 As shown in the figure, the fine-tuning training method of the image processing model may include the following steps S310 to S330:
[0033] In step S310, an image processing model is obtained.
[0034] Among them, the image processing model includes a plurality of pre-trained models connected in series; each pre-trained model is an image processing model that has been pre-trained; the pre-trained image processing model includes a single AI model for implementing any one or more image processing tasks. It should be noted that the single AI model for implementing any one or more image processing tasks means that it can be an AI model that can implement one image processing task based on a single model, or an AI model that can implement multiple processing tasks based on a single model. For example, a single AI model for single denoising processing of images; another example is a single AI model that can simultaneously perform denoising and super-resolution reconstruction on images.
[0035] In an exemplary embodiment, when obtaining the image processing model, referring to Figure 4 As shown in the figure, it may include the following steps S410 to S430:
[0036] In step S410, a plurality of models to be trained and the cascading order corresponding to the models to be trained are obtained.
[0037] Among them, each of the multiple models to be trained corresponds to each of the multiple pre-trained models respectively. After each model to be trained is trained, the image processing task to be achieved by the corresponding pre-trained model can be realized. It should be noted that the obtained serial order corresponding to the model to be trained is the same as the serial order of the corresponding pre-trained model in the image processing model, so as to ensure the execution order of each image processing task in the image processing model.
[0038] In step S420, pre-training is performed for each model to be trained respectively to obtain the pre-trained models corresponding to the respective models to be trained.
[0039] In an exemplary embodiment, after obtaining multiple models to be trained, pre-training can be performed for each model to be trained respectively to obtain the pre-trained models corresponding to each model to be trained. Among them, when performing pre-training for each model to be trained respectively, the same samples can be used for pre-training, or different samples can be used for pre-training, as long as it matches the application scenario of the image processing model. The present disclosure does not make special limitations in this regard. For example, for the model to be trained A and the model to be trained B, the same pre-training sample image 1 can be used for training, or the pre-training sample image 1 can be used to train the model to be trained A, and the pre-training sample image 2 can be used to train the model to be trained B.
[0040] In step S430, the pre-trained models are connected in series according to the serial order corresponding to the models to be trained to obtain an image processing model.
[0041] In an exemplary embodiment, after pre-training is performed for each model to be trained to obtain the corresponding pre-trained models, the pre-trained models corresponding to the models to be trained can be connected in series according to the serial order corresponding to the models to be trained to obtain an image processing model. For example, there are 3 models to be trained C, D, and E, and the corresponding serial orders are 1, 2, and 3 respectively, that is, the model to be trained C - the model to be trained D - the model to be trained E; at this time, pre-training is performed for the models to be trained C, D, and E respectively to obtain the pre-trained models C, D, and E, and they are connected in series according to the serial order corresponding to the models to be trained, and an image processing model with the serial order of the pre-trained model C - the pre-trained model D - the pre-trained model E can be obtained.
[0042] In addition, an image processing model can also be obtained by other means. For example, multiple pre-trained models and the serial order of the pre-trained models that can realize the functions required for the current scenario can be directly obtained, and the pre-trained models can be directly connected in series to obtain an image processing model.
[0043] In step S320, based on a preset sample image, a pre-trained model, and the concatenation order of the pre-trained model in the image processing model, fine-tuning training is performed on the model to be adjusted in the image processing model to obtain a fine-tuned model corresponding to the model to be adjusted.
[0044] Among them, the preset sample image can be one or more of the pre-training sample images used for pre-training the model to be trained, or other pre-training sample images suitable for the application scenario of the image processing model. The present disclosure does not make special limitations on this.
[0045] For example, when training the model to be trained A based on the pre-training sample image 1 and training the model to be trained B based on the pre-training sample image 2, the preset sample image can include the pre-training sample image 1 used for pre-training the model to be trained A, or can include the pre-training sample image 2 used for pre-training the model to be trained B.
[0046] Among them, the model to be adjusted includes at least one pre-trained model, and the model to be adjusted at least includes the pre-trained model at the end of the concatenation order in the image processing model.
[0047] Specifically, when the model to be adjusted includes only one pre-trained model, the model to be adjusted is the pre-trained model at the end of the concatenation order in the image processing model; for example, the image processing model includes 5 concatenated pre-trained models, namely pre-trained model F - pre-trained model G - pre-trained model H - pre-trained model I - pre-trained model J, then the model to be adjusted is pre-trained model J; when the model to be adjusted includes multiple pre-trained models, it can include the pre-trained model at the end of the concatenation order in the image processing model, and one or more of the other pre-trained models in the image processor. For example, for the above image processing model including 5 concatenated pre-trained models, the model to be adjusted can include pre-trained model J, and one or more of pre-trained models F, G, H, and I.
[0048] Among them, fine-tuning training refers to the process of training the pre-trained model with samples consistent with the input data of the actual application scenario, so that after fine-tuning training, the parameters of the fine-tuned model can adapt to the input data of the actual application scenario. Through fine-tuning training, the subsequent AI model in the concatenated model can be made to adapt to the output data of the previous AI model, thereby avoiding the problem that the output result of the subsequent AI model is defective due to the inconsistency between the output data of the previous AI model and the training data.
[0049] In addition, in some embodiments, when using one or more of the pre-training sample images used for pre-training the model to be trained as the preset sample image, techniques for increasing sample diversity can also be used to process the pre-training sample images first, such as image enhancement and other techniques.
[0050] For example, in the training method of supervised learning, if the pre-training sample image 1 is an input image obtained by degrading a high-definition image, the high-definition image can be degraded again by modifying the degradation parameters, and the re-degraded high-definition image can be used as the input for model training. Another example is that if the pre-training sample image 2 is an input image obtained by acquisition or synthesis, the input image can be directly degraded to obtain new input data. Among them, the degradation process can include any process for reducing the image quality. For example, it can include a process of adding noise to the image. Modifying the degradation parameters can include any method of modifying the parameters that can change the effect of reducing the image quality. For example, in the noise reduction task, modifying the parameters for increasing or decreasing the degree of image noise, modifying the parameters for changing the noise form, etc. Another example is that in the deblurring task, modifying the parameters for increasing or decreasing the degree of image blurring, modifying the parameters for changing the type of blur kernel, etc.
[0051] In an exemplary embodiment, when fine-tuning the model to be adjusted in the image processing model, based on the serial order of the pre-training models in the image processing model, in the front-to-back order of the models to be adjusted, fine-tuning training can be performed on the models to be adjusted one by one based on the preset sample images to obtain the fine-tuned models corresponding to the models to be adjusted.
[0052] Specifically, since the image processing model has a serial structure, usually the output data of the previous AI model with a prior serial order will be directly used as the input data of the subsequent AI model with a posterior serial order, that is, the subsequent AI model will perform further processing based on the output data processed by the previous AI model. At this time, in order to avoid the impact of the change in the output data after the fine-tuning training of the previous AI model on the subsequent model, based on the serial order of the pre-training models in the image processing model, fine-tuning training can be performed in the front-to-back order of the actual serial order of the models to be adjusted in the image processing model.
[0053] For example, when the serial order of the image processing model is pre-training model F - pre-training model G - pre-training model H - pre-training model I - pre-training model J, if the pre-training model J is fine-tuned first and then the pre-training model I is fine-tuned, it is very likely that the output data of the fine-tuned pre-training model I, that is, the fine-tuned model I, has changed compared with the pre-training model I. Therefore, it is still possible that the fine-tuned pre-training model J (fine-tuned model J) cannot adapt to the output data of the fine-tuned model I.
[0054] In an exemplary embodiment, when fine-tuning the models to be adjusted in the front-to-back order, the specific fine-tuning training process can refer to Figure 5 As shown, it includes steps S510 to S520:
[0055] In step S510, determine the target model to be adjusted according to the front-back order of the models to be adjusted.
[0056] In step S520, perform fine-tuning training on the target model to be adjusted.
[0057] In an exemplary embodiment, in order to perform fine-tuning training sequentially according to the front-back order, a target model to be adjusted can be first determined according to the front-back order. After the fine-tuning training for this target model to be adjusted is completed, then determine the next target model to be adjusted according to the front-back order, and then perform fine-tuning training again.
[0058] Among them, the process of fine-tuning training refers to Figure 6 as shown, and may include the following steps S610 to S630:
[0059] In step S610, read the serial order n of the target model to be adjusted in the image processing model.
[0060] Among them, n is a positive integer. Since the model to be adjusted is a pre-trained model in the image processing model, and the purpose of fine-tuning training is to enable the target model to be adjusted to adapt to the output data corresponding to its previous pre-trained model after fine-tuning training, so as to avoid the problem of defective output images of the target model to be adjusted. Therefore, the serial order n of the target model to be adjusted in the image processing model can be first obtained, so as to determine the position of the target model to be adjusted in the image processing model and the previous model corresponding to the target model to be adjusted.
[0061] In step S620, obtain the output image of the (n - 1)-th model in the image processing model based on the preset sample image, and generate a fine-tuning sample image corresponding to the target model to be adjusted based on the output image.
[0062] In an exemplary embodiment, after determining the serial order n of the target model to be adjusted in the image processing model, the (n - 1)-th model can be determined in the image processing model according to n. Then, based on the preset sample image, obtain the output image corresponding to the (n - 1)-th model, and then generate a fine-tuning sample image corresponding to the target model to be adjusted based on the output image.
[0063] Among them, since the target model to be adjusted is determined according to the front-back order of the serial order, the (n - 1)-th model may include a pre-trained model that has not undergone fine-tuning training, or may include a fine-tuned model that has been trained. The present disclosure does not make special limitations on this.
[0064] It should be noted that in an exemplary embodiment, if both the pre-training and fine-tuning training processes are performed in a supervised manner, the training samples used for pre-training and fine-tuning training may include sample data with ground-truth images (labels). Specifically, in the pre-training samples images used for pre-training, each sample includes an input image and a ground-truth image; during fine-tuning training, each sample in the preset sample images also includes an input image and a ground-truth image. Correspondingly, when generating fine-tuning sample images based on the output images, the output images can be used as the input images for each sample in the fine-tuning sample images, and the ground-truth image corresponding to the output image in the preset sample images can be used as the ground-truth image included in the sample in the fine-tuning sample images.
[0065] For example, a sample 1 in the preset sample images includes sample image 1 and ground-truth image 1. Based on sample 1, the output image 1 (the input image of the target model to be adjusted) of the (n - 1)-th model in the image processing model can be obtained. At this time, the output image 1 of the (n - 1)-th model can be used as the input image of the target model to be adjusted, and the ground-truth image 1 can be used as the ground-truth image to generate a sample in the fine-tuning sample images, that is, output image 1 - ground-truth image 1.
[0066] In step S630, using the fine-tuning sample images corresponding to the target model to be adjusted as training samples, fine-tuning training is performed on the target model to be adjusted.
[0067] In an exemplary embodiment, after obtaining the fine-tuning sample images, based on these fine-tuning sample images, fine-tuning training can be performed on the target model to be adjusted to obtain the fine-tuning model corresponding to the target model to be adjusted.
[0068] It should be noted that if the fine-tuning training process is a supervised training process, the input images in the fine-tuning sample images can be input into the target model to be adjusted, and using the ground-truth images as labels, fine-tuning training is performed on the target model to be trained.
[0069] In addition, since it is necessary to generate fine-tuning sample images based on the output images of the (n - 1)-th model, the serial order n of the target model to be adjusted read in the image processing model is usually an integer greater than 1.
[0070] In an exemplary embodiment, when obtaining the output images of the (n - 1)-th model in the image processing model based on the preset sample images, in order to obtain fine-tuning training samples with consistent output image data with the (n - 1)-th model, at least the preset sample images need to be input into the (n - 1)-th model, and after being processed by the (n - 1)-th model, the corresponding output images are obtained; it is also possible to use the fine-tuning process of the target model to be adjusted to avoid the influence of the previous multiple models on the output images of the (n - 1)-th model at one time.
[0071] Specifically, the preset sample image can be used as the input of the k-th model in the image processing model. After being processed by the k-th to the (n - 1)-th models in series, the output image of the (n - 1)-th model is obtained.
[0072] Among them, k takes a positive integer less than or equal to n - 1. That is, the preset sample image can be used as the input image of any model a (the serial order of model a is the k-th) before the n-th model in the image processing model. Then, after being processed by the serial models from model a (the k-th) to the (n - 1)-th model in the image processing model, the output image corresponding to the (n - 1)-th model is obtained. Then, a fine-tuning sample image is generated based on this output image.
[0073] For example, there are 10 pre-trained models in the image processing model. When fine-tuning and training the first target model to be adjusted (assuming the first determined target model to be adjusted is the 8th pre-trained model), the preset sample image can be used as the input of any one of the 1st to 7th pre-trained models, and the output image of the 7th pre-trained model is correspondingly obtained. For example, assuming the preset sample image is used as the input image of the 3rd pre-trained model, the preset sample image can be successively processed by the 3rd, 4th, 5th, 6th, and 7th pre-trained models in series to obtain the output image of the 7th pre-trained model to generate a fine-tuning sample image. Again, the preset sample image can be directly used as the input image of the 7th pre-trained model to directly obtain the output image of the 7th pre-trained model.
[0074] Furthermore, to avoid the influence of the fine-tuning model on the output data, after fine-tuning and training each target model to be adjusted, the obtained fine-tuning model can be used to replace the corresponding pre-trained model in the image processing model, and then the fine-tuning training process of the next target model to be adjusted is carried out. For example, in the above example, if the 8th pre-trained model is the second determined target model to be adjusted, assuming the first determined target model to be adjusted is the 4th pre-trained model, to avoid the influence of the fine-tuning model on the output data, the preset sample image can be successively processed by the 3rd pre-trained model, the fine-tuning model corresponding to the 4th pre-trained model, the 5th pre-trained model, the 6th pre-trained model, and the 7th pre-trained model, and then the output image of the 7th pre-trained model is obtained to generate a fine-tuning sample image.
[0075] In step S330, a trained image processing model is generated based on the fine-tuning model.
[0076] In an exemplary embodiment, after fine-tuning the model to be adjusted to obtain a fine-tuned model corresponding to the model to be adjusted, since the fine-tuned model can adapt to the output image of the previous pre-trained model or fine-tuned model corresponding to the model to be adjusted, a trained image processing model can be generated based on the fine-tuned model, so as to avoid the problem that the single-task AI image processing model with a later concatenation order cannot adapt to the output image of the single-task AI image processing model with an earlier concatenation order, resulting in defective image processing results.
[0077] In an exemplary embodiment, when generating a trained image processing model based on the fine-tuned model, the fine-tuned model and the pre-trained models other than the pre-trained model corresponding to the fine-tuned model can be concatenated according to the concatenation order of the pre-trained models in the image processing model to obtain the trained image processing model.
[0078] For example, the concatenation order of the image processing model in the previous example is pre-trained model F - pre-trained model G - pre-trained model H - pre-trained model I - pre-trained model J. Assuming that fine-tuning training is performed on pre-trained model G and pre-trained model J to obtain the corresponding fine-tuned models G and J, at this time, the pre-trained model F, fine-tuned model G, pre-trained model H, pre-trained model I, and fine-tuned model J can be concatenated in the order of F, G, H, I, J to obtain the trained image processing model, that is, pre-trained model F - fine-tuned model G - pre-trained model H - pre-trained model I - fine-tuned model J.
[0079] In an exemplary embodiment, when generating a trained image processing model based on the fine-tuned model, the trained image processing model can also be directly generated by replacement. Specifically, referring to Figure 7 as shown, it may include steps S710 and S720:
[0080] In step S710, determine the pre-trained models corresponding to each fine-tuned model in the image processing model.
[0081] In step S720, use the fine-tuned model to replace the pre-trained model corresponding to the fine-tuned model in the image processing model to obtain the trained image processing model.
[0082] In an exemplary embodiment, after obtaining the fine-tuned models corresponding to all the models to be adjusted, the pre-trained models corresponding to each fine-tuned model can be first determined in the original image processing model, and then each fine-tuned model is directly used to replace the pre-trained models corresponding to each fine-tuned model in the image processing model to obtain the trained image processing model.
[0083] It should be noted that when setting the serial order of the pre-trained models in the image processing model, the serial order can be determined according to the robustness of the pre-trained models on the basis of achieving the predetermined image processing effect, so as to improve the processing effect of the trained image processing model.
[0084] The following takes the image processing model including two cascaded models and the model training method being the supervised training method as an example to elaborate on the solution of the embodiments of the present disclosure in detail:
[0085] The pre-training model K is pre-trained using the pre-training sample image 3 (input 3 - ground truth 3), and the pre-training model L is pre-trained using the pre-training sample image 4 (input 4 - ground truth 4), respectively obtaining the pre-training model K and the pre-training model L. The cascading method of the image processing model is determined to be the pre-training model K - pre-training model L according to the robustness.
[0086] Embodiment 1
[0087] Refer to Figure 8 As shown, the input 3 of the pre-training sample image 3 is input into the pre-training model K, obtaining the corresponding output 3' of the pre-training model K. Based on the output 3' and the ground truth 3, a fine-tuning sample image 3' (output 3' - ground truth 3) is generated. The pre-training model L is fine-tuned using the fine-tuning sample image 3', obtaining the fine-tuned model L. The pre-training model K and the fine-tuned model L are cascaded to obtain the trained image processing model pre-training model K - fine-tuned model L.
[0088] Embodiment 2
[0089] Refer to Figure 9 As shown, the input 4 of the pre-training sample image 4 is input into the pre-training model K, obtaining the corresponding output 4' of the pre-training model K. Based on the output 4' and the ground truth 4, a fine-tuning sample image 4' (output 4' - ground truth 4) is generated. The pre-training model L is fine-tuned using the fine-tuning sample image 4', obtaining the fine-tuned model L. The pre-training model K and the fine-tuned model L are cascaded to obtain the trained image processing model pre-training model K - fine-tuned model L.
[0090] Embodiment 3
[0091] Assume that the input 3 in the pre-training sample image 3 is obtained by degrading the high-definition image 3 through the degradation parameter 1. At this time, refer to Figure 10 As shown, the high-definition image 3 can be degraded using the degradation parameter 2 different from the degradation parameter 1 to form the input 3 降质参数2 . Then the input 3 降质参数2 is input into the pre-training model K, obtaining the corresponding output 3 降质参数2 ' of the pre-training model K. Based on the output 3 降质参数2 ' and the ground truth 3, a fine-tuning sample image 3降质参数2 ’(Output 3 降质参数2 ’ - True value 3). Use the fine-tuning sample image 3 降质参数2 ’ to fine-tune the pre-trained model L to obtain the fine-tuned model L. Connect the pre-trained model K and the fine-tuned model L in series to obtain the trained image processing model Pre-trained model K - Fine-tuned model L.
[0092] Example 4
[0093] Assume that the input 3 in the pre-trained sample image 3 is obtained by acquisition or synthesis. At this time, referring to Figure 11 as shown, the input 3 can be directly degraded to form the input 3 降质 . Then input the input 3 降质 into the pre-trained model K to obtain the output 3 降质 ’ corresponding to the pre-trained model K. Based on the output 3 降质 ’ and the true value 3, generate the fine-tuning sample image 3 降质 ’(Output 3 降质 ’ - True value 3). Use the fine-tuning sample image 3 降质 ’ to fine-tune the pre-trained model L to obtain the fine-tuned model L. Connect the pre-trained model K and the fine-tuned model L in series to obtain the trained image processing model Pre-trained model K - Fine-tuned model L.
[0094] For example, when the pre-trained model K is a model for performing a deblurring task and the input 3 in the pre-trained sample image 3 is obtained by acquisition or synthesis, the input 3 降质 (as shown in Figure 12 a) can be directly generated by adding a blur kernel. After the deblurring process of the pre-trained model K, the output 3 降质 ’(as shown in Figure 12 b) is obtained. Then, based on the output 3 降质 ’(as shown in Figure 12 b) and the true value 3 (as shown in Figure 12 c), the fine-tuning sample image 3 降质 ’ is generated.
[0095] In summary, in this exemplary embodiment, a fine-tuning sample image is generated through a preset sample image (pre-training sample image 3 or pre-training sample image 4) and a pre-trained model K, and then the pre-trained model L is fine-tuned based on the fine-tuning sample image to obtain a fine-tuned model L. After that, the pre-trained model K and the fine-tuned model L are concatenated according to the concatenation order of the pre-trained model K and the pre-trained model L. The obtained trained image processing model can not only maintain the original capabilities of the pre-trained model L, but also enable the fine-tuned model L to adapt to the output image of the pre-trained model K, that is, it avoids the defects and domain gap problems caused by concatenation. At the same time, the fine-tuned model L can also repair the problem of insufficient capabilities of the pre-trained model K connected in front during real-scene prediction.
[0096] For example, in the image processing model, the pre-trained model K and the pre-trained model L are models for deblurring and detail enhancement processing of images respectively. At this time, referring to Figure 13 as shown, based on the image processing model before fine-tuning training and the image processing model after fine-tuning training (fine-tuning training is performed in the manner of Embodiment 3) respectively, Figure 13 the image shown in Figure 13 a is processed, and Figure 13 b and Figure 13 b and Figure 13 c can be correspondingly obtained. It can be clearly observed at the positions framed in Figure 13 b and
[0097] c that for the image processing model without fine-tuning training, there are defects in the image output by the pre-trained model L (
[0098] the part framed in Figure 14 b, there are black vertical stripes on the eye lenses); while for the image processing model after fine-tuning training, there are no such defects in the image output by the fine-tuned model L.
[0099] It should be noted that the above drawings are only schematic illustrations of the processing included in the method according to the exemplary embodiments of the present disclosure, rather than for limiting purposes. It is easy to understand that the processing shown in the above drawings does not indicate or limit the time sequence of these processes. In addition, it is also easy to understand that these processes can be executed synchronously or asynchronously in, for example, multiple modules.
[0100] The fine-tuning training module 1420 can be used to perform fine-tuning training on the model to be adjusted in the image processing model based on a preset sample image, a pre-trained model, and the concatenation order of the pre-trained model in the image processing model, so as to obtain a fine-tuned model corresponding to the model to be adjusted; wherein, the model to be adjusted includes at least one pre-trained model, and the model to be adjusted at least includes the pre-trained model at the end of the concatenation order in the image processing model.
[0101] The model generation module 1430 can be used to generate a trained image processing model based on the fine-tuned model.
[0102] In an exemplary embodiment, the fine-tuning training module 1420 can be used to perform fine-tuning training on the model to be adjusted in sequence based on the concatenation order of the pre-trained model in the image processing model and the preset sample image, so as to obtain a fine-tuned model corresponding to the model to be adjusted.
[0103] In an exemplary embodiment, the fine-tuning training module 1420 can be used to determine the target model to be adjusted according to the front-to-back order of the models to be adjusted; perform fine-tuning training on the target model to be adjusted; wherein the fine-tuning training includes: reading the concatenation order n of the target model to be adjusted in the image processing model; where n is a positive integer; obtaining the output image of the (n - 1)-th model in the image processing model based on the preset sample image, and generating a fine-tuning sample image corresponding to the target model to be adjusted based on the output image; using the fine-tuning sample image corresponding to the target model to be adjusted as a training sample to perform fine-tuning training on the target model to be adjusted.
[0104] In an exemplary embodiment, the fine-tuning training module 1420 can be used to use the preset sample image as the input of the k-th model in the image processing model, and after being processed by the k-th to (n - 1)-th models in series, obtain the output image of the (n - 1)-th model; where k is a positive integer less than or equal to n - 1.
[0105] In an exemplary embodiment, the model acquisition module 1410 can be used to acquire a plurality of models to be trained and the corresponding concatenation order of the models to be trained; perform pre-training on each model to be trained respectively to obtain the pre-trained models corresponding to the models to be trained; concatenate the pre-trained models according to the corresponding concatenation order of the models to be trained to obtain an image processing model.
[0106] In an exemplary embodiment, the model generation module 1430 can be used to concatenate the fine-tuned model and the pre-trained models other than the pre-trained model corresponding to the fine-tuned model according to the concatenation order of the pre-trained models in the image processing model to obtain a trained image processing model.
[0107] In an exemplary embodiment, the model generation module 1430 may be configured to determine a pre-trained model corresponding to each fine-tuning model in an image processing model; and replace the pre-trained model corresponding to the fine-tuning model in the image processing model with the fine-tuning model to obtain a trained image processing model.
[0108] The specific details of each module in the above device have been described in detail in the implementation manner of the method part. For the details not disclosed, please refer to the implementation manner of the method part, and thus will not be elaborated here.
[0109] Those skilled in the art can understand that various aspects of the present disclosure can be implemented as a system, a method, or a program product. Therefore, various aspects of the present disclosure can be specifically implemented in the following forms, namely: a complete hardware implementation manner, a complete software implementation manner (including firmware, microcode, etc.), or an implementation manner combining hardware and software aspects, which can be collectively referred to as "circuit", "module", or "system" here.
[0110] An exemplary embodiment of the present disclosure also provides an electronic device for implementing a fine-tuning training method of an image processing model. The electronic device at least includes a processor and a memory. The memory is used to store executable instructions of the processor, and the processor is configured to execute the fine-tuning training method of the image processing model by executing the executable instructions.
[0111] Next, taking the mobile terminal 1500 in Figure 15 as an example, an exemplary description of the structure of the electronic device in the embodiments of the present disclosure will be given. Those skilled in the art should understand that, except for the components specifically for mobile purposes, Figure 15 the structure in Figure 15 can also be applied to fixed-type devices. In some other embodiments, the mobile terminal 1500 may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figure can be implemented in hardware, software, or a combination of software and hardware. The interface connection relationships between the components are only schematically shown and do not constitute a limitation on the structure of the mobile terminal 1500. In some other embodiments, the mobile terminal 1500 may also adopt an interface connection manner different from
[0112] As shown in Figure 15As shown in the figure, the mobile terminal 1500 may specifically include: a processor 1510, an internal memory 1521, an external memory interface 1522, a Universal Serial Bus (USB) interface 1530, a charging management module 1540, a power management module 1541, a battery 1542, an antenna 1, an antenna 2, a mobile communication module 1550, a wireless communication module 1560, an audio module 1570, a speaker 1571, a receiver 1572, a microphone 1573, a headphone interface 1574, a sensor module 1580, a display screen 1590, a camera module 1591, an indicator 1592, a motor 1593, a button 1594, and a subscriber identification module (SIM) card interface 1595, etc. Among them, the sensor module 1580 may include a depth sensor 15801, a pressure sensor 15802, a gyroscope sensor 15803, etc.
[0113] The processor 1510 may include one or more processing units. For example, the processor 1510 may include an Application Processor (AP), a modem processor, a Graphics Processing Unit (GPU), an Image Signal Processor (ISP), a controller, a video codec, a Digital Signal Processor (DSP), a baseband processor, and / or a Neural-Network Processing Unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.
[0114] The NPU is a Neural-Network (NN) computing processor. By learning from the biological neural network structure, for example, learning from the transmission mode between human brain neurons, it can quickly process the input information and can also continuously self-learn. Through the NPU, applications such as intelligent cognition of the mobile terminal 1500 can be realized, such as: image recognition, face recognition, voice recognition, text understanding, etc. In some embodiments, the process of fine-tuning and training of the image processing model may be executed by the NPU.
[0115] A memory is provided in the processor 1510. The memory may store instructions for implementing six modular functions: detection instructions, connection instructions, information management instructions, analysis instructions, data transmission instructions, and notification instructions, and are controlled and executed by the processor 1510.
[0116] The wireless communication function of the mobile terminal 1500 can be implemented by antenna 1, antenna 2, mobile communication module 1550, wireless communication module 1560, modulation and demodulation processor, baseband processor, etc. Among them, antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals; the mobile communication module 1550 can provide solutions for wireless communication including 15G / 3G / 4G / 5G, etc. applied to the mobile terminal 1500; the modulation and demodulation processor can include a modulator and a demodulator; the wireless communication module 1560 can provide solutions for wireless communication including Wireless Local Area Networks (WLAN) (such as Wireless Fidelity (Wi-Fi) network), Bluetooth (BT), etc. applied to the mobile terminal 1500. In some embodiments, antenna 1 of the mobile terminal 1500 is coupled to the mobile communication module 1550, and antenna 2 is coupled to the wireless communication module 1560, so that the mobile terminal 1500 can communicate with the network and other devices through wireless communication technologies.
[0117] The mobile terminal 1500 implements the display function through the GPU, display screen 1590, application processor, etc. The GPU is a microprocessor for image processing, connected to the display screen 1590 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 1510 may include one or more GPUs, which execute program instructions to generate or change display information. In some embodiments, the GPU can be used to perform degradation processing or image enhancement processing on high-definition images.
[0118] The internal memory 1521 can be used to store computer-executable program code, and the executable program code includes instructions. The internal memory 1521 may include a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, image playback function, etc.). The data storage area can store data created during the use of the mobile terminal 1500 (such as audio data, phone book, etc.). In addition, the internal memory 1521 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, Universal Flash Storage (UFS), etc. The processor 1510 executes various functional applications and data processing of the mobile terminal 1500 by running the instructions stored in the internal memory 1521 and / or the instructions stored in the memory provided in the processor.
[0119] The depth sensor 15801 is used to obtain the depth information of a scene. The pressure sensor 15802 is used to sense a pressure signal and can convert the pressure signal into an electrical signal. The gyroscope sensor 15803 can be used to determine the motion posture of the mobile terminal 1500. In addition, sensors with other functions can also be set in the sensor module 1580 according to actual needs, such as a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a distance sensor, a proximity light sensor, a fingerprint sensor, a temperature sensor, a touch sensor, an ambient light sensor, a bone conduction sensor, etc.
[0120] In addition, the exemplary embodiments of the present disclosure also provide a computer-readable storage medium, on which a program product capable of implementing the above methods in this specification is stored. In some possible embodiments, various aspects of the present disclosure can also be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to cause the terminal device to execute the steps according to various exemplary embodiments of the present disclosure described in the above "Exemplary Method" section, for example, it can execute Figures 3 to 7 any one or more of the steps.
[0121] It should be noted that the computer-readable medium shown in the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0122] In this disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. And in this disclosure, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination of the above.
[0123] In addition, the program code for performing the operations of this disclosure can be written in any combination of one or more programming languages. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, by using an Internet service provider to connect through the Internet).
[0124] Those skilled in the art will readily conceive of other embodiments of this disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure, which follow the general principles of this disclosure and include known common knowledge or conventional technical means in the technical field not disclosed in this disclosure. The specification and embodiments are only considered exemplary, and the true scope and spirit of this disclosure are pointed out by the claims.
[0125] It should be understood that this disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is only limited by the appended claims.
Claims
1. A fine-tuning training method for an image processing model, characterized in that, it includes: Obtain an image processing model, where the image processing model includes a plurality of pre-trained models connected in series; Based on a preset sample image, the pre-trained model, and the serial order of the pre-trained model in the image processing model, perform fine-tuning training on the model to be adjusted in the image processing model to obtain a fine-tuned model corresponding to the model to be adjusted; wherein, the model to be adjusted includes at least one of the pre-trained models, and the model to be adjusted at least includes the pre-trained model at the end of the serial order in the image processing model; Generate a trained image processing model based on the fine-tuned model; Among them, the step of performing fine-tuning training on the model to be adjusted in the image processing model based on the preset sample image, the pre-trained model, and the serial order of the pre-trained model in the image processing model to obtain a fine-tuned model corresponding to the model to be adjusted includes: Based on the serial order of the pre-trained model in the image processing model, in the front-to-back order of the model to be adjusted, perform fine-tuning training on the model to be adjusted in turn based on the preset sample image to obtain a fine-tuned model corresponding to the model to be adjusted.
2. The method according to claim 1, characterized in that, The step of performing fine-tuning training on the model to be adjusted in turn based on the preset sample image in the front-to-back order of the model to be adjusted includes: Determine the target model to be adjusted in the front-to-back order of the model to be adjusted; Perform fine-tuning training on the target model to be adjusted; wherein the fine-tuning training includes: Read the serial order n of the target model to be adjusted in the image processing model; where n is a positive integer; Obtain the output image of the (n - 1)-th model in the image processing model based on the preset sample image, and generate a fine-tuning sample image corresponding to the target model to be adjusted based on the output image; Use the fine-tuning sample image corresponding to the target model to be adjusted as a training sample to perform fine-tuning training on the target model to be adjusted.
3. The method according to claim 2, characterized in that, The step of obtaining the output image of the (n - 1)-th model in the image processing model based on the preset sample image includes: Take the preset sample image as the input of the k-th model in the image processing model, and after being processed by the k-th to (n - 1)-th models connected in series, obtain the output image of the (n - 1)-th model; where k is a positive integer less than or equal to n - 1.
4. The method according to claim 1, characterized in that, The step of obtaining the image processing model includes: Obtain a plurality of models to be trained and the corresponding serial order of the models to be trained; Perform pre-training on each of the models to be trained respectively to obtain pre-trained models corresponding to the models to be trained; Connect the pre-trained models in series according to the serial order corresponding to the models to be trained to obtain the image processing model.
5. The method according to claim 1, characterized in that, The step of generating a trained image processing model based on the fine-tuned model includes: Connect the fine-tuning model and the pre-trained models other than the pre-trained model corresponding to the fine-tuning model in series according to the series connection order of the pre-trained models in the image processing model, to obtain a trained image processing model.
6. The method according to claim 1, wherein, the generating a trained image processing model based on the fine-tuning model includes: determining the pre-trained models corresponding to the respective fine-tuning models in the image processing model; using the fine-tuning model to replace the pre-trained model corresponding to the fine-tuning model in the image processing model, to obtain a trained image processing model.
7. A fine-tuning training device for an image processing model, wherein, it includes: a model acquisition module, configured to acquire an image processing model, where the image processing model includes a plurality of pre-trained models connected in series; a fine-tuning training module, configured to perform fine-tuning training on a model to be adjusted in the image processing model based on a preset sample image, the pre-trained model, and the series connection order of the pre-trained models in the image processing model, to obtain a fine-tuning model corresponding to the model to be adjusted; wherein, the model to be adjusted includes at least one of the pre-trained models, and the model to be adjusted includes at least the pre-trained model at the end of the series connection order in the image processing model; a model generation module, configured to generate a trained image processing model based on the fine-tuning model; wherein, the performing fine-tuning training on the model to be adjusted in the image processing model based on the preset sample image, the pre-trained model, and the series connection order of the pre-trained models in the image processing model, to obtain a fine-tuning model corresponding to the model to be adjusted includes: based on the series connection order of the pre-trained models in the image processing model, sequentially performing fine-tuning training on the model to be adjusted based on the preset sample image according to the front-back order of the model to be adjusted, to obtain a fine-tuning model corresponding to the model to be adjusted.
8. A computer-readable medium, on which a computer program is stored, wherein, the computer program, when executed by a processor, implements the method according to any one of claims 1 to 6.
9. An electronic device, wherein, it includes: a processor; and a memory, configured to store executable instructions of the processor; wherein, the processor is configured to execute the method according to any one of claims 1 to 6 by executing the executable instructions.
Citation Information
Patent Citations
Convolutional neural network structure simplification and image classification method based on evolutionary strategy
CN110427965A
Face recognition model processing method, face recognition method and face recognition device
CN112801054A