Image processing method and device, equipment and medium
Through the optimized target network model, combined with quality and consistency evaluation, the problem of poor image editing effect in the prior art is solved, and high quality and consistency of generated images in multiple dimensions are achieved.
Patent Information
- Application Number
- CN202410046849.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-11
- Publication Date
- 2025-07-11
AI Technical Summary
The existing image editing effect is poor and it is difficult to meet the needs of users in multiple dimensions. Especially when the content filling process is processed, it is difficult to take into account the consistency and quality of the generated image and the original image.
The target network model is adopted, which is based on the optimized first network model and the second network model, and optimizes the first network model through multi-dimensional evaluation. Combining the evaluation results of the quality dimension and consistency dimension, the parameters of the first network model are adjusted to improve the image editing effect.
The image editing effect is achieved to meet user needs in multiple dimensions, especially during content filling processing, the quality and consistency of generated images have been significantly improved.
Smart Images

Figure CN120298539A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing technologies, and in particular, to an image processing method, apparatus, device, and medium. Background Art
[0002] In application software such as image editing software and photo-taking software, the images input by users can be edited and processed to obtain the images required by the users. However, the existing image editing effect is not good and it is difficult to meet the user's needs. Summary of the Invention
[0003] In order to solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides an image processing method, apparatus, device, and medium.
[0004] An embodiment of the present disclosure provides an image processing method, the method including: obtaining a first image to be processed; calling a preset target network model; wherein, the target network model is obtained based on a preset first network model and a preset second network model, the first network model is used for editing and processing a sample image to obtain a generated image corresponding to the sample image; the second network model is used for performing multi-dimensional evaluation on the generated image to optimize the first network model based on the multi-dimensional evaluation results; and the target network model is the optimized first network model; using the target network model to perform editing and processing on the first image to generate a second image.
[0005] Optionally, the sample image is an image with a local content occluded, and the editing and processing includes content filling processing.
[0006] Optionally, the performing multi-dimensional evaluation on the generated image to optimize the first network model based on the multi-dimensional evaluation results includes: evaluating the generated image respectively based on a preset plurality of evaluation dimensions to obtain evaluation results corresponding to the plurality of evaluation dimensions respectively; determining the information to be optimized of the first network model based on the evaluation results corresponding to the plurality of evaluation dimensions respectively, and adjusting the parameters of the first network model based on the information to be optimized.
[0007] Optionally, the plurality of evaluation dimensions include a quality dimension and a consistency dimension, and the consistency dimension measures the consistency between the sample image and the generated image corresponding to the sample image at the pixel level.
[0008] Optionally, the quality dimension measures the quality of the generated image corresponding to the sample image based on one or more factors among the overall image, graphic-text matching, image details, image structure, image light and shadow, and image aesthetics.
[0009] Optionally, the preset second network model is obtained based on the following steps: obtaining a first training sample set corresponding to a quality dimension; the first training sample set includes a plurality of first training samples, and the first training sample set carries sorting information of quality evaluation results of the plurality of first training samples; obtaining a second training sample set corresponding to a consistency dimension; the second training sample set includes a plurality of second training samples, and the second training sample set carries sorting information of pixel evaluation results of the plurality of second training samples; based on the first training sample set and the second training sample set, training an initial network model by using a sorting loss function to obtain the preset second network model based on the trained initial network model.
[0010] Optionally, the initial network model includes a plurality of initial sub-models with the same structure, the first training sample set includes at least one first sample subset, and the sorting information carried by different first sample subsets is obtained by measuring the quality of the plurality of first training samples based on different factors. The at least one first sample subset and the second training sample set are used to train different initial sub-models.
[0011] An embodiment of the present disclosure further provides an image processing apparatus, including: an image acquisition module, configured to acquire a first image to be processed; a model calling module, configured to call a preset target network model; wherein, the target network model is obtained based on a preset first network model and a preset second network model, the first network model is configured to perform editing processing on a sample image to obtain a generated image corresponding to the sample image; the second network model is configured to perform multi-dimensional evaluation on the generated image to optimize the first network model based on the multi-dimensional evaluation result; and the target network model is the optimized first network model; an image editing module, configured to perform editing processing on the first image by using the target network model to generate a second image.
[0012] An embodiment of the present disclosure further provides an electronic device, the electronic device includes: a processor; a memory for storing executable instructions of the processor; the processor is configured to read the executable instructions from the memory and execute the instructions to implement the image processing method provided by the embodiment of the present disclosure.
[0013] An embodiment of the present disclosure further provides a computer-readable storage medium, the storage medium stores a computer program, and the computer program is used to execute the image processing method provided by the embodiment of the present disclosure.
[0014] The above technical solution provided by the embodiments of the present disclosure can call a target network model to perform editing processing on a first image to be processed. Among them, the target network model is obtained based on a preset first network model and a preset second network model. Specifically, the first network model can perform editing processing on a sample image to obtain a generated image corresponding to the sample image; the second network model can perform multi-dimensional evaluation on the generated image to optimize the first network model based on the multi-dimensional evaluation results, and use the optimized first network model as the target network model, so as to effectively ensure that the image editing effect of the target network model can meet the user's needs in multiple dimensions.
[0015] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present disclosure and used together with the specification to explain the principles of the present disclosure.
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0018] Figure 1 It is a schematic flowchart of an image processing method provided by an embodiment of the present disclosure;
[0019] Figure 2 It is a schematic diagram of optimizing a first network model provided by an embodiment of the present disclosure;
[0020] Figure 3 It is a schematic structural diagram of a second network model provided by an embodiment of the present disclosure;
[0021] Figure 4 It is a schematic structural diagram of an image processing device provided by an embodiment of the present disclosure;
[0022] Figure 5 It is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] In order to be able to more clearly understand the above objects, features, and advantages of the present disclosure, the following will further describe the solutions of the present disclosure. It should be noted that, without conflict, the embodiments of the present disclosure and the features in the embodiments can be combined with each other.
[0024] Numerous specific details are set forth in the following description to facilitate a full understanding of the present disclosure. However, the present disclosure may also be implemented in other ways different from those described herein. Obviously, the embodiments in the specification are only a part of the embodiments of the present disclosure, rather than all of the embodiments.
[0025] Figure 1 The following is a schematic flowchart of an image processing method provided by an embodiment of the present disclosure. This method can be executed by an image processing device, where the device can be implemented by software and / or hardware and is generally integrated in an electronic device. As Figure 1 shown, the method mainly includes the following steps S102 to step S106:
[0026] Step S102, obtain a first image to be processed.
[0027] The embodiments of the present disclosure do not limit the content included in the first image. For example, the first image can be a portrait image, a landscape image, etc. In addition, the embodiments of the present disclosure do not limit the acquisition method of the first image. For example, it can be an image uploaded by the user through taking a photo, an image selected by the user from the local, an image downloaded by the user through the network, or an image obtained by transmission from other devices.
[0028] Step S104, call a preset target network model; wherein, the target network model is obtained based on a preset first network model and a preset second network model. The first network model is used to perform editing processing on a sample image to obtain a generated image corresponding to the sample image; the second network model is used to perform multi-dimensional evaluation on the generated image to optimize the first network model based on the multi-dimensional evaluation results; and the target network model is the optimized first network model.
[0029] The embodiments of the present disclosure do not limit the structures of the first network model and the second network model. In practical applications, a first network model with the required editing and processing capabilities can be initially trained. The embodiments of the present disclosure do not limit the editing and processing, such as beautification processing, content filling processing, etc. In some specific implementation examples, the editing and processing include content filling processing. On this basis, the sample image is an image with partially occluded content. The embodiments of the present disclosure do not limit the position of the occluded partial content, which can be the peripheral area of the image or the internal area of the image. Therefore, this content filling processing can be an expansion and drawing process on the basis of the original image (i.e., the input image of the first network model), that is, expanding the edge content of the original image. The content filling processing can also be to complete the partially occluded or smeared areas in the original image. It can be understood that the editing and processing capabilities of the first network model may be limited. Therefore, the generated image effects obtained by editing and processing the input image may not meet the expectations. For example, in the generated image obtained by content filling for an image containing only a person, there may be low-quality images such as floating dogs, or the generated image for a billboard with a smeared local area gives a poor sense of consistency to the user. For this reason, the embodiments of the present disclosure will optimize the first network model with the help of an additional second network model. Specifically, the second network model is a model that can distinguish the advantages and disadvantages of an image in multiple dimensions, that is, it can evaluate the generated image in multiple dimensions. By evaluating the generated image output by the first network model through the second network model in multiple dimensions, the performance of the generated image in multiple dimensions can be measured relatively efficiently, and the optimization direction of the first network model can be determined. This optimization direction can be the direction that makes the multi-dimensional evaluation results of the generated image of the first network model optimal. The optimized first network model can generate images with better multi-dimensional evaluation results. Through the above method, the editing and processing capabilities of the target network model can be effectively ensured to meet the expectations in multiple dimensions.
[0030] Step S106: Use the target network model to perform editing and processing on the first image to generate a second image.
[0031] When the target network model performs content filling processing on the first image, the obtained second image not only contains the content of the first image but also performs filling processing on the content of the first image. Specifically, when it is detected that there is a target area (a local smeared area or a local occluded area) in the first image, the target area can be completed according to the existing content of the first image. When it is detected that there is no target area in the first image, the first image can be expanded peripherally according to the existing content of the first image. Since the target network model is obtained based on the optimized first network model, the multi-dimensional evaluation results are usually good, that is, the finally obtained second image can meet the user's needs in multiple dimensions.
[0032] The foregoing second network model can perform multi-dimensional evaluation on the generated image to optimize the first network model based on the multi-dimensional evaluation results. For the convenience of understanding, embodiments of the present disclosure provide an implementation example of performing multi-dimensional evaluation on the generated image to optimize the first network model, which can be executed according to the following steps A and B:
[0033] Step A: Evaluate the generated image respectively based on a variety of preset evaluation dimensions to obtain the evaluation results corresponding to the respective evaluation dimensions.
[0034] Embodiments of the present disclosure do not limit the variety of evaluation dimensions, which can be flexibly set according to requirements. In some implementation examples, the first network model is a generation model, and the editing process is a content filling process. Exemplarily, the variety of evaluation dimensions include a quality dimension and a consistency dimension, and the consistency dimension measures the consistency between the sample image and the generated image corresponding to the sample image at the pixel level. For example, on the basis that the editing process includes a content filling process, measuring the consistency between the sample image and the generated image corresponding to the sample image at the pixel level can mainly be achieved by evaluating each pixel in the image to measure the consistency between the filled area of the generated image and the original sample image in one or more aspects such as texture, layout, and color. It can be understood that the inventors have found through research that related technologies basically perform overall evaluation on images. Although the overall quality of the images generated by the model can be guaranteed to a certain extent, problems such as poor content consistency between the content filling area of the images generated by the model and the original images are likely to occur. Therefore, the generated images of the model still cannot satisfy users. For this reason, embodiments of the present disclosure enable the second network model to evaluate the first network model around two core points of quality and consistency, so as to ensure that the generated image obtained after the first network model performs editing processing on the input image not only has high quality, but also has strong consistency with the input image, and can better meet the user's needs in terms of overall quality and consistency.
[0035] Furthermore, the quality dimension can be further divided into fine-grained levels according to requirements. Exemplarily, the quality dimension measures the quality of the generated image corresponding to the sample image based on one or more factors among the overall image, graphic matching, image details, image structure, image lighting, and image aesthetics. Through the above method, the second network model can perform multi-faceted and fine-grained evaluation on the quality of the image based on a variety of factors, so as to fully ensure that the images generated by the first network model can meet multi-dimensional quality requirements.
[0036] Step B: Determine the information to be optimized of the first network model based on the evaluation results corresponding to the respective evaluation dimensions, and adjust the parameters of the first network model based on the information to be optimized.
[0037] The information to be optimized can be used to characterize the optimization direction of the first network model. Specifically, based on the evaluation results corresponding to multiple evaluation dimensions, the target dimension can be determined from multiple evaluation dimensions. The evaluation result corresponding to the target dimension does not meet the preset requirements, such as the evaluation score being lower than the preset score threshold. In other words, the target dimension is the dimension in which the generated images of the first network model perform poorly. At this time, the information to be optimized for the first network model can be determined based on the target dimension, and then the parameters of the first network model can be adjusted based on the information to be optimized to improve the image generation effect of the first network model after parameter adjustment. In addition, since adjusting the parameters of the first network model based on the target dimension will also affect the evaluation results of the generated images in other dimensions, when adjusting the parameters of the first network model, the evaluation loss (also called the reward loss) corresponding to each evaluation dimension can be calculated simultaneously, prompting each evaluation loss to be maximized, so as to maximize the sum of multiple evaluation losses, that is, to make the multi-dimensional evaluation results of the images finally output by the first network model globally optimal.
[0038] For ease of understanding, reference can be made to Figure 2 An optimization schematic diagram of a first network model as shown, which shows that an architectural image with partial occlusion is input into the first network model so that the first network model performs content filling processing based on the input image to obtain a generated image (the overall structure of the building is completed). Then the generated image is input into the second network model, and the second network model can output an image global score and a pixel score for the generated image. Among them, the global score is the evaluation result obtained based on the aforementioned quality dimension, and the pixel score is the evaluation result obtained based on the aforementioned consistency dimension. After that, the parameters of the first network model can be updated in the direction of maximizing the global score and the pixel score to achieve the optimization of the first network model. In practical applications, the global score can be further subdivided, such as being subdivided into: overall image score, text-image matching score, image detail score, image structure score, image light and shadow score, image aesthetics score, etc., so as to perform multi-dimensional evaluation of the global performance of the generated image at a fine-grained level.
[0039] In practical applications, it is also possible to input a prompt text while inputting an image into the first network model. The prompt text is an optional item and is not restricted here. In addition, the first network model in the embodiments of the present disclosure is not restricted either. Exemplarily, the first network model includes a generative model, and any generative model can be used as the first network model. In some specific implementation examples, the first network model can be a diffusion model, so as to have a stronger image generation ability and be able to generate an output image related to the input image.
[0040] Furthermore, the embodiments of the present disclosure provide a method for obtaining a preset second network model. Exemplarily, the preset second network model is obtained based on the following steps (1) to (3):
[0041] Step (1): Obtain a first training sample set corresponding to a quality dimension; the first training sample set includes multiple first training samples, and the first training sample set carries sorting information of the quality evaluation results of the multiple first training samples. Exemplarily, the sorting of the multiple first training samples can be marked manually or in other ways in the order from excellent to poor or from poor to excellent in terms of the quality dimension.
[0042] Step (2): Obtain a second training sample set corresponding to a consistency dimension; the second training sample set includes multiple second training samples, and the second training sample set carries sorting information of the pixel evaluation results of the multiple second training samples. Among them, each pixel in each second training sample corresponds to an evaluation result. Assuming that the second training sample has M * N pixels, the second training sample has M * N pixel evaluation results, and the above sorting information is the information obtained by sorting the evaluation results of each pixel at the same position in different second training samples in the order from excellent to poor or from poor to excellent.
[0043] Step (3): Based on the first training sample set and the second training sample set, use a sorting loss function to train an initial network model, so as to obtain a preset second network model based on the trained initial network model. The trained initial network model is the second network model. By using the above first training sample set, second training sample set and sorting loss function to train the initial network model, the finally trained second network model has the ability to better distinguish the quality of images and the quality of image pixels, so as to be able to reliably evaluate the images output by other models in dimensions such as quality and consistency.
[0044] In practical applications, the second network model can be a network model that can simultaneously output evaluation results corresponding to multiple dimensions; the second network model can also include multiple sub-models, and each sub-model outputs evaluation results in its corresponding dimension respectively, which is not limited here. For the convenience of processing and application, in some specific implementation examples, the initial network model includes multiple initial sub-models with the same structure, the first training sample set includes at least one first sample subset, and the sorting information carried by different first sample subsets is obtained by measuring the quality of multiple first training samples based on different factors. At least one first sample subset and the second training sample set are used to train different initial sub-models, so that the trained different initial sub-models can score in different dimensions. The second network model also includes multiple trained initial sub-models, and multiple trained initial sub-models can exist side by side and can be flexibly called according to specific needs.
[0045] For the convenience of understanding, taking the example that the second network model can simultaneously output evaluation results of the quality dimension and the consistency dimension as an example, reference can be made toFigure 3 Schematic diagram of the structure of a second network model shown. The second network model includes an image encoder, a text encoder, a multimodal fusion network, and an evaluation output network. Among them, the image encoder is used to encode the received input image to obtain image features. The text encoder is used to encode the received prompt text to obtain text features. The multimodal fusion network is used to fuse the image features and text features to obtain fused features. The evaluation output network is used to generate an evaluation result of the received image based on the fused features. For example, it outputs a global score and a pixel score. Among them, the global score is an evaluation result based on the quality dimension, and the pixel score is an evaluation result based on the consistency dimension. In specific implementation, the evaluation output network can be a multi-scale MLP (Multilayer Perceptron), or other structures can also be used, which are not limited here.
[0046] As mentioned above, in practical applications, the second network model can include multiple trained initial sub-models, and the structure of each initial sub-model can refer to Figure 3 shown, or other structures can also be used, which are not limited here. For example, in some scenarios without text input, the text encoder may not be included.
[0047] In summary, the above image processing method provided by the embodiments of the present disclosure can use the second network model to perform multi-dimensional evaluation on the image output by the first network model, so as to optimize the first network model based on the multi-dimensional evaluation results, and use the optimized first network model as the target network model, thereby effectively ensuring that the image editing effect of the target network model can meet the user's needs in multiple dimensions. Especially for the first network model (such as a diffusion model) that needs to perform content filling processing, the reliability of the generated image can be fully guaranteed from dimensions such as quality and consistency.
[0048] Corresponding to the foregoing image processing method, Figure 4 Schematic diagram of the structure of an image processing device provided by the embodiments of the present disclosure. This device can be implemented by software and / or hardware, and is generally integrated in an electronic device, such as Figure 4 shown, the image processing device includes:
[0049] An image acquisition module 402, configured to acquire a first image to be processed;
[0050] The model calling module 404 is used to call a preset target network model; wherein, the target network model is obtained based on a preset first network model and a preset second network model. The first network model is used to perform editing processing on a sample image to obtain a generated image corresponding to the sample image. The second network model is used to perform multi-dimensional evaluation on the generated image to optimize the first network model based on the multi-dimensional evaluation results. And the target network model is the optimized first network model.
[0051] The image editing module 406 is used to perform editing processing on the first image by using the target network model to generate a second image.
[0052] Through the above device, it can effectively ensure that the image editing effect of the target network model can meet the user's needs in multiple dimensions.
[0053] In some embodiments, the sample image is an image with partial content occluded, and the editing processing includes content filling processing.
[0054] In some embodiments, the model calling module 404 is specifically configured to: evaluate the generated image respectively based on a preset variety of evaluation dimensions to obtain the evaluation results corresponding to the respective evaluation dimensions; determine the information to be optimized of the first network model based on the evaluation results corresponding to the respective evaluation dimensions, and adjust the parameters of the first network model based on the information to be optimized.
[0055] In some embodiments, the variety of evaluation dimensions includes a quality dimension and a consistency dimension, and the consistency dimension measures the consistency between the sample image and the generated image corresponding to the sample image at the pixel level.
[0056] In some embodiments, the quality dimension measures the quality of the generated image corresponding to the sample image based on one or more factors among the overall image, graphic-text matching, image details, image structure, image lighting, and image aesthetics.
[0057] In some embodiments, the preset second network model is obtained based on the following steps: obtaining a first training sample set corresponding to the quality dimension; the first training sample set contains multiple first training samples, and the first training sample set carries the sorting information of the quality evaluation results of the multiple first training samples; obtaining a second training sample set corresponding to the consistency dimension; the second training sample set contains multiple second training samples, and the second training sample set carries the sorting information of the pixel evaluation results of the multiple second training samples; based on the first training sample set and the second training sample set, training an initial network model by using a sorting loss function to obtain the preset second network model based on the trained initial network model.
[0058] In some embodiments, the initial network model includes a plurality of initial sub-models with the same structure. The first training sample set includes at least one first sample subset. The sorting information carried by different first sample subsets is obtained by measuring the quality of a plurality of the first training samples based on different factors. The at least one first sample subset and the second training sample set are used to train different initial sub-models.
[0059] The image processing apparatus provided by the embodiments of the present disclosure can execute the image processing method provided by any embodiment of the present disclosure, and has functional modules and beneficial effects corresponding to the execution of the method.
[0060] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working process of the device embodiments described above can refer to the corresponding process in the method embodiments, which will not be elaborated here.
[0061] The embodiments of the present disclosure provide an electronic device, which includes: a storage device on which a computer program is stored; a processing device for executing the computer program in the storage device to implement the steps of any method in the present disclosure.
[0062] Reference is made below to Figure 5 , which shows a schematic structural diagram of an electronic device 500 suitable for implementing the embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multi-
[0063] Media Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 5 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.
[0064] As Figure 5 shown, the electronic device 500 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the electronic device 500 are also stored. The processing device 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0065] Typically, the following devices can be connected to the I / O interface 505: an input device 506 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 508 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 509. The communication device 509 can allow the electronic device 500 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 5 the electronic device 500 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. More or fewer devices can be alternatively implemented or had.
[0066] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program codes for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through the communication device 509, or installed from the storage device 508, or installed from the ROM 502. When the computer program is executed by the processing device 501, the above functions defined in the method of the embodiment of the present disclosure are executed.
[0067] In addition to the above methods and devices, an embodiment of the present disclosure can also be a computer program product, which includes computer program instructions that cause the processor to execute the image processing method provided by the embodiment of the present disclosure when the computer program instructions are run by the processor. The computer program product can be written in any combination of one or more programming languages for executing the program codes of the operations of the embodiment of the present disclosure. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program codes can be executed entirely on the user computing device, partially on the user device, executed as an independent software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0068] Furthermore, an embodiment of the present disclosure can also be a computer-readable storage medium, on which computer program instructions are stored, and the computer program instructions cause the processor to execute the image processing method provided by the embodiment of the present disclosure when the computer program instructions are run by the processor.
[0069] The computer-readable storage medium may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may include, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the readable storage medium (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0070] The embodiments of the present disclosure further provide a computer program product, including a computer program / instructions, which when executed by a processor, implement the image processing method in the embodiments of the present disclosure.
[0071] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0072] For example, when receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that performs the operations of the technical solutions of the present disclosure according to the prompt message.
[0073] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving an active request from the user may be, for example, a pop-up window manner, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0074] It can be understood that the above process of notifying and obtaining the user's authorization is only illustrative and does not constitute a limitation on the implementation manner of the present disclosure. Other manners that meet relevant laws and regulations can also be applied to the implementation manner of the present disclosure.
[0075] It should be noted that in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising said element.
[0076] The above are only specific embodiments of the present disclosure, enabling those skilled in the art to understand or implement the present disclosure. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure will not be limited to the embodiments described herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.
Claims
1. An image processing method, characterized in that, Including: Obtain a first image to be processed; Invoke a preset target network model; wherein, the target network model is obtained based on a preset first network model and a preset second network model. The first network model is used to perform editing processing on a sample image to obtain a generated image corresponding to the sample image. The second network model is used to perform multi-dimensional evaluation on the generated image to optimize the first network model based on the multi-dimensional evaluation results. And the target network model is the optimized first network model; Use the target network model to perform editing processing on the first image to generate a second image.
2. The method according to claim 1, wherein The sample image is an image with partially occluded content, and the editing processing includes content filling processing.
3. The method according to claim 1, wherein The multi-dimensional evaluation of the generated image to optimize the first network model based on the multi-dimensional evaluation results includes: Evaluate the generated image respectively based on a plurality of preset evaluation dimensions to obtain evaluation results corresponding to the respective evaluation dimensions of the plurality of evaluation dimensions; Determine the information to be optimized of the first network model based on the evaluation results corresponding to the respective evaluation dimensions of the plurality of evaluation dimensions, and adjust the parameters of the first network model based on the information to be optimized.
4. The method according to claim 3, wherein The plurality of evaluation dimensions include a quality dimension and a consistency dimension, and the consistency dimension measures the consistency between the sample image and the generated image corresponding to the sample image at the pixel level.
5. The method according to claim 4, wherein The quality dimension measures the quality of the generated image corresponding to the sample image based on one or more factors among the overall image, text-image matching, image details, image structure, image lighting, and image aesthetics.
6. The method according to claim 1, characterized in that, The preset second network model is obtained based on the following steps: Obtain a first training sample set corresponding to the quality dimension; the first training sample set contains a plurality of first training samples, and the first training sample set carries sorting information of the quality evaluation results of the plurality of first training samples; Obtain a second training sample set corresponding to the consistency dimension; The second training sample set contains a plurality of second training samples, and the second training sample set carries sorting information of the pixel evaluation results of the plurality of second training samples; Based on the first training sample set and the second training sample set, use a sorting loss function to train an initial network model to obtain a preset second network model based on the trained initial network model.
7. The method according to claim 6, characterized in that The initial network model includes a plurality of initial sub-models with the same structure. The first training sample set contains at least one first sample subset. The sorting information carried by different first sample subsets is obtained by measuring the quality of the plurality of first training samples based on different factors. The at least one first sample subset and the second training sample set are used to train different initial sub-models.
8. An image processing apparatus, characterized in that, Including: An image acquisition module for obtaining a first image to be processed; A model calling module, configured to call a preset target network model; wherein, the target network model is obtained based on a preset first network model and a preset second network model, the first network model is configured to perform editing processing on a sample image to obtain a generated image corresponding to the sample image; the second network model is configured to perform multi-dimensional evaluation on the generated image to optimize the first network model based on the multi-dimensional evaluation result; and the target network model is the optimized first network model; An image editing module, configured to use the target network model to perform editing processing on the first image to generate a second image.
9. An electronic device, characterized in that, The electronic device includes: A storage device, on which a computer program is stored; A processing device, configured to execute the computer program in the storage device to implement the steps of the image processing method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, and the computer program is configured to execute the image processing method according to any one of claims 1-7 above.