Method of image processing and electronic device
By generating a reference image using a first model and transferring its style to the original image using a second model, the problem of complex image editing in existing technologies is solved, and flexible automatic conversion and efficient editing of image styles are achieved.
Patent Information
- Application Number
- CN202311305983.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-09
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2043-10-09
AI Technical Summary
In existing technologies, users need professional knowledge and cumbersome operating procedures to adjust the style of images, resulting in a poor image editing experience.
By acquiring style conversion prompts from user input, a reference image is generated using a first model, and the style of the reference image is transferred to the original image using a second model, thus achieving automatic image style conversion and reducing manual adjustment steps.
It enables flexible switching of image styles, simplifies user operations, improves the efficiency and quality of image editing, and ensures the preservation of high-level semantic information.
Smart Images

Figure CN119850405B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of image processing, and in particular to an image processing method and an electronic device. BACKGROUND
[0002] With the popularity and development of mobile devices with cameras, people increasingly acquire images in a more convenient way. The images taken by electronic devices are affected by many external factors, resulting in images that are not so satisfactory, and manual use of software is required to modify the images. For example, the brightness and saturation of the taken images do not meet the needs of users, and users can use image editing software to adjust the color temperature, brightness, saturation, and hue of the images. If a user wants to change the background environment of an image, the user needs to use professional image editing software to manually edit the image. For example, image A includes a large tree with dense leaves; the user can use image editing software (such as Photoshop) to remove the leaves from the large tree in image A and modify the large tree with dense leaves in image A to a large tree without leaves.
[0003] However, this image editing method requires the user to have professional image editing experience to modify the image, and the image editing method requires many manual operation steps, which prevents the user from quickly modifying the image using the electronic device and affects the user's experience of editing the image. SUMMARY
[0004] To solve the above technical problems, the present application provides an image processing method and an electronic device, which quickly and accurately convert the image style of an image to the image style indicated in the prompt information according to the style conversion prompt information input by the user, the conversion image style does not require manual adjustment of the image, reduces the manual operation steps, and increases the flexibility of the electronic device in converting the image style of the image.
[0005] In a first aspect, the present application provides an image processing method, including: obtaining a first image and first prompt information input by a user, the first prompt information including text information of a target style; inputting the first image and the first prompt information into a first model to obtain a reference image of the first image, the image style of the reference image being the same as the target style, the image style including at least one of color features and texture features of the image; inputting the first image and the reference image into a second model to obtain a target image output by the second model, the image style of the target image being the same as the image style of the reference image and the high-level semantic information of the target image being the same as that of the first image.
[0006] In this way, the first model adjusts the image style of the first image according to the first prompt information, and obtains a reference image of the first image, the reference image of the first image has a target style, and the reference image of the first image has a high probability of image distortion (i.e., high-level semantic information changes) compared with the first image. The electronic device takes the reference image of the first image as a reference image for image style transfer of the first image, so that the second model can transfer the image style of the reference image to the first image to generate a target image. The target image can maintain the high-level semantic information of the first image, and ensure that the target image does not have distortion problems. The electronic device can change the image style of the first image to any image style through the first model and the second model, so that the electronic device adjusts the image style in a more flexible manner, and does not need to manually adjust the image style, thereby increasing the application scenarios of the software for adjusting the image style.
[0007] According to the first aspect, before the first image and the first prompt information are input into the first model to obtain the reference image of the first image, the method further includes: obtaining a first training sample set, the first training sample set including a first sample image, a style sample prompt information, and a second sample image corresponding to the first sample image, the image style of the second sample image being the same as the image style indicated by the style sample prompt information; inputting the first sample image and the style sample prompt information into the first training model to obtain an image output by the first training model; and verifying the image output by the first training model using the second sample image to update parameters of the first training model until the first training model converges to generate the first model. In this way, the image style of the second sample image is the same as the image style indicated by the style sample prompt information, the first sample image and the second sample image in the first training sample set are paired training data, high-precision training data improves the success rate of the first model trained based on the first training model and the first training sample set to convert the image style of the first image, and further improves the accuracy of the second model to convert the image style of the first image.
[0008] According to the first aspect, before the first image and the reference image are input into the second model, the method further includes: obtaining a second training sample set, the second training sample set including a first sample image, a first sample reference image and a second sample image, the first sample reference image being determined based on the second sample image; inputting the first sample image and the first sample reference image into a second training model to obtain an image output by the second training model; and verifying the image output by the second training model using the second sample image to update parameters of the second training model until the second training model converges, thereby generating the second model. In this way, during training of the second model, the data input into the second training model is the first sample image and the reference image of the first sample image, so that the second training model can transfer the image style of the reference image to the first sample image, instead of converting the image style of the first image to a fixed image style, thereby improving flexibility of the second model in converting the image style of the first image.
[0009] According to the first aspect, before the first image and the reference image are input into the second model, the method further includes: detecting whether the adjustment category of the first prompt information belongs to a first category, the adjustment category of the first prompt information being a category indicated by the first prompt information to instruct the first model to adjust the image style of the first image; if it is detected that the adjustment category belongs to the first category, obtaining a first candidate model as the second model, the first category being a category indicating that the first model adjusts color features in the first image; and if it is detected that the target style belongs to a second category, obtaining a second candidate model as the second model, the second category including a category indicating that the first model adjusts color and texture features in the first image, and a category indicating that the first model adjusts texture features in the first image. In this way, the electronic device selects different candidate models as the second model according to the adjustment category of the first prompt information, so that the selected candidate model processes the first image more in line with the requirements of the adjustment category, and the quality of image processing by the second model can be improved.
[0010] According to the first aspect, detecting whether the adjustment category of the first prompt information belongs to the first category comprises: if it is detected that there is information indicating an image color in the first prompt information and there is no information indicating a time, it is determined that the adjustment category of the first prompt information belongs to the first category; and if it is detected that there is information indicating a time in the first prompt information, it is determined that the adjustment category of the first prompt information belongs to the second category. In this way, the adjustment of the color in the image does not affect the texture information in the image, and the electronic device can quickly determine whether the adjustment category of the first prompt information is the first category by detecting whether there is color information in the first prompt information; and the object usually changes with time, and when the adjustment of the image involves time conversion (such as seasonal conversion, day-night conversion, etc.), it usually means that the structure of the object in the image will change (such as a tree sprouting in spring and having no leaves in winter), and the electronic device can determine whether the adjustment category of the first prompt information is the second category based on the time information. This determination method is simple and fast.
[0011] According to the first aspect, before obtaining the category of the image style adjustment indicated by the first prompt information, the method further comprises: in the case that the second training model adopts a high dynamic range (HDR) network model, obtaining a second training sample set, the second training sample set comprising a first sample image, a first sample reference image, and a second sample image, the image style of the second sample image being the same as the image style indicated by the style sample prompt information, the first sample reference image being determined based on the second sample image; the second training model performing low-resolution coefficient prediction processing on the first sample image and the second sample image to obtain a processed first sample image and a processed second sample image; the second training model performing pixel-level network fusion processing on the first sample image to generate a guided image of the first sample image; the second training model performing matrix transformation processing on the guided image, the processed first sample image, and the processed second sample image; the second training model combining the image after the matrix transformation processing with the first sample image to output a second prediction image; and verifying the second prediction image by using the second sample image to update the parameters of the second training model until the second training model converges to generate a first candidate model. In this way, the second training model adopts an HDRnet, and the first sample reference image is introduced in the training method, so that the prediction image output can obtain the image style in the first sample reference image, and the HDRnet has a good effect on the adjustment of the color of the image.
[0012] According to the first aspect, the second training model comprises a weight prediction model and N basic three-dimensional lookup tables, N is an integer greater than 1; before obtaining the category of image style adjustment indicated by the first prompt information, the method further comprises: obtaining a second training sample set, the second training sample set comprising a first sample image, a first sample reference image and a second sample image, the image style of the second sample image being the same as the image style indicated by the style sample prompt information, the first sample reference image being determined based on the second sample image; inputting the first sample image and the first sample reference image into the weight prediction model to obtain N weights corresponding to the N basic three-dimensional lookup tables output by the weight prediction model; fusing the N basic three-dimensional lookup tables according to the N weights corresponding to the N basic three-dimensional lookup tables to generate a reference style three-dimensional lookup table, the reference style three-dimensional lookup table being a color gamut mapping of a target style; transforming the image style of the first sample image by using the reference style three-dimensional lookup table to obtain a second prediction image; verifying the second prediction image by using the second sample image to iteratively update the parameters of the weight prediction model and the N basic three-dimensional lookup tables until the second training model converges to generate a first candidate model. In this way, during the process of the second training model, the electronic device continuously adjusts the N basic three-dimensional lookup tables and adjusts the weights corresponding to the N basic three-dimensional lookup tables by adjusting the parameters in the weight prediction model, so that the reference style three-dimensional lookup table obtained by fusing the N basic three-dimensional lookup tables. The electronic device can quickly and accurately convert the image style of the input image to the style in the first sample reference image through the reference style three-dimensional lookup table.
[0013] According to the first aspect, before obtaining the category of image style adjustment indicated by the first prompt information, the method further comprises: in the case that the second training model is a restoration transformer model, obtaining a second training sample set, the second training sample set comprising a first sample image, a first sample reference image and a second sample image, the image style of the second sample image being the same as the image style indicated by the style sample prompt information, the first sample reference image being determined based on the second sample image; inputting the first sample image and the first sample reference image into the second training model to obtain a second prediction image output by the second training model; verifying the second prediction image by using the second sample image to update the parameters of the second training model until the second training model converges to generate a second candidate model. In this way, the restoration transformer model has a good effect on texture adjustment of an image, and the model can better adjust the texture features of the first image.
[0014] According to the first aspect, the method further comprises: compressing the second sample image according to a preset size to generate the first sample reference image. In this way, the first sample reference image has a small image size, which reduces the complexity of training the second training model, and further reduces the magnitude of the trained second model, facilitating the deployment of the second model in a mobile terminal.
[0015] According to the first aspect, inputting the first image and the first prompt information into the first model to obtain a reference image of the first image comprises: inputting the first image and the first prompt information into the first model to obtain a first predicted image output by the first model, the image style of the first predicted image being the same as the target style; and compressing the first predicted image according to a preset size to generate the reference image of the first image. In this way, the first predicted image output by the first model is compressed, and the compressed image is used as the reference image of the first image, which can reduce the processing speed of the second model for the compressed image.
[0016] According to the first aspect, obtaining the first prompt information input by the user comprises: displaying the first image; displaying an input box in response to an operation of adjusting the image style input by the user; and obtaining prompt information input by the user in the input box as the first prompt information. In this way, after the user inputs the operation of adjusting the image style, the input box is displayed, so that the user can be prompted to input the prompt information for adjusting the image style. This way allows the user to input any information for adjusting the image style, and the electronic device can perform the image style adjustment operation on the first image according to the input information for adjusting the image style.
[0017] According to the first aspect, obtaining the first prompt information input by the user comprises: displaying the first image; displaying at least two image style controls in response to an operation of adjusting the image style input by the user, each image style control having prompt information of a corresponding image style; obtaining prompt information of a selected image style in response to a selection operation of selecting the image style from the at least two image style controls; and taking the obtained prompt information of the image style as the first prompt information. In this way, the user does not need to manually input the prompt information of the image style, reducing the operation of the user on the electronic device.
[0018] The second aspect provides an electronic device, comprising: one or more processors; a memory; and one or more computer programs, wherein the one or more computer programs are stored in the memory, and when the computer programs are executed by the one or more processors, the electronic device performs the method of image processing of the first aspect and any implementation manner of the first aspect.
[0019] The second aspect and any implementation manner of the second aspect correspond to the first aspect and any implementation manner of the first aspect, respectively. The technical effects corresponding to the second aspect and any implementation manner of the second aspect can be referred to the technical effects corresponding to the first aspect and any implementation manner of the first aspect, which will not be described here.
[0020] In a third aspect, the present application provides a computer readable medium for storing a computer program, when the computer program is run on an electronic device, the electronic device is caused to perform the image processing method of the first aspect and any one of the implementation manners of the first aspect. BRIEF DESCRIPTION OF DRAWINGS
[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed to be used in the description of the embodiments of the present application will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0022] Figure 1 is an exemplary structural schematic diagram of an electronic device;
[0023] Figure 2 is an exemplary software structural block diagram of an electronic device;
[0024] Figure 3 is an exemplary flowchart of an image processing method;
[0025] Figure 4 is an exemplary schematic diagram of adjusting the image style of a first image;
[0026] Figure 5 is an exemplary image of a large tree in different seasons;
[0027] Figure 6 is an exemplary first image, target image and reference image;
[0028] Figure 7a is an exemplary schematic diagram of a second model training;
[0029] Figure 7b is an exemplary schematic diagram of another second model training;
[0030] Figure 8 is an exemplary schematic diagram of a second candidate model training;
[0031] Figure 9 is an exemplary flowchart of an image processing method;
[0032] Figure 10 is an exemplary schematic diagram of adjusting the image style of an image;
[0033] Figure 11 is an exemplary schematic diagram of adjusting the image style of an image. DETAILED DESCRIPTION
[0034] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of the present application.
[0035] The term "and / or" in the present application is only used to describe the association relationship of the associated objects, and means that there can be three relationships, for example, A and / or B can mean that A exists alone, A and B exist together, and B exists alone.
[0036] The terms "first" and "second" and the like in the description and claims of the embodiments of the present application are used to distinguish different objects, rather than to describe a specific order of the objects. For example, the first target object and the second target object are used to distinguish different target objects, rather than to describe a specific order of the target objects.
[0037] In the embodiments of the present application, the words "exemplary" or "for example" are used to mean serving as an example, instance, or illustration. Any embodiment or design presented as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or advantageous than other embodiments or design solutions. Rather, the use of "exemplary" or "for example" is intended to present relevant concepts in a specific way.
[0038] In the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality of" is two or more. For example, a plurality of processing units means two or more processing units; a plurality of systems means two or more systems.
[0039] In some embodiments, a user can edit an acquired image through an image style editing module in an electronic device, so that the edited image meets the user's needs.
[0040] The electronic device can be a mobile phone, a computer, a digital camera, or the like. In this example, the electronic device is taken as a mobile phone for illustration. The mobile phone is installed with an image style editing module, which can adjust color features such as saturation and sharpness of an image. A user can adjust the style of an image through the image style editing module, such as Photoshop. For example, the color of the image style of an original image is a cold tone (such as winter), and the user can change the season in the image (such as changing to spring) by using Photoshop. However, this way requires the user to be proficient in using the Photoshop software, and is not suitable for every user, such as a user who does not know how to use Photoshop. That is, the electronic device cannot change the image style through a simple operation, resulting in that the electronic device is not flexible in adjusting the image style of an image and reduces the application scenarios of the electronic device in changing the image style of an image.
[0041] In an embodiment of the present application, a method for image processing is provided, which is executed by an electronic device. The electronic device can be a mobile phone, a computer, a tablet computer, a PAD, or the like. The electronic device is installed with an image style editing module. Optionally, the image style editing module can be embedded in a gallery application. For example, a user can click the gallery in the mobile phone and click a style adjustment icon in the gallery application. The mobile phone starts the image style editing module in response to the operation of the user clicking the style adjustment. Optionally, the image style editing module can also be a separate application software.
[0042] The method for image processing includes that the electronic device acquires a first image and first prompt information input by a user, the first prompt information being text information of a target style; inputs the first image and the first prompt information into a first model to obtain a reference image of the first image, the image style of the reference image being the same as the target style; inputs the first image and the reference image into a second model to obtain a target image output by the second model, the image style of the target image being the same as the image style of the reference image and the high-level semantic information of the target image being the same as the high-level semantic information of the first image.
[0043] In this example, the electronic device can adjust the image style of the first image to be any image style through the first model and the second model. Meanwhile, the high-level semantic information of the adjusted image is the same as the high-level semantic information of the original image, that is, the method for image processing of the present application can not only ensure the arbitrary conversion of the image style of the first image, but also ensure the quality of the converted image, such as not reducing the pixels and not causing distortion of objects in the image.
[0044] Figure 1 A structural schematic diagram of an electronic device 100 is shown in an embodiment of the present application. It should be understood that, Figure 1The electronic device 100 shown is only one example of an electronic device, and the electronic device 100 can have more or fewer components than shown, can combine two or more components, or can have a different configuration of components. Figure 1 The various components shown in FIG. 1 can be implemented in hardware, software, or a combination of both hardware and software.
[0045] The electronic device 100 can include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headset jack 170D, a sensor module 180, a key 190, a motor 191, an indicator 192, a camera 193, a display 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 can include a pressure sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a distance sensor, a proximity light sensor, a fingerprint sensor, a temperature sensor, a touch sensor, an ambient light sensor, a bone conduction sensor, etc.
[0046] Figure 2 is a software structure block diagram of the electronic device 100 of the embodiments of the present application.
[0047] The layered architecture of the electronic device 100 divides software into several layers, each of which has a clear role and division of labor. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into three layers, from top to bottom, the application layer, the application framework layer, and the kernel layer. It can be understood that, Figure 2 The layers in the software structure and the components contained in each layer do not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 can include more or fewer layers than shown, and each layer can include more or fewer components, which are not limited in the present application.
[0048] As Figure 2As shown, the application layer can include a series of application packages. The application packages can include gallery, social application, camera, calendar, short message, etc. The application framework layer provides application programming interface (API) and programming framework for the application of the application layer. The application framework layer includes some pre-defined functions. The image style editing module in the present application can be embedded in the gallery, social application, camera application. For example, before displaying the image in the display interface of the social application, the electronic device starts the image style editing module in response to the operation of the user starting the image style editing module in the social application, and the electronic device performs any conversion on the image style of the image to be displayed through the image style editing module, and displays the image with the converted image style in the display interface.
[0049] As shown, the application framework layer can include window manager, resource manager, content provider, view system, phone manager, notification manager, etc. Figure 2
[0050] The window manager is used to manage window programs. The window manager can obtain the size of the display screen, determine whether there is a status bar, lock the screen, and intercept the screen, etc.
[0051] The resource manager provides various resources for the application, such as localized strings, icons, pictures, layout files, video files, etc.
[0052] The content provider is used to store and obtain data, and make the data accessible to the application. The data can include video, image, audio, dialed and received phone, browsing history and bookmark, phone book, etc.
[0053] The view system includes visual controls, such as controls for displaying text, controls for displaying pictures, etc. The view system can be used to build an application. The display interface can be composed of one or more views. For example, the display interface including the short message notification icon can include a view for displaying text and a view for displaying pictures.
[0054] The phone manager is used to provide the communication function of the electronic device 100. For example, the management of the call state (including connection, hang up, etc.).
[0055] The notification manager enables the application to display notification information in the status bar, which can be used to convey the type of message, which can automatically disappear after a short stay, without user interaction. For example, the notification manager is used to inform the completion of the download, message reminder, etc.
[0056] The kernel layer is the layer between hardware and software. The kernel layer at least includes display driver, camera driver, audio driver, sensor driver.
[0057] Before describing the method of image processing in this application, some technical terms involved in this application are described.
[0058] (1) Semantic information of image: the content of image is the semantic information of image. The semantic information of image includes low-level semantic information and high-level semantic information. The low-level semantic information of image includes image style. The high-level semantic information of image refers to the closest thing expressed by the image to human understanding. The high-level semantic information of image is what people can see, for example, feature extraction is performed on a human face image, the low-level semantic information extracted includes the outline of the face, nose, eyes, etc., and the high-level semantic information extracted is the human face.
[0059] (2) Image style: image style can include texture features of image, artistic forms of image. Image style can include day style, night style, spring style, summer style, autumn style, winter style, sunny style, rainy style, meteor style, Hirakawa Hideaki style, etc. Image style can also include the style of famous paintings, such as Van Gogh's Sunflowers and Starry Night, and can also include the creation type of image, such as Chinese style, cubism, modernism, surrealism, cartoon, comic, oil painting, watercolor, ink, etc. The image style of the embodiments of the present application is not limited.
[0060] Figure 3 For the flowchart of the method of image processing shown by way of example, in this example, the electronic device is exemplified by a mobile phone. The process of image processing specifically includes:
[0061] Step 301: The electronic device acquires a first image and first prompt information input by a user.
[0062] Exemplarily, a mobile phone (such as mobile phone A) acquires an image with an image style to be adjusted as a first image. The first image can be an image sent by another device (such as mobile phone B), or an image photographed by mobile phone A.
[0063] The process of the mobile phone acquiring the first image and the first prompt information input by the user can include: in response to the operation of the user adjusting the image style of the image to be edited (such as Figure 10As shown in FIG. 10b, when the user clicks the icon of "style adjustment", the image style editing module determines that the current image to be edited is the first image, and an input box can be displayed on the display interface. The input box is used to obtain the information of the target style input by the user. Optionally, before the user inputs the text information into the input box, the text input prompt information can be displayed in the input box, for example, the text input prompt information can be "please input the information of the image style to be converted!". When the user clicks the input box, the displayed text input prompt information is eliminated, and a soft keyboard is popped up on the interface. When the image style editing module detects that the user has completed the input of the information in the input box, the information in the input box is obtained as the first prompt information, for example, the user inputs "convert to spring" in the input box. The first prompt information is "convert to spring".
[0064] Step 302: The electronic device determines the reference image of the first image according to the first image, the first prompt information and the first model.
[0065] Exemplarily, the image style editing module in the mobile phone includes a trained first model and a second model. The first model is used to output a first predicted image according to an original image and a style conversion prompt information input by the user, and the image style of the first predicted image is consistent with the target style in the style conversion prompt information.
[0066] The process of training the first model will be described below in detail. Figure 5 and Figure 6 The process of training the first model will be described below in detail.
[0067] A first training model is preset, and the network structure of the first training model can adopt the network structure of the conditional diffusion model (instructPix2Pix). A first training sample set is a data set used to train the first training model. The first training sample set includes a first sample image, a style sample prompt information and a second sample image corresponding to the first sample image, and the image style of the second sample image is the same as the image style indicated by the style sample prompt information. The first sample image and the second sample image are a pair of images.
[0068] In this example, the electronic device can collect a large number of style sample prompt information and first sample images, and the second sample image can be synthesized by a person according to the target style by using photoshop, so as to ensure that the high-level semantic information of the generated second sample image and the first sample image remains unchanged.
[0069] For example, the first sample image is Figure 5The large tree image shown in 5a is edited according to the style sample prompt information "converted into summer" to generate an image as shown in 5b. The electronic device can use the large tree image shown in 5a as the second sample image. Similarly, the large tree image shown in 5c and the large tree image shown in 5d can be synthesized according to "converted into autumn" and "converted into winter". The electronic device uses the large tree image shown in 5c and the large tree image shown in 5d as the second sample image. Referring to Figure 5 The main subject of the four images shown in 5a-5d is a large tree, and the number and shape of the branches of the large tree do not change. The leaves of the large tree have different shapes and numbers in different seasons, that is, the high-level semantic information of the large tree images shown in 5a-5d is the same. The pair of sample images can improve the accuracy of the predicted image output by the first model.
[0070] The style sample prompt information in the first training sample set can be text information with a target style, for example: background with meteor, make the tree sprout, make the tree shed leaves, convert into a painting style, make the photo look brighter, etc.
[0071] The first training model includes a text encoder and an image encoder. The text encoder can extract text features from the style sample prompt information, and the image encoder can extract image features from the first sample image. During training, the first training model fuses the text features and the image features to output a predicted image with a style prompt information indicated image style. The electronic device uses the second sample image to verify the image output by the first training model to update the parameters of the first training model until the first training model converges to generate the first model.
[0072] In this example, the image style editing module installs the trained first model. When the image style editing module obtains the first image and the first prompt information, the first image and the first prompt information are input into the first model to obtain a reference image of the first image. The first model outputs a first predicted image, and the image style of the first predicted image is the same as the target style. The main subject in the first predicted image may be distorted, incomplete, or have other problems.
[0073] For example, as shown in 6a in Figure 6 The building image shown in 6a is the first image, and the building image shown in 6b is the target image. The first prompt information is "background with meteor", and the image style editing module inputs the first image and the first prompt information into the first model to obtain a first predicted image output by the first model, as shown in 6c. Figure 6 The building in the first predicted image is distorted, and the image quality of the first predicted image is lower than that of the images shown in 6a and 6c.
[0074] The image quality of the first predicted image output by the first model is poor, but the image style of the first predicted image is consistent with the target style, so the image style editing module can take the first predicted image as the reference image of the first image.
[0075] Optionally, the image style editing module can compress the first predicted image (with a resolution of 1920*1080 pixels) according to a preset size to obtain a thumbnail image (such as a resolution of 200*200 pixels) of the first predicted image. The image style editing module takes the thumbnail image of the first predicted image as the reference image of the first image.
[0076] Step 303: The electronic device inputs the first image and the reference image into the second model to obtain a target image output by the second model, the image style of the target image being the same as that of the reference image and the high-level semantic information of the target image being the same as that of the first image.
[0077] Illustratively, the image style editing module is pre-installed with the second model. The second model is used to output a target image according to an input original image and a reference image of the original image, the image style of the target image being the same as that of the reference image and the high-level semantic information of the target image being the same as that of the first image.
[0078] The following will be described in combination with Figure 7a or Figure 7b The process of training the second model will be specifically described.
[0079] The electronic device obtains a second training set, the second training sample set including a first sample image, a first sample reference image and a second sample image, the first sample reference image being determined based on the second sample image. For example, Figure 7a As shown in Figure 7a , the original image inis the first sample image, the image after style transformation is the second sample image, and the first sample image and the second sample image are paired images. The second sample image is the same image as the second sample image used to train the first model.
[0080] The electronic device pre-sets a second training model, and the network structure of the second training model can adopt a high dynamic range network structure (High Dynamic Range Imaging Network, “HDRnet”). Figure 7a For the schematic diagram of training the second training model, as shown in Figure 7a , the electronic device inputs the first sample image and the second sample image into the second training model, and the second training model inputs the first sample image (i.e. the original image in ) and the second sample image (i.e. the image after style transformation in ) into the second training model. Figure 7a Figure 7aThe first processing is performed on the image and the style-transformed image shown in the style shown in the figure to obtain a low-resolution image 1 and an image 2. The second training model performs low-resolution coefficient prediction processing on the image 1 and the image 2 to obtain a processed first sample image (denoted as Feature Ori 1) and a processed second sample image (denoted as Feature Tar 2). The low-resolution coefficient prediction processing is used to perform bilateral filtering on the image to obtain the image features of the input image.
[0081] The second training model simultaneously performs pixel-level network fusion (Pixel-wise-network) processing on the first sample image to generate a guide image of the first sample image. The guide image is used to obtain the interaction information of the first sample image. The second training model filters (i.e., matrix transformation) the guide image of the first sample image, the first sample image (Feature Tar 1) and the second sample image (Feature Tar 2) after low-resolution coefficient prediction processing according to a preset manner to obtain a matrix-transformed image. The second training model combines the matrix-transformed image with the first sample image to generate a second prediction sample image. The electronic device obtains the difference between the second prediction sample image and the second sample image to update the parameters in the second training model until the second training model converges to generate a second model.
[0082] In some embodiments, the second training model can be based on three-dimensional look-up tables (3D Look-Up Tables, “3DLUTs”). Figure 7b A schematic diagram of the second training model trained based on 3D LUTs is shown for illustration. The second training set used to train the second training model includes a first sample image, a first sample reference image and a second sample image. The first sample image is the original image in Figure 7b , the second sample image is the style-transformed image corresponding to the original image, and the first sample reference image is a thumbnail of the style-transformed image (i.e. Figure 7b , which has a reference style). The image style of the second sample image is the same as that of the first sample reference image. In this example, the image style of the first sample reference image is referred to as the reference style.
[0083] As Figure 7bAs shown, the preset second training model can include a preset weight prediction model and N basis 3D LUTs (i.e., Basis 3D LUTs). The N basis 3D LUTs are different N basis 3D LUTs, N is an integer greater than 1, for example, N can be 3, 4, 5, etc., and in this example, N is taken as 3 for illustration, and the electronic device can pre-set three basis 3D LUTs when starting to train the second model. The weight prediction model can be a convolutional neural network (CNN), wherein the input data of the weight prediction model is the first sample image and the first sample reference image, and the output is the weight corresponding to the N basis 3D LUTs.
[0084] The electronic device inputs the first sample image and the first sample reference image into the weight prediction model, obtains three weights output by the weight prediction model, and the three weights are w1, w2 and w3, wherein w1 is the weight corresponding to the first basis 3D LUT, w2 is the weight corresponding to the second basis 3D LUT, and w3 is the weight corresponding to the third basis 3D LUT. The electronic device fuses the respective weights of the three basis 3D LUTs in the middle to obtain a 3D LUT based on the reference style (hereinafter referred to as a reference style 3D LUT). The electronic device uses the reference style 3D LUT to transform the original image to generate a style-transformed image (i.e., a second predicted sample image). The electronic device obtains the difference between the second predicted sample image and the second sample image to iteratively update the parameters in the weight prediction model and the three basis 3D LUTs until the second training model converges to generate a second model. Figure 7b
[0085] The image style editing module installs the trained second model, and after the electronic device obtains the first image and the reference image of the first image, inputs the first image and the reference image of the first image into the second model to obtain a target image output by the second model. The image style of the target image is the same as the image style of the reference image of the first image, and the high-level semantic information of the target image is the same as that of the first image.
[0086] Figure 4 A schematic diagram of the image style editing module processing the first image and the first prompt information is shown in FIG. 1. As shown in FIG. 1, the image style editing module includes a trained first model and a second model. Figure 4 As shown in FIG. 1, the image style editing module includes a trained first model and a second model. The image style editing module inputs the first image (such as the image 101 in FIG. 1) and the first prompt information (such as the image 102 in FIG. 1) into the first model to obtain a second image (such as the image 103 in FIG. 1) output by the first model. Figure 4 The image style editing module inputs the first image and the first prompt information (such as improving the brightness of the image) into the first model, and the first model outputs a first predicted image. The image style editing module compresses the first predicted image to obtain a reference image. The person in the reference image is distorted relative to the person in the original image, and the image style of the reference image is the same as the target style in the first prompt information. The image style editing module inputs the reference image and the original image into the second model, and the second model outputs a target image. The image style of the target image is consistent with the image style of the reference image, and the person in the target image is consistent with the person in the original image and is not distorted. The image style editing module realizes the purpose of arbitrarily converting the image style of the first image through the first model and the second model.
[0087] In the example, the electronic device can obtain a reference image of the first image through the first model. The image style of the reference image of the first image is the same as the target style in the first prompt information. When the first model adjusts the image style according to the first prompt information, the adjusted image may be distorted, and the image quality of the reference image of the first image is poor. Taking the reference image of the first image as the reference image of the second model adjusting the image style of the first image can migrate the image style of the reference image of the first image to the first image while keeping the high-level semantic information in the first image unchanged, so that the image quality of the target image output by the second model is high. In this way, the electronic device can change the image style of the first image through the first prompt information input by the user, simplifying the step of manually editing the image and improving the efficiency of the electronic device in processing the first image.
[0088] In one embodiment, the electronic device can pre-store a first candidate model and a second candidate model. The electronic device can select one of the first candidate model and the second candidate model as the second model according to the first prompt information input by the user. The flow of the image processing is as shown in Figure 9
[0089] Step 901: The electronic device obtains a first image and a first prompt information input by a user.
[0090] This step is similar to the related description in step 301, and will not be repeated here.
[0091] Step 902: The electronic device detects whether the adjustment category is the first category. If it is detected that it is, step 903 is performed, and if it is detected that it is not, step 904 is performed.
[0092] Exemplarily, the adjustment category of the first prompt information is a category indicating that the first model adjusts the image style of the first image, and the adjustment category includes a first category and a second category. The first category is a category indicating that the first model adjusts the color feature of the image. For example, the first prompt information is “adjust the color to be bright”, the first prompt information indicates that the first model adjusts the color feature of the first image, and the adjustment category of the first prompt information is the first category. For another example, “from fresh blue to autumn”, fresh blue indicates that the color tone of the image is blue, and autumn indicates that the color tone of the image is orange. That is, the adjustment category of the first prompt information is the first category.
[0093] The second category includes a category indicating that the first model adjusts the texture feature of the image and a category adjusting the texture and color features. For example, the first prompt information is “convert to winter”, the current image is taken in spring, the scene changes over time, that is, the texture of the image changes, and the adjustment category of the first prompt information is the second category.
[0094] In one example, since the first category is a category in which the first model adjusts the color feature of the image, the electronic device can detect whether the first prompt information carries information indicating the color and does not carry information indicating the time, and if it is detected that the first prompt information carries information indicating the color and does not carry information indicating the time, it is determined that the first prompt information indicates that the adjustment category of the first model adjusting the image style of the first image is the first category. For example, the first prompt information is “adjust the color to be bright”, the electronic device detects “bright”, which is information indicating the color, and determines that the first prompt information indicates the first category.
[0095] The electronic device can determine that the first prompt information indicates that the adjustment category of the first model adjusting the image style of the first image is the second category by detecting whether there is information indicating the time transformation in the first prompt information. For example, “convert to night”, the electronic device detects that there is night in the first prompt information, and determines that the adjustment category indicated by the first prompt information is the second category.
[0096] When the electronic device determines that the adjustment category of the first prompt information is the first category, step 903 is performed, and if it is detected that the adjustment category of the first prompt information is the second category, step 904 is performed.
[0097] Step 903: The electronic device obtains a first candidate model as a second model. After this step, step 905 is performed.
[0098] Exemplarily, the image style editing module can store the first candidate model and the second candidate model after training. The training model of the first candidate model can adopt the structure of HDRnet. The training process of the second candidate model can refer toFigure 7a or Figure 7b The related descriptions in the above steps are not repeated here.
[0099] Step 904: The electronic device obtains a second candidate model as the second model. After this step, step 905 is performed.
[0100] Exemplarily, the training model of the second candidate model can adopt a restoration transformer (Restormer) model. The electronic device obtains a second training sample set, which includes the first sample image, the first sample reference image, and the second sample image, and the first sample reference image is determined based on the second sample image. In this example, the first sample reference image is a thumbnail image of the second sample image, and the resolution of the thumbnail image is a preset resolution, such as 360*360 pixels.
[0101] Figure 8 For the process of training the second candidate model, as shown in Figure 8 , the first sample image (i.e., the original image in Figure 8 ) and the first sample reference image (i.e., the thumbnail image in Figure 8 ) are input into the second training model, a second predicted sample image output by the second training model is obtained, the electronic device obtains a difference between the second predicted sample image and the second sample image, and iteratively trains the second training model according to the obtained difference until the second training model converges, thereby obtaining the second candidate model.
[0102] When the image style editing module detects that the adjustment category of the first prompt information is the second category, the second candidate model is obtained as the second model. When the target style adjusted by the second model involves time transformation, the texture features in the image need to be adjusted, or the texture features and color features in the image need to be adjusted. The Restormer model is adopted, so that the predicted image output by the second model is more in line with the actual needs of the user, and the quality of the output image is better.
[0103] Step 905: The electronic device determines a reference image of the first image according to the first image, the first prompt information, and the first model.
[0104] This step is substantially the same as step 302, and the specific process can refer to the related descriptions in step 302.
[0105] Step 906: The electronic device inputs the first image and the reference image into the second model to obtain a target image output by the second model.
[0106] This step is substantially the same as step 303, and the specific process can refer to the related descriptions in step 302.
[0107] Figure 10 This is an illustrative diagram showing a user adjusting the image style using the image style editing module. In this example, the image style editing module stores a first model and a second model. The second model is a fixed model.
[0108] In this example, user A launches the gallery app on their phone, displaying the gallery interface. This gallery app includes an embedded image style editing module. The user then interacts with the gallery interface... Figure 10 Select an image (e.g., click on a thumbnail image in the gallery) from the image selection menu (not shown in the image). Figure 10 As shown in 10a, in response to the user's selection, the image library displays the image 1002 in interface 1001. Interface 1001 includes an edit icon 1003, a delete icon, etc. In response to the user clicking the edit icon 1003, the image library jumps to the image editing interface 1004, as shown... Figure 10 As shown in 10b, the image editing interface 1004 displays the image 1002 to be edited. The image editing interface 1004 displays a style adjustment icon 1005 and a filter icon. The image editing interface may also include icons for functions such as one-click beautification and image cropping. In response to the user clicking the style adjustment icon 1005, the gallery can jump to the style adjustment interface 1006. The style adjustment interface 1006 displays the image 1002 (i.e., the first image) whose style needs to be adjusted. The style adjustment interface 1006 also displays an input box 1007 with a first prompt message. Before the user enters information, the input box 1007 may display an input prompt message such as, "Please enter the text information of the image's target style!", to prompt the user to enter the first prompt message.
[0109] like Figure 10 As shown in 10c, the user enters the first prompt message "Convert to winter" in input box 1007. The image style editing module responds to the user clicking the "OK" control 1008, acquiring image 1002 and the first prompt message. The image style editing module inputs image 1002 and the first prompt message into the first model, acquiring the first predicted image output by the first model. The image style of the first predicted image is winter (i.e., trees without leaves). The image style editing module compresses the first predicted image (resolution 1920*1080 pixels before compression) according to a preset size, obtaining the compressed first predicted image, which is used as a reference image for the first image. The resolution of the reference image is 200*200 pixels. The image style editing module inputs the first image and the reference image into the second model, acquiring the target image output by the second model, and instructs the display interface to display the target image. Figure 10As shown in 10d, the target image 1010 is displayed in the interface 1009. The tree in the target image 1010 has no leaves, and the image was taken in winter.
[0110] This example uses the first and second models to change the texture of the trees in the image (i.e., transforming trees with leaves into trees without leaves). This image style adjustment method adjusts the image texture while ensuring that the adjusted image is not distorted, retaining the pixels of the original image, thus improving the display effect of the transformed image.
[0111] In other examples, the electronic device may pre-store multiple prompts for changing image styles, and in response to the user's selection of a prompt for changing image styles, the electronic device may obtain the user-selected prompt for changing image styles as the first prompt.
[0112] For example, the image style editing module stores a first model, a first candidate model, and a second candidate model. User B launches the gallery app on their phone, displaying the gallery interface. This gallery app embeds the image style editing module. The user then uses the gallery interface... Figure 11 When an image is selected (e.g., by clicking a thumbnail image in the gallery), the gallery responds to the user's selection by displaying an interface (which can be referenced in the image gallery). Figure 10 The image is displayed in interface 1001 (10a) of the image library. This display interface includes edit icons, delete icons, etc. In response to the user clicking the edit icon, the image library redirects to the image editing interface 1101, as shown below. Figure 11 As shown in 11a, the image editing interface 1101 includes: an image 1102 to be edited and startup icons for various image editing functions. These startup icons include: style adjustment icons, filter icons, cropping icons, and icons for adding mosaic, etc. In this example, as shown... Figure 11 As shown in 11a, the launch icons for image editing functions include: style adjustment icon 1103 and a filter icon; other launch icons are not shown in 11a. As shown in 11a, when a user clicks the style adjustment icon 1103, the user is redirected to the style adjustment interface 1104 in response to the click. Figure 11 As shown in 11b, the style adjustment interface 1104 displays image 1102 (i.e., the first image) to be styled. Figure 11As shown in FIG. 11B, the interface 1106 displays a plurality of pieces of prompt information for converting the image style in the prompt information display area 1105 of the style adjustment interface. The prompt information includes, for example, meteor, cloudy, bright, night, and the like. The user clicks the “meteor” control 11051, obtains the image 1102, and obtains the prompt information corresponding to the “meteor” control 11051 as the first prompt information. In this example, the prompt information corresponding to the “meteor” control 11051 is “convert the sky to a meteor”.
[0113] The image style editing module inputs the image 1102 and the first prompt information into the first model, obtains a first predicted image output by the first model, and the image style of the first predicted image is that the sky is a meteor. The image style editing module compresses the first predicted image (the resolution before compression is 1920*1080 pixels) according to a preset size, obtains a compressed first predicted image, and uses the compressed first predicted image as the reference image of the first image. The resolution of the reference image of the first image is 200*200 pixels. The image style editing module inputs the first image and the reference image of the first image into the second model, obtains a target image output by the second model, and instructs to display the target image on the display interface. The image 1102 is the first image, the time of the image in the first image is daytime, and the corresponding brightness value is d1. The first prompt information is “convert the sky to a meteor”. The conversion condition indicated in the first prompt information includes that the time of the image is night, and the sky of the night has a meteor. As shown in FIG. 11C, the interface 1106 displays the target image 1107. The corresponding brightness of the converted image 1107 of the image 1102 is e1, and e1 Figure 11
[0114] In one example, the user clicks the “bright” control 11052, and the image style editing module obtains the image 1102 as the first image and obtains the prompt information corresponding to the “bright” control 11052 as the first prompt information in response to the click operation of the user. In this example, the prompt information corresponding to the “bright” control is “increase the brightness value of the image to a preset value”. The image style editing module detects whether there is information indicating color transformation in the first prompt information and detects whether there is information indicating time transformation in the first prompt information. In this example, the image style editing module detects that there is information indicating color (such as brightness value) in the first prompt information and there is no time information, and obtains a first candidate model (such as a model trained based on HDRnet or a model trained based on 3D LUTs) as the second model.
[0115] The image style editing module inputs the image 1102 and the first prompt information into the first model, obtains a first predicted image output by the first model, and the image style of the first predicted image is: the luminance value is improved. The image editing software compresses the first predicted image (the resolution before compression is 1920*1080 pixels) according to a preset size, obtains a compressed first predicted image, and takes the compressed first predicted image as a reference image of the first image, and the resolution of the reference image of the first image is 200*200 pixels. The image style editing module inputs the first image and the reference image of the first image into the second model, obtains a target image output by the second model, and instructs to display the target image 1110 on the display interface. The image 1102 is the first image, and the luminance value corresponding to the image in the first image is d1. The first prompt information is "improve the luminance value of the image to a preset value". The conversion condition indicated in the first prompt information includes: increasing the luminance value to the preset value. As shown in 11d of FIG. 11, the target image 1110 is displayed in the interface 1109. The luminance of the converted image 1110 is c1, and d1<c1. It can be seen that the luminance of the image 1110 is higher than that of the image 1102. Other information in the image is not changed. Figure 11
[0116] It can be understood that, in order to implement the above functions, the electronic device comprises hardware and / or software modules corresponding to the respective functions. The algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in hardware or a combination of hardware and computer software. Whether a certain function is implemented in hardware or computer software driven hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application in conjunction with the embodiments, but such implementation should not be considered beyond the scope of the present application.
[0117] The embodiment also provides a computer storage medium, which stores computer instructions, and when the computer instructions run on an electronic device, the electronic device executes the related method steps to implement the image processing method in the above embodiment. The storage medium includes: a U disk, a mobile hard disk, a read only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various storage program codes.
[0118] The embodiment also provides a computer program product, which, when running on a computer, causes the computer to execute the related steps to implement the image processing method in the above embodiment.
[0119] The electronic device, the computer storage medium, the computer program product or the chip provided in the embodiment are used for executing the corresponding method provided in the above, and thus the beneficial effects achieved by the electronic device, the computer storage medium, the computer program product or the chip can refer to the beneficial effects of the corresponding method provided in the above, which will not be described here again.
[0120] Any content of each embodiment of the present application, and any content of the same embodiment, can be freely combined. Any combination of the above is within the scope of the present application.
[0121] The embodiments of the present application are described above in combination with the drawings, but the present application is not limited to the specific embodiments described above, which are only illustrative but not restrictive. Those skilled in the art can make many forms under the inspiration of the present application without departing from the scope of the present application and the scope protected by the claims.
Claims
1. A method of image processing, characterized by, The method comprises: obtaining a first image and first prompt information of a user input, the first prompt information comprising text information of a target style; inputting the first image and the first prompt information into a first model to obtain a reference image of the first image, the reference image having an image style same as the target style, the image style comprising at least one of color features and texture features of an image; inputting the first image and the reference image into a second model to obtain a target image output by the second model, the target image having an image style same as that of the reference image and having high-level semantic information same as that of the first image; before inputting the first image and the reference image into the second model, the method further comprises: if it is detected that the first prompt information contains information indicating an image color and does not contain information indicating a time, it is determined that an adjustment category of the first prompt information belongs to a first category; if it is detected that the first prompt information contains information indicating a time, it is determined that the adjustment category of the first prompt information belongs to a second category; the adjustment category of the first prompt information is a category indicated by the first prompt information for the first model to adjust an image style of the first image; if it is detected that the adjustment category belongs to the first category, a first candidate model is obtained as the second model, the first category being a category indicating that the first model adjusts color features in the first image; if it is detected that the adjustment category belongs to the second category, a second candidate model is obtained as the second model, the second category comprising a category indicating that the first model adjusts color and texture features in the first image, and a category indicating that the first model adjusts texture features in the first image.
2. The method of claim 1, wherein, before inputting the first image and the first prompt information into the first model to obtain the reference image of the first image, the method further comprises: obtaining a first training sample set, the first training sample set comprising a first sample image, style sample prompt information, and a second sample image corresponding to the first sample image, the second sample image having an image style same as an image style indicated by the style sample prompt information; inputting the first sample image and the style sample prompt information into the first training model to obtain an image output by the first training model; verifying the image output by the first training model by using the second sample image to update parameters of the first training model until the first training model converges, and generating the first model.
3. The method of claim 2, wherein, before inputting the first image and the reference image into the second model, the method further comprises: obtaining a second training sample set, the second training sample set comprising the first sample image, a first sample reference image, and the second sample image, the first sample reference image being determined based on the second sample image; inputting the first sample image and the first sample reference image into a second training model to obtain an image output by the second training model; The second sample image is used to verify the image output by the second training model, so as to update the parameters of the second training model until the second training model converges, and the second model is generated.
4. The method of claim 2, wherein, Before obtaining the category of image style adjustment indicated by the first prompt information, the method further includes: In the case that the second training model adopts a high dynamic range (HDR) network model, a second training sample set is obtained, the second training sample set including the first sample image, a first sample reference image, and the second sample image, the image style of the second sample image being the same as the image style indicated by the style sample prompt information, the first sample reference image being determined based on the second sample image; The second training model performs low-resolution coefficient prediction processing on the first sample image and the second sample image, to obtain a processed first sample image and a processed second sample image; The second training model performs pixel-level network fusion processing on the first sample image, to generate a guided image of the first sample image; The second training model performs matrix transformation processing on the guided image, the processed first sample image, and the processed second sample image; The second training model combines the image after the matrix transformation processing with the first sample image, and outputs a second prediction image; The second sample image is used to verify the second prediction image, so as to update the parameters of the second training model until the second training model converges, and the first candidate model is generated.
5. The method of claim 2, wherein, The second training model includes a weight prediction model and N basic three-dimensional lookup tables, N being an integer greater than 1. Before obtaining the category of image style adjustment indicated by the first prompt information, the method further includes: A second training sample set is obtained, the second training sample set including the first sample image, a first sample reference image, and the second sample image, the image style of the second sample image being the same as the image style indicated by the style sample prompt information, the first sample reference image being determined based on the second sample image; The first sample image and the first sample reference image are input into the weight prediction model, to obtain N weights corresponding to the N basic three-dimensional lookup tables output by the weight prediction model; The N basic three-dimensional lookup tables are fused according to the N weights corresponding to the N basic three-dimensional lookup tables, to generate a three-dimensional lookup table of a reference style, the three-dimensional lookup table of the reference style being a color gamut mapping of a target style; The image style of the first sample image is transformed using the three-dimensional lookup table of the reference style, to obtain a second prediction image; The second sample image is used to verify the second prediction image, so as to iteratively update the parameters of the weight prediction model and the N basic three-dimensional lookup tables until the second training model converges, and the first candidate model is generated.
6. The method of claim 2, wherein, Before obtaining the category of image style adjustment indicated by the first prompt information, the method further includes: In a case where the second training model is a restoration transformer model, a second training sample set is obtained, the second training sample set including the first sample image, a first sample reference image, and a second sample image, the second sample image having an image style same as an image style indicated by the style sample prompt information, and the first sample reference image being determined based on the second sample image; The first sample image and the first sample reference image are input into the second training model, and a second predicted image output by the second training model is obtained; The second predicted image is verified by using the second sample image, parameters of the second training model are updated until the second training model converges, and the second candidate model is generated.
7. The method according to any one of claims 4-6, characterized in that, The method further includes: The second sample image is compressed according to a preset size, and the first sample reference image is generated.
8. The method of claim 1, wherein, The method further includes: The first image and the first prompt information are input into the first model, and a first predicted image output by the first model is obtained, the first predicted image having an image style same as the target style; The first predicted image is compressed according to a preset size, and a reference image of the first image is generated.
9. The method of claim 1, wherein, The method further includes: The first image is displayed; In response to an operation of adjusting an image style input by a user, an input box is displayed; The prompt information input by the user in the input box is obtained as the first prompt information.
10. The method of claim 1, wherein, The method further includes: The first image is displayed; In response to an operation of adjusting an image style input by a user, at least two image style controls are displayed, each image style control having prompt information of a corresponding image style; In response to a selection operation of selecting an image style from the at least two image style controls, prompt information of the selected image style is obtained; The obtained prompt information of the image style is obtained as the first prompt information.
11. An electronic device, comprising: The method further includes: A memory and a processor, the memory being coupled to the processor; The memory stores program instructions, and when the program instructions are executed by the processor, the electronic device executes the image processing method in any one of claims 1 to 10.
12. A computer readable storage medium comprising a computer program, characterized in that, When the computer program runs on the electronic device, the electronic device executes the image processing method in any one of claims 1 to 10.
Citation Information
Patent Citations
Image stylization transferring method of combining deep learning and depth perception
CN107705242A
Image processing method and device, electronic equipment and storage medium
CN110516201A