Image processing method and device, electronic equipment and storage medium
Patent Information
- Application Number
- CN202310102357.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-29
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2043-01-29
AI Technical Summary
[0003]本公开提供一种图像处理方法、装置、电子设备及存储介质,以至少解决相关技术中如何提升图像风格化处理效率和精度的问题
[0109]通过显示图像风格化处理页面,并设置图像风格化处理页面展示有内容图像输入区域和风格信息输入区域;从而响应于内容图像输入区域输入的第一图像以及风格信息输入区域输入的风格表征信息,在图像风格化处理页面增量展示第一图像被处理为目标风格后的第二图像,目标风格为风格表征信息指示的风格。通过这种页面的方式进行图像风格化转换,便于用户直观地操作,也能够及时得到风格化处理结果,实现了图像风格化处理的可视化操作,更加便捷,适用范围更广;
Smart Images

Figure CN116188250B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of data processing technology, and in particular to an image processing method, apparatus, electronic device, and storage medium. Background Technology
[0002] The stylization of images is currently receiving considerable attention. For example, stylized images, such as illustration-style images or abstract-style images, are often used in short videos. Related technologies employ two approaches: using text to guide the generation of stylized images, or directly inputting an image into the model for stylization. The former requires manually providing guiding text, which is inefficient and can easily lead to inaccurate descriptions of the image's content, resulting in imprecise stylization. The latter lacks semantic guidance, performs poorly on unstructured data, and is also inefficient. Furthermore, existing image stylization processes are not convenient enough. Summary of the Invention
[0003] This disclosure provides an image processing method, apparatus, electronic device, and storage medium to at least address the problem of how to improve the efficiency and accuracy of image stylization processing in related technologies. The technical solution of this disclosure is as follows:
[0004] According to a first aspect of the present disclosure, an image processing method is provided, comprising:
[0005] The image stylization processing page displays a content image input area and a style information input area;
[0006] In response to the first image input in the content image input area and the style representation information input in the style information input area, a second image is incrementally displayed on the image stylization processing page. The content of the second image matches the content of the first image, and the style attribute of the second image is the target style, which is the style indicated by the style representation information.
[0007] In one possible implementation, the method further includes:
[0008] The first image is displayed on the image stylization page.
[0009] In one possible implementation, the incremental display of a second image on the image stylization processing page in response to the first image input in the content image input area and the style representation information input in the style information input area includes:
[0010] In response to the first image input in the content image input area and the style representation information input in the style information input area, the content text information and style description text information are incrementally displayed on the image stylization processing page; the content text information corresponds to the semantic content of the first image, and the style description text information is text information describing the target style;
[0011] In response to the stylization processing instruction, the second image is incrementally displayed on the image stylization processing page; the second image is obtained based on the content text information and the style description text information.
[0012] In one possible implementation, the incremental display of a second image on the image stylization processing page in response to the first image input in the content image input area and the style representation information input in the style information input area includes:
[0013] In response to the first image input in the content image input area and the style representation information input in the style information input area, the content text information and style description text information are incrementally displayed on the image stylization processing page; the content text information corresponds to the semantic content of the first image, and the style description text information is text information describing the target style;
[0014] Upon detecting a first adjustment operation on the content text information and / or a second adjustment operation on the style description text information, display the content adjustment information and / or the style adjustment text information;
[0015] In response to the stylization processing instruction, the second image is incrementally displayed on the image stylization processing page; the second image is obtained based on the display content adjustment information and the style adjustment text information; or the second image is obtained based on the display content adjustment information and the style description text information; or the second image is obtained based on the content text information and the style adjustment text information.
[0016] In one possible implementation, the style description text information includes at least two target styles; the second adjustment operation, which detects the style description text information, displays style adjustment text information, including:
[0017] The image stylization processing page displays style priority selection information;
[0018] In response to a priority confirmation instruction triggered based on the style priority selection information, it is determined that the second adjustment operation has been detected, and priority information corresponding to each of the at least two target styles is displayed;
[0019] Based on the priority information and the style description text information, the style adjustment text information is generated and displayed.
[0020] In one possible implementation, the style representation information is any of the following: image, text, audio, or video information.
[0021] In one possible implementation, the method further includes:
[0022] The style representation information is subjected to style extraction processing to obtain style description text information;
[0023] The first image is subjected to semantic extraction processing to obtain the content text information;
[0024] The content text information and the style description text information are fused together to obtain the target text information;
[0025] The target text information is input into a stylization processing model to obtain the second image.
[0026] In one possible implementation, the style representation information is a third image, and the style attribute of the third image is the target style; the style extraction processing of the style representation information to obtain style description text information includes:
[0027] The third image is input into the style extraction model for style extraction processing to obtain the style description text information.
[0028] In one possible implementation, the style representation information is style indicator text information; the step of performing style extraction processing on the style representation information to obtain style description text information includes:
[0029] The style instruction text information is processed by word ordering to obtain the style description text information;
[0030] Alternatively, the style instruction text information can be input into a text processing model to obtain the style description text information.
[0031] In one possible implementation, the semantic extraction processing of the first image to obtain content text information includes:
[0032] The first image is input into a semantic extraction model for semantic extraction processing to obtain the content text information.
[0033] In one possible implementation, the method further includes:
[0034] Obtain noise information;
[0035] The step of inputting the target text information into a stylization processing model to obtain the second image includes:
[0036] The noise information and the target text information are input into the stylization processing model to obtain the second image.
[0037] In one possible implementation, acquiring the noise information includes:
[0038] The first image is subjected to noise processing to obtain the noise information;
[0039] Alternatively, the noise information can be obtained by sampling a preset Gaussian noise.
[0040] According to a second aspect of the present disclosure, an image processing method is provided, comprising:
[0041] Obtain the first image to be stylized and its style representation information;
[0042] The style representation information is subjected to style extraction processing to obtain style description text information;
[0043] The first image is subjected to semantic extraction processing to obtain the content text information;
[0044] The content text information and the style description text information are fused together to obtain the target text information;
[0045] The target text information is input into a stylization processing model to obtain a second image; the content of the second image matches the content of the first image, and the style attribute of the second image is the target style, which is the style indicated by the style representation information.
[0046] In one possible implementation, the style representation information is a third image; the style extraction processing of the style representation information to obtain style description text information includes:
[0047] The third image is input into the style extraction model for style extraction processing to obtain the style description text information.
[0048] In one possible implementation, the method further includes:
[0049] Obtain noise information;
[0050] The step of inputting the target text information into a stylization processing model to obtain a second image includes:
[0051] The noise information and the target text information are input into the stylization processing model to obtain the second image.
[0052] In one possible implementation, acquiring the noise information includes:
[0053] The first image is subjected to noise processing to obtain the noise information;
[0054] Alternatively, the noise information can be obtained by sampling a preset Gaussian noise.
[0055] In one possible implementation, the method further includes:
[0056] Obtain content adjustment information obtained by performing a first adjustment operation on the content text information, and / or style adjustment text information obtained by performing a second adjustment operation on the style description text information;
[0057] The process of fusing the content text information and the style description text information to obtain the target text information includes:
[0058] The target text information is obtained based on the content adjustment information and / or the style adjustment text information.
[0059] According to a third aspect of the present disclosure, an image processing apparatus is provided, comprising:
[0060] The page display module is configured to execute the display of an image stylization processing page, which displays a content image input area and a style information input area;
[0061] The stylized image display module is configured to incrementally display a second image on the image stylization processing page in response to a first image input in the content image input area and style representation information input in the style information input area. The content of the second image matches the content of the first image, and the style attribute of the second image is a target style, which is the style indicated by the style representation information.
[0062] In one possible implementation, the device further includes:
[0063] The first display module is configured to display the first image on the image stylization processing page.
[0064] In one possible implementation, the stylized image display module includes:
[0065] The first text information display unit is configured to incrementally display content text information and style description text information on the image stylization processing page in response to the first image input in the content image input area and the style representation information input in the style information input area; the content text information corresponds to the semantic content of the first image, and the style description text information is text information describing the target style;
[0066] The first stylized image display unit is configured to execute a stylization processing instruction and incrementally display the second image on the image stylization processing page; the second image is obtained based on the content text information and the style description text information.
[0067] In one possible implementation, the stylized image display module includes:
[0068] The second text information display unit is configured to incrementally display content text information and style description text information on the image stylization processing page in response to the first image input in the content image input area and the style representation information input in the style information input area; the content text information corresponds to the semantic content of the first image, and the style description text information is text information describing the target style;
[0069] The text adjustment unit is configured to perform a first adjustment operation that detects the content text information and / or a second adjustment operation that detects the style description text information, and to display the content adjustment information and / or the style adjustment text information.
[0070] The second stylized image display unit is configured to incrementally display the second image on the image stylization processing page in response to the stylization processing instruction; the second image is obtained based on the display content adjustment information and the style adjustment text information; or the second image is obtained based on the display content adjustment information and the style description text information; or the second image is obtained based on the content text information and the style adjustment text information.
[0071] In one possible implementation, the style description text information includes at least two target styles; the text adjustment unit includes:
[0072] The priority selection subunit is configured to display style priority selection information on the image stylization processing page.
[0073] The second adjustment operation determination subunit is configured to execute a priority confirmation instruction triggered in response to the style priority selection information, determine that the second adjustment operation has been detected, and display the priority information corresponding to each of the at least two target styles.
[0074] The style adjustment text display subunit is configured to generate and display the style adjustment text information based on the priority information and the style description text information.
[0075] In one possible implementation, the style representation information is any of the following: image, text, audio, or video information.
[0076] In one possible implementation, the device further includes:
[0077] The style extraction module is configured to perform style extraction processing on the style representation information to obtain style description text information;
[0078] The semantic extraction module is configured to perform semantic extraction processing on the first image to obtain content text information;
[0079] The text fusion module is configured to perform fusion processing on the content text information and the style description text information to obtain target text information;
[0080] The stylization processing module is configured to input the target text information into the stylization processing model to obtain the second image.
[0081] In one possible implementation, the style representation information is a third image, and the style attribute of the third image is the target style; the style extraction module includes:
[0082] The first style extraction unit is configured to input the third image into the style extraction model, perform style extraction processing, and obtain the style description text information.
[0083] In one possible implementation, the style representation information is style indicator text information; the style extraction module further includes:
[0084] The second style extraction unit is configured to perform word order processing on the style indicator text information to obtain the style description text information; or, input the style indicator text information into a text processing model to obtain the style description text information.
[0085] In one possible implementation, the semantic extraction module is further configured to input the first image into a semantic extraction model, perform semantic extraction processing, and obtain the content text information.
[0086] In one possible implementation, the device further includes:
[0087] The noise acquisition module is configured to acquire noise information.
[0088] The stylization processing module is further configured to input the noise information and the target text information into the stylization processing model to obtain the second image.
[0089] In one possible implementation, the noise acquisition module is further configured to perform noise addition processing on the first image to obtain the noise information; or to sample preset Gaussian noise to obtain the noise information.
[0090] According to a fourth aspect of the present disclosure, an image processing apparatus is provided, comprising:
[0091] The acquisition module is configured to acquire the first image to be stylized and its style representation information.
[0092] The style text acquisition module is configured to perform style extraction processing on the style representation information to obtain style description text information;
[0093] The content text acquisition module is configured to perform semantic extraction processing on the first image to obtain content text information;
[0094] The target text information acquisition module is configured to perform a fusion processing on the content text information and the style description text information to obtain the target text information;
[0095] The style processing module is configured to input the target text information into a styleization processing model to obtain a second image; the content of the second image matches the content of the first image, and the style attribute of the second image is the target style, which is the style indicated by the style representation information.
[0096] In one possible implementation, the style representation information is a third image; the style text acquisition module includes:
[0097] The style text acquisition unit is configured to input the third image into the style extraction model, perform style extraction processing, and obtain the style description text information.
[0098] In one possible implementation, the device further includes:
[0099] The noise information acquisition module is configured to acquire noise information.
[0100] The style processing module is further configured to input the noise information and the target text information into the style processing model to obtain the second image.
[0101] In one possible implementation, the noise information acquisition module is further configured to perform noise addition processing on the first image to obtain the noise information; or to sample a preset Gaussian noise to obtain the noise information.
[0102] In one possible implementation, the device further includes:
[0103] The text adjustment module is configured to perform the following operations: obtain content adjustment information obtained by performing a first adjustment operation on the content text information, and / or obtain style adjustment text information obtained by performing a second adjustment operation on the style description text information.
[0104] The target text information acquisition module is further configured to perform operations based on the content adjustment information and / or the style adjustment text information to obtain the target text information.
[0105] According to a fifth aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to execute the instructions to implement the method as described in any one of the first aspects above.
[0106] According to a sixth aspect of the present disclosure, a computer-readable storage medium is provided such that, when instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform any of the methods described in the first aspect of the present disclosure.
[0107] According to a seventh aspect of the present disclosure, a computer program product is provided, including computer instructions that, when executed by a processor, cause a computer to perform the method described in any one of the first aspects of the present disclosure.
[0108] The technical solutions provided by the embodiments of this disclosure have at least the following beneficial effects:
[0109] By displaying an image stylization processing page with a content image input area and a style information input area, the image stylization processing page incrementally displays a second image after the first image has been processed to the target style, as indicated by the style representation information, in response to the first image input in the content image input area and the style representation information input in the style information input area. This page-based approach to image stylization conversion is intuitive for users, provides timely results, and enables a more convenient and widely applicable visual operation of image stylization processing.
[0110] Furthermore, by inputting the first image in the content image input area to represent the image content to be stylized, the user does not need to manually input text describing the content, and the image content can be expressed more accurately, thereby improving the accuracy of stylization processing.
[0111] In addition, by setting the content image input area and style information input area, the image content and target style for stylization processing can be obtained automatically and conveniently, making image stylization processing more accurate and efficient.
[0112] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0113] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.
[0114] Figure 1 This is a schematic diagram illustrating an application environment according to an exemplary embodiment.
[0115] Figure 2 This is a flowchart illustrating an image processing method according to an exemplary embodiment.
[0116] Figure 3a This is a schematic diagram illustrating an image stylization processing page according to an exemplary embodiment.
[0117] Figure 3b This is a schematic diagram illustrating a second image according to an exemplary embodiment.
[0118] Figure 3c This is a schematic diagram illustrating another second image according to an exemplary embodiment.
[0119] Figure 4a This is a schematic diagram illustrating another image stylization process page according to an exemplary embodiment.
[0120] Figure 4b This is a schematic diagram illustrating another second image according to an exemplary embodiment.
[0121] Figure 4c This is a schematic diagram illustrating a content text information, style description text information, and a second image according to an exemplary embodiment.
[0122] Figure 4d This is a schematic diagram illustrating the display of content text information and style description text information according to an exemplary embodiment.
[0123] Figure 4e This is a schematic diagram illustrating a style priority selection information according to an exemplary embodiment.
[0124] Figure 4f This is a schematic diagram illustrating a style-adjusted display of text information according to an exemplary embodiment.
[0125] Figure 5 This is a schematic diagram of a third image in a target style according to an exemplary embodiment.
[0126] Figure 6 This is a flowchart illustrating a stylized processing method according to an exemplary embodiment.
[0127] Figure 7 This is a flowchart illustrating another stylized processing method according to an exemplary embodiment.
[0128] Figure 8 This is a flowchart illustrating another image processing method according to an exemplary embodiment.
[0129] Figure 9 This is a block diagram of an image processing apparatus according to an exemplary embodiment.
[0130] Figure 10 This is a block diagram illustrating an electronic device for image processing according to an exemplary embodiment.
[0131] Figure 11 This is a block diagram illustrating another electronic device for image processing according to an exemplary embodiment. Detailed Implementation
[0132] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0133] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0134] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or computers-controlled machines to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. AI software technology mainly includes computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0135] In recent years, with the research and progress of artificial intelligence technology, artificial intelligence technology has been widely used in many fields. The solutions provided in the embodiments of this application involve technologies such as machine learning / deep learning, which are specifically illustrated through the following embodiments.
[0136] Please see Figure 1 , Figure 1 This is a schematic diagram illustrating an application environment according to an exemplary embodiment, such as... Figure 1 As shown, the application environment may include server 01 and terminal 02.
[0137] In an optional embodiment, server 01 can be used for image stylization processing. Specifically, server 01 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.
[0138] In an optional embodiment, terminal 02 can be used to display an image stylization processing page and show the stylized image. Specifically, terminal 02 can be, but is not limited to, electronic devices such as smartphones, desktop computers, tablets, laptops, smart speakers, digital assistants, augmented reality (AR) / virtual reality (VR) devices, and smart wearable devices. Optionally, the operating system running on the electronic device can be, but is not limited to, Android, iOS, Linux, Windows, etc.
[0139] In addition, it should be noted that, Figure 1 The illustration shown is merely one application environment of the image processing method provided in this disclosure. Optionally, image stylization processing, the display of the image stylization processing page, and the display of the stylized image can all be performed by terminal 02.
[0140] In the embodiments described in this specification, the server 01 and the terminal 02 can be directly or indirectly connected through wired or wireless communication, and this application does not impose any restrictions on this.
[0141] It should be noted that the following diagram illustrates one possible sequence of steps, and it is not strictly required to follow this order. Some steps can be performed in parallel without interdependence. The user information (including but not limited to user device information, user personal information, user behavior information, etc.) and data (including but not limited to data used for display, training data, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties.
[0142] Figure 2 This is a flowchart illustrating an image processing method according to an exemplary embodiment, which can be applied to a terminal. For example... Figure 2 As shown, the steps may include the following.
[0143] In step S201, an image stylization processing page is displayed, which can show a content image input area and a style information input area.
[0144] In practical applications, the target application can provide image stylization processing functionality. Based on this, an image stylization processing page can be set up within the target application. Entry information for accessing this image stylization processing page can then be set on the target application's homepage or a preset page. Triggering this entry information allows access to the image stylization processing page. In other words, responding to the triggering operation of this entry information displays the image stylization processing page. The target application can be a target multimedia application, a target image processing application, an image stylization processing application, etc. Target multimedia applications can include short video applications, etc., and this disclosure does not limit this. It should be noted that in the case of an image stylization processing application, launching the image stylization processing application will display the image stylization processing page, and entry information is not required.
[0145] Optionally, the image stylization processing page can be an image stylization processing webpage. Entering the corresponding image stylization processing website will display the image stylization processing webpage. This disclosure does not limit the display method or access method of the image stylization processing page.
[0146] In one example, image stylization of a page can be done as follows: Figure 3a As shown, the image stylization processing page can display a content image input area and a style information input area. The content image input area can be as follows: Figure 3a As shown in Figure 301, the style information input area can be as follows: Figure 3aAs shown in 302. In an alternative manner, a prompt message "Content to be stylized" may also be displayed for the content image input area, and a prompt message "Style representation information" may be displayed for the style information input area, to indicate the function of the input area.
[0147] The content image input area can be used to input or upload content images, providing the content of the image during the stylization process. In other words, the content of the stylized image matches the content of this content image. Alternatively, the content image can be understood as an image that provides content for the stylization process, thereby generating the content of the stylized image based on the content of this content image.
[0148] Optionally, such as Figure 3a As shown, the content image input area may include an image selection control 303, which can be triggered to select an image used to represent the content, such as a first image.
[0149] The style information input area can be used to input style representation information of the target style. As an example, such as... Figure 3a As shown in 302, the style information input area can be a text box, and correspondingly, the style representation information can be text information used to describe the target style. Thus, you can enter the desired style in this text box, such as "Let's convert it to an illustration-style image."
[0150] In another example, image stylization of a page can be done as follows: Figure 4a As shown, the image stylization processing page can display a content image input area and a style information input area. The content image input area can be as follows: Figure 4a As shown in Figure 401, the style information input area can be as follows: Figure 4a 402 shown.
[0151] The content image input area can be used to input or upload content images, that is, images that provide content for stylization processing, such as the first image, which provides the content of the image during the stylization process; that is, the content of the stylized image can match the content of this content image. Alternatively, it can be understood that the content image can refer to the image that provides content for stylization processing.
[0152] Optionally, such as Figure 4a As shown, the content image input area may include an image selection control 403, which can be triggered to select an image to represent the content, such as a first image, like XX1.jpg. The style attributes of the first image may differ from the target style; that is, the style attributes of the first image may not be the target style.
[0153] The style information input area can be used to input style representation information of the target style, such as... Figure 4a As shown in 402, the style information input area can be an image selection area used to input or upload a style image, i.e., an image of the target style, to indicate the target style. Alternatively, it can be understood that the style image can refer to an image indicating the target style for stylization processing. Accordingly, the style representation information can be a style image, such as a third image.
[0154] Optionally, such as Figure 4a As shown, the style information input area may include an image selection control 404, which can be triggered to select an image that indicates the target style, such as a third image, like T1.jpg, whose style attribute is the target style.
[0155] In step S203, in response to the first image input in the content image input area and the style representation information input in the style information input area, a second image is incrementally displayed on the image stylization processing page. The content of the second image matches the content of the first image, and the style attribute of the second image is the target style, which is the style indicated by the style representation information. The style representation information can be any of the following: image, text information, audio information, or video information, and this disclosure does not limit it.
[0156] In the embodiments of this specification, in response to the first image input in the content image input area and the style representation information input in the style information input area, a second image with the style attribute of the target style is incrementally displayed on the image stylization processing page, such as... Figure 3b The 304 error is shown. For example, in... Figure 3a After inputting the first image in 301 and style representation information in 302, in response to the first image input in the content image input area and the style representation information input in the style information input area, the following can be displayed: Figure 3b This means that a second image (304) can be incrementally displayed on the image stylization processing page. The content of the second image matches the content of the first image, and the style attribute of the second image is the target style, which can be the style indicated by style representation information. Matching the content of the second image with the content of the first image can mean that the content of the second image is the same as the content of the first image (content consistency). In this case, stylization processing can be viewed as performing a target style transformation on the first image.
[0157] Alternatively, matching the content of the second image with the content of the first image can mean that the first object in the first image and the second object in the second image are of the same type (e.g., category, variety, etc.). The first object in the first image can refer to a foreground object in the first image, and the second object in the second image can refer to a foreground object in the second image. For example, if the first object in the first image is a flower, then the second object in the second image can be a flower, and the variety of the flower in the second image is the same as the variety of the flower in the first image. The backgrounds of the first and second images can be different, and the size and quantity of the first and second objects can also be different; this disclosure does not impose any limitations on these aspects. Furthermore, the first and second objects can be any objects, and this disclosure does not impose any limitations on this either.
[0158] As an optional approach, the image stylization page may also include a style processing initiation control, such as Figure 3a The "Submit" button is shown. Based on this, the first image input in the content image input area and the style representation information input in the style information input area can refer to the triggering operation in response to this "Submit". For example, clicking "Submit" can incrementally display the second image 304 on the image stylization processing page. Figure 3b As shown. Specifically, in response to the triggering operation of the style processing start control, a first image can be extracted from the content image input area, and style representation information can be extracted from the style information input area; thus, stylization processing can be performed based on the first image and the style representation information to obtain a second image. The specific process of stylization processing can be found in the corresponding content below, and will not be repeated here.
[0159] Optionally, see Figure 3c and Figure 4b Alternatively, the first image can be displayed simultaneously, making it easier to compare style changes more intuitively. Based on this, the method can also include: displaying the first image on the image stylization processing page, that is, displaying both the first and second images on the image stylization processing page. The first image can be as follows: Figure 3c As shown in 305, or as... Figure 4b The 405 error shown is the content to be stylized, i.e., the content image: XX1.jpg.
[0160] In one example, Figure 3a After inputting the first image in 301 and style representation information in 302, it can be displayed. Figure 3c This means that the first image 305 and the second image 304 can be displayed on the image stylization processing page.
[0161] In another example, in Figure 4a After inputting the first image in 401 and the style image (i.e., the third image) in 402, it can be displayed. Figure 4b This means that the first image 405 and the second image 406 can be displayed on the image stylization processing page. The third image can be, for example... Figure 5 As shown, the style attribute of the third image is the target style, which is the illustration style here. By using a third image with an existing illustration style, the content of the first image is automatically converted into a second image with an illustration style, thereby automatically and accurately displaying the content of the first image in the target style. There is no need to manually describe the content or the style, which is efficient, convenient and accurate.
[0162] By displaying an image stylization processing page with a content image input area and a style information input area, the image stylization processing page incrementally displays a second image after the first image has been processed to the target style, as indicated by the style representation information, in response to the first image input in the content image input area and the style representation information input in the style information input area. This page-based approach to image stylization conversion is intuitive for users, provides timely results, and enables a more convenient and widely applicable visual operation of image stylization processing.
[0163] Furthermore, by inputting the first image in the content image input area to represent the image content to be stylized, the user does not need to manually input text describing the content, and the image content can be expressed more accurately, thereby improving the accuracy of stylization processing.
[0164] In addition, by setting the content image input area and style information input area, the image content and target style for stylization processing can be obtained automatically and conveniently, making image stylization processing more accurate and efficient.
[0165] The above mainly demonstrates image style transfer. The following describes the process of stylization based on style representation information and the first image. In one possible implementation, the method may further include the following steps:
[0166] Style representation information is processed by style extraction to obtain style description text information. As an example, style representation information can be a third image or style indicator text information; that is, style representation information can be either an image or text. Based on this, when the style representation information is an image, for example, a third image, the style attribute of the third image is the target style. Accordingly, processing the style representation information by style extraction to obtain style description text information can include: inputting the third image into a style extraction model, performing style extraction processing, and obtaining style description text information. By setting a style extraction model to extract the style of the third image, the accuracy and convenience of the target style can be improved. The style description text information can refer to text information describing the target style. The style description text information can be a sentence or a combination of words describing the target style; this disclosure does not limit this. The language type of the style description text information is also not limited in this disclosure. The style extraction model here can be a BLIP (Bootstrapping Language-Image Pre-training) model; this disclosure does not limit this. The BLIP model is a unified vision-language pre-training (VLP) framework that can output the style of an image based on an input image. This style extraction model can be trained using an initial style extraction model based on a first training sample. The first training sample can include multiple sample images and their corresponding style labels (e.g., annotated style text). This disclosure does not limit the training method; for example, multiple sample images can be input into the initial style extraction model to obtain style prediction text. Loss information can then be determined based on the style prediction text and style labels, and the parameters of the initial style extraction model can be adjusted based on this loss information until the training iteration conditions are met. The initial style extraction model that meets the training iteration conditions can be used as the semantic extraction model. The training iteration conditions can be a threshold for the number of training iterations, a loss threshold, etc.
[0167] When the style representation information is text, such as style indicator text information, the style representation information is processed to obtain style description text information. This can include: processing the style indicator text information to obtain style description text information. For example, the style indicator text information can be processed by word order manipulation to obtain style description text information; this can be achieved by setting a word order adjustment module. Alternatively, the style indicator text information can be input into a text processing model to obtain style description text information. The text processing model can be a text translation model, a keyword extraction model, etc., and can be pre-trained. Processing the style indicator text information by word order manipulation to obtain text information describing the target style is simpler; text processing using a text processing model can meet the needs of large-scale text processing and is more accurate.
[0168] Furthermore, semantic extraction processing can be performed on the first image to obtain content text information. This content text information can refer to text information describing the stylized processing content; it can be a complete sentence or a combination of words describing the content. The language type of the content text information is not limited in this disclosure. As an example, the first image can be input into a semantic extraction model for semantic extraction processing to obtain the content text information. Extracting the semantics of the first image using a semantic extraction model can improve the accuracy and convenience of content text information extraction. Moreover, compared to directly inputting an image into the model to output a target style image, extracting the stylized processing content using an independent semantic extraction model not only yields content text to guide the stylized processing but also extracts denser semantic features, making the content text information more comprehensive and accurate. This allows for both textual guidance and ensures the accuracy of the guidance during stylized processing, and the content text information does not require manual input, further improving the accuracy and convenience of content text information extraction.
[0169] The semantic extraction model can be a BLIP model, and this disclosure does not limit its use. The semantic extraction model can output the content of an image based on the input image. This semantic extraction model can be obtained by training an initial semantic extraction model based on a second training sample. The second training sample can include multiple sample images and corresponding content labels (e.g., sample image content text) for each sample image. This disclosure does not limit the training method; for example, multiple sample images can be input into the initial semantic extraction model to obtain predicted content text. Loss information can then be determined based on the predicted content text and content labels. The parameters of the initial semantic extraction model can then be adjusted based on this loss information until the training iteration conditions are met. The initial semantic extraction model that meets the training iteration conditions can be used as the semantic extraction model. The training iteration conditions can be a threshold for the number of training iterations, a loss threshold, etc.
[0170] Furthermore, the content text information and style description text information can be fused to obtain the target text information. The language of the target text information is not limited in this disclosure. The fusion process can be a splicing process, or it can be text deduplication and splicing. The splicing process can be based on a preset grammar, which can be the grammar of each language; this disclosure is not limited in this regard. Optionally, the fusion process can also include keyword addition processing, for example, adding keywords indicating image display metrics, such as "high definition" or "wide-angle," etc.; this disclosure is not limited in this regard.
[0171] The target text information can then be input into a stylization model to obtain a second image. This stylization model can be a text-guided generative model, such as a diffusion model or a Generative Adversarial Network (GAN). The stylization model can be trained using a third training sample.
[0172] When the stylization model is a diffusion model, the third training sample can include multiple sample texts and the image labels corresponding to each sample text; thus, the stylization model can be obtained by supervised training of the initial diffusion model based on the third training sample.
[0173] When the stylization processing model is a generative adversarial network (GAN), the third training sample can include multiple sample texts. This allows for unsupervised training of the initial GAN based on these multiple sample texts to obtain the stylization processing model. This disclosure does not limit the specific training process. Alternatively, the third training sample can include multiple sample texts and corresponding annotation information for each sample text. This annotation information can be sample images. Therefore, supervised training of the initial GAN based on multiple sample texts and annotation information can be performed to obtain the stylization processing model.
[0174] By extracting the semantics of the first image to obtain content text information, and combining it with style description text information as input to the stylization processing model, the accuracy of stylization processing can be improved. On the one hand, compared with models lacking semantic guidance, it has a holistic understanding of global semantics, performs better in stylization processing of unstructured data, and has better scalability in application scenarios. On the other hand, the automated acquisition of content text information and style description text information makes these information more accurate, thereby improving the semantic accuracy of the target text information input to the stylization processing model, which can further improve the accuracy of stylization processing.
[0175] As an example, a flowchart of the stylization process can be as follows: Figure 6 As shown, the style extraction module can be a style extraction model, a text processing model, or a word order adjustment module. A third image or style instruction text can be input into the style extraction module for style extraction processing to obtain style description text information of the target style. Conversely, a first image providing content can be input into the semantic extraction model for semantic extraction processing to obtain content text information. The style description text information and content text information can then be integrated to obtain the target text information. For example, if the content text information is "a blooming XX flower" and the style description text information is "illustration style," the resulting target text information could be "an illustration-style, blooming XX flower." Furthermore, the target text information can be input into the stylization processing model, which, guided by text, generates a second image in the illustration style, i.e., outputs a second image. The style attributes and content of this second image can be consistent with the target style indicated by the style description text information and the content indicated by the content text information, i.e., consistent with the target text information. Correspondingly, this second image can be displayed as the stylization processing result.
[0176] Alternatively, the flowchart for stylized processing can also be as follows: Figure 7 As shown, noise information can be added to the input of the stylization processing model. This not only ensures that the second image is consistent with the target text information, but also enhances the diversity of the images generated by the stylization processing. Based on this, the method may further include: obtaining noise information; for example, the first image can be noise-added, such as by adding random noise, to obtain the noise information; or, the noise information can be obtained by sampling a preset Gaussian noise. The sampling method can be random sampling, and this disclosure does not limit it.
[0177] Accordingly, inputting the target text information into the stylization processing model to obtain the second image can include: inputting noise information and target text information into the stylization processing model to obtain the second image. It should be noted that during the training of the stylization processing model, the third training sample may also include sample noise.
[0178] As an optional implementation, content text information and style description text information can also be displayed on the image stylization processing page to visually represent them. Based on this, step S203 may include: in response to a first image input in the content image input area and style representation information input in the style information input area, incrementally displaying content text information and style description text information on the image stylization processing page, such as... Figure 4c The style description text information shown in section 407 is as follows: Figure 4c As shown in 408, the content text information corresponds to the semantic content of the first image, and the style description text information is the text information describing the target style.
[0179] Furthermore, in response to stylization instructions, such as in response to Figure 4c The stylization processing instruction triggered by clicking "Submit" can incrementally display a second image on the image stylization processing page; this second image is obtained based on content text information and style description text information.
[0180] As an optional implementation, content text information and style description text information can be displayed on the image stylization processing page, and the content text information and style description text information can be adjusted to meet the flexible requirements of image processing. Based on this, step S203 above may include: in response to the first image input in the content image input area and the style representation information input in the style information input area, incrementally displaying the content text information and style description text information on the image stylization processing page, such as... Figure 4d As shown in 409 and 410, the content text information corresponds to the semantic content of the first image, and the style description text information is the text information describing the target style;
[0181] The detection of a first adjustment operation on the content text information and / or a second adjustment operation on the style description text information, for example... Figure 4d When the "Adjust" button is clicked, the content text and style description text become editable. Upon completion of editing, an edit completion command is triggered, displaying the adjusted content and / or style text information—that is, the adjusted content and style description text. As an example, if the style description text includes at least two styles, style priority selection information can be displayed on the image stylization page. Figure 4d The display of 410 can be as follows Figure 4e As shown, it may include Figure 4e As shown in 411. Based on this, in response to a priority confirmation instruction triggered based on style priority selection information, for example... Figure 4f If both Style 1 and Style 2 are selected, and the priority information for Style 1 is set to high and the priority information for Style 2 is set to low, then a second adjustment operation can be detected. This allows the display of the priority information for at least two target styles, such as... Figure 4f As shown in 412, style adjustment text information can be generated and displayed based on priority information and style description text information. This style adjustment text information can include the selected target style and its corresponding priority information, such as "The converted styles include style 1 and style 2, where style 1 has a higher priority than style 2". This is just an example and does not limit this disclosure. For example, only style 1 can be selected, in which case the priority information of style 1 is high by default.
[0182] Furthermore, in response to the stylization processing command, a second image is incrementally displayed on the image stylization processing page, as shown below. Figure 3b , Figure 3c , Figure 4b and Figure 4c Any one of the following. The second image is obtained based on content adjustment information and style adjustment text information; or the second image is obtained based on content adjustment information and style description text information; or the second image is obtained based on content text information and style adjustment text information. Specifically, in image processing methods involving a hybrid style with at least two target styles, style weights can be determined based on priority information. For example, the target style with higher priority information can be designated as the primary style with a higher weight, while the target style with lower priority information can be designated as the secondary style with a lower weight. This allows for stylizing the content to be stylized using the primary style and refining details using the secondary style. Alternatively, the target style with higher priority information can be designated as the foreground style, and the target style with lower priority information can be designated as the background style to achieve diversified style processing of the image.
[0183] Figure 8 This is a flowchart illustrating another image processing method according to an exemplary embodiment. It can be applied to a terminal or server. Figure 8 As shown, the method may include:
[0184] In step S801, the first image to be stylized and its style representation information are obtained;
[0185] In step S803, style extraction processing is performed on the style representation information to obtain style description text information;
[0186] In step S805, semantic extraction processing is performed on the first image to obtain content text information;
[0187] In step S807, the content text information and style description text information are fused to obtain the target text information;
[0188] In step S809, the target text information is input into the stylization processing model to obtain a second image; the style attribute of the second image is the target style, and the target style is the style indicated by the style representation information.
[0189] It should be noted that, when the server is executing the process, the second image can be sent to the terminal so that the terminal can display the second image on the image stylization page.
[0190] As an example, the style representation information can be a third image; accordingly, step S803 above can include: inputting the third image into the style extraction model, performing style extraction processing, and obtaining style description text information.
[0191] In an optional implementation, step S807 may further include: acquiring content adjustment information obtained by performing a first adjustment operation on the content text information, and / or style adjustment text information obtained by performing a second adjustment operation on the style description text information. Accordingly, step S807 may include: obtaining target text information based on the content adjustment information and / or the style adjustment text information. For example, the content adjustment information and the style adjustment text information may be fused to obtain the target text information; or the content adjustment information and the style description text information may be fused to obtain the target text information; or the content text information and the style adjustment text information may be fused to obtain the target text information.
[0192] Optionally, the method may further include: acquiring noise information; correspondingly, step S809 may include: inputting the noise information and the target text information into a stylization processing model to obtain a second image. Acquiring noise information may include: adding noise to the first image to obtain noise information; or, sampling preset Gaussian noise to obtain noise information.
[0193] above Figure 8 For details on the specific processing methods of the relevant steps, please refer to the processing of the corresponding steps mentioned above, which will not be repeated here.
[0194] Figure 9 This is a block diagram of an image processing apparatus according to an exemplary embodiment. (Refer to...) Figure 9 The device may include:
[0195] The page display module 901 is configured to execute the display image stylization processing page, which displays a content image input area and a style information input area;
[0196] The stylized image display module 903 is configured to incrementally display a second image on the image stylization processing page in response to a first image input in the content image input area and style representation information input in the style information input area. The content of the second image matches the content of the first image, and the style attribute of the second image is a target style, which is the style indicated by the style representation information.
[0197] By displaying an image stylization processing page with a content image input area and a style information input area, the image stylization processing page incrementally displays a second image after the first image has been processed to the target style, as indicated by the style representation information, in response to the first image input in the content image input area and the style representation information input in the style information input area. This page-based approach to image stylization conversion is intuitive for users, provides timely results, and enables a more convenient and widely applicable visual operation of image stylization processing.
[0198] Furthermore, by inputting the first image in the content image input area to represent the image content to be stylized, the user does not need to manually input text describing the content, and the image content can be expressed more accurately, thereby improving the accuracy of stylization processing.
[0199] In addition, by setting the content image input area and style information input area, the image content and target style for stylization processing can be obtained automatically and conveniently, making image stylization processing more accurate and efficient.
[0200] In one possible implementation, the stylized image display module 903 described above may include:
[0201] The first text information display unit is configured to incrementally display content text information and style description text information on the image stylization processing page in response to the first image input in the content image input area and the style representation information input in the style information input area; the content text information corresponds to the semantic content of the first image, and the style description text information is text information describing the target style;
[0202] The first stylized image display unit is configured to execute a stylization processing instruction and incrementally display the second image on the image stylization processing page; the second image is obtained based on the content text information and the style description text information.
[0203] In one possible implementation, the stylized image display module 903 described above may include:
[0204] The second text information display unit is configured to incrementally display content text information and style description text information on the image stylization processing page in response to the first image input in the content image input area and the style representation information input in the style information input area; the content text information corresponds to the semantic content of the first image, and the style description text information is text information describing the target style;
[0205] The text adjustment unit is configured to perform a first adjustment operation that detects the content text information and / or a second adjustment operation that detects the style description text information, and to display the content adjustment information and / or the style adjustment text information.
[0206] The second stylized image display unit is configured to incrementally display the second image on the image stylization processing page in response to the stylization processing instruction; the second image is obtained based on the display content adjustment information and the style adjustment text information; or the second image is obtained based on the display content adjustment information and the style description text information; or the second image is obtained based on the content text information and the style adjustment text information.
[0207] In one possible implementation, the style description text information includes at least two target styles; the aforementioned text adjustment unit may include:
[0208] The priority selection subunit is configured to display style priority selection information on the image stylization processing page.
[0209] The second adjustment operation determination subunit is configured to execute a priority confirmation instruction triggered in response to the style priority selection information, determine that the second adjustment operation has been detected, and display the priority information corresponding to each of the at least two target styles.
[0210] The style adjustment text display subunit is configured to generate and display the style adjustment text information based on the priority information and the style description text information.
[0211] In one possible implementation, the style representation information is any of the following: image, text, audio, or video information.
[0212] In one possible implementation, the device may further include:
[0213] The first display module is configured to display the first image on the image stylization processing page.
[0214] In one possible implementation, the device may further include:
[0215] The style extraction module is configured to perform style extraction processing on the style representation information to obtain style description text information;
[0216] The semantic extraction module is configured to perform semantic extraction processing on the first image to obtain content text information;
[0217] The text fusion module is configured to perform fusion processing on the content text information and the style description text information to obtain target text information;
[0218] The stylization processing module is configured to input the target text information into the stylization processing model to obtain the second image.
[0219] In one possible implementation, the style representation information is a third image, and the style attribute of the third image can be the target style; accordingly, the style extraction module may include:
[0220] The first style extraction unit is configured to input the third image into the style extraction model, perform style extraction processing, and obtain the style description text information.
[0221] In one possible implementation, the aforementioned style representation information is style indicator text information; correspondingly, the style extraction module may further include:
[0222] The second style extraction unit is configured to perform word order processing on the style indicator text information to obtain the style description text information; or, input the style indicator text information into a text processing model to obtain the style description text information.
[0223] In one possible implementation, the semantic extraction module is further configured to input the first image into the semantic extraction model, perform semantic extraction processing, and obtain the content text information.
[0224] In one possible implementation, the device may further include:
[0225] The noise acquisition module is configured to acquire noise information.
[0226] The stylization processing module is further configured to input the noise information and the target text information into the stylization processing model to obtain the second image.
[0227] In one possible implementation, the noise acquisition module is further configured to perform noise addition processing on the first image to obtain the noise information; or to sample a preset Gaussian noise to obtain the noise information.
[0228] This disclosure also provides an image processing apparatus, which can be applied to a server or terminal, and the apparatus may include:
[0229] The acquisition module is configured to acquire the first image to be stylized and its style representation information.
[0230] The style text acquisition module is configured to perform style extraction processing on the style representation information to obtain style description text information;
[0231] The content text acquisition module is configured to perform semantic extraction processing on the first image to obtain content text information;
[0232] The target text information acquisition module is configured to perform a fusion processing on the content text information and the style description text information to obtain the target text information;
[0233] The style processing module is configured to input the target text information into a styleization processing model to obtain a second image; the content of the second image matches the content of the first image, and the style attribute of the second image is the target style, which is the style indicated by the style representation information.
[0234] In one possible implementation, the style representation information is a third image; the style text acquisition module includes:
[0235] The style text acquisition unit is configured to input the third image into the style extraction model, perform style extraction processing, and obtain the style description text information.
[0236] In one possible implementation, the device further includes:
[0237] The noise information acquisition module is configured to acquire noise information.
[0238] The style processing module is further configured to input the noise information and the target text information into the style processing model to obtain the second image.
[0239] In one possible implementation, the noise information acquisition module is further configured to perform noise addition processing on the first image to obtain the noise information; or to sample a preset Gaussian noise to obtain the noise information.
[0240] In one possible implementation, the device may further include:
[0241] The text adjustment module is configured to perform the following operations: obtain content adjustment information obtained by performing a first adjustment operation on the content text information, and / or obtain style adjustment text information obtained by performing a second adjustment operation on the style description text information.
[0242] Accordingly, the aforementioned target text information acquisition module is also configured to perform operations based on the content adjustment information and / or the style adjustment text information to obtain the target text information.
[0243] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0244] Figure 10 This is a block diagram illustrating an electronic device for image processing according to an exemplary embodiment. The electronic device may be a terminal, and its internal structure diagram may be as follows: Figure 10 As shown, the electronic device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage medium. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements an image processing method. The display screen can be a liquid crystal display (LCD) or an e-ink display. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad mounted on the device's casing, or an external keyboard, touchpad, or mouse.
[0245] Those skilled in the art will understand that Figure 10 The structure shown is merely a block diagram of a portion of the structure related to the present disclosure and does not constitute a limitation on the electronic device to which the present disclosure is applied. A specific electronic device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0246] Figure 11 This is a block diagram illustrating another electronic device for image processing according to an exemplary embodiment. The electronic device may be a server, and its internal structure diagram may be as follows: Figure 11As shown, the electronic device includes a processor, memory, and a network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements an image processing method.
[0247] Those skilled in the art will understand that Figure 11 The structure shown is merely a block diagram of a portion of the structure related to the present disclosure and does not constitute a limitation on the electronic device to which the present disclosure is applied. A specific electronic device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0248] In an exemplary embodiment, an electronic device is also provided, including: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to execute the instructions to implement the image processing method as described in the embodiments of this disclosure.
[0249] In an exemplary embodiment, a computer-readable storage medium is also provided, which, when executed by a processor of an electronic device, enables the electronic device to perform the image processing method of the present disclosure. The computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, or optical data storage device, etc.
[0250] In an exemplary embodiment, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to perform the image processing method of the present disclosure embodiments.
[0251] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.
[0252] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.
[0253] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. An image processing method, characterized in that, include: The image stylization processing page displays a content image input area and a style information input area; In response to the first image input in the content image input area and the style representation information input in the style information input area, a second image is incrementally displayed on the image stylization processing page. The content of the second image matches the content of the first image, and the style attribute of the second image is the target style, which is the style indicated by the style representation information. The incremental display of a second image on the image stylization processing page, in response to the first image input in the content image input area and the style representation information input in the style information input area, includes: When the style representation information is a third image, the third image is input into a style extraction model for style extraction processing to obtain style description text information; the style attribute of the third image is the target style. In response to the first image input in the content image input area and the style representation information input in the style information input area, the content text information and style description text information are incrementally displayed on the image stylization processing page; The content text information corresponds to the semantic content of the first image, and the style description text information is text information describing the target style; Upon detecting a first adjustment operation on the content text information and / or a second adjustment operation on the style description text information, the content adjustment information and / or style adjustment text information are displayed; the first adjustment operation includes triggering the content text information to be converted to an editable state and an editing operation in the editable state; the second adjustment operation includes triggering the style description text information to be converted to the editable state and an editing operation in the editable state. In response to a stylization processing instruction, a second image is incrementally displayed on the image stylization processing page; the second image is obtained based on the content adjustment information and the style adjustment text information; or the second image is obtained based on the content adjustment information and the style description text information; or the second image is obtained based on the content text information and the style adjustment text information. The style description text information includes at least two target styles; the second adjustment operation, which detects the style description text information, displays style adjustment text information, including: The image stylization processing page displays style priority selection information; In response to a priority confirmation instruction triggered based on the style priority selection information, it is determined that the second adjustment operation has been detected, and priority information corresponding to each of the at least two target styles is displayed; Based on the priority information and the style description text information, the style adjustment text information is generated and displayed.
2. The method according to claim 1, characterized in that, The method further includes: The first image is displayed on the image stylization page.
3. The method according to claim 1, characterized in that, The style representation information can be any of the following: image, text, audio, or video.
4. The method according to claim 1 or 2, characterized in that, The method further includes: The style representation information is subjected to style extraction processing to obtain style description text information; The first image is subjected to semantic extraction processing to obtain the content text information; The content text information and the style description text information are fused together to obtain the target text information; The target text information is input into a stylization processing model to obtain the second image.
5. The method according to claim 4, characterized in that, The style representation information is style indicator text information; The process of extracting style from the style representation information to obtain style description text information includes: The style instruction text information is processed by word ordering to obtain the style description text information; Alternatively, the style instruction text information can be input into a text processing model to obtain the style description text information.
6. The method according to claim 4, characterized in that, The first image is then subjected to semantic extraction processing. The obtained content text information includes: The first image is input into a semantic extraction model for semantic extraction processing to obtain the content text information.
7. The method according to claim 4, characterized in that, The method further includes: Obtain noise information; The step of inputting the target text information into a stylization processing model to obtain the second image includes: The noise information and the target text information are input into the stylization processing model to obtain the second image.
8. The method according to claim 7, characterized in that, The acquisition of noise information includes: The first image is subjected to noise processing to obtain the noise information; Alternatively, the noise information can be obtained by sampling a preset Gaussian noise.
9. An image processing method, characterized in that, include: Obtain the first image to be stylized and its style representation information; When the style representation information is a third image, the third image is input into a style extraction model for style extraction processing to obtain style description text information; the style attribute of the third image is the target style. The first image is subjected to semantic extraction processing to obtain the content text information; The system acquires content adjustment information obtained by performing a first adjustment operation on the content text information, and / or style adjustment text information obtained by performing a second adjustment operation on the style description text information; the first adjustment operation includes triggering the content text information to be converted to an editable state and an editing operation in the editable state; the second adjustment operation includes triggering the style description text information to be converted to the editable state and an editing operation in the editable state. Based on the content adjustment information and / or the style adjustment text information, the target text information is obtained; The target text information is input into a stylization processing model to obtain a second image; the content of the second image matches the content of the first image, and the style attribute of the second image is the target style, which is the style indicated by the style representation information. The style description text information includes at least two target styles; wherein, upon detecting a second adjustment operation of the style description text information, style adjustment text information is displayed, including: The image stylization page displays information on style priority selection. In response to a priority confirmation instruction triggered based on the style priority selection information, it is determined that the second adjustment operation has been detected, and priority information corresponding to each of the at least two target styles is displayed; Based on the priority information and the style description text information, the style adjustment text information is generated and displayed.
10. The method according to claim 9, characterized in that, The method further includes: Obtain noise information; The step of inputting the target text information into a stylization processing model to obtain a second image includes: The noise information and the target text information are input into the stylization processing model to obtain the second image.
11. The method according to claim 10, characterized in that, The acquisition of noise information includes: The first image is subjected to noise processing to obtain the noise information; Alternatively, the noise information can be obtained by sampling a preset Gaussian noise.
12. An image processing apparatus, characterized in that, include: The page display module is configured to execute the display of an image stylization processing page, which displays a content image input area and a style information input area; The stylized image display module is configured to incrementally display a second image on the image stylization processing page in response to a first image input in the content image input area and style representation information input in the style information input area. The content of the second image matches the content of the first image, and the style attribute of the second image is a target style, which is the style indicated by the style representation information. The style extraction module includes a first style extraction unit, which is configured to input the third image into a style extraction model and perform style extraction processing to obtain style description text information when the style representation information is a third image; the style attribute of the third image is the target style. The stylized image display module includes: The second text information display unit is configured to incrementally display content text information and style description text information on the image stylization processing page in response to the first image input in the content image input area and the style representation information input in the style information input area; the content text information corresponds to the semantic content of the first image, and the style description text information is text information describing the target style; The text adjustment unit is configured to perform a first adjustment operation and / or a second adjustment operation on the detected content text information and / or style description text information, and to display content adjustment information and / or style adjustment text information; the first adjustment operation includes an operation to trigger the content text information to be converted to an editable state and an editing operation in the editable state; the second adjustment operation includes an operation to trigger the style description text information to be converted to the editable state and an editing operation in the editable state; The second stylized image display unit is configured to incrementally display a second image on the image stylization processing page in response to a stylization processing instruction; the second image is obtained based on the content adjustment information and the style adjustment text information; or the second image is obtained based on the content adjustment information and the style description text information; or the second image is obtained based on the content text information and the style adjustment text information. The text adjustment unit includes a priority selection subunit, configured to display style priority selection information on the image stylization processing page; The second adjustment operation determination subunit is configured to execute a priority confirmation instruction triggered in response to the style priority selection information, determine that the second adjustment operation has been detected, and display the priority information corresponding to each of the at least two target styles. The style adjustment text display subunit is configured to generate and display the style adjustment text information based on the priority information and the style description text information.
13. An electronic device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the instructions to implement the image processing method as described in any one of claims 1 to 11.
14. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device is enabled to perform the image processing method as described in any one of claims 1 to 11.
Citation Information
Patent Citations
Image animation processing method and device, intelligent equipment and storage medium
CN113393545A
Image processing method and device, electronic equipment and storage medium
CN114266840A
Real-time Intelligent Image Manipulation System
US20190026870A1