Image generation method and apparatus, and electronic device and medium

By generating emoji detail prompts based on user-input emoji descriptions using an image generation model, the problem of tedious and time-consuming emoji image production in existing technologies is solved, and emoji images that meet user needs are generated efficiently.

WO2026021384A1PCT designated stage Publication Date: 2026-01-29VIVO MOBILE COMM CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/109603
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-24
Filing Date
2025-07-21
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

The current process of creating emoji images is cumbersome and time-consuming, requiring users to input multiple times and perform multiple image processing steps.

Method used

The image generation model generates emoji detail prompts based on the user's input emoji description information and outputs emoji images that match the emoji description information. The image generation model is obtained by removing the transparency channel from the original training images.

Benefits of technology

It simplifies the process of creating emoji images, improves production efficiency, and generates emoji images that better meet user needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025109603_29012026_PF_FP_ABST
    Figure CN2025109603_29012026_PF_FP_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of artificial intelligence. Disclosed are an image generation method and apparatus, and an electronic device and a medium. The method comprises: on the basis of sticker description information input by a user, generating sticker detail prompt words, which comprise sticker detail information in at least one description dimension; and inputting the sticker detail prompt words into an image generation model, and outputting a sticker image matching the sticker description information, wherein the image generation model is obtained by training at least one model training image, and the at least one model training image is obtained by executing transparency channel removal processing on at least one original training image.
Need to check novelty before this filing date? Find Prior Art

Description

Image generation method and device, electronic device, and medium

[0001] Cross-reference to Related Applications

[0002] This application claims priority to the Chinese patent application No. 202411000034.0 filed on July 24, 2024 in China, the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD

[0003] The present application belongs to the technical field of artificial intelligence, and particularly relates to an image generation method and device, an electronic device, and a medium. BACKGROUND

[0004] In Internet socialization, the application of expression package images is becoming more and more widespread. Through expression package images, the emotional state of a user can be vividly expressed, and the communication efficiency and interest can be improved.

[0005] At present, the production of expression package images is usually based on image processing and editing technology. For example, an electronic device can first perform cropping on an image through multiple inputs of a user, then adjust the size of the cropped image through multiple inputs of the user, and finally add text in the image after size adjustment through multiple inputs of the user, so as to obtain an expression package image corresponding to the image.

[0006] However, according to the above method, in the production process of the expression package image, multiple image processing needs to be performed on an image, and each image processing needs multiple inputs of the user to complete, thus resulting in a relatively cumbersome production process and a relatively long time consumption of the expression package image. SUMMARY

[0007] The purpose of the embodiments of the present application is to provide an image generation method and device, an electronic device, and a medium, which can simplify the production process of expression package images and shorten the time consumption.

[0008] In a first aspect, the embodiments of the present application provide an image generation method, which includes: generating an expression detail prompt word according to expression description information input by a user, the expression detail prompt word including expression detail information of at least one description dimension; inputting the expression detail prompt word into an image generation model to output an expression package image matched with the expression description information; wherein the image generation model is obtained by training at least one model training image, and the at least one model training image is obtained by performing a transparency channel removal processing on at least one original training image.

[0009] In a second aspect, an embodiment of the present application provides an image generation apparatus, the apparatus comprising: a generation module configured to generate an expression detail prompt word according to expression description information input by a user, the expression detail prompt word comprising expression detail information of at least one description dimension; and a processing module configured to input the expression detail prompt word into an image generation model and output an expression package image matching the expression description information, wherein the image generation model is obtained by training at least one model training image, and the at least one model training image is obtained by performing a transparency channel removal process on at least one original training image.

[0010] In a third aspect, an embodiment of the present application provides an electronic device, comprising a processor and a memory, wherein the memory stores programs or instructions executable on the processor, and the programs or instructions are executed by the processor to implement the steps of the method according to the first aspect.

[0011] In a fourth aspect, an embodiment of the present application provides a readable storage medium, wherein the readable storage medium stores programs or instructions, and the programs or instructions are executed by a processor to implement the steps of the method according to the first aspect.

[0012] In a fifth aspect, an embodiment of the present application provides a chip, comprising a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is configured to execute programs or instructions to implement the method according to the first aspect.

[0013] In a sixth aspect, an embodiment of the present application provides a computer program stored in a storage medium, and the computer program is executed by at least one processor to implement the method according to the first aspect.

[0014] In the embodiments of the present application, on the one hand, since the expression detail prompt word input into the image generation model is generated according to the expression description information input by the user, and the expression detail prompt word comprises expression detail information of at least one description dimension, the image generation apparatus can generate an expression package image more in line with the image requirements of the user according to the personalized image requirements of the user for the expression package image; on the other hand, since the image generation model can be used to directly output an expression package image matching the expression description information after the user inputs the expression description information, the user does not need to input multiple times to perform multiple image processing on an image during the process of making the expression package image, so that the process of making the expression package image can be simplified, and the efficiency of making the expression package image can be improved. BRIEF DESCRIPTION OF DRAWINGS

[0015] FIG. 1 is a flowchart of an image generation method according to some embodiments of the present application;

[0016] FIG. 2 is a schematic diagram of an interface to which the image generation method according to some embodiments of the present application is applied;

[0017] FIG. 3 is a schematic diagram of an interface applied by the image generation method according to some embodiments of the present application;

[0018] FIG. 4 is a schematic diagram of an interface applied by the image generation method according to some embodiments of the present application;

[0019] FIG. 5 is a schematic diagram of an interface applied by the image generation method according to some embodiments of the present application;

[0020] FIG. 6 is a schematic diagram of an interface applied by the image generation method according to some embodiments of the present application;

[0021] FIG. 7 is a schematic diagram of an interface applied by the image generation method according to some embodiments of the present application;

[0022] FIG. 8 is a flowchart of the image generation method according to some embodiments of the present application;

[0023] FIG. 9 is a flowchart of the image generation method according to some embodiments of the present application;

[0024] FIG. 10 is a flowchart of the image generation method according to some embodiments of the present application;

[0025] FIG. 11 is a flowchart of the image generation method according to some embodiments of the present application;

[0026] FIG. 12 is a flowchart of the image generation method according to some embodiments of the present application;

[0027] FIG. 13 is a flowchart of the image generation method according to some embodiments of the present application;

[0028] FIG. 14 is a schematic diagram of the network structure of the VAE in the image generation method according to some embodiments of the present application;

[0029] FIG. 15 is a schematic diagram of the network structure of the image generation model in the image generation method according to some embodiments of the present application;

[0030] FIG. 16 is a schematic diagram of the image generation apparatus according to some embodiments of the present application;

[0031] FIG. 17 is a schematic diagram of the electronic device according to some embodiments of the present application;

[0032] FIG. 18 is a schematic diagram of the hardware of the electronic device according to some embodiments of the present application. DETAILED DESCRIPTION

[0033] With reference to the drawings, the technical solutions in the embodiments of the present application will be clearly described below. Obviously, the described embodiments are only some of the embodiments of the present application, but not all of them. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art are within the scope of protection of the present application.

[0034] The terms "first", "second", and the like in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than that illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally a category and do not limit the number of objects, for example, the first object can be one or more. In addition, "and / or" in the specification and claims indicates at least one of the connected objects, and the character " / ", generally indicates that the objects before and after are in an "or" relationship.

[0035] The term "indication" in the present application can be a direct indication (or explicit indication) or an indirect indication (or implicit indication). Among them, the direct indication can be understood as that the sender explicitly informs the receiver of specific information, operations to be performed or requested results, etc. in the sent indication; the indirect indication can be understood as that the receiver determines the corresponding information according to the indication sent by the sender, or judges and determines the operation to be performed or the requested result according to the judgment result.

[0036] The terms "at least one", "at least one of", and the like in the specification and claims of the present application refer to any one of the objects contained therein, a combination of any two or more than two. For example, at least one of a, b, and c can mean "a", "b", "c", "a and b", "a and c", "b and c", and "a, b, and c", wherein a, b, and c can be single or multiple. Similarly, "at least two" refers to two or more, and its meaning is similar to that of "at least one".

[0037] Some terms or phrases involved in the specification and claims of the present application will be explained below.

[0038] Artificial Intelligence (AI) is an important driving force for the new round of technological revolution and industrial transformation. It is a new technical science that studies, develops and applies systems for simulating, extending and expanding human intelligence. AI is an important part of the field of intelligence, which aims to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. AI is a very broad science, including robots, language recognition, image recognition, natural language processing, expert systems, machine learning, computer vision, etc.

[0039] Mask is a very important concept in the field of image processing and computer vision, usually used to specify a certain area of an image so that a certain operation can be performed on that area without affecting other parts of the image. Mask is usually a binary or Boolean image of the same size as the original image, where the selected area is marked as 1 and the rest is marked as 0. When applying an operation to an image, the mask can be used to limit the operation to only occur within a specific area of the image.

[0040] Transparency channel, also known as alpha channel, represents the transparency information of pixels in a digital image. White alpha pixels are used to define non-transparent color pixels, while black alpha pixels are used to define transparent pixels, and gray levels between black and white are semi-transparent parts of color images.

[0041] Text-to-image model is an advanced artificial intelligence model that can convert descriptive text into corresponding images. The core function of this model is to interpret user input natural language descriptions and generate visual content based on these descriptions. The implementation of this model usually relies on deep learning frameworks, especially generative adversarial networks or variational autoencoders.

[0042] Image super-resolution is the process of recovering a high-resolution image from a low-resolution image or image sequence. Image super-resolution technology can be divided into super-resolution restoration and super-resolution reconstruction. Image super-resolution research can be divided into three main categories: interpolation-based, reconstruction-based, and learning-based methods.

[0043] The image generation method, device, electronic equipment and medium provided by the embodiments of the present application will be described in detail below in combination with the drawings and through specific embodiments and their application scenarios.

[0044] In Internet socialization, the application of meme images is becoming more and more widespread. Compared with text and voice communication, meme images can effectively improve communication efficiency and interest. Specifically, the value of meme images can be reflected in the following seven aspects:

[0045] 1. Emotion expression enhancer: Emoticon images can vividly express the user's emotional state, such as the user's joy, sadness, surprise, etc., and can convey the user's emotional color more accurately than text;

[0046] 2. Communication efficiency improvement: When language is not enough to quickly and intuitively express a complex emotion or reaction, emoticon images can serve as a visual shorthand to quickly convey information;

[0047] 3. Social atmosphere regulator: Certain emoticon images can alleviate tense or awkward atmosphere, making conversations more relaxed and comfortable;

[0048] 4. Cross-cultural communication medium: Certain emoticon images express non-verbal meanings that are common across different cultures, helping to overcome language barriers and facilitate communication;

[0049] 5. Personalization and creative expression: Personalized emoticon images can showcase the user's style and creativity, enhancing personal characteristics;

[0050] 6. Topic guidance and switching: Using appropriate emoticon images in conversation can serve as a non-verbal cue to introduce new topics or change the direction of the conversation;

[0051] 7. Trend and sense of belonging: Following popular emoticon image trends can enhance people's sense of belonging and fashion in social interactions.

[0052] With the continuous development of AI technology, new ideas are provided to solve the above problems. First, AI can analyze different user groups' preferences for emoticon images through big data analysis, and thus design personalized emoticon images accurately; second, AI's automatic generation of emoticon images can reduce the cost of human labor in design and production; finally, AI's learning and optimization functions can continuously improve the design of emoticon images, achieving continuous improvement in quality. Based on this, the embodiments of the present application provide an image generation method, device, electronic equipment and medium. The image generation method provided by the embodiments of the present application can be applied to the production scene of emoticon images.

[0053] Illustratively, when a user chats with friends, family or colleagues through an instant messaging chat application installed in an electronic device, needs to send an emoticon image, for example, Xiaozhang and Xiaowang chat through an instant messaging application, when chatting about today Xiaowang helping Xiaozhang repair the computer, in order to thank Xiaowang for his help, Xiaozhang needs to send Xiaomimi a thank you emoticon image to Xiaowang; for example, Xiaoli and Xiaochen chat through an instant messaging application, when Xiaochen complains to Xiaoli about the injustice at work, in order to comfort Xiaochen, Xiaoli needs to send Xiaochen a hug emoticon image to Xiaobear.

[0054] It should be noted that the image generation method provided in the embodiments of the present application can be executed by an image generation device, an electronic device, or a functional module in the electronic device, etc. In some embodiments of the present application, the image generation method is executed by an electronic device as an example to illustrate the image generation method provided in the embodiments of the present application.

[0055] FIG. 1 shows a flowchart of the image generation method provided in some embodiments of the present application. As shown in FIG. 1, the image generation method provided in some embodiments of the present application can include the following steps 101 and 102.

[0056] The image generation method provided in some embodiments of the present application can be applied in the scenario of making an expression package image; for example, a user inputs text information to quickly generate an expression package image corresponding to the text information.

[0057] Step 101: The electronic device generates an expression detail prompt word according to expression description information input by a user.

[0058] The expression detail prompt word includes at least one expression detail information of a description dimension.

[0059] In some embodiments of the present application, the expression description information is information describing an expression package image that a user expects to generate, and is used to generate a corresponding expression detail prompt word. The expression description information can include, but is not limited to, at least one of the following: description information of a subject of the expression package image that the user expects to generate, description information of a style of the expression package image that the user expects to generate, description information of an expression type of the expression package image that the user expects to generate, description information of a subject behavior of the expression package image that the user expects to generate, and description information of an image element of the expression package image that the user expects to generate, etc.

[0060] In some embodiments of the present application, the expression description information can be input by a user in any possible form, such as voice input, touch input, or input by a stylus on a screen of the electronic device.

[0061] For example, taking the expression description information input by the user through touch input as an example, the user can input the expression description information by clicking a virtual key in an input method interface; the input method interface can be an input interface of any input method in the electronic device.

[0062] In some embodiments of the present application, the electronic device can first extract and generate corpus information matched with the expression description information through a large language model, and then construct the expression detail prompt word according to the corpus information.

[0063] In some embodiments of the present application, the above-mentioned corpus information can include but is not limited to at least one of the following: implied subject, emotion, object, behavior element, and whether it is a colloquial text label.

[0064] The above-mentioned subject refers to objects of different vertical categories such as characters, animals, and cute pets, which can include boy, girl, cat, dog, and other emoticon subjects. The above-mentioned emotion refers to the expression of the subject's facial emotion and expression, including basic joy, anger, sadness, and other emotions unique to the emoticon field such as gratitude, approval, fatigue, boredom, and complex contempt, and staring. The above-mentioned behavior element refers to the element information displayed in the above-mentioned expression description information, including or metaphorically expanding, such as sending flowers to express gratitude and lying down to express fatigue. The above-mentioned whether it is a colloquial text label represents whether the above-mentioned expression description information is a colloquial information, such as "thank you" which is a colloquial information and is suitable for emoticon text embedding, while "Zhang San smiles" is a description type text and is not suitable for emoticon text embedding.

[0065] For example, as shown in Table 1 below, after the user inputs the expression description information, the corresponding corpus information can be extracted by the above-mentioned language model.

[0066] Table 1

[0067] In some embodiments of the present application, the electronic device can construct the above-mentioned expression detail prompt word according to the above-mentioned corpus information.

[0068] In some embodiments of the present application, the above-mentioned expression detail prompt word is a specific format language instruction for guiding the image generation model to output the emoticon image that the user expects to generate. The expression detail prompt word can be used as the input of the image generation model, so that the image generation model generates the corresponding emoticon image.

[0069] In some embodiments of the present application, the specific format of the above-mentioned expression detail prompt word can be in the format of "{subject}, {emotion}, {behavior information}".

[0070] In some embodiments of the present application, the above-mentioned description dimension can include at least one of the following: emoticon image subject, emoticon image subject emotion, emoticon image subject behavior, emoticon image background, and emoticon image style.

[0071] Exemplarily, assuming that the expression description information input by the user is "even the kittens think it's great", the corpus information extracted by the large language model includes: the subject is a kitten, the emotion is approval, and the behavior information is thumbs up, so the expression detail prompt word can be constructed in the format of "{subject}, {emotion}, {behavior information}", and the constructed expression detail prompt word is "kitten, approval, thumbs up". If the corpus information identified does not contain a subject, an expression detail prompt word can be generated by randomly selecting a subject from a preset prompt word subject library. In addition, a style, such as a simple pen, cartoon, real, illustration, and the like, and a background word can be randomly selected from a preset prompt word style library to construct a final expression detail prompt word; for example, after adding a style word and a background word to the expression detail prompt word "kitten, approval, thumbs up", it becomes "kitten, approval, thumbs up, white background, cartoon style". In this way, the expression detail prompt word constructed can contain more dimensions and can more accurately describe the corresponding image.

[0072] In some embodiments of the present application, since the description dimension described above can include at least one of the following: expression package image subject, expression package image subject emotion, expression package image subject behavior, expression package image background, and expression package image style, the expression detail prompt word of different description dimensions can be generated according to the expression description information input by the user, so that the generated expression detail prompt word can match the information input by the user from different angles.

[0073] Step 102, the electronic device inputs the expression detail prompt word into an image generation model, and outputs an expression package image matched with the expression description information.

[0074] The image generation model is obtained by training at least one model training image, and the at least one model training image is obtained by performing a transparency channel removal process on at least one original training image.

[0075] It should be noted that the expression package image is usually an image that uses popular stars, quotes, animation, and film screenshots as materials, and matches a series of texts to express specific emotions. The expression package image essentially belongs to a kind of popular culture, relying on the continuous development of social and network, the communication between people has gradually evolved from the earliest text communication to the use of some simple symbols, expressions, and expression packages, and gradually evolved into an increasingly diversified expression culture, using some self-made popular element images for communication.

[0076] In some embodiments of the present application, the expression package image can be a static expression package image or a dynamic expression package image.

[0077] For example, the format of the sticker image that is static can include, but is not limited to, a Joint Photographic Experts Group (JPEG) format, a Portable Network Graphics (PNG) format, or a Bitmap (BMP) format, and the like; the format of the sticker image that is dynamic can include, but is not limited to, a Graphics Interchange Format (GIF) or an Animated Portable Network Graphics (APNG) format, and the like. In some embodiments of the present application, the sticker image output by the image generation model and matched with the sticker description information can be one or more.

[0078] In some embodiments of the present application, the at least one original training image described above can be a sticker image.

[0079] In some embodiments of the present application, any one of the at least one original training image described above can be a sticker image including text content, or can be a sticker image not including text content.

[0080] In some embodiments of the present application, the at least one original training image described above can be an image stored in an electronic device, or can be an image stored in another electronic device, or can be an image stored in a server, and the like.

[0081] In some embodiments of the present application, the image generation model described above can also be referred to as a text-to-sticker model.

[0082] The training method of the image generation model described above will be described in detail in the embodiments described below, and will not be described again here to avoid repetition.

[0083] In some embodiments of the present application, after the user inputs the sticker description information, the sticker image corresponding to the sticker description information can be output by the image generation model.

[0084] In some embodiments of the present application, the number of sticker images output by the image generation model and matched with the sticker description information can be one or more.

[0085] Exemplarily, in the process of chatting with friends, family or colleagues through the instant messaging chat application installed in the electronic device, an application scenario of sending an expression package image is needed, as shown in FIG. 2, Xiaozhang and Xiaowang chat through the instant messaging application, when talking about today Xiaowang helping Xiaozhang repair the computer, in order to thank Xiaowang for his help, Xiaozhang needs to send Xiaowang an expression package image of “I love you”. Xiaozhang inputs the expression description information “I love you” in the input box 11 through the input method interface, and then clicks the “expression package image generation” control 12. Then as shown in FIG. 3, the electronic device can start to generate the expression package image corresponding to the expression description information, and display the expression package image generation identifier in the preview area 13 of the input method interface. The specific generation process is as follows: the electronic device first acquires the expression description information input by the user, then generates the expression detail prompt word corresponding to the expression description information according to the expression description information, and finally inputs the expression detail prompt word into the above-mentioned image generation model to output the expression package image corresponding to the expression description information. Finally, as shown in FIG. 4, the electronic device can display the output expression package image 14, expression package image 15 and expression package image 16 in the preview area 13 of the input method interface. In this way, Xiaozhang can select the needed expression package image to send to Xiaowang to express thanks.

[0086] Exemplarily, in the process of chatting with friends, family or colleagues through the instant messaging chat application installed in the electronic device, another application scenario of sending an expression package image is needed, as shown in FIG. 5, Xiaoli and Xiaochen chat through the instant messaging application, when Xiaochen tells Xiaoli about the grievances in work, in order to comfort Xiaochen, Xiaoli needs to send Xiaochen an expression package image of “hug the little bear”. Xiaozhang inputs the expression description information “hug the little bear” in the input box 21 through the input method interface, and then clicks the “expression package image generation” control 22. Then as shown in FIG. 6, the electronic device can start to generate the expression package image corresponding to the expression description information, and display the expression package image generation identifier in the preview area 23 of the input method interface. The specific generation process is as follows: the electronic device first acquires the expression description information input by the user, then generates the expression detail prompt word corresponding to the expression description information according to the expression description information, and finally inputs the expression detail prompt word into the above-mentioned image generation model to output the expression package image corresponding to the expression description information. Finally, as shown in FIG. 7, the electronic device can display the output expression package image 24 in the preview area 23 of the input method interface. In this way, Xiaoli can send the expression package image 24 to Xiaochen to comfort Xiaochen.

[0087] In the image generation method provided by some embodiments of the present application, on the one hand, since the expression detail prompt word input into the image generation model is generated according to the expression description information input by the user, and the expression detail prompt word includes expression detail information in at least one description dimension, the image generation model can generate an expression package image that meets the image requirements of the user according to the personalized expression package image of the user; on the other hand, since the image generation model can directly output an expression package image matching the expression description information after the user inputs the expression description information, the user does not need to perform multiple image processing on an image through multiple inputs in the process of making an expression package image, so that the process of making an expression package image can be simplified, and the efficiency of making an expression package image can be improved.

[0088] In some embodiments of the present application, after the above step 102, the image generation method provided by some embodiments of the present application can further include the following step A.

[0089] Step A, the electronic device displays the expression package image in the expression package image preview area of the chat interface.

[0090] In some embodiments of the present application, the chat interface can be a chat interface between the user and any other user in the instant messaging application of the electronic device, or can be a chat interface in a video playing process.

[0091] In some embodiments of the present application, the preview area can be any area in the chat interface; for example, the preview area can be a bottom area or a top area in the chat interface, etc.

[0092] In some embodiments of the present application, the display parameter of the preview area can be any display parameter; the display parameter includes but is not limited to transparency, color, shape, size, etc.

[0093] In some embodiments of the present application, as shown in FIG. 8, before the above step 101, the image generation method provided by some embodiments of the present application can further include the following step 103 or step 104; the above step 101 can be implemented through the following step 101a, step 101b or step 101c.

[0094] Step 103, the electronic device receives expression description information input by the user in the message input area of the chat interface.

[0095] In some embodiments of the present application, the chat interface can be a chat interface between the user and any other user; for example, the chat interface can be a chat interface between the user and a contact A in the instant messaging application of the electronic device.

[0096] In some embodiments of the present application, the message input area is used for inputting a chat message.

[0097] In some embodiments of the present application, the user can input the expression description information through any possible form of input such as touch input or voice input to the message input area.

[0098] Step 104, the electronic device receives the expression description information input by the user in the bullet screen input area in the video playing interface.

[0099] In some embodiments of the present application, the video playing interface can be the playing interface of any video in a video website or a video playing application, and the playing interface includes the bullet screen input area.

[0100] In some embodiments of the present application, the bullet screen input area is used for inputting bullet screens.

[0101] In some embodiments of the present application, the user can input the expression description information through any possible form of input such as touch input or voice input to the bullet screen input area.

[0102] Step 101a, in the case that the input area of the expression description information is the message input area of the chat interface, the electronic device generates the expression detail prompt word according to the expression description information input by the user.

[0103] In some embodiments of the present application, after the user inputs the expression description information in the message input area, the electronic device can automatically generate the expression detail prompt word according to the expression description information.

[0104] For example, assuming that the user inputs the expression description information "thank you" in the message input area, the electronic device can automatically generate the expression detail prompt word corresponding to "thank you" according to the expression description information "thank you", and then generate the sticker image corresponding to "thank you" through the image generation model. In this way, the process of generating the sticker image can be simplified.

[0105] Step 101b, in the case that the input area of the expression description information is the message input area of the chat interface and the touch input of the sticker image generation control is received, the electronic device generates the expression detail prompt word according to the expression description information input by the user.

[0106] In some embodiments of the present application, the touch input described above is used to control the electronic device to generate the sticker image, and the touch input can be a touch operation. Illustratively, the touch input described above includes, but is not limited to, a touch input of the sticker image generation control by a user through a finger or a stylus and the like touch device, or a specific gesture input by the user, or a touch input of the sticker image generation control by the user through a voice assistant to control the electronic device, or other feasible inputs, which can be determined according to actual use requirements, and embodiments of the present application are not limited. The specific gesture in the embodiments of the present application can be any one of a single-click gesture, a sliding gesture, a drag gesture, a pressure recognition gesture, a long-press gesture, an area change gesture, a double-press gesture, and a double-click gesture; the click input in the embodiments of the present application can be a single-click input, a double-click input, or a click input of any number of times, and can also be a long-press input or a short-press input. For example, the touch input described above can be a click input of the sticker image generation control by the user.

[0107] In some embodiments of the present application, the sticker image generation control described above is used to trigger generation of an expression detail prompt word, and then output the sticker image.

[0108] In some embodiments of the present application, the control parameter of the sticker image generation control described above can be any parameter; the control parameter includes, but is not limited to, size, color, transparency, brightness, shape, and the like.

[0109] In some embodiments of the present application, after the user inputs the expression description information in the message input area described above, the electronic device does not automatically generate the expression detail prompt word, but generates the expression detail prompt word after receiving the touch input of the sticker image generation control described above.

[0110] Illustratively, assuming that the user inputs the expression description information "sad" in the message input area described above, the electronic device does not automatically generate the corresponding expression detail prompt word according to the expression description information "sad". When the user needs to make a sticker image corresponding to the expression description information "sad", the user can perform a touch input on the sticker image generation control described above, and the electronic device can generate the expression detail prompt word corresponding to "sad" according to the expression description information "sad" after receiving the touch input, and then generate the sticker image corresponding to "sad" through the image generation model described above. In this way, the sticker image can be generated according to the user's needs, and the flexibility of generating the sticker image is improved.

[0111] Step 101c, in the case of the input area of the expression description information being a bullet screen input area in a video playing interface, the electronic device generates an expression detail prompt word according to the expression description information input by the user.

[0112] In some embodiments of the present application, after the user inputs the expression description information in the above mentioned barrage input area, the electronic device can automatically generate the expression detail prompt word according to the expression description information.

[0113] For example, assuming that the user inputs the expression description information "wonderful" in the above mentioned barrage input area, the electronic device can automatically generate the expression detail prompt word corresponding to "wonderful" according to the expression description information "wonderful", and then generate the sticker image corresponding to "wonderful" through the above mentioned image generation model. In this way, the process of generating the sticker image can be simplified.

[0114] In some embodiments of the present application, since the user can input the expression description information in the above mentioned message input area or the above mentioned barrage input area, the electronic device can automatically or through the user's input to the function control to trigger the generation of the expression detail prompt word corresponding to the expression description information after receiving the expression description information, thereby improving the flexibility of generating the expression detail prompt word.

[0115] In some embodiments of the present application, as shown in FIG. 9, after the step 102, the image generation method provided by some embodiments of the present application can further include the following step 105.

[0116] Step 105, in the case that the expression description information includes a spoken keyword, the electronic device adds the expression description information to the sticker image to display the expression description information on the sticker image.

[0117] In some embodiments of the present application, the above mentioned spoken keyword can be "thank you" or "like" and the like which can represent spoken language.

[0118] It can be understood that if the above mentioned expression description information includes the above mentioned spoken keyword, it can be considered that the expression description information is a spoken text; if the above mentioned expression description information does not include the above mentioned spoken keyword, it can be considered that the expression description information is a non-spoken text, which can also be referred to as a written text.

[0119] In some embodiments of the present application, the electronic device can perform foreground and background separation on the transparency binary mask of the above mentioned sticker image, and determine the pixel distance Dist up and Dist down of the minimum and maximum height values of the mask pixel value from the upper and lower picture edges according to the upper and lower edge positions of the segmented mask. up <Dist down If Dist up ≥ Dist downIf the expression description information is added to the top of the sticker image, the expression description information is added to the top of the sticker image.

[0120] It should be noted that the up, down, left, right involved in some embodiments of the present application are all shown with the screen of the electronic device facing the user.

[0121] In some embodiments of the present application, the electronic device can superimpose the expression description information onto the sticker image.

[0122] For the specific method of adding the expression description information to the sticker image, please refer to the related description in the related art. To avoid repetition, it will not be described here.

[0123] In some embodiments of the present application, in the case that the expression description information does not include the spoken keyword, the electronic device can not process the sticker image.

[0124] In some embodiments of the present application, in the case that the expression description information includes the written keyword, the electronic device can convert the written keyword into a spoken keyword, and then add the expression description information to the sticker image to display the expression description information on the sticker image.

[0125] In some embodiments of the present application, since the expression description information is added to the sticker image in the case that the expression description information includes the spoken keyword, the expression description information can be added to the sticker image when the expression description information input by the user is suitable as the sticker text, thereby improving the interestingness of the sticker image.

[0126] The image generation method provided by some embodiments of the present application is an end-to-end AI transparent sticker generation scheme, which creatively proposes to directly convert user text to generate exquisite personalized sticker images with matched text and images, combines emotion recognition and natural language entity recognition to greatly improve the semantic understanding ability of user corpus, and combines a language large model to convert to an image semantic description domain and input an image generation model to realize efficient and one-key generation of AI stickers.

[0127] In some embodiments of the present application, as shown in FIG. 10, before the step 101, the image generation method provided by some embodiments of the present application can further include the following steps 106 and 107.

[0128] It should be noted that in FIG. 7, only the steps 106 and 107 described above are executed before the step 101 described above as an example. In actual implementation, the steps 106 and 107 described above can also be executed after the step 101 and before the step 102, and the embodiments of the present application are not limited thereto.

[0129] The step 106 includes that the electronic device performs image processing on the at least one original training image to obtain an image transparency mask of each original training image and at least one model training image.

[0130] The image processing includes at least removing a transparency channel.

[0131] It should be noted that in order to achieve better visualization effect when using the sticker image, the background area of the sticker image is usually not shielded. In fact, a transparency channel is added in the process of making the sticker image, and the RGB three channels of the original image are changed to RGBA four channels, where A is the transparency channel, that is, the alpha channel.

[0132] In some embodiments of the present application, the electronic device can obtain the image transparency mask of each original training image in the at least one original training image by removing the transparency channel of the at least one original training image.

[0133] In some embodiments of the present application, as shown in FIG. 11, the step 106 can be implemented by the steps 106a, 106b and 106c described below in combination with FIG. 10.

[0134] The step 106a includes that the electronic device performs image processing on the at least one original training image to obtain an image transparency mask of each original training image and a corresponding intermediate training image of each original training image.

[0135] In some embodiments of the present application, the at least one intermediate training image corresponds to the at least one original training image one by one.

[0136] It can be understood that the electronic device can obtain the image transparency mask of the original training image and the corresponding intermediate training image of the original training image based on the image processing on any one of the at least one original training image.

[0137] The step 106b includes that the electronic device determines at least one target training image from the at least one intermediate training image corresponding to the at least one original training image based on the image description information and the image parameters of each intermediate training image.

[0138] In some embodiments of the present application, the electronic device can describe each intermediate training image by using a text-image contrastive pre-training (CLIP) model or a large language and vision assistant (LLaVA) model, to obtain image description information of each intermediate training image.

[0139] In some embodiments of the present application, the image description information can include, but is not limited to, information of at least one of the following dimensions: subject, emotion, behavior, background, style, etc.

[0140] In some embodiments of the present application, the image description information can also be a prompt word.

[0141] In some embodiments of the present application, the image parameters can include, but are not limited to, at least one of the following: image resolution, image brightness, image transparency, image depth, image color temperature, image size, etc.

[0142] In some embodiments of the present application, the electronic device can score each intermediate training image according to the image description information and the image parameters of the intermediate training image, and remove the intermediate training images with scores lower than a preset threshold from the at least one intermediate training image, so as to determine the remaining intermediate training images as the at least one target training image.

[0143] For example, the electronic device can score each intermediate training image according to the matching degree of the intermediate training image and its image description information, and the background color or resolution of the intermediate training image, and then remove the intermediate training images with scores lower than a preset threshold from the at least one intermediate training image, and determine the remaining intermediate training images as the at least one target training image.

[0144] Step 106c: The electronic device performs image super-resolution processing on the at least one target training image to obtain at least one model training image corresponding to the at least one target training image.

[0145] It can be understood that the electronic device can perform image super-resolution processing on any target training image in the at least one target training image to obtain a model training image corresponding to the target training image.

[0146] In some embodiments of the present application, the resolution of each model training image in the at least one model training image is greater than the resolution of the target training image corresponding thereto.

[0147] In some embodiments of the present application, the electronic device can input the at least one target training image into the image super-resolution model to obtain the at least one model training image.

[0148] It should be noted that the image super-resolution model can increase the resolution of the input image and output the input image with adjusted resolution. For specific description of the image super-resolution model, reference can be made to the related description in the related art, which will not be repeated here.

[0149] In some embodiments of the present application, after obtaining the image transparency mask of each original training image and the at least one intermediate training image based on the image processing of the at least one original training image, the electronic device can screen the at least one intermediate training image based on the image description information and the image parameters of each intermediate training image, so as to ensure that the image quality of the screened intermediate training image is high. Then, the image super-resolution processing is performed on the screened intermediate training image to obtain the at least one model training image, so that the display effect of the obtained model training image is better.

[0150] In some embodiments of the present application, the image processing can further include text content removal processing. For example, as shown in FIG. 12, the step 106a can be implemented by the following steps 106a1 and 106a2.

[0151] Step 106a1, the electronic device performs transparency channel removal processing on the at least one original training image to obtain the image transparency mask of each original training image and the channel processing image corresponding to each original training image.

[0152] It can be understood that the electronic device performs transparency channel removal processing on any one of the at least one original training image to obtain the image transparency mask of the original training image and a channel processing image corresponding to the original training image.

[0153] For example, assuming that one of the at least one original training image is represented as A(W, H, C), where C=4; the transparency channel removal processing performed by the electronic device on the original training image A can include: filling the image area with alpha channel as 0 with the regular sticker image background color white according to whether the alpha channel of the third dimension index 3 (starting from 0) of the original training image A is 0, and the corresponding pixel value is 255. Thus, a channel processing image B corresponding to the original training image and an image transparency mask M of the original training image A can be obtained, and the specific calculation formula is shown in the following formula (1) and formula (2): M(x, y) = A(x, y, c = 3), 0≤c≤2. (2)

[0154] In step 106a2, the electronic device performs text removal processing on the image containing text in the channel-processed image corresponding to at least one original training image to obtain an intermediate training image corresponding to each original training image.

[0155] In some embodiments of the present application, the text described above can be all the text contained in the at least one channel-processed image described above.

[0156] For example, assuming that the channel-processed image 1 in the at least one channel-processed image described above contains text, the electronic device can detect the region position corresponding to the text in the channel-processed image 1 by using an optical character recognition (OCR) algorithm, and construct a whiteboard mask M according to the region position by using the following formula (3). ocr (x, y):

[0157] Then, the electronic device can input the channel-processed image 1 and M ocr (x, y) into the image implicit diffusion repair model to remove the text content, and obtain an intermediate training image B NoOCR (x, y, c) corresponding to the channel-processed image 1.

[0158] In some embodiments of the present application, since the electronic device can perform the transparency removal channel processing on the at least one original training image to obtain the image transparency mask of each original training image and the at least one channel-processed image, and perform the text removal processing on the image containing text in the at least one channel-processed image to obtain the at least one intermediate training image, the model training image can be an image not containing text content, thereby facilitating the flexible addition of text content after generating the meme image.

[0159] In step 107, the electronic device trains the initial model based on the at least one model training image, the image description information of the at least one model training image, and the at least one image transparency mask to obtain an image generation model.

[0160] The at least one image transparency mask is the image transparency mask of the original training image corresponding to the at least one model training image.

[0161] It can be understood that each image transparency mask in the at least one image transparency mask is the image transparency mask of the original training image corresponding to one model training image in the at least one model training image.

[0162] In some embodiments of the present application, the initial model described above can be a text-to-image model.

[0163] In some embodiments of the present application, the initial model described above can include an image encoder and a text encoder. For example, as shown in FIG. 13, the step 107 described above can be implemented by the following steps 107a and 107b in combination with FIG. 10.

[0164] Step 107a, the electronic device inputs at least one model training image and at least one image transparency mask into the image encoder for encoding, outputs image feature information, and inputs image description information of the at least one model training image into the text encoder for encoding, and outputs text feature information.

[0165] In some embodiments of the present application, the image encoder described above can encode the image into latent space features.

[0166] In some embodiments of the present application, the initial model described above can further include an image decoder, which can decode the image latent space features into a transparency image.

[0167] In some embodiments of the present application, the image encoder and the image decoder described above can be two modules in a variational auto-encoder (VAE).

[0168] For example, the VAE described above can increase one dimension of the image encoder input and the image decoder output convolution layer based on the multiplexing stable diffusion 1.5 encoder, adapt to the newly added mask map, change from the original (B, 3, H, W) to (B, 4, H, W), and the latent space feature dimension is (W / 8) x (H / 8) x 4, then unchanged, the original UNet weight can be reused and fine-tuned; where B is batch_size, H is image height, and W is image width. The overall network structure is shown in FIG. 14, wherein the down-sampling encoding block and the up-sampling decoding block adopt the Resnet Block residual structure, the down-sampling structure is realized by the convolution layer with stride 2, and the up-sampling is realized by the nearest neighbor up-sampling operator supported by the Pytorch framework, i.e. F. interpolate and kernel is 3x3, and the convolution layer with stride 1 is realized in series. When training the VAE with alpha channel, the reconstruction loss is defined by encoding and decoding the input image of the training, then the corresponding loss is L2 loss, and the specific calculation formula is shown in the following formula (4):

[0169] L recon = ||Bin-Brec||2. (4)

[0170] In some embodiments of the present application, the text encoder described above can encode the image prompt word into the feature space.

[0171] In step 107b, the electronic device predicts a model training loss value based on the image feature information and the text feature information, and trains the initial model based on the model training loss value to obtain the image generation model.

[0172] In some embodiments of the present application, the electronic device can update the model parameters of the initial model according to the model training loss value described above to train the initial model to obtain the image generation model described above.

[0173] For the specific method of the electronic device predicting the model training loss value based on the image feature information and the text feature information described above, please refer to the relevant description in the related art. To avoid repetition, this will not be repeated here.

[0174] Exemplarily, as shown in FIG. 15, the image generation model described above can include an image encoder, an image decoder, a text encoder, and a UNet. The electronic device can input the at least one model training image and the at least one image transparency mask into the image encoder to obtain image feature information, and input the image description information corresponding to the at least one model training image into the text encoder to obtain text feature information; and input the image feature information and the text feature information into the UNet to predict a model training loss value, and train the UNet based on the model training loss value to obtain the image generation model described above.

[0175] In some embodiments of the present application, since the electronic device can predict a model training loss value based on image feature information and text feature information, and train an initial model based on the model training loss value to obtain an image generation model, the model training loss value can optimize the accuracy of the model, thereby improving the accuracy of the image generation model obtained.

[0176] In some embodiments of the present application, since the image generation model for generating meme images can be trained based on at least one model training image and its image description information, and the image transparency mask of the original training sample image corresponding to the at least one model training image, the meme image can be directly generated through the image generation model, without the need for the user to perform multiple image processing on an image through multiple inputs, thereby simplifying the production process of the meme image and improving the production efficiency of the meme image.

[0177] From the above, the image generation model in some embodiments of the present application is actually a text-to-image sticker diffusion model with transparency, which can directly generate sticker images with an alpha channel of transparency, avoiding the labor and energy loss of secondary or manual matting. Through the process of text-to-image sticker data collection, cleaning and training, the processing of batch collecting sticker data from the network can be efficiently realized, and directly used for training of the image generation model. The overall process includes data crawling preprocessing, text removal, image-text description, image super-resolution, image-text description to obtain prompt words, and filtering low matching degree image-text pairs through image-text matching scoring, filtering low-quality pictures through aesthetic scoring, adding style words to prompt words after style classification, etc.

[0178] The image generation method provided by some embodiments of the present application is exemplarily described below.

[0179] Exemplarily, the electronic device can train the above image generation model through the following steps A to N:

[0180] Step A, crawling sticker sample images on a large scale;

[0181] Step B, extracting the transparency channel in the sticker sample image, removing the background effect of the RGB three-channel image, obtaining the sticker sample image after removing the transparency channel, and the mask information of the sticker sample image;

[0182] Step C, removing the text content in the sticker sample image after removing the transparency channel to obtain the sticker sample image after removing the text content;

[0183] Step D, obtaining the image description information, i.e. prompt words, of the sticker sample image after removing the text content through an image-text description model;

[0184] Step E, filtering the images with low image quality in the sticker sample image after removing the text content to obtain the screened sticker sample image;

[0185] Step F, performing super-resolution processing on the screened sticker sample image to enhance image clarity and increase resolution, obtaining a model training image.

[0186] Step G, taking the model training image and its image description information, and the mask information of the original sample image corresponding to the model training image as the input of a text-to-image model, training the text-to-image model to obtain the above image generation model.

[0187] Step H, the user inputs the expression description information, such as "thank you", "send you a flower" and the like;

[0188] Step I, extracting and generating corpus information of different dimensions matching the input expression description information through the LLM large language model;

[0189] Step J, in the case that the above corpus information includes a subject, constructing an expression detail prompt word corresponding to the above text corpus according to the corpus information; otherwise, proceeding to Step K;

[0190] Step K, in the case that the above corpus information does not include a subject, randomly selecting a subject from a pre-set prompt word subject library, and constructing an expression detail prompt word corresponding to the above corpus information;

[0191] Step L, randomly selecting a style from a pre-set prompt word style library, and constructing a final expression detail prompt word;

[0192] Step M, inputting the constructed expression detail prompt word into the above image generation model, and outputting an expression package image corresponding to the expression description information input by the user;

[0193] Step N, in the case that the expression description information input by the user includes a spoken keyword, adding the expression description information to the above expression package image to obtain a final expression package image.

[0194] In this way, by utilizing the natural language reasoning capability of the large model, the user input text can be converted into a prompt word input suitable for the text-to-image large model in one key, and the image generation model is used for reasoning to generate an expression package image expected to be generated by the user. Meanwhile, the image generation model has a stable diffusion structure with a transparent channel, so that a transparent expression package image can be directly generated, which can be directly applied in end-side chat or sending of bullet screen and the like. The above method can not only improve the generation efficiency of the expression package image, but also generate an expression package image more in line with the user's demand according to the user's individualized demand.

[0195] Moreover, for ordinary users, there is no need to master complex production skills, and only by inputting a self-defined text, the generation of an expression package image can be realized in one key. Compared with traditional customization tools, the process is more concise, the customization cost is lower, and the use is more convenient. The image generation method provided by some embodiments of the present application effectively solves the pain points in the prior art, and has high practical value and market prospect. The above various method embodiments, or various possible implementation manners in each method embodiment can be executed independently, or, in the absence of contradictions, can also be executed in combination, and the specific execution can be determined according to actual use requirements, which is not limited by the embodiments of the present application.

[0196] The image generation method provided in the embodiments of the present application can be executed by an image generation device. The image generation device provided in the embodiments of the present application is described by taking the image generation device executing the image generation method as an example.

[0197] As shown in FIG. 16, the embodiments of the present application provide an image generation device 160, which can include a generation module 161 configured to generate an expression detail prompt word including expression detail information of at least one description dimension according to expression description information input by a user, and a processing module 162 configured to input the expression detail prompt word into an image generation model and output an expression package image matching the expression description information. The image generation model is obtained by training at least one model training image, and the at least one model training image is obtained by performing a transparency channel removal process on at least one original training image.

[0198] In some embodiments of the present application, the description dimension can include at least one of the following: an expression package image subject, an expression package image subject emotion, an expression package image subject behavior, an expression package image background, and an expression package image style.

[0199] In some embodiments of the present application, the image generation device 160 can further include a receiving module configured to receive expression description information input by a user in a message input area of a chat interface before the generation module 161 generates the expression detail prompt word according to the expression description information input by the user, or receive expression description information input by a user in a scroll input area of a video playing interface. The generation module 161 can be specifically configured to generate the expression detail prompt word according to the expression description information input by the user in the case that the input area of the expression description information is the message input area of the chat interface, or generate the expression detail prompt word according to the expression description information input by the user in the case that the input area of the expression description information is the message input area of the chat interface and the processing module 162 receives a touch input of an expression package image generation control by the user, or generate the expression detail prompt word according to the expression description information input by the user in the case that the input area of the expression description information is the scroll input area of the video playing interface.

[0200] In some embodiments of the present application, the image generation device 160 can further include a display module configured to display the expression package image in an expression package image preview area of the chat interface after the processing module 162 inputs the expression detail prompt word into the image generation model and outputs the expression package image matching the expression description information.

[0201] In some embodiments of the present application, the processing module 162 can also be configured to input the above-mentioned expression detail prompt word into the above-mentioned image generation model, and after outputting the above-mentioned expression package image matching the above-mentioned expression description information, in the case that the expression description information includes spoken keywords, add the expression description information to the expression package image to display the expression description information on the expression package image.

[0202] In some embodiments of the present application, the processing module 162 can also be configured to perform image processing on the at least one original training image to obtain an image transparency mask of each original training image and at least one model training image; the image processing at least includes removing the transparency channel processing, each model training image corresponds to an original training image; and based on the at least one model training image, the image description information of the at least one model training image, and at least one image transparency mask, train the initial model to obtain the image generation model; the at least one image transparency mask is the image transparency mask of the original training image corresponding to the at least one model training image.

[0203] In some embodiments of the present application, the processing module 162 can be specifically configured to perform image processing on the at least one original training image to obtain an image transparency mask of each original training image and an intermediate training image corresponding to each original training image; and based on the image description information and the image parameters of each intermediate training image, determine at least one target training image from the intermediate training image corresponding to the at least one original training image; and perform image super-resolution processing on the at least one target training image to obtain at least one model training image corresponding one-to-one to the at least one target training image.

[0204] In some embodiments of the present application, the above-mentioned image processing can also include removing text content processing; the processing module 162 can be specifically configured to perform removing transparency channel processing on the at least one original training image to obtain an image transparency mask of each original training image and a channel processing image corresponding to each original training image; and perform removing text content processing on the image containing text in the channel processing image corresponding to the at least one original training image to obtain an intermediate training image corresponding to each original training image.

[0205] In some embodiments of the present application, the initial model can include an image encoder and a text encoder; the processing module 162 can be specifically configured to input the at least one model training image and the at least one image transparency mask into the image encoder for encoding, output image feature information, and input the image description information of the at least one model training image into the text encoder for encoding, output text feature information; and based on the image feature information and the text feature information, predict a model training loss value, and based on the model training loss value, train the initial model to obtain the image generation model.

[0206] In the image generation apparatus provided by the embodiments of the present application, on the one hand, since the expression detail prompt word input into the image generation model is generated according to the expression description information input by the user, and the expression detail prompt word includes expression detail information of at least one description dimension, the image generation apparatus can generate an expression package image that is more in line with the image requirement of the user according to the personalized image requirement of the user for the expression package image; on the other hand, since the image generation model can be used to directly output an expression package image matched with the expression description information after the user inputs the expression description information, the image generation apparatus can simplify the production process of the expression package image and improve the production efficiency of the expression package image, without the need for the user to input multiple times to perform multiple image processing on an image.

[0207] The image generation apparatus in the embodiments of the present application can be an electronic device or a component in an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices other than the terminal. For example, the electronic device can be a mobile phone, a tablet computer, a notebook computer, a palm computer, a vehicle-mounted electronic device, a mobile Internet device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), and the like, and can also be a server, a network attached storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, and the like, and the embodiments of the present application are not limited thereto.

[0208] The image generation apparatus in the embodiments of the present application can be an apparatus with an operating system. The operating system can be an Android operating system, an IOS operating system, or other possible operating systems, which are not limited in the embodiments of the present application.

[0209] The image generation apparatus provided in the embodiments of the present application can implement each process implemented by the method embodiments and achieve the same technical effects. To avoid repetition, details are not described herein.

[0210] As shown in FIG. 17, the embodiments of the present application further provide an electronic device 100, which includes a processor 101 and a memory 102, and the memory 102 stores programs or instructions executable on the processor 101. When the programs or instructions are executed by the processor 101, each step of the image generation method embodiments described above is implemented, and the same technical effects are achieved. To avoid repetition, details are not described herein.

[0211] It should be noted that the electronic device in the embodiments of the present application includes a mobile electronic device and a non-mobile electronic device.

[0212] FIG. 18 is a schematic diagram of a hardware structure of an electronic device for implementing the embodiments of the present application.

[0213] As shown in FIG. 18, the electronic device 1000 includes but is not limited to: a radio frequency unit 1001, a network module 1002, an audio output unit 1003, an input unit 1004, a sensor 1005, a display unit 1006, a user input unit 1007, an interface unit 1008, a memory 1009, and a processor 1010, etc.

[0214] Those skilled in the art can understand that the electronic device 1000 can further include a power supply (such as a battery) for supplying power to each component, and the power supply can be logically connected to the processor 1010 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. The structure of the electronic device shown in FIG. 18 does not constitute a limitation on the electronic device, and the electronic device can include more or fewer components than those shown, or combine certain components, or different component arrangements, which are not described herein.

[0215] The processor 1010 can be configured to generate an expression detail prompt word according to expression description information input by a user, the expression detail prompt word including expression detail information of at least one description dimension; and input the expression detail prompt word into an image generation model to output an expression package image matched with the expression description information. The image generation model is obtained by training at least one model training image, and the at least one model training image is obtained by performing a transparency channel removal process on at least one original training image.

[0216] In some embodiments of the present application, the above-mentioned description dimensions can include at least one of the following: an image subject of the sticker image, an emotion of the image subject of the sticker image, a behavior of the image subject of the sticker image, a background of the sticker image, and a style of the sticker image.

[0217] In some embodiments of the present application, the user input unit 1007 can be configured to receive the expression description information input by the user in the message input area of the chat interface before the processor 1010 generates the expression detail prompt word according to the expression description information input by the user; or, receive the expression description information input by the user in the scroll input area of the video playing interface. The processor 1010 can be specifically configured to: in the case that the input area of the expression description information is the message input area of the chat interface, generate the expression detail prompt word according to the expression description information input by the user; or, in the case that the input area of the expression description information is the message input area of the chat interface and the touch input of the user to the sticker image generation control is received, generate the expression detail prompt word according to the expression description information input by the user; or, in the case that the input area of the expression description information is the scroll input area of the video playing interface, generate the expression detail prompt word according to the expression description information input by the user.

[0218] In some embodiments of the present application, the display unit 1006 can be configured to input the expression detail prompt word into the image generation model and output the sticker image matched with the expression description information, and then display the sticker image in the sticker image preview area of the chat interface.

[0219] In some embodiments of the present application, the processor 1010 can be further configured to, after inputting the expression detail prompt word into the image generation model and outputting the sticker image matched with the expression description information, in the case that the expression description information includes a spoken keyword, add the expression description information to the sticker image to display the expression description information on the sticker image.

[0220] In some embodiments of the present application, the processor 1010 can be further configured to perform image processing on at least one original training image to obtain an image transparency mask of each original training image and at least one model training image; the image processing at least includes removing a transparency channel processing, each model training image corresponds to an original training image; and based on the at least one model training image, image description information of the at least one model training image, and at least one image transparency mask, train an initial model to obtain an image generation model; the at least one image transparency mask is the image transparency mask of the original training image corresponding to the at least one model training image.

[0221] In some embodiments of the present application, the processor 1010 can be specifically configured to perform image processing on at least one original training image to obtain an image transparency mask of each original training image and a corresponding intermediate training image of each original training image; and determine at least one target training image from the corresponding intermediate training image of the at least one original training image based on image description information and image parameters of each intermediate training image; and perform image super-resolution processing on the at least one target training image to obtain at least one model training image corresponding to the at least one target training image.

[0222] In some embodiments of the present application, the above-mentioned image processing can further include text content removal processing; the processor 1010 can be specifically configured to perform transparency channel removal processing on at least one original training image to obtain an image transparency mask of each original training image and a corresponding channel processing image of each original training image; and perform text content removal processing on an image containing text in the corresponding channel processing image of the at least one original training image to obtain a corresponding intermediate training image of each original training image.

[0223] In some embodiments of the present application, the above-mentioned initial model can include an image encoder and a text encoder; the processor 1010 can be specifically configured to input the at least one model training image and the at least one image transparency mask into the image encoder for encoding to output image feature information, and input image description information of the at least one model training image into the text encoder for encoding to output text feature information; and predict a model training loss value based on the image feature information and the text feature information, and train the above-mentioned initial model based on the model training loss value to obtain an image generation model.

[0224] In the electronic device provided in the embodiments of the present application, on the one hand, since the expression detail prompt word input into the image generation model is generated according to the expression description information input by the user, and the expression detail prompt word includes expression detail information of at least one description dimension, the expression package image that is more in line with the image requirements of the user can be generated according to the personalized image requirements of the user for the expression package image; on the other hand, since the expression package image matched with the expression description information can be directly output by the image generation model after the user inputs the expression description information, the user does not need to perform multiple image processing on one image through multiple inputs in the process of making the expression package image, so that the making process of the expression package image can be simplified, and the making efficiency of the expression package image can be improved.

[0225] It should be understood that in the embodiments of the present application, the input unit 1004 can include a graphics processor (GPU) 10041 and a microphone 10042. The graphics processor 10041 processes image data of a still picture or a video obtained by an image capture device (such as a camera) in a video capture mode or an image capture mode. The display unit 1006 can include a display panel 10061, which can be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 1007 includes at least one of a touch panel 10071 and other input devices 10072. The touch panel 10071 is also referred to as a touch screen. The touch panel 10071 can include two parts of a touch detection device and a touch controller. The other input devices 10072 can include, but are not limited to, a physical keyboard, function keys (such as volume control keys, on-off keys, etc.), a trackball, a mouse, a joystick, and the like, which will not be described here.

[0226] The memory 1009 can be used to store software programs and various data. The memory 1009 can mainly include a first storage area storing programs or instructions and a second storage area storing data, wherein the first storage area can store an operating system, application programs or instructions required by at least one function (such as a sound playing function, an image playing function, etc.), etc. In addition, the memory 1009 can include a volatile memory or a non-volatile memory, or the memory 1009 can include both volatile and non-volatile memories. The non-volatile memory can be a Read-Only Memory (ROM), a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically EPROM (EEPROM), or a flash memory. The volatile memory can be a Random Access Memory (RAM), a Static RAM (SRAM), a Dynamic RAM (DRAM), a Synchronous DRAM (SDRAM), a Double Data Rate SDRAM (DDR SDRAM), an Enhanced SDRAM (ESDRAM), a Synch link DRAM (SLDRAM), and a Direct Rambus RAM (DRRAM). The memory 1009 in the embodiments of the present application includes but is not limited to these and any other suitable types of memories.

[0227] The processor 1010 can include one or more processing units; optionally, the processor 1010 integrates an application processor and a modem processor, wherein the application processor mainly processes operations related to an operating system, a user interface, and an application program, and the modem processor mainly processes wireless communication signals, such as a baseband processor. It can be understood that the above-mentioned modem processor can also not be integrated into the processor 1010.

[0228] The embodiments of the present application also provide a readable storage medium, the readable storage medium stores programs or instructions, the programs or instructions are executed by a processor to realize various processes of the above-mentioned image generation method embodiments, and the same technical effects can be achieved. To avoid repetition, details are not described here.

[0229] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes a computer readable storage medium, such as a computer readable only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0230] The embodiment of the present application further provides a chip, which comprises a processor and a communication interface, the communication interface is coupled with the processor, the processor is used for running programs or instructions to realize the processes of the above image generation method embodiments and achieve the same technical effects. To avoid repetition, details are not described herein.

[0231] It should be understood that the chip mentioned in the embodiment of the present application can also be referred to as a system level chip, a system chip, a chip system or a system on chip, etc.

[0232] The embodiment of the present application provides a computer program / program product, which is stored in a storage medium, and is executed by at least one processor to realize the processes of the above image generation method embodiments and achieve the same technical effects. To avoid repetition, details are not described herein.

[0233] It should be noted that in this document, the term "comprising" or "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. Without more limitations, the element defined by the statement "comprising a" does not exclude the presence of other identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the method and device in the embodiment of the present application is not limited to the order of performing the functions as shown or discussed, but can also include performing the functions in a substantially simultaneous manner or in a reverse order according to the functions involved, for example, the described method can be performed in an order different from the described order, and various steps can also be added, omitted or combined. In addition, the features described with reference to some examples can be combined in other examples.

[0234] Through the above description of the embodiments, those skilled in the art can clearly understand that the above-mentioned example methods can be realized by means of software and a necessary general hardware platform, and of course, can also be realized by hardware, but in many cases, the former is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a computer software product in essence or in the form of a part that contributes to the prior art, which is stored in a storage medium (such as a ROM / RAM, a magnetic disk, or an optical disk) and includes a plurality of instructions for causing a terminal (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present application.

[0235] The embodiments of the present application are described above in combination with the drawings, but the present application is not limited to the above-mentioned specific embodiments, and the above-mentioned specific embodiments are only illustrative and not restrictive. Those skilled in the art can make many forms under the inspiration of the present application without departing from the scope of the present application and the scope protected by the claims.

Claims

1. An image generation method, the method comprising: generating an expression detail prompt word according to expression description information input by a user, the expression detail prompt word comprising expression detail information of at least one description dimension; inputting the expression detail prompt word into an image generation model to output an expression package image matching the expression description information; wherein the image generation model is obtained by training at least one model training image, the at least one model training image being obtained by performing a transparency channel removal process on at least one original training image.

2. The method of claim 1, wherein, The description dimension comprises at least one of the following: expression package image subject, expression package image subject emotion, expression package image subject behavior, expression package image background, and expression package image style.

3. The method of claim 1, wherein, Before the expression detail prompt word is generated according to the expression description information input by the user, the method further comprises: receiving expression description information input by a user in a message input area of a chat interface; or, receiving expression description information input by a user in a scroll input area of a video playing interface; The expression detail prompt word is generated according to the expression description information input by the user, comprising: in the case that the input area of the expression description information is a message input area of a chat interface, generating an expression detail prompt word according to the expression description information input by the user; or, in the case that the input area of the expression description information is a message input area of a chat interface and a touch input of an expression package image generation control by the user is received, generating an expression detail prompt word according to the expression description information input by the user; or, in the case that the input area of the expression description information is a scroll input area of a video playing interface, generating an expression detail prompt word according to the expression description information input by the user.

4. The method of claim 1, wherein, After the expression detail prompt word is input into the image generation model to output an expression package image matching the expression description information, the method further comprises: displaying the expression package image in an expression package image preview area of a chat interface.

5. The method of claim 1, wherein, After the expression detail prompt word is input into the image generation model to output an expression package image matching the expression description information, the method further comprises: in the case that the expression description information comprises a colloquial keyword, adding the expression description information to the expression package image to display the expression description information on the expression package image.

6. The method of claim 1, wherein, The method further comprises: performing image processing on at least one original training image to obtain an image transparency mask of each original training image and at least one model training image; the image processing at least comprises a transparency channel removal process, and each model training image corresponds to an original training image; training an initial model based on the at least one model training image, image description information of the at least one model training image, and at least one image transparency mask to obtain an image generation model; the at least one image transparency mask is an image transparency mask of an original training image corresponding to the at least one model training image.

7. The method of claim 6, wherein, The image processing on at least one original training image to obtain an image transparency mask of each original training image and at least one model training image comprises: perform image processing on the at least one original training image to obtain an image transparency mask of each original training image and a corresponding intermediate training image of each original training image; determine at least one target training image from the at least one corresponding intermediate training image of the at least one original training image based on image description information and image parameters of each intermediate training image; perform image super-resolution processing on the at least one target training image to obtain at least one model training image corresponding to the at least one target training image.

8. The method of claim 7, wherein, The image processing further includes removing text content processing. The image processing further includes removing text content processing. The image processing further includes removing text content processing. The initial model includes an image encoder and a text encoder.

9. The method of claim 6, wherein, The initial model includes an image encoder and a text encoder. The initial model includes an image encoder and a text encoder. The initial model includes an image encoder and a text encoder.

10. An image generation device, the device comprising: a generation module configured to generate an expression detail prompt word based on expression description information input by a user, the expression detail prompt word including expression detail information of at least one description dimension; a processing module configured to input the expression detail prompt word into an image generation model to output an expression package image matching the expression description information; wherein the image generation model is obtained by training at least one model training image, and the at least one model training image is obtained by performing removing transparency channel processing on at least one original training image. The description dimension includes at least one of the following: expression package image subject, expression package image subject emotion, expression package image subject behavior, expression package image background, and expression package image style.

11. The apparatus of claim 10, wherein, The device further comprises:

12. The apparatus of claim 10, wherein, a receiving module configured to receive expression description information input by a user in a message input area of a chat interface before the generation module generates the expression detail prompt word based on the expression description information input by the user; or, receive expression description information input by a user in a scroll input area in a video playing interface. ​ The generation module is specifically configured to: in a case where the input area of the expression description information is a message input area of a chat interface, generate an expression detail prompt word according to the expression description information input by the user; or in a case where the input area of the expression description information is the message input area of the chat interface and touch input of an expression package image generation control by the user is received, generate the expression detail prompt word according to the expression description information input by the user; or in a case where the input area of the expression description information is a scroll input area in a video playing interface, generate the expression detail prompt word according to the expression description information input by the user.

13. The apparatus of claim 10, wherein, The device further includes: The display module is configured to input the expression detail prompt word into an image generation model, and after outputting an expression package image matched with the expression description information, display the expression package image in an expression package image preview area of a chat interface.

14. The apparatus of claim 10, wherein, The processing module is further configured to, after inputting the expression detail prompt word into the image generation model and outputting the expression package image matched with the expression description information, in a case where the expression description information includes a colloquial keyword, add the expression description information to the expression package image to display the expression description information on the expression package image.

15. The apparatus of claim 10, wherein, The processing module is further configured to perform image processing on at least one original training image to obtain an image transparency mask of each original training image and at least one model training image; the image processing at least includes removing a transparency channel processing, and each model training image corresponds to an original training image. And based on the at least one model training image, image description information of the at least one model training image, and at least one image transparency mask, the processing module is configured to train an initial model to obtain an image generation model; the at least one image transparency mask is an image transparency mask of the original training image corresponding to the at least one model training image.

16. The apparatus of claim 15, wherein, The processing module is specifically configured to perform image processing on at least one original training image to obtain an image transparency mask of each original training image and an intermediate training image corresponding to each original training image; and based on image description information and image parameters of each intermediate training image, the processing module is configured to determine at least one target training image from the intermediate training image corresponding to the at least one original training image. And the processing module is configured to perform image super-resolution processing on the at least one target training image to obtain at least one model training image corresponding one by one to the at least one target training image.

17. The apparatus of claim 16, wherein, The image processing further includes removing text content processing; the processing module is specifically configured to perform removing transparency channel processing on at least one original training image to obtain an image transparency mask of each original training image and a channel processing image corresponding to each original training image; and perform removing text content processing on an image containing text in the channel processing image corresponding to the at least one original training image to obtain an intermediate training image corresponding to each original training image.

18. The apparatus of claim 15, wherein, The initial model comprises an image encoder and a text encoder; the processing module is specifically configured to input the at least one model training image and the at least one image transparency mask into the image encoder for encoding, output image feature information, and input image description information of the at least one model training image into the text encoder for encoding, output text feature information; and predict a model training loss value based on the image feature information and the text feature information, and train the initial model based on the model training loss value to obtain an image generation model. 19.An electronic device comprising a processor and a memory, the memory storing programs or instructions executable on the processor, the programs or instructions, when executed by the processor, implement steps of the image generation method according to any one of claims 1-9. 20.A readable storage medium, the readable storage medium storing programs or instructions, the programs or instructions, when executed by a processor, implement steps of the image generation method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Image generation method and device, electronic equipment and storage medium

    CN117173497A

  • Weak supervision semantic segmentation method based on background prior

    CN118015282A

  • Model fine tuning method, model fine tuning device and electronic equipment

    CN118153638A

  • Model training method and device, equipment, storage medium and program product

    CN118378633A

  • Image generation method and device, electronic equipment and medium

    CN118864637A