Method, device, apparatus and storage medium for generating an image

By processing input text to generate material description text and material images, and combining glyph images, the problem of restricted generation of static two-dimensional images and character borders is solved, and a diverse art word image generation is achieved.

CN117197292BActive Publication Date: 2025-05-02BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311270742.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-27
Publication Date
2025-05-02
Estimated Expiration
2043-09-27

AI Technical Summary

Technical Problem

Traditional image generation methods only support the generation of static two-dimensional images, and are prone to limited character borders, affecting the effect of the art word image.

Method used

By obtaining the input text, the first model process is used to determine the material description text, a material image is obtained, and a target image is generated based on the material image and the glyph image.

Benefits of technology

It realizes the generation of target images corresponding to characters based on the input text, allowing users to freely edit the input text to generate diversified target images, meeting users' diverse image generation needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117197292B_ABST
    Figure CN117197292B_ABST
Patent Text Reader

Abstract

According to an embodiment of the present disclosure, a method, apparatus, device and storage medium for generating an image are provided. The method includes: obtaining input text, the input text indicating generation of an image corresponding to at least one character; processing the input text using a first model to determine a material description text corresponding to at least one character; obtaining a material image generated based on the material description text; and generating a target image corresponding to at least one character based on the material image and a glyph image corresponding to at least one character. In this way, a target image corresponding to a character can be generated based on the input text, and the user is allowed to freely edit the input text to generate a variety of target images, thereby meeting the user's diverse image generation needs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Example embodiments of the present disclosure relate generally to information processing, and more particularly, to methods, apparatuses, devices, and computer-readable storage media for generating images. Background Art

[0002] With the development of machine learning technology, machine learning models can be used to perform tasks in a variety of application environments. Model-based visual tasks are used to process visual data, such as images, videos, etc. Examples of visual tasks include but are not limited to image generation, image classification, object detection, semantic segmentation, optical character recognition (OCR), etc., among which image generation tasks are important tasks in visual tasks. Artistic word image generation in image generation has received more and more attention due to its wide application, and has gradually become an important task in image generation tasks. Summary of the invention

[0003] In a first aspect of the present disclosure, a method for generating an image is provided. The method comprises: obtaining input text, the input text indicating generation of an image corresponding to at least one character; processing the input text using a first model to determine a material description text corresponding to the at least one character; obtaining a material image generated based on the material description text; and generating a target image corresponding to the at least one character based on the material image and a glyph image corresponding to the at least one character.

[0004] In a second aspect of the present disclosure, a device for generating an image is provided. The device includes: a text acquisition module configured to acquire input text, the input text indicating generation of an image corresponding to at least one character; a text determination module configured to process the input text using a first model to determine a material description text corresponding to at least one character; an image acquisition module configured to acquire a material image generated based on the material description text; and an image generation module configured to generate a target image corresponding to at least one character based on the material image and a glyph image corresponding to at least one character.

[0005] In a third aspect of the present disclosure, an electronic device is provided. The electronic device includes at least one processing unit; and at least one memory, the at least one memory is coupled to the at least one processing unit and stores instructions for execution by the at least one processing unit. When the instructions are executed by the at least one processing unit, the electronic device executes the method according to the first aspect of the present disclosure.

[0006] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein a computer program is stored on the computer-readable storage medium, and the computer program can be executed by a processor to perform the method according to the first aspect of the present disclosure.

[0007] It should be understood that the content described in this section is not intended to limit the key features or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] In the following, in conjunction with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages and aspects of various implementations of the present disclosure will become more apparent. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:

[0009] Figure 1 A schematic diagram showing an example environment in which embodiments of the present disclosure can be implemented;

[0010] Figure 2 A flowchart showing a process of generating an image according to some embodiments of the present disclosure;

[0011] Figure 3 A schematic diagram showing an example architecture for generating an image according to some embodiments of the present disclosure;

[0012] Figure 4 A schematic diagram showing an example of loop filling according to some embodiments of the present disclosure;

[0013] Figure 5 A schematic diagram showing an example of blur processing according to some embodiments of the present disclosure;

[0014] Figure 6 A schematic diagram showing an example of adding noise according to some embodiments of the present disclosure;

[0015] Figure 7 A schematic structural block diagram of an apparatus for generating an image according to some embodiments of the present disclosure is shown; and

[0016] Figure 8 A block diagram of an electronic device that can be used to implement some embodiments of the present disclosure is shown. DETAILED DESCRIPTION

[0017] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.

[0018] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below.

[0019] The term "in response to" means that the corresponding event occurs or the condition is satisfied. It will be understood that the timing of the execution of the subsequent action executed in response to the event or condition is not necessarily strongly related to the time when the event occurs or the condition is satisfied. In some cases, the subsequent action may be executed immediately when the event occurs or the condition is satisfied; in other cases, the subsequent action may be executed some time after the event occurs or the condition is satisfied.

[0020] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and relevant provisions.

[0021] It is understandable that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, scope of use, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0022] For example, in response to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested to be performed will require obtaining and using the user's personal information, so that the user can independently choose whether to provide personal information to software or hardware such as electronic devices, applications, servers or storage media that execute operations of the technical solution of the present disclosure based on the prompt message.

[0023] As an optional but non-limiting implementation, in response to receiving an active request from the user, the prompt information is sent to the user in a manner such as a pop-up window, in which the prompt information can be presented in text form. In addition, the pop-up window can also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0024] It is understandable that the above notification and the process of obtaining user authorization are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that meet the relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0025] As used herein, the term "model" can learn the association between the corresponding input and output from the training data, so that after the training is completed, the corresponding output can be generated for a given input. The generation of the model can be based on machine learning technology. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multi-layer processing units. A neural network model is an example of a model based on deep learning. In this article, "model" may also be referred to as "machine learning model", "learning model", "machine learning network" or "learning network", and these terms are used interchangeably in this article.

[0026] A "neural network" is a machine learning network based on deep learning. A neural network is capable of processing inputs and providing corresponding outputs, and typically includes an input layer and an output layer and one or more hidden layers between the input layer and the output layer. Neural networks used in deep learning applications typically include many hidden layers, thereby increasing the depth of the network. The layers of a neural network are connected in sequence so that the output of the previous layer is provided as input to the next layer, where the input layer receives the input of the neural network and the output of the output layer serves as the final output of the neural network. Each layer of a neural network includes one or more nodes (also called processing nodes or neurons), each of which processes input from the previous layer.

[0027] Generally, machine learning can be roughly divided into three stages, namely the training stage, the testing stage, and the application stage (also called the inference stage). In the training stage, a given model can be trained using a large amount of training data, and the parameter values ​​are continuously updated iteratively until the model can obtain consistent inferences that meet the expected goals from the training data. Through training, the model can be considered to be able to learn the association from input to output (also called the mapping of input to output) from the training data. The parameter values ​​of the trained model are determined. In the testing stage, the test input is applied to the trained model to test whether the model can provide the correct output, thereby determining the performance of the model. In the application stage, the model can be used to process the actual input based on the parameter values ​​obtained from the training to determine the corresponding output.

[0028] Figure 1 1 is a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. Figure 1 As shown, environment 100 may include electronic device 110 .

[0029] The electronic device 110 can generate a target image 112 corresponding to the input text 102 based on the input text 102. That is, the electronic device 110 can perform an image generation task based on the input text 102 to generate the target image 112. The electronic device 110 can obtain the input text 102 in any appropriate manner. For example, the electronic device 110 can determine that the input text is obtained in response to detecting the user's input in the input box. For example, the electronic device 110 can receive the user's voice and convert it into input text. The input text 102 here can be a text sequence of any appropriate language and any number of words. For example, the electronic device 110 can perform an image generation task based on the input text 102 in Chinese to generate its corresponding target image 112. In the case where the input text 102 indicates the generation of an image corresponding to at least one character, the electronic device 110 can perform an artistic word image generation task based on such input text 102, and the electronic device 110 can generate an artistic word image, that is, a target image 112.

[0030] The electronic device 110 can, for example, use the model 120 to perform the task of generating an artistic word image. The model 120 can include, for example, but is not limited to any appropriate model such as a Transformer model, a LORA model, a convolutional neural network (CNN), a recurrent neural network (RNN), a deep neural network (DNN), a generative adversarial network (GAN), etc. The model 120 can be a model local to the electronic device 110, or a model installed in other electronic devices 110 (for example, installed in a remote device). The model 120 can include multiple models, for example, a model for text question and answer (also known as a text question and answer model), a model for generating images based on text (also known as a text-generated image model), a model for generating images based on images (also known as an image-generated image model), and the like.

[0031] The electronic device 110 may include any computing system with computing capabilities, such as various computing devices / systems, terminal devices, server devices, etc. The terminal device may be any type of mobile terminal, fixed terminal or portable terminal, including a mobile phone, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a handheld computer, a portable game terminal, a VR / AR device, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a game device, or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof.

[0032] The server-side device can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks, and big data and artificial intelligence platforms. Server-side devices can include computing systems / servers such as mainframes, edge computing nodes, computing devices in cloud environments, and so on.

[0033] It should be understood that the structure and function of the various elements in the environment 100 are described for exemplary purposes only and do not imply any limitation on the scope of the present disclosure.

[0034] As mentioned above, image generation tasks are important tasks in visual tasks. The generation of word art images in image generation has gradually become an important task in image generation tasks. Traditionally, an image generation method has been proposed, which supports input text and generates word art images based on the input text. However, the traditional image generation method only supports the generation of static two-dimensional images, and as the number of training times increases, it is easy for the character borders in the word art image to be limited. This will affect the effect of the final generated word art image.

[0035] To this end, an embodiment of the present disclosure proposes a scheme for generating an image. According to the scheme, an input text indicating the generation of an image corresponding to at least one character is obtained. The input text is processed using a first model to determine a material description text corresponding to the at least one character. A material image generated based on the material description text is obtained. Based on the material image and a glyph image corresponding to the at least one character, a target image corresponding to the at least one character is generated.

[0036] According to the image generation scheme disclosed in the present invention, a target image corresponding to a character can be generated based on an input text, and a user is allowed to freely edit the input text to generate a variety of target images, thereby meeting the user's diverse image generation needs.

[0037] Some example embodiments of the present disclosure will be described below with continued reference to the accompanying drawings.

[0038] Figure 2 1 shows a flowchart of a process 200 for generating an image according to some embodiments of the present disclosure. The process 200 may be implemented at the electronic device 110. For ease of discussion, reference will be made to Figure 1 The process 200 is described with reference to the environment 100 of FIG.

[0039] In block 210 , the electronic device 110 obtains input text, the input text indicating generation of an image corresponding to at least one character.

[0040] In some embodiments, the electronic device 110 can obtain an indication of user input to generate a text sequence of images corresponding to at least one character, and the text sequence is the user's input text. The electronic device 110 may, for example, include a touch display screen, and the electronic device 110 may obtain the input text in response to detecting a preset operation of the user on an operating control (such as an input box) in the touch display screen. The electronic device 110 may, for example, also collect user voice through an audio collection device and convert the collected voice into input text. It is understood that the input text may be a text sequence of any language (such as Chinese, English, etc.) and any number of words. It is understood that the electronic device 110 may also obtain the input text in any other appropriate manner.

[0041] The input text acquired by the electronic device 110 may be, for example, "I need a 3D wooden-style building that looks like the letters 'ABCD', which is relatively old, brown, and has a lot of moss on it." For the convenience of description, the following exemplary description is given by taking this sample text as an example unless otherwise specified. It can be understood that the input text here is only an example, and the electronic device 110 can acquire any appropriate input text.

[0042] At block 220 , the electronic device 110 processes the input text using the first model to determine material description text corresponding to the at least one character.

[0043] The first model here may be a model included in the model 120, which may be, for example, a text question-answering model. The first model may, for example, be a model installed locally on the electronic device 110, and the electronic device 110 may directly use the local first model to process the input text. The first model may, for example, also be a model installed on other electronic devices, and the electronic device 110 may send the input text to other electronic devices through a communication connection with other electronic devices to process the input text using the first model, and the electronic device 110 may then directly obtain the processing result, i.e., the material description text. For the sake of convenience, the following illustrative description is given by taking the example that the first model is installed locally on the electronic device 110 and the electronic device 110 may directly use the first model.

[0044] The electronic device 110 may generate first input information of the first model based on the input text. The first input information may include, for example, preset constraint information for constraining the output generation of the first model. Such preset constraint information may indicate constraints on the capabilities or roles of the first model, which may also be referred to as capability constraint information, role constraint information, system-level setting information, etc. For example, if the preset constraint information indicates that the model is an artist, then after the first model receives the preset constraint information, it will assume that it is an artist to generate the output result, thereby improving the quality of the model-generated image. The first input information may also include, for example, a set of reference material description texts. In some embodiments, the electronic device 110 may obtain a training set for training a material generation model (i.e., the second model, the text image model, etc. described below), and the training set includes a plurality of material image-description text pairs. The description text in each material image-description text pair is used to describe the corresponding material image. The electronic device 110 may obtain a set of material image-description text pairs from such a training set, and use a set of description texts in a set of material image-description text pairs as a set of reference material description texts.

[0045] The electronic device 110 then generates the first input information of the first model based on the input text, the preset constraint information and a set of reference material description texts. Exemplarily, the first input information generated by the electronic device 110 based on the input text may indicate that the task to be processed is "generate a description text about the material", indicate that the preset constraint information is "You are a rendering expert and you know the material very well. Output example: rough ground covered with leaves", and the output example is "rough ground covered with leaves". The output example here is the reference material description text, and the first model can generate a material description text with a language structure similar to that of the output example based on the output example. It can be understood that the first input information here is only an example, and the electronic device 110 can generate any appropriate first input information based on any appropriate input text obtained.

[0046] Figure 3 A schematic diagram of an example architecture 300 for generating an image according to some embodiments of the present disclosure is shown. The architecture 300 may include, for example, a text question answering model 310 (ie, a first model).

[0047] The electronic device 110 may provide the acquired input text 102 to the text question answering model 310. In some embodiments, the electronic device 110 generates first input information based on the input text 102, and provides the generated first input information to the text question answering model 310. The text question answering model 310 may output a material description text 312 corresponding to the input text 102. Taking the above-mentioned first input information as an example, the material description text generated by the first model may be, for example, "brown old wood covered with moss".

[0048] In block 230 , the electronic device 110 obtains a material image generated based on the material description text.

[0049] In some embodiments, the electronic device 110 may acquire a material image library including a large number of material images. The electronic device 110 may acquire a material image matching the material description text from the material image library based on a predetermined mapping relationship or matching rule.

[0050] In some embodiments, the electronic device 110 can also provide second input information to the second model to generate a material image based on the material description text using the second model. Specifically, the electronic device 110 can generate the second description information based on the material description text. The electronic device 110 then provides the second description information to the second model and obtains the material image generated by the second model. Similarly, the second model can be a model included in the model 120, which can be, for example, a Wenshengtu model. The second model can be a model installed locally on the electronic device 110, or it can be a model installed on other electronic devices. For the convenience of description, the following also takes the second model installed locally on the electronic device 110, and the electronic device 110 can directly use the second model as an example for exemplary description.

[0051] Continue to refer Figure 3 ,like Figure 3 As shown, the architecture 300 may also include a Wensheng graph model 320 (i.e., a second model). The material description text 312 is provided to the Wensheng graph model 320. In some embodiments, the electronic device 110 generates second description information based on the material description text 312, and the second description information is provided to the Wensheng graph model 320. The Wensheng graph model 320 is configured to generate a material image 322 that matches the material description text 312. The Wensheng graph model 320 may include, for example, a plurality of convolutional layers. At least one of the plurality of convolutional layers may generate a first feature map corresponding to a first size, and fill the first feature map to obtain a second feature map corresponding to a second size. The second feature map will be used as an input to at least one convolutional layer. The second size is greater than the first size. Exemplarily, the convolutional layer may generate a first feature map of size 7*7, and the convolutional layer may fill the first feature map of 7*7 to obtain a second feature map of 8*8.

[0052] In some embodiments, in order to ensure that the generated material image 322 is tileable, the convolution layer in the Vincent graph model 320 is configured to generate a feature image by cyclic filling. Specifically, after the convolution layer generates a first feature image of a first size, the first feature image is cyclically filled to generate a second feature image of a second size, and the second size is larger than the first size. Specifically, the convolution layer can determine a first edge position in the first feature image. The first edge position can be any edge position of the first feature image, for example, it can be a position on the left edge, right edge, upper edge or lower edge of the first feature image.

[0053] Figure 4 Schematic diagram 400 showing an example of loop filling according to some embodiments of the present disclosure. Figure 4 The first feature map 410 and the second feature map 420 generated by filling the first feature map are included. Figure 4 As shown, if the first edge position is the position of the upper left corner of the first feature map, the corresponding value at the first edge position is a value 401 (for example, 5). The electronic device 110 can determine a second edge position symmetrical to the first edge position (i.e., the position of the upper left corner) in the target direction based on the first feature map 410. The target direction includes at least one of the horizontal direction, the vertical direction, or the diagonal direction. Exemplarily, for the first edge position, the position symmetrical to it in the horizontal direction at the upper right corner, the position symmetrical to it in the vertical direction at the lower left corner, and the position symmetrical to it in the diagonal direction at the lower right corner can be obtained. These positions can all be second edge positions. Then, the value of the second edge position can be used to fill the adjacent position of the first edge pixel in the target direction, and the adjacent position is outside the first feature map. Exemplarily, the value corresponding to the position in the upper right corner is a value 404 (for example, 1), and the value 404 is filled in the adjacent position in the horizontal direction of the value 401. The value corresponding to the position in the lower left corner is a value 402 (for example, 7), and the value 402 is filled in the adjacent position in the vertical direction of the value 401. The value corresponding to the position at the lower right corner is the value 403 (for example, 0), and the value 403 is filled in the adjacent position in the diagonal direction of the value 401.

[0054] In this way, the convolution layer that receives the second feature map can obtain the information of the opposite sides of the image in the vector space. The structure of such a second feature map can be approximately regarded as a ring structure, and the convolution layer can obtain the information of its opposite sides to eliminate the seams generated when multiple feature maps are spliced, which can improve the quality of the generated target image.

[0055] Return to reference Figure 3In some embodiments, the architecture 300 further includes a fine-tuning model 330. The fine-tuning model 330 may be, for example, a LORA model. The fine-tuning model 330 is configured to adjust parameters of some layers in the text graph model 320 and store these parameters. It can help the text graph model 320 establish a relationship between an image and text to generate a material image 322.

[0056] Continue to refer Figure 2 In block 240, the electronic device 110 generates a target image corresponding to at least one character based on the material image and the glyph image corresponding to the at least one character. The target image may correspond to an artistic representation of the at least one character, that is, the target image is an artistic word image including the at least one character. The artistic word image here may be any appropriate image, for example, a two-dimensional image, a three-dimensional image, a static image, a dynamic image, and the like.

[0057] In some embodiments, the electronic device 110 may obtain a set of preset glyph libraries including a large number of glyphs. A set of preset glyph libraries may include, for example, a font library, which may include a large number of font files (e.g., ttf files). The electronic device 110 may determine the target glyph from a set of preset glyph libraries based on the input text. Specifically, the electronic device 110 may, for example, process the input text to obtain the target glyph indicated by the input text. For example, if the input text includes content such as "the letters 'ABCD' in XX font", the electronic device 110 may identify the input text to determine that the target glyph is the glyph corresponding to the "XX font". The electronic device 110 then obtains the font file of the "XX font" from a set of preset glyph libraries and determines the corresponding target glyph.

[0058] In some embodiments, the electronic device 110 may also determine the target glyph based on the glyph selection information. For example, the electronic device 110 may also present a glyph (or font) selection interface to the user, and the glyph selection interface may include at least some glyphs (or fonts) included in a set of preset glyph libraries (or font libraries). The electronic device 110 further determines the selected glyph (font) as the target glyph in response to the user's selection operation on a glyph (or font).

[0059] After the target glyph is determined, the electronic device 110 may generate a glyph image corresponding to at least one character indicated by the input text based on the target glyph. The glyph image includes at least one character, and the glyph of the at least one character is the target glyph. Figure 3 ,like Figure 3As shown, in some embodiments, the electronic device 110 may obtain character text 302 based on the input text 102. The character text 302 may be, for example, text associated with a character in the input text 102. The electronic device 110 may obtain a glyph image 306 from a glyph library 304 based on the character text 302.

[0060] The electronic device 110 can then generate a guide image based on the glyph image 306 and the material image 322. The glyph part of the guide image is filled based on the material image 322. The glyph image 306 here can, for example, indicate mask information, and the electronic device 110 can then fill the material image 322 based on the mask information so that the glyph part is filled. Exemplarily, at least one character in the glyph image 306 can be, for example, white, and the area outside at least one character can be, for example, black, wherein only the white area is set to be filled. The electronic device 110 can superimpose the glyph image 306 and the material image 322, wherein the white area in the glyph image 306 can be superimposed by the material image 322, and the black area cannot be superimposed by the material image 322. The image obtained after superposition is the guide image.

[0061] In some embodiments, before generating the guide image, in order to make the target image 112 generated finally more creative, the electronic device 110 may process the glyph image 306 to weaken the edge of the glyph image 306. Such processing may include blurring, for example. The electronic device 110 then generates the guide image based on the processed glyph image 306 and the material image 322. Figure 5 Schematic diagram 500 showing an example of blur processing according to some embodiments of the present disclosure. Figure 5 As shown, the electronic device 110 may perform blur processing on the glyph image 510 to weaken the edge of the glyph image 510 , thereby obtaining a processed glyph image 520 .

[0062] like Figure 3As shown, the architecture 300 also includes a graph model 340 (i.e., a third model). After the electronic device 110 obtains the guide image, it can provide the guide image and the input text 102 to the graph model 340 to obtain the target image 112 generated by the graph model 340. In order to make the target image 112 generated finally more creative, in some embodiments, the electronic device 110 can also provide a first control parameter to the graph model 340. The control parameter is used to indicate the intensity of the noise added to the guide image. The graph model 340 can add noise to the guide image based on the noise intensity indicated by the first control parameter, and then generate the target image 112 based on the guide image and the input text 102 after adding the noise. Alternatively or additionally, the noise can also be added by the electronic device 110 before the electronic device 110 provides the guide image to the third model. Specifically, the electronic device 110 can generate an intermediate image based on the glyph image and the material image, and the electronic device 110 then generates the guide image by adding noise corresponding to the second control parameter to the intermediate image. That is, the guide image generated by the electronic device 110 is an image to which noise has been added.

[0063] Figure 6 Schematic diagram 600 showing an example of adding noise according to some embodiments of the present disclosure. Figure 6 As shown, the electronic device 110 can generate an image 610 based on the glyph image 306 and the material image 322, and the image 610 can be a guide image or an intermediate image. The electronic device 110 and / or the graph generation model 340 can add noise to the image 610 to obtain the image 610 after adding noise. According to the intensity of the added noise, the image 610 after adding noise can be, for example, an image 620, an image 630 or an image 640. Take the intensity of the noise as 0-1 as an example, where 0 means no noise is added and 1 means the highest level of noise intensity (e.g., full). Image 620, image 630 and image 640 respectively show images with different noise intensity, wherein the intensity of the noise added in image 620 is low and the intensity of the noise added in image 640 is high. The lower the intensity of the added noise, the higher the similarity between the target image 112 and the image 610, and the higher the intensity of the added noise, the higher the creativity of the target image 112, that is, the lower the similarity with the image 610.

[0064] Return to reference Figure 3In some embodiments, the electronic device 110 may also obtain a second input text 308. The electronic device 110 may provide the target image 112 and the second input text 308 to the image generation model 340, so that the image generation model 340 may output an image based on the target image 112 and the second input text 308. The newly output image is generated based on the target image 112. In this way, the target image 112 may be processed multiple times to improve the creativity and richness of the finally generated image.

[0065] In summary, according to the image generation scheme disclosed in the present invention, it is possible to generate a target image corresponding to a character based on the input text, and allow the user to freely edit the input text to generate a variety of target images, thereby meeting the user's diverse image generation needs.

[0066] According to some embodiments of the present disclosure, a device for generating an image is also provided.

[0067] Figure 7 A schematic structural block diagram of an apparatus 700 for generating an image according to some embodiments of the present disclosure is shown. The apparatus 700 may be implemented as or included in the electronic device 110. Each module / component in the apparatus 700 may be implemented by hardware, software, firmware or any combination thereof.

[0068] As shown in the figure, the device 700 includes a text acquisition module 710, which is configured to acquire input text, and the input text indicates the generation of an image corresponding to at least one character. The device 700 also includes a text determination module 720, which is configured to process the input text using a first model to determine the material description text corresponding to the at least one character. The device 700 also includes an image acquisition module 730, which is configured to acquire a material image generated based on the material description text. The device 700 also includes an image generation module 740, which is configured to generate a target image corresponding to at least one character based on the material image and the glyph image corresponding to the at least one character.

[0069] In some embodiments, the text determination module 720 includes: a first information generation module, configured to generate first input information of a first model based on input text; and a first input providing module, configured to provide the first input information to the first model to obtain a material description text generated by the first model.

[0070] In some embodiments, the first input information includes: preset constraint information for constraining output generation of the first model; and / or a set of reference material description texts.

[0071] In some embodiments, the image acquisition module 730 includes: a second input providing module configured to provide second input information to the second model to obtain a material image generated by the second model, wherein the second input information is generated based on the material description text.

[0072] In some embodiments, the second model includes multiple convolutional layers, and at least one of the multiple convolutional layers is configured to: generate a first feature map corresponding to a first size; and pad the first feature map to obtain a second feature map corresponding to a second size as an input to at least one convolutional layer, the second size being larger than the first size.

[0073] In some embodiments, filling the first feature map includes: determining a first edge position in the first feature map; based on the first feature map, determining a second edge position symmetrical to the first edge position in the target direction; and using the value of the second edge position to fill an adjacent position of the first edge pixel in the target direction, the adjacent position being outside the first feature map.

[0074] In some embodiments, the first edge position includes a position on a left edge, a right edge, an upper edge, or a lower edge of the first feature graph, and the target direction includes at least one of a horizontal direction, a vertical direction, or a diagonal direction.

[0075] In some embodiments, the apparatus 700 further includes: a glyph determination module configured to determine a target glyph from a set of preset glyph libraries based on input text and / or glyph selection information; and a glyph image generation module configured to generate a glyph image corresponding to at least one character based on the target glyph.

[0076] In some embodiments, the image generation module 740 includes: a guide image generation module, configured to generate a guide image based on a glyph image and a material image, so that the glyph portion of the guide image is filled based on the material image; and a third input providing module, configured to provide the guide image and input text to a third model to obtain a target image generated by the third model.

[0077] In some embodiments, the guide image generation module includes: a processing module configured to process a glyph image to weaken an edge of the glyph image; and a generation module configured to generate a guide image based on the processed glyph image and a material image.

[0078] In some embodiments, the image generation module 740 further includes: a parameter providing module configured to provide a first control parameter to the third model, where the control parameter is used to indicate the intensity of noise added to the guidance image.

[0079] In some embodiments, the guide image generation module includes: an intermediate image generation module, configured to generate an intermediate image based on a glyph image and a material image; and a noise adding module, configured to generate a guide image by adding noise corresponding to a second control parameter to the intermediate image.

[0080] In some embodiments, the input text is a first input text, and the device 700 further includes: a second text acquisition module configured to acquire a second input text; and an image providing module configured to provide at least one image generated based on the target image and the second input text.

[0081] In some embodiments, the target image corresponds to an artistic representation of at least one character.

[0082] The units and / or modules included in the device 700 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units and / or modules can be implemented using software and / or firmware, such as machine executable instructions stored on a storage medium. In addition to or as an alternative to machine executable instructions, some or all of the units and / or modules in the device 700 can be implemented at least in part by one or more hardware logic components. As an example and not limitation, exemplary types of hardware logic components that can be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.

[0083] Figure 8 A block diagram of an electronic device 800 is shown in which one or more embodiments of the present disclosure may be implemented. It should be understood that Figure 8 The electronic device 800 shown is merely exemplary and should not constitute any limitation on the functionality and scope of the embodiments described herein. Figure 8 The electronic device 800 shown can be used to implement Figure 1 Electronic device 110 or Figure 7 Device 700.

[0084] like Figure 8 As shown, the electronic device 800 is in the form of a general electronic device. The components of the electronic device 800 may include, but are not limited to, one or more processors or processing units 810, a memory 820, a storage device 830, one or more communication units 840, one or more input devices 850, and one or more output devices 860. The processing unit 810 may be an actual or virtual processor and is capable of performing various processes according to a program stored in the memory 820. In a multi-processor system, multiple processing units execute computer executable instructions in parallel to improve the parallel processing capability of the electronic device 800.

[0085] The electronic device 800 typically includes a plurality of computer storage media. Such media can be any accessible media that can be obtained by the electronic device 800, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 820 can be a volatile memory (e.g., a register, a cache, a random access memory (RAM)), a non-volatile memory (e.g., a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 830 can be a removable or non-removable medium, and can include a machine-readable medium, such as a flash drive, a disk, or any other medium, which can be used to store information and / or data and can be accessed within the electronic device 800.

[0086] The electronic device 800 may further include additional removable / non-removable, volatile / non-volatile storage media. Figure 8 As shown in , a disk drive for reading or writing from a removable, non-volatile disk (e.g., a "floppy disk") and an optical drive for reading or writing from a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to the bus (not shown) by one or more data media interfaces. The memory 820 may include a computer program product 825 having one or more program modules that are configured to perform various methods or actions of various embodiments of the present disclosure.

[0087] The communication unit 840 implements communication with other electronic devices through a communication medium. Additionally, the functions of the components of the electronic device 800 can be implemented with a single computing cluster or multiple computing machines that can communicate through a communication connection. Therefore, the electronic device 800 can operate in a networked environment using a logical connection with one or more other servers, a network personal computer (PC), or another network node.

[0088] The input device 850 may be one or more input devices, such as a mouse, a keyboard, a tracking ball, etc. The output device 860 may be one or more output devices, such as a display, a speaker, a printer, etc. The electronic device 800 may also communicate with one or more external devices (not shown) through the communication unit 840 as needed, such as a storage device, a display device, etc., communicate with one or more devices that allow a user to interact with the electronic device 800, or communicate with any device that allows the electronic device 800 to communicate with one or more other electronic devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).

[0089] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above.

[0090] Various aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of the methods, devices, equipment, and computer program products implemented according to the present disclosure. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer-readable program instructions.

[0091] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device that implements the functions / actions specified in one or more boxes in the flowchart and / or block diagram is generated. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause the computer, programmable data processing device, and / or other equipment to work in a specific manner, so that the computer-readable medium storing the instructions includes a manufactured product, which includes instructions for implementing various aspects of the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0092] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, so that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0093] The flow chart and block diagram in the accompanying drawings show the possible architecture, function and operation of the system, method and computer program product according to multiple implementations of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a part of 1 module, program segment or instruction, and a part of the module, program segment or instruction includes one or more executable instructions for realizing the specified logical function. In some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous square boxes can actually be executed substantially in parallel, and they can sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs the specified function or action, or can be implemented with a combination of special hardware and computer instructions.

[0094] The above descriptions of various implementations of the present disclosure are exemplary, non-exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described implementations. The selection of terms used herein is intended to best explain the principles of the implementations, practical applications, or improvements to the technology in the market, or to enable other persons of ordinary skill in the art to understand the various implementations disclosed herein.

Claims

1. A method for generating an image, comprising: Obtaining input text, wherein the input text indicates generating an image corresponding to at least one character; Processing the input text using a first model to determine material description text corresponding to the at least one character; Obtaining a material image generated based on the material description text; as well as generating a target image corresponding to the at least one character based on the material image and the glyph image corresponding to the at least one character, Wherein generating the target image corresponding to the at least one character comprises: generating a guide image based on the glyph image and the material image, so that the glyph portion of the guide image is filled based on the material image; And based on the guide image and the input text, a target image corresponding to the at least one character is generated.

2. The method of claim 1, wherein processing the input text using the first model comprises: Based on the input text, generating first input information of the first model; as well as The first input information is provided to the first model to obtain the material description text generated by the first model.

3. The method according to claim 2, wherein the first input information comprises: Preset constraint information, used to constrain output generation of the first model; and / or A set of reference material description text.

4. The method according to claim 1, wherein obtaining a material image generated based on the material description text comprises: Second input information is provided to a second model to obtain the material image generated by the second model, wherein the second input information is generated based on the material description text.

5. The method of claim 4, wherein the second model comprises a plurality of convolutional layers, and at least one of the plurality of convolutional layers is configured as: generating a first feature map corresponding to the first size; and The first feature map is padded to obtain a second feature map corresponding to a second size as an input of the at least one convolutional layer, wherein the second size is larger than the first size.

6. The method of claim 5, wherein filling the first feature map comprises: Determine a first edge position in the first feature map; Based on the first feature map, determining a second edge position symmetrical to the first edge position in a target direction; as well as The adjacent position of the first edge pixel in the target direction is filled by using the value of the second edge position, where the adjacent position is outside the first feature map.

7. The method according to claim 6, wherein the first edge position comprises a position on a left edge, a right edge, an upper edge or a lower edge of the first feature map, and the target direction comprises at least one of a horizontal direction, a vertical direction or a diagonal direction.

8. The method according to claim 1, further comprising: Determine a target glyph from a set of preset glyph libraries based on the input text and / or glyph selection information; as well as Based on the target glyph, the glyph image corresponding to the at least one character is generated.

9. The method according to claim 1, wherein generating a target image corresponding to the at least one character further comprises: The guide image and the input text are provided to a third model to obtain the target image generated by the third model.

10. The method of claim 1, wherein generating a guidance image comprises: Processing the glyph image to weaken the edge of the glyph image; as well as The guide image is generated based on the processed glyph image and the texture image.

11. The method according to claim 9, further comprising: A first control parameter is provided to the third model, the control parameter being used to indicate the intensity of noise added to the guidance image.

12. The method of claim 1, wherein generating a guidance image comprises: Generate an intermediate image based on the glyph image and the texture image; as well as The guide image is generated by adding noise corresponding to the second control parameter to the intermediate image.

13. The method according to claim 1, wherein the input text is a first input text, the method further comprising: Get the second input text; as well as At least one image generated based on the target image and the second input text is provided.

14. The method of claim 1, wherein the target image corresponds to an artistic representation of the at least one character.

15. An apparatus for generating an image, comprising: A text acquisition module, configured to acquire input text, wherein the input text indicates generation of an image corresponding to at least one character; a text determination module configured to process the input text using a first model to determine a material description text corresponding to the at least one character; An image acquisition module is configured to acquire a material image generated based on the material description text; as well as an image generation module, configured to generate a target image corresponding to the at least one character based on the material image and the glyph image corresponding to the at least one character, The image generation module is further configured to: generate a guide image based on the glyph image and the material image, so that the glyph part of the guide image is filled based on the material image; and generate a target image corresponding to the at least one character based on the guide image and the input text.

16. An electronic device, comprising: at least one processing unit; as well as At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to any one of claims 1 to 14 when executed by the at least one processing unit.

17. A computer-readable storage medium having a computer program stored thereon, wherein the computer program can be executed by a processor to implement the method according to any one of claims 1 to 14.

Citation Information

Patent Citations

  • Image generation method and device, readable medium and electronic equipment

    CN115908640A

  • Method for extracting attribute value of description object and related equipment

    CN116127080A