Method and apparatus for generating image, and device and storage medium

By acquiring text and character features and using the model to generate image content, the problem of traditional models being unable to render specific characters is solved, and high-quality image generation of multilingual characters is achieved.

WO2026158076A1PCT designated stage Publication Date: 2026-07-30BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
BEIJING ZITIAO NETWORK TECH CO LTD
Filing Date
2026-01-12
Publication Date
2026-07-30

AI Technical Summary

Technical Problem

Traditional models struggle to generate images with specific characters, failing to meet users' rendering needs.

Method used

By acquiring text content, determining text features and character graphic features, and using a model to process these features to generate image content, including graphic representations of characters.

Benefits of technology

It enables the generation of images with specific characters, supports character rendering tasks in multiple languages, and improves the quality and diversity of image generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2026072024_30072026_PF_FP_ABST
    Figure CN2026072024_30072026_PF_FP_ABST
Patent Text Reader

Abstract

The embodiments of the present disclosure relate to a method and apparatus for generating an image, and a device and a computer-readable storage medium. The method provided herein includes: acquiring text content, wherein the text content indicates at least one character to be rendered; determining a text feature corresponding to the text content and a character glyph feature corresponding to the at least one character; and processing the text feature and the character glyph feature by using a model, so as to generate image content, wherein the image content comprises a graphical representation corresponding to the at least one character.
Need to check novelty before this filing date? Find Prior Art

Description

Methods, apparatus, devices and storage media for generating images

[0001] This application claims priority to Chinese Patent Application No. 202510122515.7, filed on January 24, 2025, entitled “Method, Apparatus, Device and Storage Medium for Generating Images”, the entire contents of which are incorporated herein by reference. Technical Field

[0002] The exemplary embodiments disclosed herein relate generally to the field of computers, and more particularly to methods, apparatus, devices, and computer-readable storage media for generating images. Background Technology

[0003] With the development of computer technology, some models can now generate corresponding images based on user input text. Traditional models typically only generate images with styles or objects that correspond to the text, making it difficult to support character rendering requirements. For example, a user's request to render specific characters within text may not be responsive. Summary of the Invention

[0004] In a first aspect of this disclosure, a method for generating an image is provided. The method includes: acquiring text content indicating at least one character to be rendered; determining text features corresponding to the text content and character graphic features corresponding to the at least one character; and processing the text features and character graphic features using a model to generate image content, the image content including a graphic representation corresponding to the at least one character.

[0005] In a second aspect of this disclosure, an apparatus for generating an image is provided. The apparatus includes: an acquisition module configured to acquire text content indicating at least one character to be rendered; a determination module configured to determine text features corresponding to the text content and character graphic features corresponding to the at least one character; and a generation module configured to process the text features and character graphic features using a model to generate image content, the image content including a graphic representation corresponding to the at least one character.

[0006] In a third aspect of this disclosure, an electronic device is provided. The device includes at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor. When executed by the at least one processor, the instructions cause the electronic device to perform the method of the first aspect.

[0007] In a fourth aspect of this disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores computer-executable instructions that, when executed by a processor, implement the method of the first aspect.

[0008] In a fifth aspect of this disclosure, a computer program product is provided. The computer program product includes computer-executable instructions, wherein when executed by a processor, the computer-executable instructions implement the method according to a first aspect of this disclosure.

[0009] It should be understood that the content described in this content section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0010] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0011] Figure 1 shows a schematic diagram of an example environment in which embodiments of the present disclosure may be implemented;

[0012] Figure 2 shows a flowchart of an example process for generating an image according to some embodiments of the present disclosure;

[0013] Figure 3 shows an architecture diagram of an example model of an image generated according to some embodiments of the present disclosure;

[0014] Figure 4 shows a schematic structural block diagram of an example apparatus for generating images according to some embodiments of the present disclosure; and

[0015] Figure 5 shows a block diagram of an electronic device capable of implementing several embodiments of the present disclosure. Detailed Implementation

[0016] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0017] It should be noted that the headings of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and embodiments of any type may be included under any section / subsection. Furthermore, embodiments described in any section / subsection may be combined in any way with any other embodiments described in the same section / subsection and / or different sections / subsections.

[0018] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0019] The embodiments of this disclosure may involve user data, data acquisition, and / or use. All of these aspects comply with applicable laws, regulations, and relevant provisions. In the embodiments of this disclosure, all data collection, acquisition, processing, manipulation, forwarding, and use are conducted with the user's knowledge and confirmation. Accordingly, in implementing the embodiments of this disclosure, the type, scope of use, and usage scenarios of any data or information that may be involved should be communicated to the user and their authorization obtained in accordance with relevant laws and regulations through appropriate means. The specific methods of notification and / or authorization may vary depending on the actual situation and application scenario, and the scope of this disclosure is not limited in this respect.

[0020] In this specification and the embodiments, any processing of personal information will be carried out only under the premise of legality (such as obtaining the consent of the personal information subject, or being necessary for the performance of a contract), and will only be carried out within the scope stipulated or agreed upon. A user's refusal to process personal information other than that necessary for basic functions will not affect the user's use of basic functions.

[0021] As mentioned above, traditional models can typically only generate images with styles or objects corresponding to the text, making it difficult to support character rendering requirements. For example, a user's request to render a specific character in text may not be able to be responded to.

[0022] Embodiments of this disclosure propose a scheme for generating an image. The scheme includes: acquiring text content indicating at least one character to be rendered; determining text features corresponding to the text content and character graphic features corresponding to the at least one character; and processing the text features and character graphic features using a model to generate image content, the image content including a graphic representation corresponding to the at least one character.

[0023] In this way, embodiments of the present disclosure can extract text features related to semantic information and character graphic features related to the characters to be rendered, and use such features to jointly control the generation of the image, thereby generating an image with graphic representations corresponding to the characters.

[0024] The following section provides a detailed description of various example implementations of this scheme, with reference to the accompanying drawings.

[0025] Example Environment

[0026] Figure 1 illustrates a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. As shown in Figure 1, the example environment 100 may include an electronic device 110.

[0027] In this example environment 100, the electronic device 110 can acquire text content 130 and can use model 120 to generate image content 140 corresponding to the text content 130.

[0028] In some embodiments, text content 130 may indicate at least one character to be rendered. Taking FIG1 as an example, text content 130 may include "an A-style image with the words 'BCD' written on it". Accordingly, text content 130 may indicate that the characters to be rendered include "BCD".

[0029] As will be detailed below, electronic device 110 can use model 120 to generate image content 140, which may include a graphic representation 150 corresponding to the characters “BCD” to be rendered.

[0030] The detailed process of generating image content 140 using model 120 will be described in detail below with reference to Figures 2 and 3.

[0031] Electronic device 110 can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, handheld computers, portable gaming terminals, VR / AR devices, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, electronic device 110 can also support any type of user-facing interface (such as "wearable" circuitry).

[0032] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.

[0033] The following description will continue with reference to the accompanying drawings, which will provide some exemplary embodiments of this disclosure.

[0034] Example process

[0035] Figure 2 shows a flowchart of an example process 200 for generating an image according to some embodiments of the present disclosure. Process 200 can be implemented at electronic device 110.

[0036] As shown in the figure, in box 210, electronic device 110 can acquire text content 130, which indicates at least one character to be rendered.

[0037] In some embodiments, the electronic device 110 can determine at least one character to be rendered from the text content 130 based on the grammatical information of the text content 130. As an example, the electronic device 110 can determine that the character enclosed in quotation marks is the character to be rendered. Accordingly, the electronic device 110 can determine the character "BCD" to be rendered by detecting quotation mark elements in the text content 130.

[0038] In some embodiments, considering that the user's input may not meet such grammatical rules, the electronic device 110 can acquire the user's input and rewrite the input to determine the text content 130.

[0039] As an example, electronic device 110 can acquire a prompt input by the user and rewrite the prompt using a language model to obtain text content. For example, the prompt input by the user could include "an A-style image with BCD written on it". Furthermore, the language model can rewrite the prompt and add quotation marks to the characters to be rendered.

[0040] In some embodiments, the electronic device 110 may also provide text content 130 to the language model to determine at least one character to be rendered based on the processing results of the language model.

[0041] In box 220, electronic device 110 determines text features corresponding to text content and character graphic features corresponding to at least one character.

[0042] The following description, with further reference to Figure 3, illustrates an example process for generating images using a model. Figure 3 shows an example model architecture according to some embodiments of this disclosure.

[0043] As shown in Figure 3, the first encoding unit 320 can be used to process the text content 130 to generate text features corresponding to the text content 130. In some embodiments, the first encoding unit 320 can be implemented using a language model or other suitable encoding model, for example.

[0044] Furthermore, the second encoding unit 322 can be used to process the character "BCD" to be rendered in the text content 130 and determine the character graphic features corresponding to the character "BCD". In some embodiments, the second encoding unit 322 can be a graphic encoding model associated with multiple languages, which can process characters associated with multiple languages ​​and generate character graphic features independent of the font, color, size and other attributes of the characters.

[0045] As an example, the graph encoding model can support the processing of Chinese or English characters to generate the corresponding character graph features.

[0046] In box 230, electronic device 110 uses a model to process text features and character graphic features to generate image content 140, the image content including a graphic representation 150 corresponding to at least one character.

[0047] As shown in Figure 3, model 120 may include, for example, a diffusion model based on a transformer, which may include, for example, multiple processing blocks (also known as DiT Blocks) 330.

[0048] In some embodiments, during the inference phase of the diffusion model, the diffusion model can acquire input features constructed based on text features and character graphic features. As an example, mapping unit 324 can be used to map the character graphic features of the character reminder to a feature space corresponding to the text features, thereby determining the mapped character graphic features. Furthermore, the input features of the model can be constructed by concatenating the text features and the mapped character graphic features.

[0049] As an example, mapping unit 324 can be implemented using a multilayer perceptron (MLP), which can map character graphic features to the same feature dimension as text features. Furthermore, complete text features, i.e., text embedding 326, are obtained by concatenating text features and the mapped character graphic features.

[0050] In some embodiments, before being provided to the processing block 330, a position code 328 may be added to the text embedding 326, which may indicate the position of the character to be rendered in the text content 130.

[0051] In addition, the diffusion model can also acquire the initial noise and perform noise reduction based on the input features to generate image content 140.

[0052] As shown in Figure 3, each processing block 330 in model 120 may include an attention unit 338, which may include, for example, acquiring the corresponding input image features 332 and input text features 334, and may update the input image features and input text features based on the attention mechanism.

[0053] In this way, embodiments of this disclosure can implement cross-attention mechanisms in both text and image modalities, thereby improving the quality of generated images.

[0054] Furthermore, as shown in Figure 3, the updated input image features are provided to the first feature transfer unit 340, and the updated input text features are provided to the second feature transfer unit 342. The first and second feature transfer units are independent of each other and may, for example, have independently trained model parameters.

[0055] As an example, the first feature transfer unit 340 and the second feature transfer unit 342 may include two independently trained MLP units.

[0056] In some embodiments, scaling control parameters can also be determined based on the resolution of the image to be generated. As an example, different resolutions can correspond to different scaling control parameters. The scaling control parameters can, for example, be provided to the RoPE (Rotary Positional Embedding) unit 336 to add control features corresponding to the scaling control parameters to the input image features.

[0057] In this way, embodiments of the present disclosure can control the model to generate image content corresponding to the resolution by injecting scaling control parameters into the model. Thus, the model can generalize to generate images at resolutions or scales not present during the training phase.

[0058] In this way, embodiments of the present disclosure can extract text features related to semantic information and character graphic features related to the characters to be rendered, and use such features to jointly control the generation of the image, thereby generating an image with graphic representations corresponding to the characters.

[0059] The following section will further describe the specific training process of the model.

[0060] In some embodiments, during the training of model 120, sample images 302 may be acquired. For example, such sample images 302 may include sample graphic representations corresponding to sample characters. For instance, sample image 302 may include a graphic representation corresponding to the characters "BCD".

[0061] Furthermore, a visual language model or other suitable model can be used to generate descriptive text for sample image 302. As an example, such descriptive text can also be referred to as the caption of sample image 302.

[0062] In some embodiments, the descriptive text may indicate one or more rendering features of the sample character. For example, the descriptive text may indicate the font in the sample image 302 of the sample character "BCD". As another example, the descriptive text may indicate the color in the sample image 302 of the sample character "BCD". As yet another example, the descriptive text may indicate the size in the sample image 302 of the sample character "BCD". As yet another example, the descriptive text may indicate the position in the sample image 302 of the sample character "BCD".

[0063] Furthermore, based on the Optical Character Recognition (OCR) results of the sample image 302, sample characters, such as "BCD", can be determined in the sample image 302.

[0064] Accordingly, sample text 316 corresponding to sample image 302 can be constructed based on descriptive text and sample characters. Furthermore, model 120 can be trained using sample image 302 and corresponding sample text 316.

[0065] As shown in Figure 3, the encoder 308 can be used to encode the sample image 302, and noise 310 can be added to the image features to obtain noisy image features. Further, the switching processing unit 312 can be used to segment the noisy image features to obtain multiple image tokens corresponding to different image segments, i.e., image embeddings 314. Image embeddings 314 are also called a set of training image features.

[0066] On the other hand, similar to the inference stage of the model, the encoding unit 320 can be used to process the sample image 316, and the encoding unit 322 and the mapping unit 324 can be used to process the sample characters, thereby obtaining the text embedding 326 of the training stage, that is, a set of training image features.

[0067] As can be seen, the text embedding 326 may include a first part (also called rendering feature) corresponding to the sample text 316 and a second part (i.e., character graphic feature, also called glyph feature) corresponding to the sample character “BCD”.

[0068] Furthermore, model 120 can be trained based on image embedding 314 and text embedding 326, thereby enabling model 120 to generate images containing character graphics based on text content.

[0069] Furthermore, as shown in Figure 3, during the training process, the scaling control parameter 306 can be determined based on the resolution of the sample image 302. This scaling control parameter 306 can be used to control the processing of the RoPE unit 336 during the training phase. As an example, the scaling control parameter 306 can indicate the position representation corresponding to the image center region of the sample image 302. In this way, the embodiments of this disclosure can improve the stability of model training and accelerate the model training convergence.

[0070] In some embodiments, after the pre-training of model 120 is completed, one or more post-training procedures may be performed. As an example, the post-training procedure may include continued training, which may, for example, improve the average aesthetic level of the training data, and may add parameters related to the aesthetic level, thereby improving the aesthetic level of the image content generated by the model.

[0071] As another example, high-quality training datasets can be used for supervised fine-tuning of the model to further improve the quality of the image content generated by the model. As yet another example, the model can be further fine-tuned based on user feedback.

[0072] Based on the process described above, on the one hand, embodiments of this disclosure can support the generation of images with graphical representations of characters; on the other hand, by utilizing graphical encoding models associated with multiple languages, embodiments of this disclosure can support character graphical rendering tasks for multiple languages.

[0073] Example devices and equipment

[0074] Embodiments of this disclosure also provide corresponding apparatus for implementing the methods or processes described above. FIG4 shows a schematic structural block diagram of an example apparatus 400 for generating images according to certain embodiments of this disclosure. Apparatus 400 may be implemented as or included in electronic device 110. The various modules / components in apparatus 400 may be implemented by hardware, software, firmware, or any combination thereof.

[0075] As shown in Figure 4, the device 400 includes: an acquisition module 410 configured to acquire text content indicating at least one character to be rendered; a determination module 420 configured to determine text features corresponding to the text content and character graphic features corresponding to at least one character; and a generation module 430 configured to process the text features and character graphic features using a model to generate image content, the image content including a graphic representation corresponding to at least one character.

[0076] In some embodiments, the apparatus 400 further includes a character determination module configured to: determine at least one character to be rendered from text content based on grammatical information of the text content; or provide text content to a language model to determine at least one character to be rendered.

[0077] In some embodiments, the acquisition module 410 is further configured to: acquire user input content; and determine text content by rewriting the input content.

[0078] In some embodiments, the apparatus 400 further includes a feature construction module configured to: map character graphic features to a feature space corresponding to text features to determine the mapped character graphic features; and construct input features of the model by concatenating text features and mapped character graphic features.

[0079] In some embodiments, the model is trained based on the following process: generating descriptive text for sample images, the sample images including sample graphic representations corresponding to sample characters; determining sample characters based on optical character recognition (OCR) results of the sample images; constructing sample text corresponding to the sample images based on the descriptive text and the sample characters; and training the model using the sample images and the sample text.

[0080] In some embodiments, training a model using sample images and sample text includes: constructing a set of training image features based on the sample images; constructing a set of training text features based on the sample text, the set of training text features including a first part corresponding to the sample text and a second part corresponding to the sample characters, the second part including the character graphic features of the sample characters; and training a model based on the set of training image features and the set of training text features.

[0081] In some embodiments, the description text indicates the rendering characteristics of the sample characters, and the rendering characteristics include at least one of the following: the font of the sample characters; the color of the sample characters; the size of the sample characters; and the position of the sample characters.

[0082] In some embodiments, the apparatus 400 further includes a scaling control module configured to: determine scaling control parameters based on the resolution of the image to be generated; and provide scaling control parameters to the model to control the model to generate image content corresponding to the resolution.

[0083] In some embodiments, the model includes an attention unit configured to: acquire input image features and input text features; and update the input image features and input text features based on an attention mechanism.

[0084] In some embodiments, updated input image features are provided to a first feature transfer unit, and updated input text features are provided to a second feature transfer unit, wherein the first feature transfer unit is independent of the second feature transfer unit.

[0085] In some embodiments, the determining module 420 is further configured to provide at least one character to a graphical encoding model associated with multiple languages ​​in order to determine the character graphical features corresponding to the at least one character.

[0086] Figure 5 shows a block diagram of an electronic device 500 in which one or more embodiments of the present disclosure may be implemented. It should be understood that the electronic device 500 shown in Figure 5 is merely exemplary and should not constitute any limitation on the functionality and scope of the embodiments described herein. The electronic device 500 shown in Figure 5 can be used to implement the electronic device 110 of Figure 1.

[0087] As shown in Figure 5, the electronic device 500 is in the form of a general-purpose electronic device. Components of the electronic device 500 may include, but are not limited to, one or more processing units or processors 510, memory 520, storage devices 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. The processor 510 may be a physical or virtual processor and is capable of performing various processes according to the programs stored in the memory 520. In a multiprocessor system, multiple processors execute computer-executable instructions in parallel to improve the parallel processing capability of the electronic device 500.

[0088] Electronic device 500 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 500, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 520 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 530 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data and can be accessed within electronic device 500.

[0089] Electronic device 500 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 5, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks may be provided. In these cases, each drive may be connected to a bus (not shown) via one or more data media interfaces. Memory 520 may include computer program product 525 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.

[0090] Communication unit 540 enables communication with other electronic devices via a communication medium. Additionally, the functionality of components of electronic device 500 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, electronic device 500 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.

[0091] Input device 550 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 560 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 500 can also communicate with one or more external devices (not shown) via communication unit 540 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 500, or with any device that enables electronic device 500 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).

[0092] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.

[0093] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0094] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0095] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0096] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0097] Various implementations of this disclosure have been described above. The foregoing description is exemplary and not exhaustive, nor is it limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A method for generating an image, comprising: Obtain text content, which indicates at least one character to be rendered; Determine the text features corresponding to the text content and the character graphic features corresponding to the at least one character; as well as The text features and character graphic features are processed using a model to generate image content, the image content including a graphic representation corresponding to the at least one character.

2. The method according to claim 1, further comprising: Based on the grammatical information of the text content, determine the at least one character to be rendered from the text content; or The text content is provided to the language model to determine the at least one character to be rendered.

3. The method according to claim 1 or 2, wherein obtaining the text content includes: Get the user's input; as well as The text content is determined by rewriting the input content.

4. The method according to any one of claims 1 to 3, further comprising: The character graphic features are mapped to the feature space corresponding to the text features to determine the mapped character graphic features; as well as The input features of the model are constructed by concatenating the text features and the mapped character graphic features.

5. The method according to any one of claims 1 to 4, wherein the model is trained based on the following process: Based on the sample image, a descriptive text for the sample image is generated, wherein the sample image includes a sample graphic representation corresponding to the sample characters; Based on the optical character recognition (OCR) results of the sample image, the sample character is determined; Based on the descriptive text and the sample characters, construct sample text corresponding to the sample image; and The model is trained using the sample images and the sample text.

6. The method of claim 5, wherein training the model using the sample image and the sample text comprises: Based on the sample images, a set of training image features is constructed; Based on the sample text, a set of training text features is constructed. The set of training text features includes a first part corresponding to the sample text and a second part corresponding to the sample character. The second part includes the character graphic features of the sample character. as well as The model is trained based on the set of training image features and the set of training text features.

7. The method of claim 5, wherein the descriptive text indicates the rendering features of the sample characters, and the rendering features include at least one of the following: The font of the sample characters; The color of the sample character; The size of the sample character; The position of the sample character.

8. The method according to any one of claims 1 to 7, further comprising: Determine the scaling control parameters based on the resolution of the image to be generated; as well as The scaling control parameters are provided to the model to control the model to generate image content corresponding to the resolution.

9. The method according to any one of claims 1 to 8, wherein the model includes an attention unit, the attention unit being configured to: Obtain input image features and input text features; and The input image features and the input text features are updated based on an attention mechanism.

10. The method of claim 9, wherein the updated input image features are provided to a first feature transfer unit, the updated input text features are provided to a second feature transfer unit, and the first feature transfer unit is independent of the second feature transfer unit.

11. The method according to any one of claims 1 to 10, wherein determining the character graphic feature corresponding to the at least one character comprises: The at least one character is provided to a graphical encoding model associated with multiple languages ​​to determine the graphical features of the character corresponding to the at least one character.

12. An apparatus for generating an image, comprising: The acquisition module is configured to acquire text content indicating at least one character to be rendered; The determining module is configured to determine text features corresponding to the text content and character graphic features corresponding to the at least one character; as well as The generation module is configured to process the text features and the character graphic features using a model to generate image content, the image content including a graphic representation corresponding to the at least one character.

13. An electronic device, comprising: At least one processor; as well as At least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions causing the electronic device to perform the method according to any one of claims 1 to 11 when executed by the at least one processor.

14. A computer-readable storage medium having stored thereon computer-executable instructions that can be executed by a processor to implement the method according to any one of claims 1 to 11.

15. A computer program product comprising computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the method according to any one of claims 1 to 11.