Image generation method and device, electronic equipment and storage medium
By acquiring background image description information and the text to be displayed from the target text, semantic understanding is used to determine the text to be displayed and synthesize the image. This solves the problem of inaccurate text display in existing technologies, realizes the generation of images containing specific text, and meets the diverse needs of users.
Patent Information
- Application Number
- CN202411155324.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-21
- Publication Date
- 2026-03-06
AI Technical Summary
Existing image generation models cannot accurately represent specific text in the generated images, resulting in unsatisfactory text display effects.
By acquiring background image description information and the text to be displayed from the target text, semantic understanding is used to determine the text to be displayed. After generating the first image, it is combined with the text to be displayed to ensure that the text is accurately displayed in the image.
It enables the generation of images to include user-specified specific text, meeting diverse image generation needs and improving the accuracy and consistency of text display.
Smart Images

Figure CN121616686A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to an image generation method, apparatus, electronic device, and storage medium. Background Technology
[0002] In the digital age, image generation technology, as a significant breakthrough in the field of artificial intelligence, has greatly enriched users' content creation experience and brought unprecedented innovation to numerous industries. However, it still has certain limitations and cannot fully meet the needs of all image generation scenarios. For example, when users want to include specific text in the generated image, existing image generation models cannot accurately represent this text in the generated image, resulting in unsatisfactory text display effects. Summary of the Invention
[0003] In order to solve the above-mentioned technical problems, or at least partially solve the above-mentioned technical problems, this disclosure provides an image generation method, apparatus, electronic device, and storage medium.
[0004] In a first aspect, this disclosure provides an image generation method, including:
[0005] Obtain the target text; the target text includes descriptive information about the background image and the text to be displayed;
[0006] Based on the target text, determine the text to be displayed;
[0007] Based on the target text, generate a first image;
[0008] The text to be displayed is combined with the first image to obtain a target image, which includes the text to be displayed.
[0009] Secondly, this disclosure also provides an image generation method, including:
[0010] An acquisition module is used to acquire target text; the target text includes descriptive information about the background image and text to be displayed; a determination module is used to determine the text to be displayed based on the target text;
[0011] A generation module is used to generate a first image based on the target text;
[0012] A compositing module is used to combine the text to be displayed with the first image to obtain a target image, wherein the target image includes the text to be displayed.
[0013] Thirdly, this disclosure also provides an electronic device, the electronic device comprising:
[0014] One or more processors;
[0015] Storage device for storing one or more programs;
[0016] When the one or more programs are executed by the one or more processors, the one or more processors implement the image generation method as described above.
[0017] Fourthly, this disclosure also provides a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the image generation method described above.
[0018] The technical solution provided in this disclosure has the following advantages compared with the prior art:
[0019] The technical solution provided in this disclosure involves setting and obtaining target text; the target text includes descriptive information about a background image and text to be displayed; determining the text to be displayed based on the target text; generating a first image based on the target text; and combining the text to be displayed with the first image to obtain a target image, which includes the text to be displayed. Essentially, it provides a method to generate images containing specific text, satisfying diverse image generation needs of users. It can be applied to scenarios involving the creation of wallpapers or posters. Attached Figure Description
[0020] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0021] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 A flowchart of an image generation method provided in this disclosure embodiment;
[0023] Figure 2 A schematic diagram of a first image provided for an embodiment of this disclosure;
[0024] Figure 3 A schematic diagram of a target image provided in an embodiment of this disclosure;
[0025] Figure 4 This is a schematic diagram of the structure of an image generation device according to an embodiment of the present disclosure;
[0026] Figure 5 This is a schematic diagram of the structure of an electronic device according to an embodiment of this disclosure. Detailed Implementation
[0027] To better understand the above-mentioned objectives, features, and advantages of this disclosure, the solutions disclosed herein will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.
[0028] Numerous specific details are set forth in the following description in order to provide a full understanding of this disclosure, but this disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some, and not all, of the embodiments of this disclosure.
[0029] Figure 1 This is a flowchart illustrating an image generation method provided in this embodiment. This embodiment is applicable to situations where image generation is performed on a client side. The method can be executed by an image generation device, which can be implemented in software and / or hardware. This device can be configured in an electronic device, such as a terminal, including but not limited to smartphones, PDAs, tablets, wearable devices with displays, desktop computers, laptops, all-in-one computers, smart home devices, etc. Alternatively, this embodiment can be applied to situations where image generation is performed on a server side. The method can be executed by an image generation device, which can be implemented in software and / or hardware. This device can be configured in an electronic device, such as a server.
[0030] like Figure 1 As shown, the method may specifically include:
[0031] S110. Obtain the target text; the target text includes descriptive information about the background image and the text to be displayed.
[0032] The target text can be, for example, text input by the user. Subsequently, the target text can be directly used as a prompt for the image generation model, inputting into the model to generate an image.
[0033] In this application, the desired final image is the target image. The target image includes a background image and text information (also referred to as text) superimposed on the background image. The descriptive information of the background image in the target text is used to guide the image generation model to generate the background image, which is the first image mentioned below. The text to be displayed in the target text is the text information that the user wants to superimpose on the background image.
[0034] In practice, this application does not restrict the descriptive information of the background image or the order of the text to be displayed in the target text.
[0035] For example, suppose the target text is "There is a person playing frisbee with a puppy on the grass, with the words 'Happy times on the grass' written above it," where "There is a person playing frisbee with a puppy on the grass" is descriptive information about the background image, and "Happy times on the grass" is the text to be displayed.
[0036] S120. Based on the target text, determine the text to be displayed.
[0037] Although the target text includes the text to be displayed, the electronic device does not "know" which characters or characters belong to the text to be displayed. This step essentially involves parsing the target text so that the electronic device understands which characters or characters belong to the text to be displayed.
[0038] There are multiple ways to implement this step, and this application does not limit this one. For example, the implementation method of this step includes: performing semantic understanding on the target text to obtain the semantic understanding result; and obtaining the text to be displayed based on the semantic understanding result.
[0039] Semantic understanding of target text can be achieved using models with semantic understanding capabilities. These models can perform comprehensive and in-depth analysis of the target text, clarifying its overall meaning, the meaning of each word, and the relationships between entities within the text. By understanding the semantics of the target text, electronic devices can determine which characters(s) the user intends to display, thus improving the accuracy of text selection.
[0040] Furthermore, the target text can be configured to include a location identifier, which indicates the position of the text to be displayed within the target text. Obtaining the text to be displayed based on the semantic understanding results can include: determining the text to be displayed within the target text based on the semantic understanding results and the location identifier.
[0041] Location identifiers can be pre-defined, and prompts can guide users to use them when they enter target text. This application does not limit the specific type of identifier a location identifier can be. For example, a location identifier can be a specific punctuation mark, such as “”. Or, a location identifier can be a specific phrase or word, such as “writing…”.
[0042] For example, if the target text is "There is a person playing frisbee with a puppy on the grass, with the words 'Happy times on the grass' written on it", the location identifier is ''. Based on the location identifier, "Happy times on the grass" can be obtained as the text to be displayed.
[0043] Optionally, based on the semantic understanding results and location identifiers, determining the text to be displayed in the target text may include: performing semantic understanding on the target text to obtain a first text and its confidence level; determining a second text in the target text based on the location identifiers, and determining the confidence level of the second text. Based on the confidence levels of the first and second texts, the text to be displayed is determined from both the first and second texts. The first and second texts are determined from the target text using different methods and may be characters or strings used as the text to be displayed. Optionally, if the confidence level of the first text is higher than that of the second text, the first text is determined to be the text to be displayed. If the confidence level of the second text is higher than that of the first text, the second text is determined to be the text to be displayed. This setting can make the obtained text to be displayed more accurate.
[0044] S130. Generate the first image based on the target text.
[0045] There are multiple ways to implement this step, and this step does not limit this one. For example, the implementation method of this step may include: inputting the target text into the image generation module (such as a text-to-image model or an image-to-image model) so that the image generation model generates a first image.
[0046] S140. Combine the text to be displayed with the first image to obtain a target image, which includes the text to be displayed.
[0047] There are multiple ways to implement this step, and this application does not limit this one. For example, the text to be displayed can be used as a top-level element and superimposed on the first image to obtain the target image.
[0048] For example, suppose the target text is "A person is playing frisbee with a puppy on the grass, with the words 'Happy times on the grass' written on it." Based on this target text, the text to be displayed can be "Happy times on the grass." The first image generated based on this target text is as follows: Figure 2 As shown, the "Happy Time on the Grass" image is composited with the first image to obtain the target image. The resulting target image is as follows. Figure 3 As shown, the target image contains the text "Happy times on the grass".
[0049] Existing image generation models, especially diffusion models, typically involve progressively denoising an initial image with random noise to generate the final image. While these models can understand and reflect the provided prompts to some extent, in practice, due to their complexity and uncertainty, they cannot accurately represent the text to be displayed in the generated image.
[0050] The above technical solution involves setting and obtaining target text; the target text includes descriptive information about the background image and text to be displayed; based on the target text, the text to be displayed is determined; based on the target text, a first image is generated; and the text to be displayed is combined with the first image to obtain a target image, which includes the text to be displayed. Essentially, it provides a method to generate images containing specific text, satisfying diverse image generation needs of users. This technical solution can be applied to scenarios involving the creation of wallpapers or posters.
[0051] It is also important to emphasize that, since the target text explicitly contains text that needs to be displayed in the image, the text to be displayed is directly extracted from the target text and overlaid directly onto the generated image. This ensures that the text displayed in the target image is consistent with the text to be displayed in the target text. In other words, the text included in the target image generated using the technical solution provided in this application is user-specified and conforms to the user's needs, rather than being randomly generated.
[0052] In the above technical solution, S140 may include: determining a rendering scheme corresponding to the text to be displayed; rendering the text to be displayed based on the rendering scheme corresponding to the text to be displayed; and combining the rendered text to be displayed with the first image to obtain the target image.
[0053] The rendering scheme corresponding to the text to be displayed can be, for example, a rendering scheme suitable for the text to be displayed, which may include one or more aspects such as the display font, display size, display position, and display color of the text to be displayed.
[0054] There are various specific methods for "determining the rendering scheme corresponding to the text to be displayed," and this application does not limit this. In practice, the rendering scheme corresponding to the text to be displayed can be determined through user interaction. For example, if the rendering scheme includes the display position of the target text, determining the rendering scheme corresponding to the text to be displayed may include: displaying a first pattern and a region selection tool, the region selection tool being used to assist the user in selecting the display position of the text to be displayed in the first pattern; in response to the use of the region selection tool, the region indicated by the region selection tool is taken as the display position corresponding to the text to be displayed.
[0055] If the rendering scheme includes one or more of the display font, display size, and display color of the target text, determining the rendering scheme corresponding to the text to be displayed may include: displaying options related to the rendering scheme, wherein at least one of the display font, display size, and display color corresponding to the options is different; in response to the selection operation of one of the options related to the rendering scheme, the rendering scheme corresponding to the selected option is used as the rendering scheme corresponding to the text to be displayed.
[0056] In some scenarios, the target text may include user requirements regarding the specific rendering scheme for the text to be displayed. For example, in such cases, "determining the rendering scheme corresponding to the text to be displayed" may include: performing semantic understanding on the target text to obtain the semantic understanding result; and based on the semantic understanding result, obtaining the rendering scheme corresponding to the text to be displayed.
[0057] For example, the target text is "A person is playing frisbee with a puppy on the grass, and the words 'Happy times on the grass' are written in the sky." In this example, the target text specifies the display location of "Happy times on the grass" (i.e., the text to be displayed)—in the sky. By semantically understanding the target text, it can be determined that the user wants the text to be displayed to be shown in the sky.
[0058] Furthermore, in order to ensure that the text to be displayed can be properly displayed in the target image so that the text to be displayed can be well integrated with the background image in the target image, a rendering scheme corresponding to the text to be displayed can be set based on the semantic understanding results, including: determining the rendering scheme corresponding to the text to be displayed based on the semantic understanding results and the feature information of the first image; the feature information of the first image includes at least one of the size, color and content of the first image.
[0059] For example, the size of the text to be displayed can be determined based on the size of the first image. For instance, the upper and lower limits of the text size can be determined based on the size of the first image, thereby ensuring that the size of the text to be displayed is greater than a set minimum size threshold and smaller than the size of the first image. This prevents situations where the text size is too small, making it difficult for the user to quickly observe, or where the text size is too large, resulting in incomplete display in the first image.
[0060] The content of the first image may include information such as the position of objects within the first image. Objects in the first image can be things in the image, such as people, animals, plants, buildings, objects, the sky, the ground, etc. For example, the text to be displayed is set not to overlap with the main object in the first image. The main object can be, for example, an object in the first image that needs to be emphasized or highlighted. For example, the main object can be a person, animal, plant, building, or object. There can be one or more main objects. Setting the text to be displayed not to overlap with the main object in the first image aims to prevent the target text from obscuring the main object in the first image.
[0061] The color of the first image may be, for example, the main color tone of the first image, and / or the color of a local area in the first image.
[0062] Optionally, the display color of the text to be displayed can be determined based on the color of the first image. This setting serves two purposes: firstly, it ensures that the rendered color tone of the text to be displayed is similar to the overall color tone of the first image, allowing the text to blend naturally and seamlessly into the first image. Secondly, it ensures that the color of the rendered text to be displayed differs from the color of the area in the first image where the text is displayed, allowing the first image to be distinguished from the text by color, making the text easier to observe. This maintains overall visual harmony while ensuring high readability and prominence of the text.
[0063] Furthermore, a color reference area can be determined in the first image, and the display color of the target text can be determined based on the average pixel color in the color reference area. Optionally, the entire first image can be used as the color reference area, or the color reference area can be determined with the center point of the display area of the first image as the center point and a preset distance as the radius.
[0064] Optionally, if the display color of the text to be displayed is determined based on the color of a local area in the first image, a color reference area can be determined with the display position of the first image as the center point and a preset distance as the radius, and the display color of the text to be displayed can be determined based on the color of the color reference area.
[0065] Based on the above technical solution, the method may optionally include: recognizing the text in the target image to obtain a first text recognition result; if the first text recognition result is consistent with the text to be displayed, outputting the target image.
[0066] In practice, if the target text is used directly as the prompt information, the generated first image may also contain text, but this text may differ from the text to be displayed. By setting the text in the target image to be recognized, a first text recognition result is obtained; if the first text recognition result matches the text to be displayed, the target image is output. The purpose is to strictly control the target image, filtering out target images containing text that differs from the text to be displayed, thereby improving the image generation experience for users with specific needs.
[0067] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0068] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0069] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0070] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0071] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, because according to the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.
[0072] Figure 4 This is a schematic diagram of an image generation apparatus according to an embodiment of this disclosure. The image generation apparatus provided in this embodiment can be configured in a client or in a server. See also Figure 4 The image generation device specifically includes:
[0073] The acquisition module 310 is used to acquire target text; the target text includes descriptive information about the background image and text to be displayed.
[0074] The determining module 320 is used to determine the text to be displayed based on the target text;
[0075] The generation module 330 is used to generate a first image based on the target text;
[0076] The compositing module 340 is used to combine the text to be displayed with the first image to obtain a target image, wherein the target image includes the text to be displayed.
[0077] Furthermore, module 320 is defined as being used for:
[0078] Perform semantic understanding on the target text to obtain the semantic understanding result;
[0079] Based on the semantic understanding results, the text to be displayed is obtained.
[0080] Further, the target text includes a location identifier, which is used to indicate the position of the text to be displayed within the target text. The determining module 320 is used to:
[0081] Based on the semantic understanding results and the location identifier, the text to be displayed is determined in the target text.
[0082] Furthermore, the synthesis module 340 is used for:
[0083] Determine the rendering scheme corresponding to the text to be displayed; the rendering scheme includes at least one of the target text's display font, display size, display position, and display color;
[0084] The text to be displayed is rendered based on the rendering scheme corresponding to the text to be displayed.
[0085] The rendered text to be displayed is combined with the first image to obtain the target image.
[0086] Furthermore, the synthesis module 340 is used for:
[0087] Perform semantic understanding on the target text to obtain the semantic understanding result;
[0088] Based on the semantic understanding results, a rendering scheme corresponding to the text to be displayed is obtained.
[0089] Furthermore, the synthesis module 340 is used for:
[0090] Based on the semantic understanding results and the feature information of the first image, a rendering scheme corresponding to the text to be displayed is determined; the feature information of the first image includes at least one of the size, color, and content of the first image.
[0091] Furthermore, the device also includes an output module for:
[0092] The text in the target image is recognized to obtain a first text recognition result;
[0093] If the first character recognition result matches the text to be displayed, the target image is output.
[0094] The image generation apparatus provided in this disclosure can execute the steps performed by the client or server in the image generation method provided in this disclosure, and has the execution steps and beneficial effects, which will not be described in detail here.
[0095] Figure 5 This is a schematic diagram of the structure of an electronic device according to an embodiment of this disclosure. See below for details. Figure 5 The diagram illustrates a structural schematic suitable for implementing the electronic device 1000 in the embodiments of this disclosure. The electronic device 1000 in the embodiments of this disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), wearable electronic devices, etc., as well as fixed terminals such as digital TVs, desktop computers, smart home devices, etc. Figure 5 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0096] like Figure 5 As shown, the electronic device 1000 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 1001, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1008 into a random access memory (RAM) 1003 to implement the image generation method as described in the embodiments of this disclosure. The RAM 1003 also stores various programs and information required for the operation of the electronic device 1000. The processing device 1001, ROM 1002, and RAM 1003 are interconnected via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0097] Typically, the following devices can be connected to the I / O interface 1005: input devices 1006 including, for example, a touchscreen, touchpad, keyboard, mouse, camera, microphone, accelerometer, gyroscope, etc.; output devices 1007 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1008 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. Communication device 1009 allows electronic device 1000 to exchange information with other devices wirelessly or via wired communication. Although Figure 5 An electronic device 1000 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0098] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts, thereby implementing the image generation method as described above. In such embodiments, the computer program can be downloaded and installed from a network via communication device 1009, or installed from storage device 1008, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of embodiments of this disclosure.
[0099] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include information signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated information signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0100] In some implementations, clients and servers may communicate using any known or future network protocol, such as HTTP (Hypertext Transfer Protocol), and may interconnect with any form or medium of digital information communication (e.g., a communication network). Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any known or future network.
[0101] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0102] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to:
[0103] Obtain the target text; the target text includes descriptive information about the background image and the text to be displayed;
[0104] Based on the target text, determine the text to be displayed;
[0105] Based on the target text, generate a first image;
[0106] The text to be displayed is combined with the first image to obtain a target image, which includes the text to be displayed.
[0107] Optionally, when one or more of the above-described procedures are executed by the electronic device, the electronic device may also perform other steps described in the above embodiments.
[0108] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0109] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0110] The units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units are not, in some cases, intended to limit the specific unit.
[0111] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0112] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0113] According to one or more embodiments of this disclosure, this disclosure provides an electronic device, including:
[0114] One or more processors;
[0115] Memory, used to store one or more programs;
[0116] When the one or more programs are executed by the one or more processors, the one or more processors implement any of the image generation methods provided in this disclosure.
[0117] According to one or more embodiments of the present disclosure, the present disclosure provides a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements an image generation method as described in any of the present disclosure.
[0118] This disclosure also provides a computer program product, which includes a computer program or instructions that, when executed by a processor, implement the image generation method described above.
[0119] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0120] The above description is merely a specific embodiment of this disclosure, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An image generation method characterized by, The method comprises: obtaining target text; the target text comprises description information of a background image and to-be-displayed text; determining to-be-displayed text based on the target text; generating a first image based on the target text; synthesizing the to-be-displayed text and the first image to obtain a target image, the target image comprising the to-be-displayed text.
2. The method of claim 1, wherein, The method further comprises: performing semantic understanding on the target text to obtain a semantic understanding result; obtaining to-be-displayed text based on the semantic understanding result.
3. The method of claim 2, wherein, The target text comprises a position identifier used to indicate a position of the to-be-displayed text in the target text, and the method further comprises: obtaining to-be-displayed text in the target text based on the semantic understanding result and the position identifier.
4. The method of claim 1, wherein, The method further comprises: determining a rendering scheme corresponding to the to-be-displayed text; the rendering scheme comprises at least one of a display font, a display size, a display position, and a display color of the target text; rendering the to-be-displayed text based on the rendering scheme corresponding to the to-be-displayed text; synthesizing the rendered to-be-displayed text and the first image to obtain a target image.
5. The method of claim 1, wherein, The method further comprises: performing semantic understanding on the target text to obtain a semantic understanding result; obtaining a rendering scheme corresponding to the to-be-displayed text based on the semantic understanding result.
6. The method of claim 5, wherein, The method further comprises: obtaining a rendering scheme corresponding to the to-be-displayed text based on the semantic understanding result and feature information of the first image; the feature information of the first image comprises at least one of a size, a color, and content of the first image.
7. The method of claim 1, wherein, The method further comprises: recognizing text in the target image to obtain a first text recognition result; if the first text recognition result is consistent with the to-be-displayed text, outputting the target image.
8. An image generation method characterized by, The method comprises: an obtaining module configured to obtain target text; the target text comprises description information of a background image and to-be-displayed text; a determining module configured to determine to-be-displayed text based on the target text; a generating module configured to generate a first image based on the target text; a synthesizing module configured to synthesize the to-be-displayed text and the first image to obtain a target image, the target image comprising the to-be-displayed text.
9. An electronic device, comprising: The electronic device comprises: one or more processors; a storage device configured to store one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-7.
10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The program is executed by the processor to implement the method according to any one of claims 1-7.