Image Generation Method, Device, Storage Medium, and Program Product
By displaying visual identification of text and non-text media in the display interface and generating metadata, the problem of user interaction mode solidification in the prior art is solved, and more flexible and accurate image generation is achieved.
Patent Information
- Application Number
- CN202510429901.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-08
AI Technical Summary
In the existing image generation technology, users need to enter text descriptions and upload media files in preset independent areas, which lacks flexibility and freedom, resulting in the semantic correlation of multimodal elements being separated, making it difficult to accurately express the user's complex intentions and reduce the accuracy of image generation.
Displays the visual identity of the text content input by the user and non-text media in the display interface, allowing the user to mix and arrange text and visual identity, generate metadata bound to the target visual identity, and use the multimedia input content and metadata to control the image generation model to generate images.
It improves the user's sense of participation and flexibility, enhances the richness and hierarchy of multimedia input content, improves the accuracy of image generation, and ensures that the generated images meet the user's image generation needs.
Smart Images

Figure CN119941910B_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of this specification relate to the field of image generation technologies, and in particular, to an image generation method, an electronic device, a computer-readable storage medium, and a computer program product. Background Art
[0002] Currently, deep learning-based image generation technologies have been widely applied in fields such as art creation, advertising design, and virtual scene construction. In a related image generation method, an image generation model can be controlled to generate an image through a natural language description input by a user and other media files uploaded by the user.
[0003] However, this image generation method has the defect of a fixed interaction mode. The user needs to submit a text description and upload other media files separately in a preset independent input area, lacking freedom and flexibility. Moreover, this interaction mode splits the semantic relevance of multimodal elements (text, image, video, etc.), making it difficult to accurately express the complex intentions of the user and reducing the accuracy of image generation. Summary of the Invention
[0004] In view of this, one or more embodiments of this specification provide an image generation method, an electronic device, a computer-readable storage medium, and a computer program product.
[0005] To achieve the above object, one or more embodiments of this specification provide the following technical solutions:
[0006] According to a first aspect of one or more embodiments of this specification, an image generation method is proposed, including:
[0007] Displaying the text content input by the user on a display interface, and displaying, on the display interface, a visualization identifier corresponding to at least one non-text medium uploaded by the user;
[0008] If a mixing event for at least a part of the text content and any one of the visualization identifiers is monitored, mixing at least a part of the text content and the target visualization identifier indicated by the mixing event to form multimedia input content, and generating metadata bound to the target visualization identifier, where the metadata is at least used to record the target position of the target visualization identifier in the multimedia input content, the storage path and type of the non-text medium indicated by the target visualization identifier;
[0009] If an image generation instruction is received, using the multimedia input content and the metadata bound to the target visualization identifier to control a trained image generation model to generate an image.
[0010] According to a second aspect of the embodiments of the present specification, there is provided an electronic device, including:
[0011] A processor;
[0012] A memory for storing instructions executable by the processor;
[0013] Wherein, when the processor executes the executable instructions, it is used to implement the method described in the first aspect.
[0014] According to a third aspect of the embodiments of the present specification, there is provided a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps of the method described in the first aspect.
[0015] According to a fourth aspect of the embodiments of the present specification, there is provided a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the steps of the method described in the first aspect.
[0016] The technical solutions provided by the embodiments of the present specification may include the following beneficial effects:
[0017] In the embodiments of the present specification, by displaying the text content input by the user and the visual identifier corresponding to the non-text media in the display interface, the user can intuitively see the text content and the non-text media, and support the mixed arrangement of at least part of the text of the text content and any visual identifier. This interaction method can improve the user's sense of participation and flexibility, making the formed multimedia input content richer and more hierarchical, helping the user to accurately express their image generation intention and improving the expression effect. And the device automatically records the storage path, type, and target position of the non-text media indicated by the target visual identifier to be mixed and arranged through metadata. When receiving an image generation instruction, it uses the multimedia input content and metadata to control the image generation model to generate an image. The content recorded by the metadata can reflect the semantic relevance between the text content and the non-text media, so as to help the image generation model accurately analyze the combination logic of the multimedia input content, making the generated image further meet the user's image generation needs and improving the accuracy of image generation.
[0018] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present specification. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 is a schematic diagram of the architecture of an image generation service system provided by an exemplary embodiment.
[0020] Figure 2 is a flowchart of an image generation method provided by an exemplary embodiment.
[0021] Figure 3 It is a schematic diagram showing text content on a display interface provided by an exemplary embodiment.
[0022] Figure 4 It is a schematic diagram showing a visual identifier displayed on the left side of a display interface provided by an exemplary embodiment.
[0023] Figure 5 It is an interaction schematic diagram related to selecting non-text media provided by an exemplary embodiment.
[0024] Figure 6A It is an interaction schematic diagram of inserting a visual identifier of non-text media into text content provided by an exemplary embodiment.
[0025] Figure 6B It is another interaction schematic diagram of inserting a visual identifier of non-text media into text content provided by an exemplary embodiment.
[0026] Figure 7 It is a schematic diagram of the finally generated image provided by an exemplary embodiment.
[0027] Figure 8 It is a schematic diagram of the structure of an electronic device provided by an exemplary embodiment. Detailed implementation manners
[0028] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all the implementation manners consistent with one or more embodiments of this specification. On the contrary, they are only examples of the devices and methods consistent with some aspects of one or more embodiments of this specification as detailed in the appended claims.
[0029] It should be noted that: In other embodiments, the steps of the corresponding methods are not necessarily executed in the order shown and described in this specification. In some other embodiments, the steps included in the method may be more or less than those described in this specification. In addition, a single step described in this specification may be decomposed into multiple steps for description in other embodiments; and multiple steps described in this specification may also be combined into a single step for description in other embodiments.
[0030] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this specification are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards in relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or reject.
[0031] Based on the problems in the related art, the embodiments of this specification provide an image generation method. A user can freely combine the visual identifiers of non-text media (such as images, audio, etc.) with text content on the display interface, and can mix and arrange at least part of the text content and any visual identifier, so as to form a multimedia input content that can accurately express the complex intentions of the user, providing a more free and flexible interaction method; and in order to support this interaction method, metadata bound to the target visual identifier to be mixed and arranged can be generated, and then the multimedia input content and the metadata are used to control the trained image generation model to generate an image. The content recorded in the metadata can reflect the semantic relevance between the text content and the non-text media, thereby helping the image generation model to accurately analyze the combination logic of the multimedia input content, so that the generated image can further meet the user's image generation requirements and improve the accuracy of image generation.
[0032] Figure 1 is a schematic structural diagram of an image generation service system provided by an exemplary embodiment. As Figure 1 shown, the system may include a server 11, a network 12, and several user terminals, such as a PC (Personal Computer) 13, a mobile phone 14, etc.
[0033] The server 11 may be a physical server including an independent host, or the server 11 may be a virtual server hosted by a host cluster. During operation, the server 11 may run the server-side program of the image generation service to implement the corresponding image generation service platform.
[0034] PC 13 and mobile phone 14 are only some types of user terminals that users can use. In fact, users can obviously also use user terminals of the following types: tablet devices, laptop computers, personal digital assistants (PDAs), wearable devices (such as smart glasses, smart watches, etc.). One or more embodiments of this specification do not limit this. During operation, the user terminal can run a program on the client side of the image generation service and can be implemented as the client of the image generation service. Among them, the application program of the client of the above image generation service can be started and run on the user terminal. The program on the client side can be a native application installed on the user terminal, or the program on the client side can be a mini program, a fast application, or other similar forms. Of course, when using web technologies such as HTML5 or similar, relevant functions can be implemented through the page displayed by the browser. The browser here can be an independent browser application or a browser module embedded in some applications.
[0035] For the network 12 for interaction between user terminals such as PC 13 and mobile phone 14 and the server 11, it can be specifically selected to use a wired or wireless network to implement communication based on the communication methods supported by the corresponding user terminals. This specification does not limit this. For example, if PC 13 supports both wired and wireless communication, then wired or wireless network can be used to implement communication according to needs, while mobile phone 14 usually only supports wireless communication, then wireless network can be used to implement communication.
[0036] Among them, the image generation method provided in the embodiments of this specification can be executed by the user terminal, or by the server 11, or a part of the image generation method can be executed by the user terminal and another part by the server 11. This embodiment does not make any restrictions on this.
[0037] Exemplarily, if the image generation method provided in the embodiments of this specification is executed by the user terminal. In one case, if the image generation model and other related models are deployed on the user terminal, the user terminal does not need to interact with the server during the execution of the image generation method provided in the embodiments of this specification. In another case, if the image generation model and other related models are deployed in the server 11, the user terminal needs to send the relevant input data of the model (such as multimedia input content and metadata bound to the target visualization identifier to be mixed) to the server 11 during the execution of the image generation method provided in the embodiments of this specification, so that the server 11 returns the output result of the model.
[0038] Exemplarily, if the image generation method provided in the embodiments of this specification is executed by a server, and the image generation model and other related models are deployed in server 11, then server 11 can control the user terminal to display text content, visual identifiers, multimedia input content, etc. on its display interface, and perform other operations.
[0039] Alternatively, a part of the image generation method provided in the embodiments of this specification is executed by the user terminal and another part is executed by server 11. For example, the user terminal can display text content, visual identifiers, and multimedia input content on its display interface based on user operations, and can generate metadata bound to the target visual identifier to be mixed. Finally, in response to receiving an image generation instruction, the multimedia input content and the metadata bound to the target visual identifier can be sent to server 11, so that server 11 can use the multimedia input content and the metadata bound to the target visual identifier to control the trained image generation model to generate an image.
[0040] In some embodiments, please refer to Figure 2 , which shows a schematic flowchart of an image generation method. Applying an electronic device (such as the above-mentioned user terminal, server, or other devices), the method includes:
[0041] In S201, display the text content input by the user on the display interface, and display the visual identifiers corresponding to at least one non-text medium uploaded by the user on the display interface.
[0042] Exemplarily, the user can input text content according to actual needs. The text content can describe the user's requirements for generating an image. Then, the electronic device can display the text content input by the user on its display interface. For example, please refer to Figure 3 , which shows the following text content displayed on the display interface: "Generate a picture according to the content of the book, 1:1, with the style referring to this picture."
[0043] It can be understood that the input method of the text content can be keyboard input (including but not limited to physical keyboards, virtual keyboards, etc.), voice input (voice-to-text), handwriting input (such as writing with a finger or a stylus on a touch screen), or scanning input (such as taking a picture of a paper document through a scanner, a smartphone camera, etc., and then using optical character recognition technology to recognize and convert the text in the picture into editable text), etc., but not limited to this.
[0044] Exemplarily, the user can also upload non-text media according to actual needs. Different types of non-text media include at least one of the following: document files, images, audio, video, and image generation models; but not limited to this.
[0045] Document files usually contain not only plain text but also other types of data. For example, document file formats include: (1) PDF files, which contain text content as well as non-text data such as images, tables, and charts. (2) Spreadsheets (such as Excel files), which include information such as charts, formulas, and graphics in addition to text data. (3) PowerPoint presentations, which contain elements such as images, charts, and animations and are usually used together with text to convey information.
[0046] Images are a type of medium that conveys information through static visual representations. Audio is a type of medium that transmits information through sound. Video combines images and audio and is a multimedia form used to convey information through continuous images and sounds. Image generation models refer to technologies that generate images through artificial intelligence (especially deep learning techniques). These models usually create new images by learning the characteristics of a large amount of image data.
[0047] Based on the user's upload operation, the electronic device can display visual identifiers corresponding to at least one non-text medium uploaded by the user in its display interface. These visual identifiers represent the non-text medium and can be in visual forms such as icons, thumbnails, etc., to help the user better understand and operate the non-text medium they uploaded. The display styles of the visual identifiers corresponding to different types of non-text media are different, making it convenient for the user to distinguish. The visual identifiers corresponding to at least one non-text medium uploaded by the user are displayed in a floating manner in the display interface, but are not limited to this. For example, please refer to Figure 4 which shows the visual identifier corresponding to the picture 1 uploaded by the user and the visual identifier corresponding to the PDF file uploaded by the user floating and displayed on the left side of the display interface.
[0048] By intuitively displaying the text content and the visual identifiers of non-text media in the display interface, users can conveniently view and manage the content of multiple media forms they input, enhancing the friendliness and interactivity of the interface operation. Users do not need to switch between different input boxes or interfaces and can directly process text content and non-text media in one interface, which can reduce the number of operation steps and improve efficiency.
[0049] In a possible implementation, a non-text media upload control can be displayed in the display interface. If the electronic device receives a trigger operation from the user on the non-text media upload control displayed in the display interface, it can display multiple candidate non-text media included in the non-text media library in the display interface. Subsequently, in response to receiving a selection operation from the user on at least one of the multiple candidate non-text media, a visualization identifier corresponding to the selected non-text media is displayed in the display interface, which helps the user clearly understand the selected media content. This enables the user to quickly identify the selected content when uploading multiple non-text media, avoiding confusion and errors. Different types of non-text media correspond to different non-text media upload controls, facilitating the user to upload various types of non-text media in a targeted manner. In this embodiment, by providing a dedicated non-text media upload control in the display interface, the user can upload different types of non-text media (such as pictures, audio, videos, etc.) more intuitively and conveniently. The user can quickly access multiple candidate non-text media through a simple trigger operation, reducing complex operation steps and improving usage efficiency.
[0050] For example, please refer to Figure 5 , after the user triggers the "Album" ( Figure 5 the trigger process is not shown), the electronic device displays multiple candidate images in the album in the display interface. The user selects Picture 1. Then, the user triggers the "File" ( Figure 5 the trigger process is not shown), and the electronic device displays multiple candidate document files in the "File" in the display interface. The user selects one of the PDF files. Subsequently, the electronic device displays the visualization identifier corresponding to Picture 1 and the visualization identifier corresponding to the PDF file on the left side of the display interface. It can be understood that Figure 5 the "Album" (i.e., the image upload control), "File" (i.e., the document file upload control), etc. shown are only for illustrative purposes and do not constitute a limitation on the non-text media upload control. The setting of the non-text media upload control and the related upload mechanism can be specifically set according to the actual application scenario.
[0051] It can be understood that this embodiment of the specification does not impose any restrictions on the sequence between the input of text content and the upload of non-text media. The text content can be input first, and then the non-text media can be uploaded; or the non-text media can be uploaded first, and then the text content can be input; it can be customized based on user requirements.
[0052] In S202, if a mixing event of at least part of the text content and any visual identifier is monitored, at least part of the text content and the target visual identifier indicated by the mixing event are mixed and arranged to form multimedia input content, and metadata bound to the target visual identifier is generated. The metadata is at least used to record the target position of the target visual identifier in the multimedia input content, the storage path and type of the non-text media indicated by the target visual identifier.
[0053] Among them, the mixing event includes at least one of the following: an event of dragging any visual identifier to the front or back position of any character in the text content, and an event of inserting at least part of the text content before or after any visual identifier.
[0054] In the event of dragging any visual identifier to the front or back position of any character in the text content, the user directly drags the visual identifier (such as an icon or thumbnail) representing the non-text media to a specific position in the text content through a drag-and-drop operation. Specifically, the user can place the identifier in front of or behind any character in the text, indicating that the user hopes to integrate the non-text media element with the text content to form multimedia input content.
[0055] In the event of inserting at least part of the text content before or after any visual identifier, the user inserts at least part of the text content before or after the originally existing visual identifier. That is to say, instead of moving the visual identifier, the selected at least part of the text content is embedded before or after the visual identifier, or text is input before or after the target visual identifier, indicating that the user hopes to mix the text content with the non-text media element visually and semantically, so that the two are more closely combined.
[0056] In short, both of these mixing events aim to achieve the integration of text and non-text media through different user interaction methods, so as to form a rich multimedia input content. The electronic device can accurately identify the position and attributes of the non-text media indicated by the visual identifier according to the generated metadata, which is convenient for subsequent processing.
[0057] Exemplarily, the user can, according to actual needs, drag any visual identifier to the position before any character in the text content, and / or drag any visual identifier to the position after any character in the text content. Through the drag-and-drop operation, the user can more intuitively control the relationship between the text content and the non-text media, providing higher degrees of freedom and flexibility and supporting complex input methods. The user can accurately position the visual identifiers corresponding to different types of non-text media at specific positions in the text content according to their own needs, which helps the user accurately convey their intentions.
[0058] When the electronic device monitors that the user drags and drops a visual identifier to a specific position in the text content, it inserts the target visual identifier into the corresponding position of the text content, thus forming multimedia input content. Moreover, the electronic device generates metadata bound to the target visual identifier. The metadata is at least used to record the storage path, type of the non-text media indicated by the target visual identifier, and the target position of the target visual identifier in the multimedia input content. The storage path records the actual storage position of the non-text media indicated by the target visual identifier, ensuring that the non-text media can be found and accessed during subsequent image generation. The type helps the electronic device perform different processing and rendering according to the characteristics of the media type. The target position records the insertion position of the target visual identifier in the text content, indicating which part of the text the target visual identifier should be displayed in, and can accurately reflect the semantic relevance between the text content and the non-text media. In short, by providing accurate storage paths, type information, and target positions, the metadata can help the image generation model correctly understand the multimedia input content, ensuring the accuracy of the generated image and meeting the user's needs. Of course, the metadata can further record other information, such as additional descriptions of the non-text media by the user, size or resolution information of the non-text media, upload time of the non-text media, source or copyright information of the non-text media, etc. This embodiment does not impose any restrictions on this.
[0059] In one example, please refer to Figure 6A and Figure 6B , Figure 6A which shows the process of the user performing a drag-and-drop operation on the visual identifier corresponding to the PDF file. Figure 6B which shows the process of the user performing a drag-and-drop operation on the visual identifier corresponding to Picture 1. Finally, the visual identifier corresponding to the PDF file is dragged and dropped behind "this" in the text content, and the visual identifier corresponding to Picture 1 is dragged and dropped behind "piece" in the text content, forming the multimedia input content: "Generate a picture according to the content of the book [PDF icon], one-to-one, with the style referring to this [Picture 1 icon] picture."
[0060] In S203, if an image generation instruction is received, the trained image generation model is controlled to generate an image by using the multimedia input content and the metadata bound to the target visual identifier.
[0061] In this step, when the electronic device receives an image generation instruction, the electronic device passes the multimedia input content and the metadata bound to the visual identifier to the trained image generation model, and the image generation model generates an image that meets the user's needs based on these inputs. By combining the multimedia input content with the metadata, the image generation model can better understand the semantic associations between text and non-text media, thereby generating images that better meet the user's needs. Since users can freely combine text content and non-text media, the generated images can be more personalized and reflect the user's unique intentions.
[0062] Continuing with the above example, based on Figure 6A and Figure 6B the multimedia input content and related metadata shown, a target image can be generated, as shown in Figure 7 The target image is displayed on the display interface, and the style of the target image is the same as that of Picture 1.
[0063] In some embodiments, the electronic device can generate a multimodal hybrid prompt based on the multimedia input content and the metadata bound to the target visual identifier, and then control the trained image generation model to generate an image according to the multimodal hybrid prompt.
[0064] The following is an exemplary description of the generation process of the multimodal hybrid prompt: The electronic device can input the multimedia input content and the metadata bound to the target visual identifier into a pre-trained language model, so that the pre-trained language model can perform intention recognition on the multimedia input content and the metadata bound to the target visual identifier, thereby determining at least one constraint condition for the image to be generated and the processing content corresponding to each constraint condition, and generating a processing task based on the processing content corresponding to each constraint condition. Among them, the processing content corresponding to each constraint condition is at least part of the multimedia input content and the metadata bound to the target visual identifier. In this embodiment, through the intention recognition of the pre-trained language model, accurate intentions can be extracted from various input forms (text, image, metadata, etc.), ensuring that the generated images meet the constraint conditions specified by the user and avoiding inaccurate or deviated generation results.
[0065] Among them, the pre-trained language model, such as the Large Language Model (LLM), refers to an artificial intelligence model based on deep learning technology, especially trained using a large corpus, aiming to understand and generate text similar to human language, and has powerful natural language understanding and generation capabilities. The goal of the LLM is to achieve various applications through natural language processing capabilities, such as text generation, translation, summarization, question answering, and dialogue systems, so as to help improve the efficiency of human-computer interaction and the degree of automation.
[0066] After obtaining the processing tasks corresponding to the respective constraint conditions output by the pre-trained language model, the electronic device may execute the processing tasks corresponding to the respective constraint conditions to obtain processing results. Among them, if the types of media elements included in the processing content corresponding to the respective constraint conditions are different, the processing tasks corresponding to the constraint conditions are also different. For example, the processing tasks include but are not limited to at least one of the following: key information extraction tasks for text content and / or document files, content analysis and / or style extraction tasks for images, speech recognition and analysis tasks for audio, deployment tasks for the first image generation model uploaded by the user, and so on.
[0067] Exemplarily, in the process of the electronic device executing the processing tasks corresponding to the respective constraint conditions, there are the following several situations:
[0068] In the first situation, if the processing content corresponding to the constraint condition includes at least one of text content and document files, the electronic device may use the trained natural language processing model to perform semantic analysis on at least one of the text content and document files to extract key information. In this embodiment, no restrictions are imposed on the natural language processing model, and the trained natural language processing model in the related technology can be directly applied.
[0069] In the second situation, if the processing content corresponding to the constraint condition includes an image, the electronic device may use the trained image processing model to perform at least one of content analysis and style extraction on the image to obtain an image processing result. In this embodiment, no restrictions are imposed on the image processing model, and the trained image processing model in the related technology can be directly applied.
[0070] In the third situation, if the processing content corresponding to the constraint condition includes audio, the electronic device may use the trained speech recognition model to perform speech recognition and analysis on the audio to obtain an audio processing result. In this embodiment, no restrictions are imposed on the speech recognition model, and the trained speech recognition model in the related technology can be directly applied.
[0071] In the fourth situation, if the processing content corresponding to the constraint condition includes the first image generation model uploaded by the user, the electronic device may deploy the first image generation model in the image generation model deployment environment.
[0072] The electronic device can process various types of input content (text, images, audio, etc.) and perform corresponding processing tasks according to different types of content, so as to more comprehensively understand the user's needs and generate results that meet the requirements. For example, text content can be used to extract key information, image content can be used for style extraction, and audio content can be used for speech recognition. The combination of these tasks can improve the processing ability of multi-modal data. By using the trained models (such as natural language processing models, image processing models, speech recognition models, etc.), the electronic device can efficiently perform semantic analysis, content analysis, and style extraction on the input content. These models have been trained with a large amount of data, have high precision and reliability, can automatically perform complex analysis tasks, avoid manual intervention, and improve the processing efficiency and accuracy.
[0073] Finally, after obtaining the processing results corresponding to each constraint condition, the electronic device can generate a multi-modal hybrid prompt based on the multimedia input content and the processing results corresponding to each constraint condition. Based on the multimedia input content and the processing results corresponding to each constraint condition, the generated multi-modal hybrid prompt not only integrates text information but also includes the characteristics of other media information such as images and audio. In this way, the image generation model can receive a multi-level and all-round description and generate an image according to this information. The multi-modal hybrid prompt can include details of text descriptions, requirements for image styles, and prompts for audio emotions, etc.
[0074] Among them, the multimedia input content is the main source for the user to convey intentions. The text content describes the basic requirements or scenarios for image generation, while non-text media (such as images, audio) provide direct visual or sound information about image details, styles, emotions, etc. The multimedia input content provides a preliminary description framework and specific details for image generation, ensuring that the basic direction and content of image generation meet the user's expectations.
[0075] The processing results corresponding to each constraint condition (such as text analysis results, image content analysis, speech recognition results, etc.) can more precisely constrain and guide the image generation process. These processing results may contain key information, style requirements, semantic relationships, etc., further refining the user's needs. For example, the text analysis result may extract specific elements (such as people, scenery, actions, etc.) that should be included in the image, while the image processing result provides the style or visual effect requirements of the image (such as realism, abstract art, etc.). These processing results help the model understand the semantic relationship between text and non-text content, so as to generate an image that more conforms to the user's intention. For example, the speech recognition result can help the system understand the semantic information in the audio and integrate this information into the image generation task to generate an image with more emotion and context.
[0076] In a possible implementation, the electronic device can directly splice the multimedia input content and the processing results corresponding to each constraint condition to obtain a multi-modal hybrid prompt. In this embodiment, the splicing process is relatively straightforward. It only needs to combine the multimedia input content and the processing results according to certain rules, without involving excessive calculations, and can quickly generate a multi-modal hybrid prompt, which is suitable for scenarios with high efficiency requirements.
[0077] In another possible implementation, to further improve the accuracy of the multi-modal hybrid prompt, the electronic device can input the multimedia input content and the processing results corresponding to each constraint condition into a pre-trained language model, so that the pre-trained language model can organize the multimedia input content and the processing results corresponding to each constraint condition to obtain a multi-modal hybrid prompt. In this way, the pre-trained language model will organize and restructure the input information to generate a multi-modal hybrid prompt that is more in line with language logic, semantically accurate and clear. The pre-trained language model has powerful semantic analysis capabilities, can understand the relationships between different modalities and optimize them. Through the organization of the pre-trained language model, a multi-modal hybrid prompt that is more in line with the context logic can be generated to ensure that each input information is reasonably interpreted and integrated.
[0078] For example, the user hopes to generate a cartoon-style seaside sunset image. The user embeds non-text media (such as a wave texture map, a sailboat map, an audio file, a custom image generation model) into a specified position in the text content by dragging and dropping the visualization identifier corresponding to the non-text media, forming the multimedia input content: "A peaceful [wave icon] seaside sunset, with a [sailboat icon] sailboat in the distance, and the sound of [audio icon] ocean waves playing in the background, using [model icon] my cartoon rendering model."
[0079] The metadata bound to the visualization identifier includes: (1) Embedded rich media metadata: Wave icon: Located after the 5th character of the text, type is image, path is / media / waves_texture.png. (2) Sailboat icon: Located after the 15th character of the text, type is image, path is / media / sailboat.png. (3) Audio icon: Located after the 25th character of the text, type is audio, path is / audio / ocean_waves.mp3. (4) Model icon: Located after the 35th character of the text, type is a custom image generation model, path is / models / cartoon_style.ckpt.
[0080] The electronic device inputs the multimedia input content and the metadata bound to the target visual identifier into the pre-trained language model. The pre-trained language model combines the anchor words in the text content (such as "seaside", "sailboat", "sound of ocean waves") with the type of the embedded non-text media to identify the following intention: "Generate a cartoon-style seaside sunset scene, which needs to integrate dynamic ocean wave textures, sailboats, and make the rhythm of the ocean wave fluctuations match the background audio." Then, based on this intention, the following constraints, processing content, and processing tasks are determined: (1) Ocean wave texture; Processing content: The image indicated by the ocean wave icon; Processing task: Act on the "seaside" area, which needs to be dynamic and the frequency is synchronized with the audio. (2) Sailboat; Processing content: The image indicated by the sailboat icon; Processing task: Need to perform content analysis on the image indicated by the sailboat icon. (3) Audio; Processing content: The audio indicated by the audio icon; Processing task: Extract the voiceprint spectrum to drive the rhythm of the ocean wave fluctuations. (4) Custom image generation model, Processing content: The image generation model indicated by the model icon; Processing task: Deploy the custom image generation model.
[0081] The electronic device can execute the processing tasks corresponding to each constraint to obtain the processing results, and finally integrate the multimedia input content and the processing results corresponding to each constraint to obtain a multi-modal mixed prompt, such as "(1) Cartoon-style seaside sunset scene (apply model cartoon_style.ckpt); (2) Sea area: Use dynamic wave texture, texture path waves_texture.png; (3) Distant view object: Cartoon sailboat; (4) Global rhythm: The ocean wave fluctuations are synchronized with the soothing voiceprint of the audio ocean_waves.mp3."
[0082] Exemplarily, multiple image generation models can be pre-deployed in the image generation model deployment environment. When the user does not upload a custom image generation model, the image generation models in the image generation model deployment environment can be directly used. Then, in the process of generating an image according to the multi-modal mixed prompt, there are the following situations:
[0083] In the first situation, if the processing task includes the task of deploying the first image generation model uploaded by the user, and the electronic device has completed the deployment process of the first image generation model, the electronic device can input the multi-modal mixed prompt into the deployed first image generation model, so that the deployed first image generation model generates an image according to the multi-modal mixed prompt. The custom image generation model uploaded by the user can meet specific requirements or preferences. By using the custom image generation model uploaded by the user, images that meet the unique requirements of the user or industry needs can be generated. This method provides great flexibility and freedom for the user and can generate images according to specific styles, content, or other customization requirements.
[0084] In the second case, if the text content includes a second image generation model specified by the user and the second image generation model is deployed in the image generation model deployment environment, the electronic device can input the multi-modal mixed prompt into the second image generation model, so that the second image generation model generates an image according to the multi-modal mixed prompt. In this case, the user explicitly specifies the second image generation model, usually to utilize the specific capabilities of the model (such as processing specific types of content, styles, languages, etc.). By using the image generation model specified by the user, the electronic device can generate images that better meet the user's requirements and enhance the user experience. The user can choose their preferred generation model instead of relying passively on the default model, which provides the user with a variety of image generation model options to meet a wider range of application requirements.
[0085] In the third case, if the user does not upload the first image generation model, or the second image generation model specified by the user does not exist in the image generation model deployment environment, or the user does not specify the second image generation model, the electronic device can input the multi-modal mixed prompt into the default third image generation model deployed in the image generation model deployment environment, so that the third image generation model generates an image according to the multi-modal mixed prompt. In this case, even if the user does not upload a custom image generation model or does not specify the second image generation model, it can automatically fallback to the default third image generation model to ensure that the image generation process will not be interrupted.
[0086] It can be understood that the "first image generation model", "second image generation model", and "third image generation model" mentioned above are used to distinguish the model usage in different scenarios, and "first, second, third" themselves do not have other special meanings.
[0087] In some embodiments, in addition to the multi-modal prompt, preset parameters can also be added. The preset parameters include at least one of the following: parameters related to the image, parameters related to the image generation model. Parameters related to the image are, for example, image resolution, image ratio, hue and contrast, light and shadow, etc. These parameters directly affect the visual characteristics of the generated image, such as appearance, style, content, etc. Parameters related to the image generation model can be parameters used to control the internal algorithm of the image generation model. For example, (1) the classifier free guidance weight, which is used to adjust the compliance of the image generation model to the input prompt. The higher the weight, the more strictly the generated result conforms to the prompt, but the diversity may be reduced; (2) the number of iterations (or sampling steps), which is used to control the number of iterations in the image generation process. More iterations may result in a more refined image, but it requires more computing time; but not limited to this.
[0088] The electronic device inputs the multi-modal mixed prompt words and preset parameters into the trained image generation model, so that the trained image generation model (such as any one of the above-mentioned first image generation model, second image generation model, and third image generation model) generates an image based on the multi-modal mixed prompt words and preset parameters. By presetting these parameters, the electronic device can reduce manual operations and adjustments during the generation process and provide more automated, fast, and accurate generation results. For example, after setting the resolution and style of the generated image, the electronic device will automatically control the image generation model to generate an image according to these requirements without the user manually adjusting each time.
[0089] Exemplarily, the electronic device can check whether the multi-modal mixed prompt words already contain key information related to the image, such as resolution, style, lighting, etc. If this information has been clearly specified, there is no need to repeat setting these contents through additional preset parameters. For example, if the multi-modal mixed prompt words already clearly require "generate an oil painting-style image with a resolution of 1920x1080", the system can directly utilize this information without the user providing additional "style" and "resolution" parameters again.
[0090] Similarly, the electronic device will also detect whether the multi-modal mixed prompt words already involve parameters related to the image generation model, such as the number of iterations, etc. If the multi-modal mixed prompt words already contain parameters related to the image generation model, it means that the model can automatically execute the image generation process according to this information without repeated addition.
[0091] If the multi-modal mixed prompt words do not contain the necessary preset parameters, the electronic device can automatically supplement these missing parameters. For example, if the user does not specify the image resolution or style in the multi-modal mixed prompt words, the electronic device can automatically fill in these parameters according to the default settings or according to the task requirements to ensure the smooth progress of the image generation task. This method enables the electronic device to have sufficient flexibility when processing different types of inputs, while avoiding the trouble for the user to provide all details every time they input. Through intelligent judgment and automatic supplementation, the user operation is simplified and the generation efficiency is improved.
[0092] For example, assume the multi-modal mixed prompt words are: "generate an image of a city landscape on a sunny day". At this time, the electronic device will check the multi-modal mixed prompt words, determine that the resolution or style is not mentioned in the multi-modal mixed prompt words, and the electronic device automatically uses the default resolution (such as 1024x1024) and style (such as realistic style); and determine that the relevant image generation model is not mentioned in the multi-modal mixed prompt words, and the electronic device then uses the default third image generation model.
[0093] Suppose the multi-modal mixed prompt is: "Generate an image of a city landscape on a sunny day with a resolution of 1920x1080 and in the watercolor painting style". After the electronic device checks the multi-modal mixed prompt and determines that it contains the number of acceptances related to the image, it directly uses the information in the multi-modal mixed prompt to generate the image.
[0094] It can be understood that the relevant models mentioned in this specification, such as image generation models, pre-trained language models, natural language processing models, image processing models, and speech recognition models, etc., can be trained based on the training methods in related technologies. This embodiment does not impose any restrictions on the model training process; and the model architectures of the relevant models mentioned in this specification can directly adopt the model architectures mentioned in related technologies. For example, the image generation model can be a diffusion model (such as Stable Diffusion), a generative adversarial network (GAN), a variational autoencoder (VAE), an autoregressive model, a Transformer architecture, etc. This embodiment does not impose any restrictions on this.
[0095] The various technical features in the above embodiments can be combined arbitrarily as long as there is no conflict or contradiction between the features. However, due to space limitations, they are not described one by one. Therefore, any combination of the various technical features in the above embodiments also belongs to the scope disclosed in this specification.
[0096] In some embodiments, this embodiment of the specification also provides an electronic device, including: a processor; a memory for storing processor-executable instructions; wherein, the processor realizes the method described in any one of the above by running the executable instructions.
[0097] Figure 8 It is a schematic structural diagram of a device provided by an exemplary embodiment. Please refer to Figure 8 , at the hardware level, this device includes a processor 802, an internal bus 804, a network interface 806, a memory 808, and a non-volatile memory 810. Of course, there may also be other hardware required for other functions. One or more embodiments of this specification can be implemented in a software manner. For example, the processor 802 reads the corresponding computer program from the non-volatile memory 810 into the memory 808 and then runs it. Of course, in addition to the software implementation manner, one or more embodiments of this specification do not exclude other implementation manners, such as a logic device or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, and can also be hardware or a logic device.
[0098] Exemplarily, the image generation device can be applied to a device as shown in Figure 8 to implement the technical solution of this specification. Among them, the image generation device may include:
[0099] A display module, configured to display the text content input by the user on the display interface, and display a visual identifier corresponding to at least one non-text medium uploaded by the user on the display interface;
[0100] A mixing event processing module, configured to, if a mixing event for at least part of the text content and any visual identifier is detected, mix and arrange at least part of the text content and the target visual identifier indicated by the mixing event to form a multimedia input content, and generate metadata bound to the target visual identifier, where the metadata is at least used to record the target position of the target visual identifier in the multimedia input content, the storage path and type of the non-text medium indicated by the target visual identifier.
[0101] An image generation module, configured to, if an image generation instruction is received, use the multimedia input content and the metadata bound to the target visual identifier to control a trained image generation model to generate an image.
[0102] In one implementation, the mixing event includes at least one of the following: an event of dragging any of the visual identifiers to a position before or after any character in the text content, and an event of inserting at least part of the text content before or after any of the visual identifiers.
[0103] In one implementation, the visual identifier corresponding to at least one non-text medium uploaded by the user is displayed in a floating manner on the display interface.
[0104] In one implementation, the display styles of the visual identifiers corresponding to different types of non-text media are different.
[0105] In one implementation, different types of non-text media include at least one of the following: document files, images, audio, video, and image generation models.
[0106] In one implementation, the display module is specifically configured to, if a trigger operation on a non-text media upload control displayed on the display interface is received by the user, display a plurality of candidate non-text media included in the non-text media library on the display interface; in response to receiving a selection operation by the user on at least one of the plurality of candidate non-text media, display a visual identifier corresponding to the selected non-text media on the display interface.
[0107] In one implementation, the image generation module is specifically configured to generate a multi-modal hybrid prompt based on the multimedia input content and the metadata bound to the target visualization identifier; and control the trained image generation model to generate an image according to the multi-modal hybrid prompt.
[0108] In one implementation, the image generation module is specifically configured to input the multimedia input content and the metadata bound to the target visualization identifier into a pre-trained language model, so that the pre-trained language model performs intent recognition on the multimedia input content and the metadata bound to the target visualization identifier, thereby determining at least one constraint condition for the image to be generated and the processing content corresponding to each constraint condition, and generating a processing task based on the processing content corresponding to each constraint condition; execute the processing tasks corresponding to each constraint condition to obtain a processing result; and generate a multi-modal hybrid prompt based on the multimedia input content and the processing results corresponding to each constraint condition.
[0109] In one implementation, the image generation module is specifically configured to, if the processing content corresponding to the constraint condition includes text content and / or a document file, use a trained natural language processing model to perform semantic analysis on the text content and / or the document file to extract key information; if the processing content corresponding to the constraint condition includes an image, use a trained image processing model to perform content analysis and / or style extraction on the image to obtain an image processing result; if the processing content corresponding to the constraint condition includes audio, use a trained speech recognition model to perform speech recognition and analysis on the audio to obtain an audio processing result; if the processing content corresponding to the constraint condition includes a first image generation model uploaded by a user, deploy the first image generation model in an image generation model deployment environment.
[0110] In one implementation, the image generation module is specifically configured to input the multimedia input content and the processing results corresponding to each constraint condition into a pre-trained language model, so that the pre-trained language model organizes the multimedia input content and the processing results corresponding to each constraint condition to obtain the multi-modal hybrid prompt.
[0111] In one implementation, the image generation module is specifically configured to, if the processing task includes a task of deploying a first image generation model uploaded by a user, input the multi-modal mixed prompt into the deployed first image generation model, so that the deployed first image generation model generates an image according to the multi-modal mixed prompt; if the text content includes a second image generation model specified by the user and the second image generation model is deployed in the image generation model deployment environment, input the multi-modal mixed prompt into the second image generation model, so that the second image generation model generates an image according to the multi-modal mixed prompt; if the user does not upload a first image generation model, or the second image generation model specified by the user does not exist in the image generation model deployment environment, or the user does not specify a second image generation model, input the multi-modal mixed prompt into a third image generation model deployed in the image generation model deployment environment, so that the third image generation model generates an image according to the multi-modal mixed prompt.
[0112] In one implementation, the image generation module is specifically configured to input the multi-modal mixed prompt and preset parameters into a trained image generation model, so that the trained image generation model generates an image based on the multi-modal mixed prompt and the preset parameters; wherein, the preset parameters include at least one of the following: parameters related to the image, parameters related to the image generation model.
[0113] The implementation processes of the functions and roles of each module in the above device are specifically described in the implementation processes of the corresponding steps in the above method, and will not be elaborated here.
[0114] Based on the same concept as the above method, this specification also provides an electronic device, including: a processor; a memory for storing executable instructions of the processor; wherein, the processor runs the executable instructions to implement the steps of the method as described in any one of the above embodiments.
[0115] Based on the same concept as the above method, this specification also provides a computer-readable storage medium, on which computer instructions are stored, and when the instructions are executed by a processor, the steps of the method as described in any one of the above embodiments are implemented.
[0116] A computer-readable medium includes both permanent and non-permanent, removable and non-removable media and can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, disk storage, quantum memory, graphene-based storage media, or other magnetic storage devices, or any other non-transitory media that can be used to store information that can be accessed by a computing device. As defined herein, a computer-readable medium does not include transitory computer-readable media, such as modulated data signals and carrier waves.
[0117] Based on the same concept as the above method, this specification also provides a computer program product, including a computer program / instructions, which when executed by a processor, implement the steps of the method described in any of the above embodiments.
[0118] The above are only the preferred embodiments of one or more embodiments of this specification and are not intended to limit one or more embodiments of this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of one or more embodiments of this specification shall be included within the scope of protection of one or more embodiments of this specification.
Claims
1. An image generation method, comprising: Displaying the text content input by the user on a display interface, and displaying, on the display interface, visual identifiers corresponding to at least one non-text medium uploaded by the user; If a mixing event for at least part of the text content and any one of the visual identifiers is monitored, mixing and arranging at least part of the text content and the target visual identifier indicated by the mixing event to form multimedia input content, and generating metadata bound to the target visual identifier, where the metadata is at least used to record the target position of the target visual identifier in the multimedia input content, the storage path and type of the non-text medium indicated by the target visual identifier; If an image generation instruction is received, using the multimedia input content and the metadata bound to the target visual identifier to control a trained image generation model to generate an image; Wherein, the using the multimedia input content and the metadata bound to the target visual identifier to control a trained image generation model to generate an image includes: generating a multimodal mixed prompt word based on the multimedia input content and the metadata bound to the target visual identifier; controlling the trained image generation model to generate an image according to the multimodal mixed prompt word.
2. The method according to claim 1, wherein the mixing event includes at least one of the following: an event of dragging any one of the visual identifiers to a position before or after any character in the text content, and an event of inserting at least part of the text content before or after any one of the visual identifiers; and / or The visual identifiers corresponding to at least one non-text medium uploaded by the user are displayed in a floating manner on the display interface.
3. The method according to claim 1, wherein the display styles of the visual identifiers corresponding to different types of non-text media are different; and / or, different types of non-text media include at least one of the following: document files, images, audio, video, and image generation models.
4. The method according to claim 1, wherein displaying, on the display interface, visual identifiers corresponding to at least one non-text medium uploaded by the user includes: If a trigger operation on a non-text medium upload control displayed on the display interface is received, displaying a plurality of candidate non-text media included in a non-text medium library on the display interface; In response to receiving a selection operation by the user on at least one of the plurality of candidate non-text media, displaying visual identifiers corresponding to the selected non-text media on the display interface.
5. The method according to claim 1, wherein the generating a multimodal mixed prompt word based on the multimedia input content and the metadata bound to the target visual identifier includes: Input the multimedia input content and the metadata bound to the target visualization identifier into a pre-trained language model, so that the pre-trained language model performs intent recognition on the multimedia input content and the metadata bound to the target visualization identifier, thereby determining at least one constraint condition for the image to be generated and the processing content corresponding to each of the constraint conditions, and generating a processing task based on the processing content corresponding to each of the constraint conditions; Execute the processing tasks corresponding to each of the constraint conditions to obtain processing results; Generate a multi-modal hybrid prompt based on the multimedia input content and the processing results corresponding to each of the constraint conditions.
6. The method according to claim 5, wherein the executing the processing tasks corresponding to each of the constraint conditions to obtain processing results includes: If the processing content corresponding to the constraint condition includes text content and / or a document file, use a trained natural language processing model to perform semantic analysis on the text content and / or the document file to extract key information; If the processing content corresponding to the constraint condition includes an image, use a trained image processing model to perform content analysis and / or style extraction on the image to obtain an image processing result; If the processing content corresponding to the constraint condition includes audio, use a trained speech recognition model to perform speech recognition and analysis on the audio to obtain an audio processing result; If the processing content corresponding to the constraint condition includes a first image generation model uploaded by a user, deploy the first image generation model in an image generation model deployment environment.
7. The method according to claim 5, wherein the generating a multi-modal hybrid prompt based on the multimedia input content and the processing results corresponding to each of the constraint conditions includes: Input the multimedia input content and the processing results corresponding to each of the constraint conditions into a pre-trained language model, so that the pre-trained language model organizes the multimedia input content and the processing results corresponding to each of the constraint conditions to obtain the multi-modal hybrid prompt.
8. The method according to claim 6, wherein the controlling the trained image generation model to generate an image according to the multi-modal hybrid prompt includes: If the processing task includes a task of deploying a first image generation model uploaded by a user, input the multi-modal hybrid prompt into the deployed first image generation model, so that the deployed first image generation model generates an image according to the multi-modal hybrid prompt; If the text content includes a second image generation model specified by a user and the second image generation model is deployed in the image generation model deployment environment, input the multi-modal hybrid prompt into the second image generation model, so that the second image generation model generates an image according to the multi-modal hybrid prompt; If the user does not upload the first image generation model, or the second image generation model specified by the user does not exist in the image generation model deployment environment, or the user does not specify the second image generation model, the multi-modal mixed prompt is input into the third image generation model deployed in the image generation model deployment environment, so that the third image generation model generates an image according to the multi-modal mixed prompt.
9. The method according to claim 1, wherein controlling the trained image generation model to generate an image according to the multi-modal mixed prompt comprises: Inputting the multi-modal mixed prompt and preset parameters into the trained image generation model, so that the trained image generation model generates an image based on the multi-modal mixed prompt and the preset parameters; wherein the preset parameters include at least one of the following: parameters related to the image, parameters related to the image generation model.
10. An electronic device, comprising: Processor; A memory for storing instructions executable by the processor; wherein the processor realizes the steps of the method according to any one of claims 1-9 by running the executable instructions.
11. A computer-readable storage medium, on which computer instructions are stored, and when the instructions are executed by a processor, the steps of the method according to any one of claims 1-9 are realized.
12. A computer program product, comprising a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1-9 are realized.
Citation Information
Patent Citations
Image generation method and device and storage medium
CN117475031A
Text-based image generation method and device, electronic equipment and storage medium
CN117493599A