Image generation method and device, storage medium and program product
By displaying visual identifiers of text content and non-text media in the display interface, and supporting users to mix arrangements to generate metadata bound to the target visual identifier, the problem of interaction mode solidification in the prior art is solved, and the accuracy and flexibility of image generation are improved.
Patent Information
- Application Number
- CN202510429901.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-08
AI Technical Summary
The existing image generation technology has solidification in the interactive mode and lacks flexibility and freedom, making it difficult for users to express complex intentions accurately and reduce the accuracy of image generation.
By displaying the text content input by the user and visual identification of non-text media in the display interface, and supporting the user to mix and arrange the text content with the visual identification to form multimedia input content. At the same time, metadata bound to the target visual identification is generated, and its position and attributes in the multimedia input content are recorded so as to subsequently control the image generation model to generate images.
It improves user's sense of interaction and flexibility, makes multimedia input content richer and more layered, helps users accurately express image generation intentions, and improves the accuracy of image generation.
Smart Images

Figure CN119941910A_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of the present specification relate to the field of image generation technology, and in particular, to an image generation method, an electronic device, a computer-readable storage medium, and a computer program product. Background Art
[0002] Currently, image generation technology based on deep learning has been widely used in the fields of art creation, advertising design, virtual scene construction, etc. In a related technology, an image generation method can control the image generation model to generate images through natural language descriptions input by users and other uploaded media files.
[0003] However, this image generation method has the defect of fixed interaction mode. Users need to submit text descriptions and upload other media files separately in preset independent input areas, which lacks freedom and flexibility. In addition, this interaction mode breaks the semantic relevance of multimodal elements (text, images, videos, etc.), making it difficult to accurately express the user's complex intentions and reducing the accuracy of image generation. Summary of the invention
[0004] In view of this, one or more embodiments of the present specification provide an image generating method, an electronic device, a computer-readable storage medium, and a computer program product.
[0005] To achieve the above objectives, one or more embodiments of this specification provide the following technical solutions: According to a first aspect of one or more embodiments of this specification, a method for generating an image is provided, comprising: Displaying text content input by the user on a display interface, and displaying a visual identifier corresponding to at least one non-text media uploaded by the user on the display interface; If a shuffling event for at least part of the text content and any of the visual identifiers is monitored, at least part of the text content and the target visual identifier indicated by the shuffling event are mixed and arranged to form multimedia input content, and metadata bound to the target visual identifier is generated, where the metadata is used to record at least a target position of the target visual identifier in the multimedia input content and a storage path and type of the non-text media indicated by the target visual identifier; If an image generation instruction is received, the multimedia input content and the metadata bound to the target visualization identifier are used to control the trained image generation model to generate an image.
[0006] According to a second aspect of the embodiments of this specification, an electronic device is provided, including: processor; a memory for storing processor-executable instructions; When the processor executes the executable instructions, it is used to implement the method described in the first aspect.
[0007] According to a third aspect of the embodiments of this specification, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the steps of the method described in the first aspect are implemented.
[0008] According to a fourth aspect of the embodiments of this specification, a computer program product is provided, comprising a computer program, which implements the steps of the method described in the first aspect when executed by a processor.
[0009] The technical solutions provided by the embodiments of this specification may have the following beneficial effects: In the embodiments of the present specification, by displaying the text content input by the user and the visual identifier corresponding to the non-text media in the display interface, the user can intuitively see the text content and the non-text media, and supports the mixed arrangement of at least part of the text content and any visual identifier. This interactive method can improve the user's sense of participation and flexibility, making the formed multimedia input content richer and more layered, helping the user to accurately express his image generation intention and improve the expression effect. And the device automatically records the storage path, type and target location of the non-text media indicated by the mixed target visual identifier through metadata. When receiving the image generation instruction, the multimedia input content and metadata are used to control the image generation model to generate the image. The content recorded by the metadata can reflect the semantic correlation between the text content and the non-text media, so as to help the image generation model accurately parse the combination logic of the multimedia input content, so that the generated image can further meet the user's image generation needs and improve the accuracy of image generation.
[0010] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present specification. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 It is a schematic diagram of the architecture of an image generation service system provided by an exemplary embodiment.
[0012] Figure 2 It is a flowchart of an image generating method provided by an exemplary embodiment.
[0013] Figure 3 It is a schematic diagram of displaying text content in a display interface provided by an exemplary embodiment.
[0014] Figure 4 It is a schematic diagram of displaying a visual identifier on the left side of a display interface provided by an exemplary embodiment.
[0015] Figure 5 It is a schematic diagram of an interaction involving selection of non-text media provided by an exemplary embodiment.
[0016] Fig. 6A An exemplary embodiment provides an interactive schematic diagram of inserting a visual identifier of a non-text media into text content.
[0017] Figure 6B It is another interactive schematic diagram of inserting a visual identifier of non-text media into text content provided by an exemplary embodiment.
[0018] Figure 7 is a schematic diagram of a finally generated image provided by an exemplary embodiment.
[0019] Figure 8 It is a schematic structural diagram of an electronic device provided by an exemplary embodiment. DETAILED DESCRIPTION
[0020] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with one or more embodiments of this specification. Instead, they are merely examples of devices and methods consistent with some aspects of one or more embodiments of this specification as detailed in the appended claims.
[0021] It should be noted that: in other embodiments, the steps of the corresponding method are not necessarily performed in the order shown and described in this specification. In some other embodiments, the steps included in the method may be more or less than those described in this specification. In addition, a single step described in this specification may be decomposed into multiple steps for description in other embodiments; and multiple steps described in this specification may be combined into a single step for description in other embodiments.
[0022] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this manual are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0023] Based on the problems in the related art, the embodiments of this specification provide an image generation method, and the user can freely combine the visual identifiers of non-text media (such as images, audio, etc.) with text content on the display interface, and can mix and arrange at least part of the text content and any visual identifier, so as to form multimedia input content that can accurately express the user's complex intentions, providing a more free and flexible interaction method; and in order to support this interaction method, metadata bound to the mixed target visual identifier can be generated, and then the multimedia input content and metadata are used to control the trained image generation model to generate images. The content recorded by the metadata can reflect the semantic correlation between the text content and the non-text media, so as to help the image generation model accurately parse the combination logic of the multimedia input content, so that the generated image can further meet the user's image generation needs and improve the accuracy of image generation.
[0024] Figure 1 FIG. 1 is a schematic diagram of the architecture of an image generation service system provided by an exemplary embodiment. Figure 1 As shown, the system may include a server 11, a network 12, and several user terminals, such as a PC (Personal Computer) 13, a mobile phone 14, and the like.
[0025] The server 11 may be a physical server including an independent host, or the server 11 may be a virtual server carried by a host cluster. During operation, the server 11 may run a server-side program of the image generation service to realize a corresponding image generation service platform.
[0026] PC13 and mobile phone 14 are only some types of user terminals that can be used by users. In fact, users can obviously also use user terminals such as the following types: tablet devices, laptops, PDAs (Personal Digital Assistants), wearable devices (such as smart glasses, smart watches, etc.), etc., and one or more embodiments of this specification do not limit this. During operation, the user terminal can run the program on the client side of the image generation service, and can be implemented as a client of the image generation service. Among them, the application program of the client of the above-mentioned image generation service can be started and run on the user terminal. The program on the client side can be a native application installed on the user terminal, or the program on the client side can be a small program, a quick application or other similar forms. Of course, when using web page technologies such as HTML5 or similar, the relevant functions can be implemented through the page displayed by the browser. The browser here can be an independent browser application or a browser module embedded in certain applications.
[0027] As for the network 12 for interaction between user terminals such as PC 13 and mobile phone 14 and server 11, it is possible to select a wired or wireless network to achieve communication based on the communication mode supported by the corresponding user terminal, and this specification does not limit this. For example, PC 13 can support both wired and wireless communication, so it is possible to use a wired or wireless network to achieve communication as needed, while mobile phone 14 usually only supports wireless communication, so it is possible to use a wireless network to achieve communication.
[0028] Among them, the image generation method provided in the embodiment of this specification can be executed by a user terminal or by a server 11, or part of the image generation method can be executed by a user terminal and the other part can be executed by a server 11. This embodiment does not impose any restrictions on this.
[0029] Exemplarily, if the image generation method provided in the embodiment of this specification is executed by a user terminal. In one case, if the user terminal is deployed with an image generation model and other related models, the user terminal does not need to interact with the server during the execution of the image generation method provided in the embodiment of this specification. In another case, if the image generation model and other related models are deployed in the server 11, the user terminal needs to send the relevant input data of the model (such as multimedia input content and metadata bound to the mixed target visualization identifier) to the server 11 during the execution of the image generation method provided in the embodiment of this specification, so that the server 11 returns the output result of the model.
[0030] For example, if the image generation method provided in the embodiment of this specification is executed by the server, the image generation model and other related models are deployed in the server 11, then the server 11 can control the user terminal to display text content, visual identification and multimedia input content on its display interface, and perform other operations.
[0031] Alternatively, part of the image generation method provided in the embodiment of the present specification is executed by the user terminal and the other part is executed by the server 11. For example, the user terminal can display text content, visual identifiers and multimedia input content in its display interface based on user operations, and can generate metadata bound to the mixed target visual identifier. Finally, in response to receiving an image generation instruction, the multimedia input content and the metadata bound to the target visual identifier can be sent to the server 11, so that the server 11 can use the multimedia input content and the metadata bound to the target visual identifier to control the trained image generation model to generate an image.
[0032] In some embodiments, see Figure 2, shows a flow chart of an image generation method, using an electronic device (such as the above-mentioned user terminal, server or other device), the method includes: In S201, the text content input by the user is displayed on the display interface, and the visual identifier corresponding to at least one non-text media uploaded by the user is displayed on the display interface.
[0033] For example, the user may input text content according to actual needs, and the text content may describe the user's need to generate an image, and then the electronic device may display the text content input by the user in its display interface, as shown in FIG. Figure 3 , showing that the following text content is displayed in the display interface: "Generate a picture based on the content of the book, one-to-one, and refer to this picture for style." It is understandable that the input method of text content can be keyboard input (including but not limited to physical keyboard, virtual keyboard, etc.), voice input (voice to text), handwriting input (such as writing with fingers or stylus on a touch screen) or scanning input (such as taking pictures of paper documents with devices such as scanners and smartphone cameras, and then using optical character recognition technology to recognize the text in the picture and convert it into editable text), etc., but not limited to these.
[0034] Exemplarily, the user may also upload non-text media according to actual needs, and different types of non-text media include at least one of the following: document files, images, audio, video, and image generation models; but not limited thereto.
[0035] Document files usually contain not only plain text, but also other types of data. For example, document file formats include: (1) PDF files, which contain text content, as well as non-text data such as images, tables, and charts. (2) Spreadsheets (such as Excel files), which contain not only text data but also charts, formulas, graphics, and other information. (3) PowerPoint presentations, which contain elements such as images, charts, and animations and are usually used together with text to convey information.
[0036] Image is a type of media that conveys information through static visual representation. Audio is a type of media that conveys information through sound. Video combines images and audio and is a form of multimedia used to convey information through continuous images and sounds. Image generation models refer to technologies that generate images through artificial intelligence (especially deep learning technology). These models usually create new images by learning the characteristics of a large amount of image data.
[0037] Based on the user's upload operation, the electronic device can display a visual identifier corresponding to at least one non-text media uploaded by the user in its display interface. These visual identifiers represent the non-text media and can be in a visual form such as an icon or thumbnail to help the user better understand and operate the non-text media uploaded by the user. The display styles of the visual identifiers corresponding to different types of non-text media are different, so that it is convenient for the user to distinguish them. The visual identifier corresponding to at least one non-text media uploaded by the user is displayed in a floating manner in the display interface, but is not limited to this. For example, please refer to Figure 4 , showing that a visual identifier corresponding to the picture 1 uploaded by the user and a visual identifier corresponding to the PDF file uploaded by the user are suspended and displayed on the left side of the display interface.
[0038] By intuitively displaying the visual identifiers of text content and non-text media in the display interface, users can easily view and manage the input content in various media forms, enhancing the friendliness and interactivity of the interface operation. Users do not need to switch between different input boxes or interfaces, and can directly process text content and non-text media in one interface at the same time, which can reduce the operation steps and improve efficiency.
[0039] In a possible implementation, a non-text media upload control can be displayed in the display interface. If the electronic device receives a trigger operation from the user for the non-text media upload control displayed in the display interface, multiple candidate non-text media included in the non-text media library can be displayed in the display interface; and then in response to receiving a selection operation from the user for at least one of the multiple candidate non-text media, a visual identifier corresponding to the selected non-text media is displayed in the display interface, which helps the user to clearly understand the selected media content, so that when uploading multiple non-text media, the user can quickly identify the selected content to avoid confusion and errors. Different types of non-text media correspond to different non-text media upload controls, so that users can upload various types of non-text media in a targeted manner. In this embodiment, by providing a special non-text media upload control in the display interface, users can upload different types of non-text media (such as pictures, audio, video, etc.) more intuitively and conveniently. Users can quickly access multiple candidate non-text media through simple trigger operations, which can reduce complex operation steps and improve usage efficiency.
[0040] See also Figure 5 , the user triggers the "Album" ( Figure 5 After the triggering process is not shown), the electronic device displays multiple candidate images in the album on the display interface, and the user selects picture 1. Then, the user triggers the "file" ( Figure 5The triggering process is not shown), the electronic device displays multiple candidate document files in "File" in the display interface, and the user selects one of the PDF files, and then the electronic device displays the visual identifier corresponding to the picture 1 and the visual identifier corresponding to the PDF file on the left side of the display interface. It can be understood that Figure 5 The "Album" (i.e., image upload control) and "File" (i.e., document file upload control) shown are only examples and do not constitute a limitation on non-text media upload controls. The settings of non-text media upload controls and related upload mechanisms can be specifically set according to actual application scenarios.
[0041] It is understandable that the embodiments of this specification do not impose any restrictions on the order of inputting text content and uploading non-text media. You can input text content first and then upload non-text media; you can also upload non-text media first and then input text content; it can be customized based on user needs.
[0042] In S202, if a mixing event for at least a portion of the text content and any visual identifier is monitored, at least a portion of the text content and the target visual identifier indicated by the mixing event are mixed and arranged to form multimedia input content, and metadata bound to the target visual identifier is generated. The metadata is used to record at least the target position of the target visual identifier in the multimedia input content, and the storage path and type of the non-text media indicated by the target visual identifier.
[0043] The mixing event includes at least one of the following: an event of dragging and dropping any visual marker to the front or back position of any word in the text content, and an event of inserting at least part of the words in the text content to the front or back position of any visual marker.
[0044] In the event of dragging and dropping any visual marker to the position before or after any word in the text content, the user directly drags the visual marker (such as an icon or thumbnail) representing the non-text media to a specific position in the text content through the drag-and-drop operation. Specifically, the user can place the marker before or after any character in the text, indicating that the user wants to merge the non-text media element with the text content to form multimedia input content.
[0045] In the event of inserting at least part of the text content before or after any visual identifier, the user inserts at least part of the text content before or after the originally existing visual identifier, that is, instead of moving the visual identifier, the user embeds at least part of the selected text content before or after the visual identifier, or enters text before or after the target visual identifier, indicating that the user wishes to visually and semantically mix the text content with the non-text media elements so that the two are more closely integrated.
[0046] In summary, both types of mixing events are aimed at achieving the integration of text and non-text media through different user interaction methods, thereby forming a rich multimedia input content. Electronic devices can accurately identify the location and attributes of non-text media indicated by visual markers based on the generated metadata, which is convenient for subsequent processing.
[0047] Exemplarily, the user can drag and drop any visual marker to a position before any word in the text content, and / or drag and drop any visual marker to a position after any word in the text content according to actual needs. Through the drag-and-drop operation, the user can more intuitively manipulate the relationship between the text content and the non-text media, providing a higher degree of freedom and flexibility, and supporting complex input methods. The user can accurately position the visual markers corresponding to different types of non-text media at specific locations in the text content according to their needs, which helps the user to accurately convey their intentions.
[0048] When the electronic device monitors that the user drags and drops the visual identifier to a specific position in the text content, the target visual identifier is inserted into the corresponding position of the text content, thereby forming multimedia input content. In addition, the electronic device generates metadata bound to the target visual identifier, and the metadata is at least used to record the storage path, type, and target position of the non-text media indicated by the target visual identifier in the multimedia input content. The storage path records the actual storage location of the non-text media indicated by the target visual identifier, ensuring that the non-text media can be found and accessed in the subsequent image generation process. The type helps the electronic device to perform different processing and rendering according to the characteristics of the media type. The target position records the insertion position of the target visual identifier in the text content, indicating which part of the text the target visual identifier should be displayed in, and can accurately reflect the semantic relevance between the text content and the non-text media. In short, metadata can help the image generation model correctly understand the multimedia input content by providing accurate storage path, type information, and target location, ensuring the accuracy of the generated image and meeting user needs. Of course, metadata can further record other information, such as the user's additional description of the non-text media, the size or resolution information of the non-text media, the time when the non-text media was uploaded, the source or copyright information of the non-text media, etc. This embodiment does not impose any restrictions on this.
[0049] For an example, see Fig. 6A and Figure 6B , Fig. 6A The process of a user dragging and dropping a visual identifier corresponding to a PDF file is shown. Figure 6BShows the process of the user performing a drag-and-drop operation on the visual identifier corresponding to Picture 1, and finally dragging the visual identifier corresponding to the PDF file behind the "book" in the text content, and dragging the visual identifier corresponding to Picture 1 behind the "picture" in the text content, forming the multimedia input content: "Generate a picture according to the content of the book [PDF icon], one-to-one, with the style referring to this [Picture 1 icon] picture."
[0050] In S203, if an image generation instruction is received, use the multimedia input content and the metadata bound to the target visual identifier to control the trained image generation model to generate an image.
[0051] In this step, when the electronic device receives an image generation instruction, the electronic device passes the multimedia input content and the metadata bound to the visual identifier to the trained image generation model, and the image generation model generates an image that meets the user's needs based on these inputs. By combining the multimedia input content with the metadata, the image generation model can better understand the semantic relationship between the text and non-text media, and thus generate an image that better meets the user's needs. Since the user can freely combine the text content and non-text media, the generated image can be more personalized and can reflect the user's unique intention.
[0052] Continuing with the above example, based on Fig. 6A and Figure 6B the multimedia input content and related metadata shown, a target image can be generated, as Figure 7 shown, the target image is displayed on the display interface, and the style of the target image is the same as the style of Picture 1.
[0053] In some embodiments, the electronic device can generate a multi-modal hybrid prompt based on the multimedia input content and the metadata bound to the target visual identifier, and then control the trained image generation model to generate an image according to the multi-modal hybrid prompt.
[0054] The following is an exemplary description of the generation process of multimodal mixed prompt words: the electronic device can input multimedia input content and metadata bound to the target visual identifier into a pre-trained language model, so that the pre-trained language model can perform intent recognition on the multimedia input content and metadata bound to the target visual identifier, thereby determining at least one constraint condition of the image to be generated and the processing content corresponding to each constraint condition, and generating a processing task based on the processing content corresponding to each constraint condition. Among them, the processing content corresponding to each constraint condition is at least part of the multimedia input content and the metadata bound to the target visual identifier. In this embodiment, through the intent recognition of the pre-trained language model, accurate intent can be extracted from a variety of input forms (text, image, metadata, etc.), ensuring that the generated image meets the constraints specified by the user, and avoiding inaccurate or deviated generation results from user needs.
[0055] Among them, pre-trained language models, such as large language models (LLM), refer to artificial intelligence models based on deep learning technology, especially those trained with a large corpus, which are designed to understand and generate text similar to human language and have strong natural language understanding and generation capabilities. The goal of LLM is to achieve a variety of applications through natural language processing capabilities, such as text generation, translation, summarization, question-answering, and dialogue systems, thereby helping to improve the efficiency and automation of human-computer interaction.
[0056] After obtaining the processing tasks corresponding to the constraints output by the pre-trained language model, the electronic device can execute the processing tasks corresponding to the constraints to obtain the processing results. Wherein, the types of media elements contained in the processing contents corresponding to the constraints are different, and the processing tasks corresponding to the constraints are also different. For example, the processing tasks include but are not limited to at least one of the following: key information extraction tasks for text content and / or document files, content analysis and / or style extraction tasks for images, speech recognition and analysis tasks for audio, deployment tasks for the first image generation model uploaded by the user, etc.
[0057] For example, when the electronic device is executing the processing tasks corresponding to the various constraint conditions, there are the following situations: In the first case, if the processing content corresponding to the constraint condition includes at least one of text content and document files, the electronic device can use a trained natural language processing model to perform semantic analysis on at least one of the text content and document files to extract key information. This embodiment does not impose any restrictions on the natural language processing model, and a trained natural language processing model in the relevant technology can be directly applied.
[0058] In the second case, if the processing content corresponding to the constraint condition includes an image, the electronic device can use the trained image processing model to perform at least one of content analysis and style extraction on the image to obtain an image processing result. This embodiment does not impose any restrictions on the image processing model, and the trained image processing model in the related art can be directly applied.
[0059] In the third case, if the processing content corresponding to the constraint condition includes audio, the electronic device can use the trained speech recognition model to perform speech recognition and analysis on the audio to obtain an audio processing result. This embodiment does not impose any restrictions on the speech recognition model, and the trained speech recognition model in the relevant technology can be directly applied.
[0060] In the fourth case, if the processing content corresponding to the constraint condition includes a first image generation model uploaded by a user, the electronic device may deploy the first image generation model in the image generation model deployment environment.
[0061] Electronic devices can process multiple types of input content (text, images, audio, etc.) and perform corresponding processing tasks according to different types of content, so as to more comprehensively understand user needs and generate results that meet the needs. For example, text content can be used to extract key information, image content can be used for style extraction, and audio content can be used for speech recognition. The combination of these tasks can improve the processing capabilities of multimodal data. By using trained models (such as natural language processing models, image processing models, speech recognition models, etc.), electronic devices can efficiently perform semantic analysis, content analysis, and style extraction on input content. These models have been trained with a large amount of data and have high precision and reliability. They can automatically perform complex analysis tasks, avoid human intervention, and improve processing efficiency and accuracy.
[0062] Finally, after obtaining the processing results corresponding to each constraint, the electronic device can generate a multimodal mixed prompt word based on the multimedia input content and the processing results corresponding to each constraint. The generated multimodal mixed prompt word based on the multimedia input content and the processing results corresponding to each constraint not only integrates text information, but also includes the features of other media information such as images and audio. In this way, the image generation model can receive a multi-level, all-round description and generate an image based on this information. The multimodal mixed prompt word can include details of the text description, image style requirements, audio emotional prompts, etc.
[0063] Among them, multimedia input content is the main source of user intention. The text content describes the basic requirements or scenarios of image generation, while non-text media (such as images, audio) provides direct visual or sound information about image details, style, emotions, etc. Multimedia input content provides a preliminary description framework and specific details for image generation, ensuring that the basic direction and content of image generation meet user expectations.
[0064] The processing results corresponding to each constraint (such as text analysis results, image content analysis, speech recognition results, etc.) can more accurately constrain and guide the image generation process. These processing results may contain key information, style requirements, semantic relationships, etc., to further refine user needs. For example, text analysis results may extract specific elements that should be included in the image (such as people, scenery, actions, etc.), while image processing results provide image style or visual effect requirements (such as realism, abstract art, etc.). These processing results help the model understand the semantic relationship between text and non-text content, thereby generating images that are more in line with user intentions. For example, speech recognition results can help the system understand the semantic information in the audio, integrate this information into the task of image generation, and generate images with more emotional and contextual sense.
[0065] In a possible implementation, the electronic device can directly splice the multimedia input content and the processing results corresponding to each constraint condition to obtain a multi-modal mixed prompt word. In this embodiment, the splicing processing method is relatively direct, and only the multimedia input content and the processing results need to be combined according to certain rules, without involving too much calculation, and the multi-modal mixed prompt word can be quickly generated, which is suitable for scenarios with high efficiency requirements.
[0066] In another possible implementation, in order to further improve the accuracy of multimodal mixed prompt words, the electronic device can input the multimedia input content and the processing results corresponding to each constraint condition into a pre-trained language model, so that the pre-trained language model can organize the multimedia input content and the processing results corresponding to each constraint condition to obtain multimodal mixed prompt words. In this way, the pre-trained language model will organize and reorganize the input information to produce multimodal mixed prompt words that are more consistent with language logic, semantically accurate and clear. The pre-trained language model has a strong semantic analysis capability, can understand the relationship between different modes and optimize it, and through the organization of the pre-trained language model, it can generate multimodal mixed prompt words that are more consistent with the context logic, ensuring that each input information is reasonably interpreted and integrated.
[0067] For example, a user wants to generate a cartoon-style seaside sunset image. The user can embed non-text media (such as wave texture images, sailboat images, audio files, and custom image generation models) into the specified position of the text content by dragging and dropping the corresponding visual logo of the non-text media to form multimedia input content: "A tranquil [wave icon] seaside sunset, with a [sailboat icon] sailboat in the distance, [audio icon] sea wave sound playing in the background, using [model icon] my cartoon rendering model." The metadata associated with the visual logo include: (1) Embedded rich media metadata: Wave icon: located after the 5th character of the text, type is image, path is / media / waves_texture.png. (2) Sailboat icon: located after the 15th character of the text, type is image, path is / media / sailboat.png. (3) Audio icon: located after the 25th character of the text, type is audio, path is / audio / ocean_waves.mp3. (4) Model icon: located after the 35th character of the text, type is custom image generation model, path is / models / cartoon_style.ckpt.
[0068] The electronic device inputs the multimedia input content and the metadata bound to the target visual identifier into the pre-trained language model. The pre-trained language model combines the anchor words in the text content (such as "seaside", "sailboat", "wave sound") and the type of embedded non-text media to identify the following intent: "Generate a cartoon-style seaside sunset scene, which requires the integration of dynamic wave textures and sailboats, and the wave rhythm to match the background audio." Based on this intent, the following constraints, processing content and processing tasks are determined: (1) Wave texture; processing content: the image indicated by the wave icon; processing task: acting on the "seaside" area, it needs to be dynamic and the frequency is synchronized with the audio. (2) Sailboat; processing content: the image indicated by the sailboat icon; processing task: the image indicated by the sailboat icon needs to be analyzed. (3) Audio; processing content: the audio indicated by the audio icon; processing task: extract the voiceprint spectrum to drive the wave rhythm. (4) Customized image generation model, processing content: the image generation model indicated by the model icon, processing task: deploy the customized image generation model.
[0069] The electronic device can execute the processing tasks corresponding to each constraint condition to obtain the processing results, and finally integrate the multimedia input content and the processing results corresponding to each constraint condition to obtain a multimodal mixed prompt word, such as "(1) Cartoon-style seaside sunset scene (application model cartoon_style.ckpt); (2) Sea surface area: using dynamic wave texture, texture path waves_texture.png; (3) Distant object: cartoon sailboat; (4) Global rhythm: wave fluctuations are synchronized with the soothing soundprint of the audio ocean_waves.mp3." For example, multiple image generation models may be pre-deployed in the image generation model deployment environment. When the user does not upload a custom image generation model, the image generation model in the image generation model deployment environment may be directly used. The process of generating images according to multi-modal mixed prompt words has the following situations: In the first case, if the processing task includes the task of deploying a first image generation model uploaded by a user, and the electronic device has completed the deployment process of the first image generation model, the electronic device can input the multimodal mixed prompt words into the deployed first image generation model so that the deployed first image generation model can generate an image according to the multimodal mixed prompt words. User-uploaded customized image generation models can meet specific needs or preferences. By using the image generation model uploaded by the user, images that meet the user's unique requirements or industry needs can be generated. This method provides users with great flexibility and freedom, and can generate images according to specific styles, content or other customized needs.
[0070] In the second case, if the text content includes a second image generation model specified by the user, and the second image generation model is deployed in the image generation model deployment environment, the electronic device can input the multimodal mixed prompt words into the second image generation model so that the second image generation model can generate an image according to the multimodal mixed prompt words. In this case, the user explicitly specifies the second image generation model, usually to utilize the specific capabilities of the model (such as processing specific types of content, style, language, etc.). By using the image generation model specified by the user, the electronic device can generate images that better meet the user's requirements and improve the user experience. Users can choose their preferred generation model instead of passively relying on the default model. This method provides users with a variety of image generation models to choose from to meet a wider range of application needs.
[0071] In the third case, if the user does not upload the first image generation model, or the second image generation model specified by the user does not exist in the image generation model deployment environment, or the user does not specify the second image generation model, the electronic device can input the multimodal mixed prompt words into the default third image generation model deployed in the image generation model deployment environment, so that the third image generation model generates an image according to the multimodal mixed prompt words. In this case, even if the user does not upload a custom image generation model or does not specify a second image generation model, it can automatically fall back to the default third image generation model to ensure that the image generation process is not interrupted.
[0072] It can be understood that the “first image generation model”, “second image generation model” and “third image generation model” mentioned above are to distinguish the use of models in different situations, and “first, second and third” themselves do not have any other special meanings.
[0073] In some embodiments, in addition to the multimodal prompt words, preset parameters may be added, and the preset parameters include at least one of the following: parameters related to the image, parameters related to the image generation model. Parameters related to the image are, for example, image resolution, image ratio, hue and contrast, lighting and shadow, etc. These parameters directly affect the appearance, style, content and other visual features of the generated image. Parameters related to the image generation model may be parameters used to control the internal algorithm of the image generation model, such as (1) classifier free bootstrap weights, which are used to adjust the compliance of the image generation model with the input prompt words. The higher the weight, the more strictly the generated result fits the prompt words, but the diversity may be reduced; (2) the number of iterations (or sampling steps), which is used to control the number of iterations of the image generation process. More iterations may result in a more refined image, but require more computing time; but not limited to this.
[0074] The electronic device inputs the multimodal mixed prompt words and preset parameters into the trained image generation model, so that the trained image generation model (such as any one of the first image generation model, the second image generation model, and the third image generation model) generates an image based on the multimodal mixed prompt words and preset parameters. By presetting these parameters, the electronic device can reduce manual operations and adjustments during the generation process, and provide more automated, fast, and accurate generation results. For example, after setting the resolution and style of the generated image, the electronic device will automatically control the image generation model to generate the image according to these requirements, without the need for the user to manually adjust it each time.
[0075] Exemplarily, the electronic device can check whether the multimodal mixed prompt word contains key information related to the image, such as resolution, style, lighting, etc. If this information has been clearly specified, there is no need to repeatedly set these contents through additional preset parameters. For example, if the multimodal mixed prompt word clearly requires "generate an oil painting style image with a resolution of 1920x1080", the system can directly use this information without the user providing additional "style" and "resolution" parameters again.
[0076] Similarly, the electronic device will also detect whether the multimodal mixed prompt word already involves parameters related to the image generation model, such as the number of iterations, etc. If the multimodal mixed prompt word already contains the relevant parameters of the image generation model, it means that the model can automatically perform the image generation process based on this information without repeated addition.
[0077] If the multimodal mixed prompt does not contain the necessary preset parameters, the electronic device can automatically supplement these missing parameters. For example, if the user does not specify the image resolution or style in the multimodal mixed prompt, the electronic device can automatically fill in these parameters according to the default settings or according to the task requirements to ensure that the image generation task proceeds smoothly. This method enables electronic devices to have sufficient flexibility in handling different types of inputs, while avoiding the trouble of users having to provide all the details every time they input. Through intelligent judgment and automatic supplementation, user operations are simplified and generation efficiency is improved.
[0078] For example, assuming that the multimodal mixed prompt is: "Generate an image of a city landscape on a sunny day", the electronic device will check the multimodal mixed prompt, determine that the multimodal mixed prompt does not mention the resolution or style, and automatically use the default resolution (such as 1024x1024) and style (such as realistic style); and determine that the multimodal mixed prompt does not mention the relevant image generation model, the electronic device uses the default third image generation model.
[0079] Assume that the multimodal mixed prompt word is: "Generate an image of a city landscape on a sunny day, with a resolution of 1920x1080 and a watercolor style". After the electronic device checks the multimodal mixed prompt word and determines that it contains the adoption number related to the image, it directly uses the information in the multimodal mixed prompt word to generate the image.
[0080] It is understandable that the related models mentioned in this specification, such as image generation models, pre-trained language models, natural language processing models, image processing models, and speech recognition models, can be trained based on the training methods in the relevant technologies, and this embodiment does not impose any restrictions on the model training process; and the model architecture of the related models mentioned in this specification can directly adopt the model architecture mentioned in the relevant technologies. For example, the image generation model can be a diffusion model (such as Stable Diffusion), a generative adversarial network (GAN, Generative Adversarial Network), a variational autoencoder (VAE, Variational Autoencoder), an autoregressive model (Autoregressive Model), a Transformer architecture, etc., and this embodiment does not impose any restrictions on this.
[0081] The various technical features in the above embodiments can be arbitrarily combined as long as there is no conflict or contradiction between the combinations of features. However, due to space limitations, they are not described one by one. Therefore, any combination of the various technical features in the above embodiments also falls within the scope of this specification.
[0082] In some embodiments, the embodiments of this specification also provide an electronic device, including: a processor; a memory for storing processor executable instructions; wherein the processor implements any of the above methods by running the executable instructions.
[0083] Figure 8 is a schematic structural diagram of a device provided by an exemplary embodiment. Figure 8 At the hardware level, the device includes a processor 802, an internal bus 804, a network interface 806, a memory 808, and a non-volatile memory 810, and may also include hardware required for other functions. One or more embodiments of this specification may be implemented based on software, such as the processor 802 reading the corresponding computer program from the non-volatile memory 810 into the memory 808 and then running it. Of course, in addition to the software implementation, one or more embodiments of this specification do not exclude other implementations, such as logic devices or a combination of software and hardware, etc., that is, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0084] Exemplarily, the image generating device can be applied to Figure 8 The device shown in the figure is used to implement the technical solution of this specification. The image generating device may include: A display module, configured to display text content input by a user on a display interface, and to display a visual identifier corresponding to at least one non-text media uploaded by a user on the display interface; A shuffling event processing module is used to, if a shuffling event for at least a portion of the text content and any visual identifier is monitored, shuffle at least a portion of the text content and a target visual identifier indicated by the shuffling event to form multimedia input content, and generate metadata bound to the target visual identifier, the metadata being used to record at least a target position of the target visual identifier in the multimedia input content, and a storage path and type of the non-text media indicated by the target visual identifier.
[0085] The image generation module is used to control the trained image generation model to generate an image by using the multimedia input content and the metadata bound to the target visualization identifier when an image generation instruction is received.
[0086] In one implementation, the shuffling event includes at least one of the following: an event of dragging and dropping any of the visual identifiers to a position before or after any word in the text content, and an event of inserting at least part of the words in the text content to a position before or after any of the visual identifiers.
[0087] In one implementation, the visual identifier corresponding to the at least one non-text media uploaded by the user is displayed in the display interface in a suspended manner.
[0088] In one implementation, different types of non-text media correspond to different display styles of visual identifiers.
[0089] In one implementation, the different types of non-text media include at least one of the following: a document file, an image, an audio, a video, and an image generation model.
[0090] In one implementation, the display module is specifically used to display multiple candidate non-text media included in the non-text media library in the display interface if a trigger operation from the user is received for the non-text media upload control displayed in the display interface; in response to receiving a selection operation from the user for at least one of the multiple candidate non-text media, display a visual identifier corresponding to the selected non-text media on the display interface.
[0091] In one implementation, the image generation module is specifically used to generate multimodal mixed prompt words based on the multimedia input content and the metadata bound to the target visualization identifier; and control the trained image generation model to generate an image according to the multimodal mixed prompt words.
[0092] In one implementation, the image generation module is specifically used to input the multimedia input content and the metadata bound to the target visualization identifier into a pre-trained language model, so that the pre-trained language model can perform intent recognition on the multimedia input content and the metadata bound to the target visualization identifier, thereby determining at least one constraint condition of the image to be generated and the processing content corresponding to each of the constraints, and generating a processing task based on the processing content corresponding to each of the constraints; executing the processing tasks corresponding to each of the constraints to obtain a processing result; and generating a multimodal mixed prompt word based on the multimedia input content and the processing results corresponding to each of the constraints.
[0093] In one implementation, the image generation module is specifically used to, if the processing content corresponding to the constraint condition includes text content and / or document files, use a trained natural language processing model to perform semantic analysis on the text content and / or document files to extract key information; if the processing content corresponding to the constraint condition includes images, use a trained image processing model to perform content analysis and / or style extraction on the images to obtain image processing results; if the processing content corresponding to the constraint condition includes audio, use a trained speech recognition model to perform speech recognition and analysis on the audio to obtain audio processing results; if the processing content corresponding to the constraint condition includes a first image generation model uploaded by a user, deploy the first image generation model in an image generation model deployment environment.
[0094] In one implementation, the image generation module is specifically used to input the multimedia input content and the processing results corresponding to each of the constraints into a pre-trained language model, so that the pre-trained language model can organize the multimedia input content and the processing results corresponding to each of the constraints to obtain the multimodal mixed prompt words.
[0095] In one implementation, the image generation module is specifically used to, if the processing task includes a task of deploying a first image generation model uploaded by a user, input the multimodal mixed prompt word into the deployed first image generation model, so that the deployed first image generation model generates an image according to the multimodal mixed prompt word; if the text content includes a second image generation model specified by a user, and the second image generation model is deployed in the image generation model deployment environment, input the multimodal mixed prompt word into the second image generation model, so that the second image generation model generates an image according to the multimodal mixed prompt word; if the user does not upload the first image generation model, or the second image generation model specified by the user does not exist in the image generation model deployment environment, or the user does not specify the second image generation model, input the multimodal mixed prompt word into a third image generation model deployed in the image generation model deployment environment, so that the third image generation model generates an image according to the multimodal mixed prompt word.
[0096] In one implementation, the image generation module is specifically used to input the multimodal mixed prompt words and preset parameters into a trained image generation model, so that the trained image generation model generates an image based on the multimodal mixed prompt words and the preset parameters; wherein the preset parameters include at least one of the following: parameters related to the image, parameters related to the image generation model.
[0097] The implementation process of the functions and effects of each module in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, which will not be repeated here.
[0098] Based on the same concept as the above method, this specification also provides an electronic device, including: a processor; a memory for storing processor executable instructions; wherein the processor implements the steps of the method described in any of the above embodiments by running the executable instructions.
[0099] Based on the same concept as the above method, this specification also provides a computer-readable storage medium on which computer instructions are stored. When the instructions are executed by a processor, the steps of the method described in any of the above embodiments are implemented.
[0100] Computer-readable media include permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, disk storage, quantum memory, graphene-based storage media or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include temporary computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0101] Based on the same concept as the above method, this specification also provides a computer program product, including a computer program / instruction, which implements the steps of the method described in any of the above embodiments when executed by a processor.
[0102] The above description is merely a preferred embodiment of one or more embodiments of the present specification and is not intended to limit one or more embodiments of the present specification. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of one or more embodiments of the present specification shall be included in the scope of protection of one or more embodiments of the present specification.
Claims
1. A method for generating an image, comprising: Displaying text content input by the user on a display interface, and displaying a visual identifier corresponding to at least one non-text media uploaded by the user on the display interface; If a shuffling event for at least part of the text content and any of the visual identifiers is monitored, at least part of the text content and the target visual identifier indicated by the shuffling event are mixed and arranged to form multimedia input content, and metadata bound to the target visual identifier is generated, where the metadata is used to record at least a target position of the target visual identifier in the multimedia input content and a storage path and type of the non-text media indicated by the target visual identifier; If an image generation instruction is received, the multimedia input content and the metadata bound to the target visualization identifier are used to control the trained image generation model to generate an image.
2. According to the method of claim 1, the shuffling event comprises at least one of the following: an event of dragging and dropping any of the visual markers to a position before or after any word in the text content, and an event of inserting at least part of the words in the text content to a position before or after any of the visual markers; and / or, The visual identifier corresponding to the at least one non-text media uploaded by the user is displayed in the display interface in a suspended manner.
3. According to the method of claim 1, different types of non-text media have different display styles for the visual identifiers; And / or, the different types of non-text media include at least one of the following: document files, images, audio, video, and image generation models.
4. The method according to claim 1, wherein displaying a visual identifier corresponding to at least one non-text media uploaded by the user on the display interface comprises: If a trigger operation of the user for the non-text media upload control displayed in the display interface is received, a plurality of candidate non-text media included in the non-text media library is displayed in the display interface; In response to receiving a user's selection operation on at least one of the plurality of candidate non-text media, a visual identifier corresponding to the selected non-text media is displayed on the display interface.
5. The method according to any one of claims 1 to 4, wherein the step of controlling a trained image generation model to generate an image by using the multimedia input content and metadata bound to the target visualization identifier comprises: Generate a multimodal mixed prompt word based on the multimedia input content and the metadata bound to the target visual identifier; The trained image generation model is controlled to generate an image according to the multimodal mixed prompt words.
6. The method according to claim 5, wherein generating a multimodal mixed prompt word based on the multimedia input content and metadata bound to the target visual identifier comprises: Inputting the multimedia input content and the metadata bound to the target visualization identifier into a pre-trained language model, so that the pre-trained language model performs intent recognition on the multimedia input content and the metadata bound to the target visualization identifier, thereby determining at least one constraint condition of the image to be generated and processing content corresponding to each constraint condition, and generating a processing task based on the processing content corresponding to each constraint condition; Execute the processing tasks corresponding to the respective constraint conditions to obtain processing results; Based on the multimedia input content and the processing results corresponding to the respective constraint conditions, a multimodal mixed prompt word is generated.
7. The method according to claim 6, wherein the step of executing the processing tasks corresponding to the respective constraint conditions to obtain the processing results comprises: If the processing content corresponding to the constraint condition includes text content and / or document files, a trained natural language processing model is used to perform semantic analysis on the text content and / or document files to extract key information; If the processing content corresponding to the constraint condition includes an image, using the trained image processing model to perform content analysis and / or style extraction on the image to obtain an image processing result; If the processing content corresponding to the constraint condition includes audio, use the trained speech recognition model to perform speech recognition and analysis on the audio to obtain an audio processing result; If the processing content corresponding to the constraint condition includes a first image generation model uploaded by a user, the first image generation model is deployed in the image generation model deployment environment.
8. The method according to claim 6, wherein generating a multimodal mixed prompt word based on the multimedia input content and the processing results corresponding to each of the constraint conditions comprises: The multimedia input content and the processing results corresponding to each of the constraints are input into a pre-trained language model, so that the pre-trained language model can organize the multimedia input content and the processing results corresponding to each of the constraints to obtain the multimodal mixed prompt words.
9. The method according to claim 7, wherein the step of controlling the trained image generation model to generate an image according to the multimodal mixed prompt words comprises: If the processing task includes a task of deploying a first image generation model uploaded by a user, inputting the multimodal mixed prompt word into the deployed first image generation model, so that the deployed first image generation model generates an image according to the multimodal mixed prompt word; If the text content includes a second image generation model specified by a user, and the second image generation model is deployed in the image generation model deployment environment, inputting the multimodal mixed prompt word into the second image generation model so that the second image generation model generates an image according to the multimodal mixed prompt word; If the user has not uploaded the first image generation model, or the second image generation model specified by the user does not exist in the image generation model deployment environment, or the user has not specified the second image generation model, the multimodal mixed prompt words are input into the third image generation model deployed in the image generation model deployment environment, so that the third image generation model generates an image according to the multimodal mixed prompt words.
10. The method according to claim 5, wherein the step of controlling the trained image generation model to generate an image according to the multimodal mixed prompt words comprises: Inputting the multimodal mixed prompt words and preset parameters into a trained image generation model, so that the trained image generation model generates an image based on the multimodal mixed prompt words and the preset parameters; The preset parameters include at least one of the following: parameters related to the image and parameters related to the image generation model.
11. An electronic device, comprising: processor; A memory for storing processor-executable instructions; wherein the processor implements the steps of the method according to any one of claims 1 to 10 by executing the executable instructions.
12. A computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the steps of the method according to any one of claims 1 to 10.
13. A computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Image generation method and device and storage medium
CN117475031A
Text-based image generation method and device, electronic equipment and storage medium
CN117493599A
Figure generation method, apparatus and device, and computer readable storage medium
CN118918414A
Multimedia processing method, system, device, equipment, medium and product
CN119416889A
Image generation method and device, equipment and storage medium
CN119625102A
Cited By
Picture processing method and device, electronic equipment and storage medium
CN121502021A