Image generation method, apparatus, device, medium, and product
Patent Information
- Application Number
- CN202610772664.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-29
- Publication Date
- 2026-08-18
AI Technical Summary
[0003]在此提供一种图像生成方法、装置、设备、介质以及产品,解决了对图像生成或编辑时,机器学习模型的输出结果与需求不适配的问题,提高了图像生成或图像编辑的精准度
[0008] In another scenario, this document also provides a computer program product, including a computer program that, when executed by a processor, implements the image generation method as described herein.
Smart Images

Figure CN122597573A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer processing technology, and in particular to an image generation method, apparatus, device, medium, and product. Background Technology
[0002] Image generation or editing tasks commonly employ machine learning models. However, when using machine learning models, there is a problem where the output results do not match the requirements. Summary of the Invention
[0003] This invention provides an image generation method, apparatus, device, medium, and product that solves the problem of mismatch between the output of machine learning models and requirements during image generation or editing, thereby improving the accuracy of image generation or image editing.
[0004] In one scenario, this paper provides an image generation method, which includes: Based on the data unit and the first content, a second content is obtained, wherein the first content is used to indicate the first planning layout information of the first media element, and the second content is used to characterize the spatial location information of the first media element. Based on the second content, a third content is obtained, which is provided by a first machine learning model and is used to indicate the output image corresponding to the first content.
[0005] In one instance, this document also provides an image generation apparatus, which includes: The second content acquisition module is used to acquire second content based on the data unit and the first content, wherein the first content is used to indicate the first planning layout information of the first media element, and the second content is used to characterize the spatial location information of the first media element. The third content acquisition module is used to obtain third content based on the second content, wherein the third content is provided by the first machine learning model and is used to indicate the output image corresponding to the first content.
[0006] In one instance, this document also provides an electronic device comprising: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the image generation method as described herein.
[0007] In one instance, this document also provides a storage medium containing computer-executable instructions that, when executed by a computer processor, are used to perform an image generation method as described herein.
[0008] In another scenario, this document also provides a computer program product, including a computer program that, when executed by a processor, implements the image generation method as described herein.
[0009] The beneficial effects of the aforementioned image generation method are as follows: Based on data units and first content, second content is obtained. The first content is used to indicate at least the first planning layout information of the first media element, and the second content is used to represent the spatial position information of the first media element. Based on the second content, third content is obtained. The third content is provided by a first machine learning model and is used to indicate the output image corresponding to the first content. This solves the problem of mismatch between the output results of the machine learning model and the requirements during image generation or editing, and realizes refined and precise control over the spatial position of image content during image generation or image editing, thereby improving the accuracy of image generation or image editing. Attached Figure Description
[0010] The above and other features, advantages, and aspects of the embodiments described herein will become more apparent when taken in conjunction with the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0011] Figure 1 This is a schematic diagram of the architecture of an exemplary system in one scenario. Figure 2 This is a flowchart illustrating an image generation method for one scenario. Figure 3 This is a flowchart illustrating an image generation method for one scenario. Figure 4 An example diagram illustrating an image generation method in one scenario; Figure 5 This is a flowchart illustrating an image generation method for one scenario. Figure 6 This is a flowchart illustrating an image generation method for one scenario. Figure 7 Here is an example image of the output image under one condition; Figure 8 This is a flowchart illustrating an image generation method for one scenario. Figure 9 This is a flowchart illustrating an image generation method for one scenario. Figure 10 This is a schematic diagram of an image generation device in one scenario. Figure 11 This is a schematic diagram of the structure of an electronic device under one specific scenario. Detailed Implementation
[0012] The embodiments will now be described in more detail with reference to the accompanying drawings. While some embodiments are shown in the drawings, it should be understood that the technical solutions can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the technical solutions herein. It should be understood that the illustrated drawings and embodiments are for illustrative purposes only and are not intended to limit the scope of protection of the technical solutions.
[0013] It should be understood that the steps described in the method implementation may be performed in different orders and / or in parallel. Furthermore, the method implementation may include additional steps and / or omit the steps shown. The scope of this document is not limited in this respect.
[0014] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one situation" means "at least one situation"; the term "another situation" means "at least one additional situation"; the term "some situations" means "at least some situations". Definitions of other terms will be given in the following description.
[0015] It should be noted that the concepts of "first" and "second" mentioned are only used to distinguish different devices, modules or units, and are not used to limit the order of the functions performed by these devices, modules or units or their interdependencies.
[0016] It should be noted that the terms "one" and "more" used in this document are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0017] The names of messages or information exchanged between multiple devices in this document are for illustrative purposes only and are not intended to limit the scope of these messages or information.
[0018] It is understood that before using the technical solutions disclosed in the various embodiments of this document, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this document in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0019] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as electronic devices, applications, servers, or storage media, that perform the operations described herein, based on the prompt message.
[0020] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0021] It is understood that the above notification and user authorization process is merely illustrative and does not limit the implementation method described in this article. Other methods that comply with relevant laws and regulations may also be applied to the implementation method described in this article.
[0022] It is understood that the data involved in the technical solutions in this article (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0023] In this paper, we can first explain the application scenarios of this technical solution: it can be applied to any scenario that requires generating corresponding images based on input information. For example, the input information can be requirement description information, and the output image can be an image adapted based on the requirement description information.
[0024] In one scenario, a first machine learning model processes first content to obtain an output image corresponding to that content. Here, the first content primarily refers to user-edited or other intelligent system-output descriptive content more suited to the user's needs. Typically, the descriptive content may include the planned layout information of the required content within the image, such as the specific content and its location within the image.
[0025] In another scenario, the scene that generates the corresponding image can include two types: image generation and image editing. An image generation scene can be understood as input consisting only of a requirement description; an image editing scene can be understood as input including not only the requirement description but also the input image that needs to be edited based on that requirement description.
[0026] In some cases, the provided image generation method can be applied to Figure 1The image generation system shown may include a client 101 and a server 102. The client 101 may include, but is not limited to, browsers, applications (Apps), HyperText Markup Language (HTML) applications, lightweight applications (also known as mini-programs), or cloud applications. The client 101 may be deployed on an electronic device and relies on the operation of that device or certain applications within the device to implement its functions. The electronic device may be, for example, a device with a display screen that supports information browsing, such as a smartphone, tablet, personal computer, or other client terminal. For ease of understanding, Figure 1 The client is primarily represented in the form of a device. Other types of applications can also be configured on the electronic device, such as media content publishing applications, session applications, etc. Server 102 can be one or more servers providing various services. That is, it can be implemented as a distributed server cluster composed of multiple servers, or as a single server; furthermore, it can be a server for a distributed system, a server integrating blockchain technology, a cloud server, or an intelligent cloud computing server or intelligent cloud host deployed with machine learning models, etc.
[0027] The image generation method described herein allows interaction between client 101 and server 102, such as receiving or sending messages. For example, in this paper, server 102 can receive first content sent by client 101 based on an information carrier, and send third content corresponding to the first content to client 101 for display on the display interface.
[0028] It should be noted that the image generation method can be executed by client 101. In this case, the first machine learning model is deployed on client 101, which determines the second content based on the data unit and the first content, and processes the second content based on the first machine learning model deployed on client 101 to output an image. Alternatively, the image generation method can be executed by client 101 and server 102, with different functional parts of the corresponding image generation device deployed on client 101 and server 102 respectively. In one scenario, the second content acquisition module and the third content acquisition module of the device are deployed on server 102. In this case, client 101 receives the first content and sends it to the second content acquisition module. Based on the data unit in the second content acquisition module and the first content sent by client 101, the second content is determined, and based on the third content acquisition module, the third content corresponding to the second content is obtained to provide feedback on the third content. In another scenario, the second content acquisition module of the device is deployed on client 101, and the third content acquisition module is deployed on server 102. In this case, the second content acquisition module of client 101 obtains the second content based on the data unit and the first content, and sends the second content to server 102. Based on the third content acquisition module in server 102, the third content corresponding to the second content is obtained, and the third content is fed back to client 101.
[0029] Client 101 and server 102 communicate over a network to exchange data and coordinate functions. It should be understood that... Figure 1 The number of clients and servers shown is for illustrative purposes only. Any number of clients and servers can be configured to meet specific implementation requirements.
[0030] Figure 2 This is a flowchart illustrating an image generation method for one scenario, applicable to situations requiring the generation of images that match the specified requirements. This method can be executed by an image generation device, which can be implemented in software and / or hardware, optionally through an electronic device, such as a client. For a detailed description of the specific implementation, please refer to this scenario.
[0031] like Figure 2 As shown, the image generation method may specifically include: S210. Based on the data unit and the first content, obtain the second content, wherein the first content is used to indicate the first planning layout information of the first media element, and the second content is used to characterize the spatial location information of the first media element.
[0032] The first content can be a description of the requirements corresponding to the actual needs. This first content can be text information obtained through user editing or text information generated by an intelligent system. Alternatively, the first content can be text information converted from input audio or requirement information expressed through video.
[0033] It should be noted that the presentation format of the first content is not limited in this case and can be adjusted according to actual needs, as long as it can express the current needs.
[0034] Furthermore, the first content includes at least: a first media element and corresponding first planning and layout information. The first media element can be a visual basic material used to constitute the image. The first media element can specifically refer to graphics, people, text, and scenery, etc.
[0035] The first layout information can be used to indicate the overall arrangement of the first media element in the image. Optionally, the first layout information can be characterized by text description information. The text description information can represent the element position information, element size information, element stacking level information, and element density arrangement information of the first media element, etc.
[0036] It should be noted that the requirement description information is usually discrete text information. In order to make the output results of the subsequent machine learning model more compatible with the actual requirements, data units are introduced to realize the association between text information and image space through data units.
[0037] Data units can be MetaPoint Tokens (MPTs). A MetaPoint Token can be understood as a sparse, semantic, and controllable visual basic unit. Through MetaPoint Tokens, the position of the first media element can be controlled; that is, the spatial position of each first media element can be pinpointed, allowing independent control of its pose, position, size, and layout during image output. This avoids the adhesion or misalignment of first media elements, solving problems such as inaccurate positioning and difficult layout control, and achieving fine-grained control over the position of media elements.
[0038] In one scenario, the first content is processed using data units to obtain the spatial location information of a first media element. Correspondingly, the spatial location information of the first media element determined by the data units is used as the second content. Optionally, the spatial location information can be represented in the form of a feature vector.
[0039] Specifically, upon receiving the first content edited by the user or the first content generated by the intelligent system, the data unit processes the first planning layout information of the first media element to determine the spatial location information of the first media element.
[0040] S220. Based on the second content, obtain the third content, which is provided by the first machine learning model and is used to indicate the output image corresponding to the first content.
[0041] The first machine learning model can be a pre-trained generative model used for image generation or image editing. The first machine learning model can employ a diffusion model architecture. Under this architecture, the first machine learning model processes the noisy image or the input image based on the second input content to output an image that corresponds to the actual requirement.
[0042] The third content can be understood as the result obtained by the first machine learning model processing the second content. The third content is the output image corresponding to the first content. For example, if the first content is: the arrangement information of the first egg, the arrangement information of the second egg, ..., the arrangement information of the tenth egg, then the third content can be the corresponding output image, which includes image information of 10 eggs. The arrangement of each egg in the output image is compatible with the arrangement information set in the first content, and also compatible with the spatial position information corresponding to the second content.
[0043] Specifically, the first machine learning model obtains an output image corresponding to the first content based on the spatial location information of the first media element represented by the second content, and outputs the output image.
[0044] The beneficial effects of the aforementioned image generation method are as follows: Based on data units and first content, second content is obtained. The first content is used to indicate at least the first planning layout information of the first media element, and the second content is used to represent the spatial position information of the first media element. Based on the second content, third content is obtained. The third content is provided by a first machine learning model and is used to indicate the output image corresponding to the first content. This solves the problem of mismatch between the output results of the machine learning model and the requirements during image generation or editing, and realizes refined and precise control over the spatial position of image content during image generation or image editing, thereby improving the accuracy of image generation or image editing.
[0045] Figure 3 This is a flowchart illustrating an image generation method under one scenario. Based on the above, the first content can be obtained in the following manner, the specific implementation of which can be found in the detailed description of this scenario. Technical features that are the same as or similar to those described above will not be repeated here.
[0046] like Figure 3 As shown, the image generation method includes the following steps: S310. Receive first input and obtain first content. The first input includes at least content editing instructions composed of natural language, and the first content is provided by the intelligent system.
[0047] The first input can be descriptive information related to user needs. Optionally, the descriptive information can be presented in the form of text, audio, video, etc.
[0048] In one scenario, the initial input can be obtained through the first page. The first page can be the page triggered during image generation or editing. If only one page is associated with image generation or editing, the first page is that single page. If multiple pages or external applications need to be invoked during image generation or editing, all displayed pages can be considered the first page. For example, if an intelligent system is invoked during image editing or generation, the page presented by the intelligent system is also considered the first page.
[0049] Users can interact with page elements on the first page using input tools such as a mouse and / or keyboard. When the function plugin is activated, all user interactions with the first page can be recorded to obtain the first input.
[0050] The first input includes at least content editing instructions in natural language. Content editing instructions are descriptive information corresponding to the actual needs. For example, a content editing instruction could be "Generate an image containing 10 eggs".
[0051] The intelligent system can be an intelligent agent or a Large Language Model (LLM). After analyzing and processing the first input, the intelligent system can generate structured text. This structured text can be clear, readable, reusable, and standardized. The structured text can be edited by the user or the intelligent system, and the edited structured text can be retained. This structured text includes the first planning and layout information of the first media element, i.e., the first content.
[0052] Specifically, after detecting a trigger operation on the corresponding control, the first page can be displayed. The user can edit the first input into the input box on the first page and click the corresponding send control to allow the intelligent system to receive the first input. The intelligent system processes the first input, which includes at least a content editing instruction, to determine the first planning layout information of the first media element corresponding to the content editing instruction, i.e., the first content.
[0053] For example, see Figure 4On the first page, the user can input the content editing command "generate an image containing 10 eggs". When the first page receives this content editing command, it uses the command as the first input and enables the intelligent system to process it to obtain the arrangement information of each egg, i.e., the first planning layout information. This first planning layout information is used as the first content for subsequent processing.
[0054] Optionally, in response to an adjustment operation on the first content, the adjusted first content is obtained. The adjustment operation includes one of the following: receiving an input editing operation on the first content; receiving a natural language instruction to adjust the first content.
[0055] The input editing operation can be used to represent the user's editing and input of the first content. The natural language instruction to adjust the first content can be an instruction received by the intelligent system to adjust the first content.
[0056] Specifically, to make the initial content more closely match actual needs, the initial content output by the intelligent system can be adjusted. There are two main adjustment methods: one is for the user to input and edit the initial content to obtain the adjusted content; the other is to send natural language instructions for adjusting the initial content to the initial page based on the initial content presented by the intelligent system, so that the intelligent system can adjust the initial content according to the instructions, obtaining the adjusted initial content. Subsequently, an output image corresponding to the adjusted initial content can be obtained.
[0057] S320. Based on the data unit and the first content, obtain the second content, wherein the first content is used to indicate the first planning layout information of the first media element, and the second content is used to characterize the spatial location information of the first media element.
[0058] S330. Based on the second content, obtain the third content, which is provided by the first machine learning model and is used to indicate the output image corresponding to the first content.
[0059] Based on the examples above, see further. Figure 4 By building a bridge between the first content and space based on data units, the planning layout of the first media element "egg" and its spatial position in the image can be established, and the spatial position information of each egg in the image can be obtained.
[0060] Accordingly, the first machine learning model, based on the spatial location information of each egg in the image, combines the content editing instruction "generate an image containing 10 eggs" and outputs the image. The output image includes the image information of these 10 eggs, and the arrangement of the eggs matches the determined spatial location information.
[0061] The beneficial effects corresponding to the above image generation method are as follows: receiving a first input and processing the first input based on an intelligent system to obtain a first content, where the first input at least includes a content editing instruction composed of natural language. Based on a data unit and the first content, a second content is obtained. The first content is at least used to indicate the first planning layout information of a first media element, and the second content is used to characterize the spatial position information of the first media element. Based on the second content, a third content is obtained. The third content is provided by a first machine learning model and is used to indicate an output image corresponding to the first content. By processing the first input through the intelligent system, the determination of the spatial position information of the media element is achieved, fundamentally ensuring the accuracy of the position control of the first media element, and achieving the technical effect of improving the accuracy of image generation or editing. In addition, the user can still interact through the content editing instruction, ensuring the user experience.
[0062] Figure 5 It is a schematic flowchart of an image generation method in a certain situation. On the above basis, the step of "receiving a first input and obtaining a first content" is further refined. The specific implementation situation can be seen in the detailed description of this situation. Among them, the same or similar technical features as the foregoing are not described herein again.
[0063] As Figure 5 shown, the image generation method includes the following steps: S410. Receive a content editing instruction and a first image to obtain a first content; or, receive a content editing instruction to obtain a first content.
[0064] Among them, the first inputs corresponding to different actual requirements are different. Optionally, the actual requirements may include: image generation requirements and image editing requirements.
[0065] In the case of image generation requirements, the first input includes a content editing instruction composed of natural language. For example, the content editing instruction may be "generate an image of the character 'King' composed of 13 watermelons".
[0066] In the case of image editing requirements, the first input may include a content editing instruction composed of natural language and a first image corresponding to the content editing instruction. The first image may be an input image transmitted by the user clicking a corresponding control on the first page. For example, the first image may be an image including multiple books, and the content editing instruction may be "delete all books".
[0067] It should be noted that there are two ways to obtain the first image corresponding to the content editing command: Firstly, the first image can be a frame from a real-time captured image or video. When the user triggers the corresponding camera control on the first page, the camera is invoked to capture an image or video of the surrounding environment, and a frame from the acquired image or video is used as the first image. Secondly, the first image can be a frame from a pre-captured image or video, transmitted to the first page via the upload control on the first page, and used as the first image.
[0068] To better match the output of the intelligent system with actual needs, prompts can be configured for the system. These prompts represent a series of processing operations performed on the input to the intelligent system. For inputs under different actual needs, the intelligent system can use two types of prompts for targeted processing.
[0069] The first type of prompt information corresponds to the first input under the image editing requirement. That is, the first type of prompt information corresponding to the image editing requirement can be called the first prompt information. The second type of prompt information corresponds to the first input under the image generation requirement. That is, the second type of prompt information corresponding to the image generation requirement can be called the second prompt information.
[0070] Accordingly, under image editing requirements, the first content output by the intelligent system represents the result obtained by analyzing the content editing instructions and the first image based on the first prompt information. Under image generation requirements, the first content output by the intelligent system represents the result obtained by analyzing the content editing instructions based on the second prompt information.
[0071] Specifically, the following describes how the intelligent system processes the first input in conjunction with corresponding prompts under different practical needs. When the practical need is image editing, the first prompt is obtained. Based on the first prompt, the intelligent system analyzes the received content editing instructions and the first image to obtain the first content. When the practical need is image generation, the second prompt is obtained. Based on the second prompt, the intelligent system analyzes the received content editing instructions to obtain the first content.
[0072] In one scenario, the first content is used at least to indicate first planning and layout information for the first media element. The first planning and layout information is presented in a first format file; the first planning and layout information includes the coordinates of the first media element and processing instructions.
[0073] The first format file can be a file in a preset format. Optionally, to make the first content presentation more organized and improve the accuracy of subsequent processing, the first format can be a JSON file. It should be noted that the first planning layout information can also be presented in a regular text format, but to ensure the accuracy of the output image, the first planning layout information is usually presented in a JSON format file.
[0074] The first planning layout information includes two aspects. First, the coordinates of the first media element. These coordinates represent the arrangement position of the first media element. Typically, to achieve precise positioning of the first media element, the coordinates can be the center coordinates or corner coordinates of the image area to which the first media element belongs. Alternatively, to improve the accuracy of the subsequently output image, multiple boundary coordinates corresponding to the element's shape can be selected as coordinates. It should be noted that the coordinates of the first media element in the first planning layout information can be represented by pixel coordinates or UV coordinates, as long as the spatial position of the first media element can be clearly defined.
[0075] Second, there are the processing instructions for the first media element. These instructions characterize the specific operations performed on the first media element. It should be noted that there are corresponding processing instructions for each first media element.
[0076] Specifically, when the actual need is image editing, a first prompt is obtained. Based on the intelligent system's analysis of the received content editing instructions and the first image using the first prompt, at least one first media element and a corresponding processing instruction for each first media element are obtained. The at least one first media element and the corresponding processing instruction for each media element are used as the first content, and the first content is presented in the form of a first format file.
[0077] When the actual requirement is image generation, a second prompt is obtained. Based on the intelligent system and the second prompt, the received content editing instructions are analyzed to obtain at least one first media element and its corresponding processing instructions, i.e., the first content is obtained. The first content is then presented in a first format file.
[0078] For example, the following examples illustrate the processing procedures of the intelligent system under image generation and image editing requirements, respectively.
[0079] When the actual requirement is an image generation requirement, take the content editing instruction "generate an image of the character 'King' composed of 13 watermelons" as an example for illustration. Based on the intelligent system combined with the second prompt information, analyze the received content editing instruction, obtain the coordinates corresponding to 13 first media elements (the elements corresponding to "watermelon"), and assign a processing instruction of "generate one watermelon" to each first media element. Output the 13 coordinates and the processing instruction corresponding to each coordinate in the form of a JSON list.
[0080] When the actual requirement is an image editing requirement, take the first image as an image including multiple books, and the content editing instruction is "delete all books" as an example for illustration. The intelligent system combines the first prompt information to process the received first image and content editing instruction, obtains the coordinates corresponding to all "books" in the first image, and assigns a processing instruction of "delete books" to each first media element. Output the coordinates corresponding to all "books" and the corresponding processing instruction in the form of a JSON list.
[0081] It should be noted that to ensure the position accuracy of the generated image or the edited image at different resolutions, the planned layout of the first media element can be determined on a preset normalized canvas first, and the coordinates can be scaled according to the expected output image resolution to obtain the coordinates in the first planned layout information.
[0082] Take "generate an image including 10 eggs" as an example for illustration. Determine the image area where eggs need to be generated (that is, a total of 10 areas) according to the preset normalized canvas, and use the center point coordinates of each image area as the coordinates to be scaled. According to the expected generated image resolution, scale the coordinates to be scaled to obtain the coordinates corresponding to each image area that match the image resolution.
[0083] In addition, it should also be noted that the first content can be selected to be visible to the user according to the actual requirement. [[ID=...]]
[0084] S420. Obtain the second content based on the data unit and the first content. The first content is at least used to indicate the first planned layout information of the first media element, and the second content is used to represent the spatial position information of the first media element.
[0085] S430. Obtain the third content based on the second content. The third content is provided by the first machine learning model and is used to indicate the output image corresponding to the first content.
[0086] The beneficial effects of the above method are as follows: First content is obtained by receiving content editing instructions and a first image; or, second content is obtained by receiving content editing instructions and obtaining the first content. Based on the data unit and the first content, second content is obtained, where the first content at least indicates the first planning layout information of the first media element, and the second content represents the spatial position information of the first media element. Third content is obtained based on the second content. By processing different first inputs separately to obtain corresponding first content, the above method can meet the image generation or image editing needs in practical application scenarios, thus improving the applicability of the scenarios. Simultaneously, by using an intelligent system to perform targeted processing on different first inputs, the arrangement position control of the first media element is achieved, improving the accuracy of subsequent image output.
[0087] Figure 6 This is a flowchart illustrating an image generation method under one scenario. Building upon the above, and assuming the first planning layout information in the first content includes the coordinates of the first media element, the step "obtaining the second content based on the data unit and the first content" is further refined. For a detailed explanation of its specific implementation, please refer to the detailed description of this scenario. Technical features identical or similar to those described above will not be repeated here.
[0088] like Figure 6 As shown, the image generation method includes the following steps: S510. Based on the coordinates of the first media element, obtain the image position code.
[0089] Image location encoding can be understood as a spatial feature vector corresponding to coordinates that can be recognized by the first machine learning model.
[0090] Alternatively, one way to convert coordinates into image location encoding is: using The coordinates of the first media element can be determined in the following way. This is mapped to a d-dimensional spatial feature vector, i.e., image location encoding. Specifically, the first d / 2 dimensions are used to encode u, and the last d / 2 dimensions are used to encode v. , in, K is a preset base number, such as K can be 10000.
[0091] Accordingly, based on the above-described method, v is encoded to obtain the image position encoding representation corresponding to the coordinates as follows: ; in, This represents the encoded information obtained by encoding u. This represents the encoded information obtained by encoding v. express and Obtained by splicing, with coordinates The corresponding image location encoding.
[0092] Specifically, for the coordinates of the first media element in the first planning layout information, the coordinates of the first media element are encoded to determine the image position code corresponding to the first media element.
[0093] S520: Based on image location encoding and data units, the second content is obtained.
[0094] To achieve rapid identification or precise positioning of media elements, image location encoding can be concatenated or bound to data units. Data units can be MetaPoint Tokens (MPTs). Specifically, a MetaPoint Token is a basic token used to represent key points or regions in an image, and the image location encoding is determined based on the coordinates of the first media element. By overlaying and fusing the image location encoding and MetaPoint Tokens, a feature vector with spatial location information, i.e., the second content, is obtained.
[0095] The second content is used to represent the spatial location information of the first media element. The first machine learning model analyzes the second content, establishes a correlation between actual needs and image spatial location, and achieves accurate positioning of the first media element.
[0096] Specifically, feature overlay processing is performed through data units and image location encoding to obtain the second content, which is then processed by the first machine learning model to improve positioning accuracy.
[0097] The above processing reuses the position encoding logic during image segmentation, that is, assigning a two-dimensional position encoding logic to each image block during image segmentation without adding new encoding logic, and binding the data unit with the image position encoding. This improves the processing speed of the subsequent first machine learning model while achieving fine and precise control over the spatial position of the first media element.
[0098] For example, see Figure 7 If the content editing instruction is "generate an image of a red aluminum can", the intelligent system processes the instruction to obtain at least one coordinate and its corresponding processing instruction. Connecting these coordinates yields... Figure 7 (a) or Figure 7 The blue rectangle shown in (b).
[0099] If based on the first machine learning model, the content editing instructions and Figure 7(a) The coordinates corresponding to the blue rectangle are processed, and the output image is usually as follows: Figure 7 As shown in (a). To improve the accuracy of controlling the spatial position of media elements, one can... Figure 7 (b) The coordinates corresponding to the blue rectangle shown are used for position encoding to obtain the image position code. The MetaPoint Token is then concatenated with the two-dimensional image position code to obtain structured information containing both the MetaPoint Token and the image position code. This structured information can be found in [reference needed]. Figure 7 (b) The coordinates shown below. Based on the first machine learning model, this structured information is processed, and the output is similar to... Figure 7 (b) shows the output image.
[0100] S530. Based on the second content, obtain the third content, which is provided by the first machine learning model and is used to indicate the output image corresponding to the first content.
[0101] The beneficial effects of the above method are as follows: Based on the coordinates of the first media element, an image position code is obtained. According to the image position code and data units, the second content is obtained. The second content is processed using a first machine learning model to obtain an output image corresponding to the first content. This solves the problem of inaccurate description of image numbers, quantities, and complex spatial structures by traditional machine learning models, achieving refined and precise control over the spatial position of media elements during image generation or editing, thus improving the accuracy of image generation or editing.
[0102] Figure 8 This is a flowchart illustrating an image generation method in one scenario. Based on the above, the step "obtaining the third content based on the second content" is further refined. For a detailed explanation of its implementation, please refer to the detailed description of this scenario. Technical features that are the same as or similar to those described above will not be repeated here.
[0103] like Figure 8 As shown, the image generation method includes the following steps: S610. Based on the data unit and the first content, obtain the second content, wherein the first content is used to indicate the first planning layout information of the first media element, and the second content is used to characterize the spatial location information of the first media element.
[0104] S620. Based on the second content, the first content, and the first input, obtain the second input, which is used at least to indicate the task that the first machine learning model needs to perform.
[0105] Where the first input differs according to the actual needs, the corresponding second input will also differ. The following explains how to obtain the second input under image generation and image editing needs, respectively.
[0106] Specifically, in cases where the actual need is for image generation, the second input is obtained based on the second content, the content editing instructions of the first input, and the coordinates and processing instructions of the first media element in the first content, so as to use the second input as the input of the first machine learning model.
[0107] In cases where the actual need is for image editing, the second input is obtained based on the second content, the content editing instructions of the first input, the coordinates and processing instructions of the first media element in the first image and the first content, so as to use the second input as the input of the first machine learning model.
[0108] In one scenario, the second input is obtained by: obtaining the second input based on the second content, the processing instructions of the first media element in the first content, and the second image.
[0109] It should be noted that, in cases where the actual requirement is image generation, a noisy image can be obtained before generating the output image. This noisy image can then be progressively denoised using a first machine learning model to obtain the output image. That is, the second image includes one of the following: the first image in the first input; or a noisy image obtained when the first input does not include the first image. In other words, for image generation, the second image is a noisy image; for image editing, the second image is the first image as input.
[0110] Specifically, in cases where the actual requirement is image generation, the processing instructions for the second content, the first media element in the first content, and the noisy image are used as the second input.
[0111] In cases where the actual requirement is image editing, the processing instructions for the second content, the first media element in the first content, and the first image are used as the second input.
[0112] Optionally, in order to improve the processing speed of the first machine learning model, the obtained second input can be adjusted using a preset template to standardize the second input.
[0113] S630: Based on the second input, obtain the third content.
[0114] Specifically, when the second input differs depending on the specific needs, the processing of the second input by the first machine learning model will also differ. The following describes the processing of the second input by the first machine learning model under image generation and image editing needs, respectively.
[0115] For image generation, the second content, the processing instructions for the first media element, and the noisy image are used as the second input and fed into the first machine learning model. This allows the first machine learning model to locate the corresponding first media element and determine its processing operation based on the second content and the processing instructions for the first media element from the second input. The noisy image is then iteratively denoised according to this processing operation to obtain the output image, i.e., the third content.
[0116] For image editing needs, the second content, the processing instructions for the first media element, and the first image are used as the second input and fed into the first machine learning model. This allows the first machine learning model to locate the first media element in the first image and determine its processing operation based on the second content and the processing instructions for the first media element. The first media element in the first image is then processed according to this processing operation to obtain the output image, i.e., the third content.
[0117] The beneficial effects of the above method are as follows: Based on the data unit and the first content, the second content is obtained. Based on the second content, the first content, and the first input, the second input is obtained. Based on the second input and the first machine learning model, the third content is obtained. This optimizes the input information of the first machine learning model, avoiding the problem of mismatch between the output results and actual needs caused by directly inputting content editing instructions into the machine learning model, thus improving the accuracy of image generation or editing. Furthermore, by optimizing the model input rather than modifying the model architecture, the technical implementation threshold and cost can be reduced, improving the efficiency of technology application.
[0118] Figure 9 This is a flowchart illustrating the training process of the first machine learning model under one scenario. Based on the above scenario, to ensure that the first machine learning model can accurately output the image corresponding to the content editing instructions, it can be trained first. The specific training process for the first machine learning model can be found in the detailed description of this scenario. Technical features that are the same as or similar to those described above will not be repeated here.
[0119] like Figure 9 As shown, the image generation method includes the following steps: S710. Obtain multiple samples, including the third input and the first result.
[0120] To improve the accuracy of machine learning models, it is important to obtain as many and varied samples as possible. Each sample can include a third input and a first result.
[0121] The third input is used to indicate at least the second planning layout information. The second planning layout information indicates the overall arrangement of the media elements corresponding to the sample. The second planning layout information may include the coordinates of the media elements in the sample and processing instructions. The second planning layout information can be user-edited or generated based on an intelligent system. The first result indicates the expected output of the third input. This can be understood as the first result characterizing the desired output image.
[0122] Specifically, to ensure the diversity of input samples, multiple samples can be obtained under different practical needs, and the model can be trained based on these samples. The following describes the sample acquisition under different practical needs.
[0123] In cases where the actual requirement is image generation, sample content editing instructions corresponding to image generation are obtained. These instructions are then processed by an intelligent system to obtain second planning layout information, i.e., the third input, used to indicate sample media elements. Finally, a first result matching this third input is determined.
[0124] In cases where the actual need is image editing, sample content editing instructions and sample input images corresponding to the image editing are obtained. The intelligent system processes the sample input image and sample content editing instructions to obtain second planning layout information, i.e., the third input, used to indicate media elements in the sample input image. A first result matching this third input is then determined.
[0125] S720 processes the third input based on the second machine learning model and outputs the prediction result.
[0126] The second machine learning model is a pre-built generative model used for image generation or editing. Optionally, the second machine learning model can be a diffusion model architecture. Under the diffusion model architecture, the second machine learning model processes information from a noisy image or a sample input image based on a third input to output a prediction result. The prediction result is used to indicate the output image of the second machine learning model.
[0127] Specifically, based on the data units and the third input, the spatial location information of media elements in the sample is obtained. The spatial location information of the sample media elements is then processed using a second machine learning model to obtain prediction results.
[0128] S730. Based on the prediction results and the first result, adjust the model parameters in the second machine learning model.
[0129] Generally, the parameters of the second machine learning model are usually the initial parameters or default parameters. When training the second machine learning model, the various model parameters can be adjusted based on the prediction results and the loss value determined by the first result. The loss value is used to characterize the degree of difference between the prediction results and the first result.
[0130] Specifically, loss processing is performed based on the prediction result and the first result to determine the loss value. The model parameters of the second machine learning model are then adjusted based on the loss value to obtain a trained second machine learning model.
[0131] When using the loss value to refine the model parameters in the second machine learning model, the convergence of the loss function can be used as a training objective. This could be achieved by checking if the training error is less than a preset error, if the error trend stabilizes, or if the current number of iterations equals a preset number. If convergence is achieved—for example, if the training error of the loss function is less than the preset error, or if the error trend stabilizes—it indicates that the second machine learning model has been successfully trained, and iterative training can be stopped. If convergence has not yet been achieved, additional samples can be acquired to continue training the second machine learning model until the training error of the loss function is within a preset range. When the training error of the loss function converges, the successfully trained second machine learning model is obtained.
[0132] S740, in response to the completion of training of the second machine learning model, uses the second machine learning model as the first machine learning model.
[0133] Specifically, the trained second machine learning model is used as the first machine learning model, and based on the first machine learning model, an output image matching the content editing instructions is output.
[0134] The beneficial effects of the above method are as follows: By acquiring multiple samples, the third input is processed based on the second machine learning model to output a prediction result. Based on the prediction result and the first result, the model parameters in the second machine learning model are adjusted. In response to the completion of the training of the second machine learning model, the second machine learning model is used as the first machine learning model. By training the second machine learning model and using the trained second machine learning model as the first machine learning model, support is provided for the application of the first machine learning model, ensuring the accuracy of image editing or image generation based on the first machine learning model.
[0135] Figure 10 This is a schematic diagram of the structure of an image generation device in one scenario, such as... Figure 10 As shown, the device includes a second content acquisition module 810 and a third content acquisition module 820.
[0136] The second content acquisition module 810 is used to acquire second content based on the data unit and the first content, wherein the first content is used to indicate the first planning layout information of the first media element and the second content is used to characterize the spatial location information of the first media element; the third content acquisition module 820 is used to acquire third content based on the second content, wherein the third content is provided by the first machine learning model and is used to indicate the output image corresponding to the first content.
[0137] Optionally, in one scenario, the device further includes: a first content acquisition module, configured to receive a first input and acquire first content, wherein the first input includes at least content editing instructions composed of natural language, and the first content is provided by an intelligent system.
[0138] In another scenario, optionally, the first content acquisition module is configured to receive a content editing instruction and a first image, and acquire the first content, wherein the first content is provided by an intelligent system and is used to characterize the result obtained by analyzing the content editing instruction and the first image based on a first prompt information; or, to receive a content editing instruction and acquire the first content, wherein the first content is provided by an intelligent system and is used to characterize the result obtained by analyzing the content editing instruction based on a second prompt information.
[0139] In another scenario, the first planning layout information is presented in a first format file, which includes the coordinates of the first media element and processing instructions.
[0140] In another scenario, optionally, the first planning layout information includes the coordinates of the first media element, and the second content acquisition module is used to obtain an image location code based on the coordinates of the first media element; and to obtain the second content based on the image location code and the data unit.
[0141] In another scenario, optionally, the third content acquisition module includes: a second input acquisition unit, configured to acquire a second input based on the second content, the first content, and the first input, wherein the second input is at least used to indicate the task to be performed by the first machine learning model; and a third content acquisition unit, configured to acquire the third content based on the second input.
[0142] In another scenario, optionally, the second input obtaining unit is configured to obtain the second input based on the second content, the processing instructions of the first media element in the first content, and the second image; wherein the second image includes one of the following: the first image in the first input; or a noisy image obtained when the first input does not include the first image.
[0143] In another optional embodiment, the above apparatus further includes: a model training module for acquiring multiple samples, the samples including a third input and a first result, the third input being used to indicate at least second planning layout information, and the first result being used to indicate the expected output of the third input; processing the third input based on a second machine learning model to output a prediction result; adjusting the model parameters in the second machine learning model based on the prediction result and the first result; and, in response to the completion of training of the second machine learning model, using the second machine learning model as the first machine learning model.
[0144] The beneficial effects of the aforementioned device are as follows: Based on the data unit and the first content, second content is obtained. The first content is used to indicate at least the first planning layout information of the first media element, and the second content is used to characterize the spatial position information of the first media element. Based on the second content, third content is obtained. The third content is provided by the first machine learning model and is used to indicate the output image corresponding to the first content. This solves the problem that the output results of the machine learning model do not match the requirements during image generation or editing, and realizes refined and precise control over the spatial position of image content during image generation or image editing, thereby improving the accuracy of image generation or image editing.
[0145] The image generation apparatus provided in this paper can execute any of the image generation methods provided in this paper, and has the corresponding functional modules and beneficial effects of executing the methods.
[0146] It is worth noting that the various units and modules included in the above-mentioned device are divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the scope of protection of this document.
[0147] The following is for reference. Figure 11 This document illustrates a schematic diagram of an electronic device (e.g., a terminal device or server) 900 suitable for implementing the above-described methods. The terminal device referred to herein may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital radio receivers, personal digital assistants (PDAs), tablet computers (PADs), portable multimedia players (PMPs), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital televisions and desktop computers. Figure 11 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments described herein.
[0148] like Figure 11 As shown, the electronic device 900 may include a processing unit (e.g., a central processing unit, a graphics processor, etc.) 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage device 908 into a random access memory (RAM) 903. The RAM 903 also stores various programs and data required for the operation of the electronic device 900. The processing unit 901, ROM 902, and RAM 903 are interconnected via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0149] Typically, the following devices can be connected to I / O interface 905: input devices 906 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 907 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 908 including, for example, magnetic tapes, hard disks, etc.; and communication devices 909. Communication device 909 allows electronic device 900 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 11 An electronic device 900 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0150] In particular, according to embodiments of this document, the processes described in the above-referenced flowcharts can be implemented as computer software programs. For example, the technical solutions of this document include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device 909, or installed from a storage device 908, or installed from a ROM 902. When the computer program is executed by a processing device 901, it performs the functions defined in the methods of the embodiments of this document.
[0151] The names of messages or information exchanged between multiple devices in this document are for illustrative purposes only and are not intended to limit the scope of these messages or information.
[0152] The electronic device provided in this embodiment and the image generation method provided in the above technical solutions belong to the same inventive concept. Technical details not described in detail in this document can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.
[0153] This article provides a computer storage medium on which a computer program is stored, which, when executed by a processor, implements the image generation method provided in the above embodiments.
[0154] It should be noted that the computer-readable medium mentioned above can be a computer-readable signal medium, a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory (EPROM, also known as flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this document, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. The transmitted data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0155] Based on one or more scenarios described herein, an image generation method is provided, comprising: Based on the data unit and the first content, a second content is obtained, wherein the first content is used to indicate the first planning layout information of the first media element, and the second content is used to characterize the spatial location information of the first media element. Based on the second content, a third content is obtained, which is provided by a first machine learning model and is used to indicate the output image corresponding to the first content.
[0156] According to one or more scenarios described herein, Example 2 provides a method for image generation, the method further comprising: receiving a first input and obtaining first content, wherein the first input includes at least content editing instructions composed of natural language, and the first content is provided by an intelligent system.
[0157] Based on one or more scenarios described herein, Example 3 provides a method for image generation, wherein receiving a first input and obtaining first content includes: The system receives a content editing instruction and a first image to obtain the first content, which is provided by an intelligent system and is used to characterize the result obtained by analyzing the content editing instruction and the first image based on a first prompt information. The system receives a content editing instruction and obtains the first content, which is provided by the intelligent system and is used to represent the result obtained by analyzing the content editing instruction based on the second prompt information.
[0158] According to one or more scenarios described herein, Example 4 provides an image generation method, wherein the first planning layout information is presented in a first format file, and the first planning layout information includes the coordinates of the first media element and processing instructions.
[0159] Based on one or more scenarios described herein, Example 5 provides an image generation method, wherein the first planning layout information includes the coordinates of the first media element, and the step of obtaining the second content based on the data unit and the first content includes: Based on the coordinates of the first media element, the image position code is obtained; The second content is obtained based on the image location encoding and the data unit.
[0160] Based on one or more scenarios described herein, Example Six provides an image generation method, wherein obtaining third content based on the second content includes: Based on the second content, the first content, and the first input, a second input is obtained, which is at least used to indicate the task that the first machine learning model needs to perform. Based on the second input, the third content is obtained.
[0161] Based on one or more scenarios described herein, Example 7 provides an image generation method, wherein obtaining the second input based on the second content, the first content, and the first input includes: The second input is obtained based on the second content, the processing instructions of the first media element in the first content, and the second image; The second image includes one of the following: the first image in the first input; or a noisy image obtained when the first input does not include the first image.
[0162] Based on one or more scenarios described herein, Example 8 provides an image generation method, wherein the first machine learning model is obtained based on the following: Multiple samples are obtained, the samples including a third input and a first result, the third input being used to indicate at least the second planning layout information, and the first result being used to indicate the expected output of the third input; The third input is processed based on the second machine learning model, and a prediction result is output. Based on the prediction results and the first results, adjust the model parameters in the second machine learning model; In response to the completion of training of the second machine learning model, the second machine learning model is used as the first machine learning model.
[0163] According to one or more of the provisions of this article, Example 9 provides an image generation apparatus, comprising: The second content acquisition module is used to acquire second content based on the data unit and the first content, wherein the first content is used to indicate the first planning layout information of the first media element, and the second content is used to characterize the spatial location information of the first media element. The third content acquisition module is used to obtain third content based on the second content, wherein the third content is provided by the first machine learning model and is used to indicate the output image corresponding to the first content.
[0164] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol, such as HTTP (Hypertext Transfer Protocol), and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include Local Area Networks (LANs), Wide Area Networks (WANs), the Internet (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0165] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0166] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: Based on the data unit and the first content, a second content is obtained, wherein the first content is used to indicate the first planning layout information of the first media element, and the second content is used to characterize the spatial location information of the first media element. Based on the second content, a third content is obtained, which is provided by a first machine learning model and is used to indicate the output image corresponding to the first content.
[0167] Computer program code for performing the operations described herein can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" programming language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including local area networks (LANs) or wide area networks (WANs), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0168] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this document. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0169] The modules or units described herein can be implemented in software or hardware. The names of modules or units do not necessarily limit the functionality of the module or unit itself; for example, a second result acquisition unit can also be described as a "second result receiving unit".
[0170] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include at least one of the following: Field-Programmable Gate Array (FPGA), Application-Specific Integrated Circuit (ASIC), Application-Specific Standard Product (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), etc.
[0171] In the context of this document, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory (flash memory), optical fibers, portable compact disk read-only memory, optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0172] The above description is merely a preferred embodiment and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of disclosure herein is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed herein that have similar functions.
[0173] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be interpreted as limiting the scope of this document. Certain features described in the context of individual implementations may also be implemented in combination in a single implementation. Conversely, various features described in the context of a single implementation may also be implemented individually or in any suitable sub-combination in multiple implementations.
[0174] Although the subject matter has been described using a programming language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative forms of implementing the claims.
Claims
1. An image generation method, comprising: Based on the data unit and the first content, a second content is obtained, wherein the first content is used to indicate the first planning layout information of the first media element, and the second content is used to characterize the spatial location information of the first media element. Based on the second content, a third content is obtained, which is provided by a first machine learning model and is used to indicate the output image corresponding to the first content.
2. The image generation method according to claim 1, further comprising: Receive a first input and obtain first content, wherein the first input includes at least content editing instructions composed of natural language, and the first content is provided by an intelligent system.
3. The image generation method according to claim 2, wherein receiving the first input and obtaining the first content includes: The system receives a content editing instruction and a first image to obtain the first content, which is provided by an intelligent system and is used to characterize the result obtained by analyzing the content editing instruction and the first image based on a first prompt information. The system receives a content editing instruction and obtains the first content, which is provided by the intelligent system and is used to represent the result obtained by analyzing the content editing instruction based on the second prompt information.
4. The image generation method according to claim 1, wherein the first planning layout information is presented in a first format file, and the first planning layout information includes the coordinates of the first media element and processing instructions.
5. The image generation method according to claim 1, wherein the first planning layout information includes the coordinates of the first media element, and the step of obtaining the second content based on the data unit and the first content includes: Based on the coordinates of the first media element, the image position code is obtained; The second content is obtained based on the image location encoding and the data unit.
6. The image generation method according to claim 1, wherein obtaining the third content based on the second content includes: Based on the second content, the first content, and the first input, a second input is obtained, which is at least used to indicate the task that the first machine learning model needs to perform. Based on the second input, the third content is obtained.
7. The image generation method according to claim 6, wherein obtaining the second input based on the second content, the first content, and the first input includes: The second input is obtained based on the second content, the processing instructions of the first media element in the first content, and the second image; The second image includes one of the following: the first image in the first input; or a noisy image obtained when the first input does not include the first image.
8. The image generation method according to claim 1, wherein the first machine learning model is obtained based on the following method: Multiple samples are obtained, the samples including a third input and a first result, the third input being used to indicate at least the second planning layout information, and the first result being used to indicate the expected output of the third input; The third input is processed based on the second machine learning model, and a prediction result is output. Based on the prediction results and the first results, adjust the model parameters in the second machine learning model; In response to the completion of training of the second machine learning model, the second machine learning model is used as the first machine learning model.
9. An image generation apparatus, comprising: The second content acquisition module is used to acquire second content based on the data unit and the first content, wherein the first content is used to indicate the first planning layout information of the first media element, and the second content is used to characterize the spatial location information of the first media element. The third content acquisition module is used to obtain third content based on the second content, wherein the third content is provided by the first machine learning model and is used to indicate the output image corresponding to the first content.
10. An electronic device, the electronic device comprising: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the image generation method as described in any one of claims 1-8.
11. A storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the image generation method as described in any one of claims 1-8.
12. A computer program product comprising a computer program that, when executed by a processor, implements the image generation method as described in any one of claims 1-8.