Method and apparatus for generating image, device, medium and program product
By generating content description text and template elements, and combining them with machine learning models, the problem of low efficiency and poor relevance in user-generated content image generation has been solved, achieving more efficient and accurate image generation and improving the user experience.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-05
- Publication Date
- 2026-03-12
AI Technical Summary
In existing technologies, the image generation process for user-generated content is inefficient and lacks personalization, resulting in a poor user experience. Fixed templates and materials have poor correlation, failing to meet diverse needs.
By generating content description text for user-generated content, using machine learning models to generate template elements, and combining them with user-generated content, a target composite image is generated, improving relevance and efficiency.
The generated target combination images are more accurate and highly correlated, improving image generation efficiency and enhancing user experience.
Smart Images

Figure CN2024117276_12032026_PF_FP_ABST
Abstract
Description
Method, device, equipment, medium and program product for generating image TECHNICAL FIELD
[0001] Embodiments of the present disclosure generally relate to the field of image processing, and in particular, to a method, device, equipment, medium and program product for generating image. BACKGROUND
[0002] At present, machine learning is becoming more and more important in people's daily life and work, and gradually becomes an indispensable important tool for people. More and more work begins to use machine learning models to process. For example, word processing work, picture processing work, video processing work, etc. begin to use machine learning models to process. Especially when processing multi-modal type data, machine learning models with multi-modal data processing capability have obvious advantages.
[0003] With the rapid development of machine learning technology, when processing various multi-modal data, the processing process has become faster and more accurate. For example, when performing text and image processing work, a multi-modal machine learning model can be used to help users process text and image related operations. In addition, in order to meet the development needs of text and image processing technology, machine learning models are also used more and more in image generation.
[0004] SUMMARY
[0005] Embodiments of the present disclosure provide a method, device, equipment, medium and program product for generating image.
[0006] According to a first aspect of the present disclosure, a method for generating image is provided. The method comprises generating a content description text for user-generated content based on the user-generated content. The method further comprises generating a set of template elements of a template for the user-generated content based on the content description text. The method further comprises generating a target combined image based on the set of template elements and the user-generated content.
[0007] In a second aspect of the present disclosure, a device for generating image is provided. The device comprises a content description text generation module configured to generate a content description text for user-generated content based on the user-generated content; a set of template elements generation module configured to generate a set of template elements of a template for the user-generated content based on the content description text; and a target combined image generation module configured to generate a target combined image based on the set of template elements and the user-generated content.
[0008] In a third aspect of the present disclosure, an electronic device is provided, comprising at least one processor; and a storage device for storing at least one program, which when executed by the at least one processor, causes the at least one processor to implement the method according to the first aspect of the present disclosure.
[0009] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, having stored thereon a computer program, which when executed by a processor, implements the method according to the first aspect of the present disclosure.
[0010] In a fifth aspect of the present disclosure, a computer program product is provided. The computer program product comprises a computer program which, when executed by a processor, implements the method according to the first aspect of the present disclosure.
[0011] It should be understood that the contents described in this section are not intended to limit key or important features of the embodiments of the present disclosure, nor to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0012] The above and other objects, features and advantages of the present disclosure will become more apparent from the following description when taken in conjunction with the accompanying drawings, in which like reference characters designate the same components in several views.
[0013] FIG. 1 illustrates a schematic diagram of an example environment in which devices and / or methods of some embodiments of the present disclosure can be implemented;
[0014] FIG. 2 illustrates a schematic diagram of an example method for generating an image according to some embodiments of the present disclosure;
[0015] FIG. 3 illustrates a schematic diagram of a flowchart of a flow for generating an image according to some embodiments of the present disclosure;
[0016] FIG. 4 illustrates a schematic diagram of one example for generating an image according to some embodiments of the present disclosure;
[0017] FIG. 5 illustrates a schematic diagram of another example for generating an image according to some embodiments of the present disclosure;
[0018] FIG. 6 illustrates a schematic diagram of yet another example for generating an image according to some embodiments of the present disclosure;
[0019] FIG. 7 illustrates a schematic diagram of still another example for generating an image according to some embodiments of the present disclosure;
[0020] FIG. 8 illustrates a schematic diagram of a combined image containing identification information for generating an image according to some embodiments of the present disclosure;
[0021] FIG. 9 illustrates a schematic diagram of one example of a user-generated content and a set of template materials for generating an image, according to some embodiments of the present disclosure;
[0022] FIG. 10 illustrates a schematic diagram of one specific embodiment for generating an image, according to some embodiments of the present disclosure;
[0023] FIG. 11 illustrates a schematic diagram of an example for generating a summary description text, according to some embodiments of the present disclosure;
[0024] FIG. 12 illustrates a schematic diagram of an example of a training process of an image generation model for generating an image, according to some embodiments of the present disclosure;
[0025] FIG. 13 illustrates a schematic block diagram of an apparatus for generating an image, according to some embodiments of the present disclosure;
[0026] FIG. 14 illustrates a schematic block diagram of an example device suitable for implementing various embodiments of the present disclosure.
[0027] In the various drawings, same or corresponding reference numbers denote same or corresponding parts. DETAILED DESCRIPTION
[0028] It can be understood that the data involved in the technical solutions of the present disclosure (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the relevant laws and regulations and relevant provisions.
[0029] It can be understood that, before using the technical solutions disclosed in various embodiments of the present disclosure, the type of personal information involved in the present disclosure, the scope of use, the scenario of use, etc. should be informed to the user and the authorization of the user should be obtained through appropriate means according to relevant laws and regulations.
[0030] For example, when receiving an active request of a user, a prompt information is sent to the user to explicitly prompt the user that the operation requested to be executed will need to acquire and use the personal information of the user. Thus, the user can voluntarily choose whether to provide the personal information to the software or hardware such as an electronic device, an application program, a server or a storage medium, etc. that executes the operation of the technical solutions of the present disclosure according to the prompt information.
[0031] As an optional but non-limiting implementation manner, in response to receiving an active request of a user, the manner of sending a prompt information to the user may, for example, be a pop-up window manner, in which the prompt information can be presented in a text manner. In addition, the pop-up window can also carry a selection control for the user to select "agree" or "disagree" to provide the personal information to the electronic device.
[0032] It can be understood that the above notification and user authorization obtaining process is only illustrative and does not limit the implementation of the present disclosure, and other ways that meet relevant laws and regulations can also be applied to the implementation of the present disclosure.
[0033] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms, and should not be interpreted as being limited to the embodiments set forth herein, but rather, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes, and are not intended to limit the scope of protection of the present disclosure.
[0034] In the description of embodiments of the present disclosure, the term "comprising" and similar terms are understood to encompass open-ended inclusion, i.e., "comprising but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", and the like can refer to different or the same objects. Other explicit and implicit definitions can also be included below.
[0035] In the process of generating images, there are still many problems to be solved. For example, usually when a user wants to further generate personalized content creation according to existing information (such as user-generated content, which is multi-modal, usually video or image), the user usually needs to collect relevant information by himself, process the existing information, and sometimes may also need to write relevant scripts by himself. However, this way requires a lot of time and resources, and is low in efficiency, which affects the user experience.
[0036] For example, in a traditional scheme, a user can mine different types of user-generated content (UGC), such as video creation content, image creation content, audio creation content, and text creation content. Then, the user further processes the materials in various types of user-generated content, and splices the processed materials into new materials that can be put out. In this scheme, the user content usually needs to be mined manually, and then the mined user-generated content is refined to find suitable and available materials, and the materials are placed in a fixed template that has been made in advance.
[0037] Generally, the fixed template contains background content, product logo, and text, etc. In addition, it also leaves a blank "slot" for the replacement material in a fixed position, and the position in the fixed template is also fixed. Since the fixed template is manually made, the number of fixed templates is limited, which cannot meet the use demand of a large number of material delivery, and thus appears relatively single and lacks attraction to other users. In addition, the generated fixed template generally has no relevance to the material mined, and various styles and different types of materials often appear in the same fixed template, which leads to a sense of fragmentation between the fixed template and the material in the fixed "slot", and also cannot provide personalized customization for users, greatly reducing the user experience.
[0038] At least to solve the above and other potential problems, embodiments of the present disclosure propose a method for generating an image. In the method, a content description text for user-generated content can be first generated at a computing device. The user-generated content is created by a user. For example, the user-generated content is a video or an image, or a combination of a video and an image. The content description text is obtained by applying the user-generated content to a machine learning model. Then, the computing device further processes the content description text to generate a set of template elements of a template for the user-generated content. Finally, the computing device generates a target combined image using the generated set of template elements and the user-generated content. Through the method, since the content created by the user is used to generate a set of template elements associated with the user-generated content, and the associated user-generated content and the set of template elements are further combined, the target combined image generated is more accurate and has high relevance, and the image generation efficiency is improved and the user experience is improved.
[0039] Embodiments of the present disclosure will be described in detail below with further reference to the accompanying drawings. FIG. 1 shows an example environment in which devices and / or methods of embodiments of the present disclosure can be implemented. In the environment 100, a computing device 102 first generates a content description text 106 for user-generated content 104. The user-generated content 104 is created by a user. For example, the user-generated content 104 is a video or an image, or a combination of a video and an image. The content description text 106 is obtained by processing the user-generated content 104. For example, the user-generated content 104 is applied to a machine learning model to obtain the content description text 106. Then, the computing device 102 generates a set of template elements 108 of a template for the user-generated content 104 using the content description text 106. Finally, after determining the set of template elements 108, the computing device 102 further generates a target combined image 110 in combination with the user-generated content 104.
[0040] Examples of the computing device 102 include, but are not limited to, a personal computer, a server computer, a handheld or laptop device, a mobile device (such as a mobile phone, a personal digital assistant (PDA), a media player, etc.), a multiprocessor system, a consumer electronic product, a minicomputer, a mainframe computer, a distributed computing environment including any of the above systems or devices, etc.
[0041] As shown in FIG. 1, the computing device 102 can first generate a content description text 106 for a user-generated content 104. Wherein, the user-generated content 104 is created by a user. For example, the user-generated content 104 is a video or an image, or a combination of a video and an image. In one example, the user-generated content 104 can be a video, for example, a 10-second video created by a user. In another example, the user-generated content 104 can be an image, for example, an image created by a user through collecting, cropping, splicing, etc. In yet another example, the user-generated content 104 can be a combination of a video and an image, for example, a user-generated content 104 created by a user through splicing a video and an image by using a video editing software. Alternatively, the user-generated content can be an audio or a combination of an audio and a video.
[0042] In some embodiments, the content description text 106 contains information extracted from the user-generated content 104, which can contain one or more of the following: what the user-generated content 104 mainly describes, what the style of the user-generated content 104 is, and what the theme color of the user-generated content 104 is, etc. Additionally, the theme of the user-generated content 104 can be further extracted and determined from the content description text 106, for example, the theme of the user-generated content 104 can be determined as a singing live, a football event, a game entertainment, etc.
[0043] In some embodiments, the content description text 106 is obtained by applying the user-generated content 104 to a machine learning model. For example, the content description text 106 is obtained by applying the user-generated content 104 to a visual model. Additionally, the content description text 106 is obtained by applying the user-generated content 104 to a visual large model. It can be understood that the examples are only used to describe the present disclosure, and are not specific limitations of the present disclosure.
[0044] The computing device 102 then utilizes the content description text 106 to generate a set of template elements 108 for a template of the user-generated content 104. In some embodiments, the set of template elements 108 only contains a background image. In other embodiments, the set of template elements 108 contains a background image and a sticker. In other embodiments, the set of template elements 108 contains a background image and a summary description text. In yet other embodiments, the set of template elements 108 contains a background image, a sticker, and a summary description text. Additionally, the set of template elements 108 can also contain a product logo, a company logo, or a username watermark, etc. For convenience of description and ease of understanding, the summary description text is also referred to as a script.
[0045] In some embodiments, the computing device can utilize an image generation model to process the content description text 106 to generate the background image, the sticker, etc. in the set of template elements 108. In other embodiments, the computing device can utilize a large language model to process the content description text 106 to generate the script for the user-generated content 104.
[0046] Finally, after determining the set of template elements 108, the computing device 102 further generates a target combined image 110 in combination with the user-generated content 104. In some embodiments, the computing device generates more than one sticker and more than one script, and combines the multiple stickers and the multiple scripts according to a predetermined position or a preset rule to generate the set of template elements 108.
[0047] By this method, since the content generated by the user is utilized, a set of template elements associated with the user-generated content is generated, and further the associated user-generated content and the set of template elements are combined to generate a target combined image, which is more accurate and highly relevant, and improves the image generation efficiency and the user experience.
[0048] The above describes a schematic diagram of an example environment in which the device and / or method of some embodiments of the present disclosure can be implemented in conjunction with FIG. 1, and the following describes a schematic diagram of an example method for generating an image according to some embodiments of the present disclosure in conjunction with FIG. 2. The method of FIG. 2 can be performed by the computing device 102 or any suitable computing device in FIG. 1.
[0049] As shown in FIG. 2, in the example method 200, at block 202, the computing device 102 generates a content description text 106 for the user-generated content 104 based on the user-generated content 104. The user-generated content 104 is generated by a user. The user-generated content 104 can be a video or an image, or a combination of a video and an image.
[0050] In some embodiments, the content description text 106 is obtained by applying the user generated content 104 to a machine learning model. For example, the content description text 106 is obtained by applying the user generated content 104 to a visual model. Additionally, the content description text 106 is obtained by applying the user generated content 104 to a visual large model. In some embodiments, there is a predetermined mapping relationship between user generated content and content description text. After obtaining the user generated content 104, the content description text 106 corresponding to the user generated content can be obtained from the predetermined mapping relationship. The above examples are only used to describe the present disclosure, but not specific limitations of the present disclosure.
[0051] Then, at block 204, the computing device 102 generates a set of template elements 108 of a template for the user generated content 104 based on the content description text 106. In order to better display the user generated content, a template combined with the user generated content needs to be determined according to the content description text 106 determined according to the user generated content, and the template is composed of a set of template elements.
[0052] In some embodiments, the computing device can utilize an image generation model to process the content description text 106 to generate a background image, a sticker, and the like in the set of template elements 108. Additionally, the computing device utilizes a machine learning model to process the content description text 106 to generate prompt information for the background image and prompt information for the sticker. For example, the machine learning model is a large language model. Then, the prompt information for the background image and the prompt information for the sticker are input into the image generation model to generate the background image and the sticker. Alternatively, the computing device 102 can also pre-obtain a mapping relationship between description text and background image and a mapping relationship between description text and sticker. Then, after obtaining the content description text 106, the computing device 102 obtains the background image corresponding to the content description text 106 according to the mapping relationship between the description text and the background image. The computing device 102 can also obtain the sticker corresponding to the content description text 106 according to the mapping relationship between the description text and the sticker. The above examples are only used to describe the present disclosure, but not specific limitations of the present disclosure.
[0053] In some embodiments, the computing device can utilize a large language model to process the content description text 106 to generate a script for the user generated content 104. In some embodiments, the computing device 102 can obtain a mapping relationship between description text and script. After obtaining the content description text 106, the computing device 102 can utilize the mapping relationship to find the script corresponding to the description text. The above examples are only used to describe the present disclosure, but not specific limitations of the present disclosure.
[0054] Finally, at block 206, the computing device 102 generates the target combined image 110 based on the set of template elements 108 and the user-generated content 104. After obtaining the set of template elements 108, the set of template elements 108 and the user-generated content 104 can be further combined to generate the target combined image.
[0055] In some embodiments, the target combined image 110 only contains the background image and the user-generated content 104, and the user-generated content 104 is in front of the background image. In this case, the computing device can adjust the size of the user-generated content 104 according to a preset ratio, for example, adjust the user-generated content 104 to occupy 40% of the size of the background image. Additionally, at this time the computing device can place the user-generated content at any suitable position in the background image.
[0056] In some embodiments, the target combined image 110 contains the background image, the sticker, and the user-generated content 104, and the user-generated content 104 and the sticker are also in front of the background image. In this case, the computing device can place the background image and the sticker according to the preset position information of the set of template elements 108, and generate a combined image with the user-generated content 104. In one example, the sticker has no contact with the user-generated content 104 in the background image. In another example, the sticker has partial contact with the user-generated content 104 in the background image, for example, a part of the sticker covers the user-generated content 104 and is displayed in front of the user-generated content 104. Additionally, the computing device can also place the background image and the sticker according to the preset rules of the set of template elements 108, for example, the user can set the sticker to cover the user-generated content 104, such as setting a threshold of 20% of the sticker to cover the user-generated content 104.
[0057] In some embodiments, the target combined image 110 contains the background image, the summary description text, and the user-generated content 104, and the user-generated content 104 and the summary description text are also in front of the background image. In this case, the computing device can place the background image and the summary description text according to the preset position information of the set of template elements 108, and generate a combined image with the user-generated content 104. It can be understood that there can be multiple summary description texts, and they are placed in front of the background image. In one example, the summary description text has no contact with the user-generated content. In another example, the summary description text has partial contact with the user-generated content. Additionally, the computing device can also place the background image and the summary description text according to the preset rules of the set of template elements 108, for example, the user can set the summary description text to cover the user-generated content 104, such as setting a threshold of 10% of the summary description text to cover the user-generated content 104.
[0058] In some embodiments, the target combined image 110 includes a background image, a sticker, a text, and user-generated content, and the user-generated content 104, the sticker, and the text are also located in front of the background image. In this case, the computing device can arrange the user-generated content 104, the sticker, and the text according to the preset position information of the set of template elements 108 and / or preset rules. In this case, the computing device can place the user-generated content 104, the sticker, and the text at the preset position information. The computing device can also place the user-generated content 104, the sticker, and the text according to the preset rules, for example, the user-generated content 104, the sticker, and the text can be placed in contact with each other according to the preset proportion. Additionally, the user can also set the display priority of the user-generated content 104, the sticker, and the text, for example, set the display priority of the text as the highest, in which case the text will not be in contact with the user-generated content 104 and / or the sticker at all times and will always be displayed in the frontmost position of the combined image. The above is only an example and is not a limitation of the present application.
[0059] In some embodiments, the set of template elements 108 includes identification information in addition to the user-generated content 104, the sticker, and the text, and the identification information includes product identification, company identification, or username watermark, etc. In one example, the user can set the transparency of the identification information, for example, set the transparency of the identification information to 50%. In another example, the user can set the identification information to be bold or highlighted to highlight the identification information. In yet another example, the user can set the position information of the identification information, for example, set the identification information to be located in the four corners of the combined image. Additionally, the user can set a range centered on the identification information within which no other elements except the background image are allowed, for example, set a range of 50% outward diffusion of the identification information within which no other elements except the background image are allowed.
[0060] By this method, since the content generated by the user is utilized, a set of template elements associated with the user-generated content is generated, and the associated user-generated content and the set of template elements are further combined, so that the target combined image generated is also more accurate and highly relevant, and the image generation efficiency is improved and the user experience is improved.
[0061] The above describes an example method for generating an image according to some embodiments of the present disclosure in conjunction with FIG. 2. The following describes a schematic diagram of a flowchart of a flow for generating an image according to some embodiments of the present disclosure in conjunction with FIG. 3. The example of FIG. 3 can be performed by the computing device 102 shown in FIG. 1 or any suitable device.
[0062] As shown in the example 300 of FIG. 3, the computing device 102 first generates the content description text 304 for the user-generated content 302. The user-generated content 302 is created by a user, and the user-generated content 302 is a video or an image, or a combination of a video and an image. The content description text 304 is obtained by applying the user-generated content 302 to a machine learning model. Additionally, the machine learning model is a visual large model.
[0063] In some embodiments, after the computing device applies the user-generated content 302 to the visual large model, the computing device obtains the distilled information for the user-generated content 302. For example, in some embodiments, the content description text 304 contains the distilled information for the user-generated content 302, which contains at least one of the following: what the user-generated content 302 mainly describes, what the style of the user-generated content 302 is, and what the theme color of the user-generated content 302 is, and so on. Additionally, the theme of the user-generated content 302 can be further distilled and determined from the content description text 304, for example, the theme of the user-generated content 302 can be determined to be a singing live broadcast, a football event, a game entertainment, and so on.
[0064] After obtaining the above-mentioned information, the computing device further generates the summary description text 306 according to the information in the content description text 304. For the sake of description, the summary description text 306 is also referred to as a script, which is generated by distilling the text information for the user-generated content in the content description text. The computing device can use a large language model to generate the text information. Additionally, the large language model is part of the visual large model, and the summary description text 306 is one of a set of template elements.
[0065] Meanwhile, the computing device also generates the image prompt information 308 according to the content description text 304. The image prompt information 308 includes first image prompt information and second image prompt information.
[0066] In some embodiments, the first image prompt information includes text information such as text for a main description of the user-generated content 302, text for a style of the user-generated content 302, and text for a theme color of the user-generated content 302. Additionally, text for a region of the user-generated content 302 or text for a date can also be obtained. The second image prompt information includes text for a theme of the user-generated content 302. For example, the theme of the user-generated content 302 can be determined to be a singing live broadcast, a football event, a game entertainment, and so on.
[0067] The computing device then applies the image prompt information 308 to an image generation model 310 to further generate a background image 312 and a sticker 318. Both the background image 312 and the sticker 318 are elements in a set of template elements. Additionally, the image generation model 310 is a diffusion model, and further, the image generation model 310 is a Stable Diffusion model. It is understood that this is merely an example and not a limitation of the present disclosure.
[0068] In some embodiments, the background image 312 is generated by the computing device according to the first image prompt information. The background image 312 is always at the back of the set of template elements and the user-generated content 302, and the background image 312 is an indispensable part of the final combined image. That is, the combined image is at least composed of the background image and the user-generated content 302. Additionally, the background image and the user-generated content are strongly associated.
[0069] In some embodiments, the sticker 318 is generated by the computing device according to the second image prompt information. Specifically, the sticker 318 is generated according to the text in the second image prompt information that is related to the theme of the user-generated content 302. In one example, when it is determined that the theme of the user-generated content 302 is a singing live broadcast, the sticker can be an image related to singing or music. In another example, when it is determined that the theme of the user-generated content 302 is a football event, the sticker can be an image related to football, events, etc. In yet another example, when it is determined that the theme of the user-generated content 302 is game entertainment, the sticker can be an image related to games, e-sports, etc.
[0070] After the computing device determines the summary description text 306, the background image 312, and the sticker 318, the above three elements are then determined as a set of template elements. Additionally, the computing device can also determine other template elements as needed and add them to the above set of template elements. Then, the computing device can apply the above set of template elements to the layout calculation 314 to determine a plurality of candidate combined images
[0071] In some embodiments, the above three elements can be placed according to preset position information, such as a first plurality of positions. And the user-generated content also has preset position information, such as a second plurality of positions. The computing device places the three elements and the user-generated content according to the respective position information.
[0072] In some embodiments, the above three elements can also be placed according to preset rules. For example, the size of the user-generated content 104 is adjusted according to a preset ratio, such as adjusting the user-generated content 302 to occupy 40% of the size of the background image. Additionally, at this time, the computing device can place the user-generated content at any suitable position in the background image
[0073] In some embodiments, a portion of the map 318 masks the user-generated content 302 and is displayed in front of the user-generated content 302. Additionally, the user can set the proportion of the map that masks the user-generated content 302, for example, the user can set a threshold proportion of 20% of the map that masks the user-generated content 104.
[0074] In some embodiments, the computing device determines a set of candidate combined images of the user-generated content 302 and the three elements according to a plurality of preset position information and a plurality of rules, for example, the placement positions and the placement rules described above. The set of candidate combined images has a plurality of candidate combined images, and each of the set of candidate combined images can be determined according to a scoring model to determine a respective score.
[0075] In some embodiments, the user can set a threshold score, and in this case, after the layout calculation 314, if the score of a candidate combined image reaches or exceeds the threshold score, the corresponding candidate combined image is determined as the target combined image 316. Additionally, when there are a plurality of candidate combined images whose scores reach or exceed the threshold score, the candidate combined image with the highest score is determined as the target combined image 316.
[0076] Through this method, since the content generated by the user is utilized, a set of template elements associated with the user-generated content is generated, and the associated user-generated content and the set of template elements are further combined, so that the target combined image generated is also more accurate and has high relevance, and the image generation efficiency is improved, and the user experience is improved.
[0077] The above describes a schematic diagram of a flowchart of a flow for generating an image according to some embodiments of the disclosure in conjunction with FIG. 3. The following describes a schematic diagram of an example of generating an image according to some embodiments of the disclosure in conjunction with FIG. 4.
[0078] In the example 400, the background image 402 and the user-generated content 404 constitute a target combined image. In this example, the target combined image only contains the background image 402 and the user-generated content 404.
[0079] In some embodiments, the user-generated content 404 is generated by the user, and the user-generated content 404 is a video or an image, or a combination of a video and an image. Additionally, the user-generated content 404 can be a video generated by live recording.
[0080] In some embodiments, the computing device can adjust the size of the user-generated content 404 according to a preset proportion, for example, adjust the user-generated content 404 to occupy 40% of the size of the background image. Additionally, at this time, the computing device can place the user-generated content at any suitable position in the background image.
[0081] The above describes one example of a schematic diagram for generating an image according to some embodiments of the present disclosure in connection with FIG. 4. The following describes another example of a schematic diagram for generating an image according to some embodiments of the present disclosure in connection with FIG. 5.
[0082] In example 500, on the basis of the previous example 400, on the basis of the background image 502 and the user-generated content 508, the decal 504 and the decal 506 are added. The size and shape of the decal 504 and the decal 506 are set by the user, and both are placed in front of the background image.
[0083] In some embodiments, there is at least one decal, and the decal can be placed according to a preset position. Additionally, the decal can also be placed according to a preset rule. In one example, the decal has no contact with the user-generated content in the background image. In another example, the decal has partial contact with the user-generated content in the background image, for example, a part of the decal masks the user-generated content and is displayed in front of the user-generated content. The user can also set the proportion of the decal masking the user-generated content 104, for example, the user can set a threshold proportion of the decal masking the user-generated content to be 20%.
[0084] The above describes another example of a schematic diagram for generating an image according to some embodiments of the present disclosure in connection with FIG. 5. The following describes yet another example of a schematic diagram for generating an image according to some embodiments of the present disclosure in connection with FIG. 6.
[0085] In example 600, on the basis of the previous example 400, on the basis of the background image 602 and the user-generated content 606, the summary description text 604 and the summary description text 608 are added. For the convenience of description, the summary description text 604 is referred to as Script 1, and the summary description text 608 is referred to as Script 2.
[0086] In some embodiments, Script 1 and Script 2 have no contact with the user-generated content, and are prevented from being in the background according to a predetermined position.
[0087] In some embodiments, Script 1 and / or Script 2 have contact with the user-generated content, and Script 1 and / or Script 2 partially mask the user-generated content. The computing device can place the script according to a preset rule, for example, the user can set the proportion of the script masking the user-generated content, for example, the user can set a threshold proportion of the decal masking the user-generated content to be 10%.
[0088] The above describes yet another example of a schematic diagram for generating an image according to some embodiments of the present disclosure in connection with FIG. 6. The following describes still another example of a schematic diagram for generating an image according to some embodiments of the present disclosure in connection with FIG. 7.
[0089] In example 700, on the basis of previous examples 400, 500 and 600, on the basis of background image 702 and user-generated content 710, summary description text 704, sticker 706, sticker 708 and summary description text 712 are placed in front of background image 702. For the convenience of description, summary description text 704 is referred to as text 1, and summary description text 712 is referred to as text 2.
[0090] In some embodiments, the computing device can place user-generated content 710, sticker 706, sticker 708, text 1 and text 2 according to predetermined position information. The computing device can also place user-generated content, stickers and texts according to predetermined rules, for example, user-generated content, stickers and texts can be placed according to preset rules, such as thickening the edges of stickers, highlighting texts, etc. Additionally, the user can also set the display priority of user-generated content, stickers and texts, for example, set the display priority of texts to the highest, in which case the texts will never be in contact with user-generated content and / or stickers, and will always be displayed in the front of the combined image. The above is only an example, and is not a limitation of the present application.
[0091] The above describes another example of a schematic diagram for generating an image according to some embodiments of the present disclosure in conjunction with FIG. 7. The following describes a schematic diagram for generating a combined image containing identification information of an image according to some embodiments of the present disclosure in conjunction with FIG. 8.
[0092] In example 800, on the basis of previous example 700, in addition to background image 802, summary description text 806, sticker 808, sticker 810, user-generated content 812 and summary description text 814, identification 804 is also contained. The identification 804 includes product identification, company identification or username watermark, etc.
[0093] In some embodiments, the user can set the transparency of the identification information, for example, set the transparency of the identification information to 50%. In another example, the user can set the identification information to be bold or highlighted to highlight the identification information. In yet another example, the user can set the position information of the identification information, for example, set the identification information in the four corners of the combined image. Additionally, the user can set to center the identification information, and diffuse a certain range in which no other elements except the background image are allowed, for example, set the identification information to diffuse 50% outwardly in a range in which no other elements except the background image are allowed to exist.
[0094] In some embodiments, the computing device sets the position of the logo 804 as the highest priority, that is, no other element in the combined image can obstruct the logo 804, so that the logo 804 can always appear in the front of the combined image.
[0095] The above describes a schematic diagram for generating a combined image of an image containing logo information according to some embodiments of the present disclosure in connection with FIG. 8. The following describes a schematic diagram of one position example of user-generated content and a set of template materials for generating an image according to some embodiments of the present disclosure in connection with FIG. 9.
[0096] In this example 900, the example 900 is composed of a background image 902, a logo 904, a sticker 906, a summary description text 908, a sticker 910, user-generated content 912, a sticker 914, and a summary description text 916.
[0097] In some embodiments, the position of the logo 904 is fixed, and is placed by a preset position range. In one example, the computing device places the logo 904 in a range of 5% of the upper left corner of the background image 902 according to a preset rule. In another example, the logo 904 is placed without contacting any edge of the background image after placement.
[0098] In some embodiments, the sticker 906 is placed according to a preset position, and is displayed in front of the background image. In some embodiments, the summary description text 908 and the sticker 910 are placed according to a preset rule, for example, a part of the summary description text 908 will obscure the sticker 910, and the summary description text 908 is always displayed in front of the sticker 910. Additionally, the proportion of the summary description text 908 obscuring the sticker 910 does not exceed a predetermined threshold, for example, does not exceed 30% of the display area of the sticker 910.
[0099] In some embodiments, the sticker 914 can also obscure the user-generated content 912, and similarly, the proportion of the sticker 914 obscuring the user-generated content 912 also does not exceed a predetermined threshold, for example, does not exceed 5% of the display area of the user-generated content.
[0100] In some embodiments, the text content of the summary description text 908 and the summary description text 916 can be the same, and is displayed in different display styles. In some embodiments, the summary description text 908 is displayed in the form of bold highlighting, and the summary description text 916 is displayed in artistic font.
[0101] In some embodiments, the text content of the summary description text 908 and the summary description text 916 are not the same. Additionally, the content of the summary description text 908 and the summary description text 916 should have strong relevance, for example, the summary of the content should revolve around the same topic, or the same keywords appear.
[0102] The above describes one example of a schematic diagram of a location for generating user-generated content and a set of template materials for an image according to some embodiments of the present disclosure in connection with FIG. 9. The following describes a schematic diagram of one specific embodiment for generating an image according to some embodiments of the present disclosure in connection with FIG. 10.
[0103] In the example 1000, the background image 1002 and the user-generated content 1014 are included. Among them, the user-generated content 1014 is a piece of content in which the host explains specific matters about the music festival activity, and the host mentions that there will be a mysterious guest appearance at the music festival activity, and that it is sponsored and held by the sponsor and the organizer. In addition, the background image 1002 is generated according to the first image prompt information for the user-generated content 1014.
[0104] In some embodiments, the identification 1004 is the identification of the sponsor and the organizer. In other embodiments, the identification 1004 is a watermark of the username of the host.
[0105] In some embodiments, the transparency of the identification 1004 can be set, for example, the transparency of the identification 1004 is set to 50%. The identification 1004 can also be set to bold or highlighted to highlight the identification 1004. The position information of the identification 1004 can also be set, for example, the identification 1004 is set at the four corners of the combined image. Additionally, the user can set the identification information as the center, and diffuse a certain range within which no other elements except the background image are allowed, for example, set the identification information to diffuse 50% outwardly within which no other elements except the background image are allowed to exist.
[0106] In some embodiments, the summary description information 1008 is “Mysterious guest, surprise appearance! Please look forward to it!”, which is generated by summarizing the user-generated content using a large language model. Similarly, the summary description information 1018 “Refreshing music festival!” is also generated based on the same large language model. And the summary description information 1008 and the summary description information 1018 have relevance, and can be directly or indirectly derived from the user-generated content.
[0107] In some embodiments, the sticker 1016 is a CD, the sticker 1010 is a speaker, the sticker 1012 is a shuffle mark, and the sticker 1016 is a musical note. The stickers are generated according to the second image prompt information for the user-generated content 1014, and further, are generated according to theme information in the second image prompt information, which includes “music festival”, “music”, and the like.
[0108] The above describes, in combination with FIG. 10, a schematic diagram of one specific embodiment of generating an image according to some embodiments of the present disclosure. The following describes, in combination with FIG. 11, a schematic diagram of an example of generating an image according to some embodiments of the present disclosure, which generates corresponding summary description text according to user-generated content.
[0109] In the example 1100, the scene content of the user-generated content 1102 is that a pet dog is having a birthday, and the pet dog is smiling at the camera.
[0110] In some embodiments, the computing device obtains the content description text for the user-generated content 1102 by applying the user-generated content 1102 to the visual large model, and then the computing device understands the user-generated content 1102 by using the visual large model, obtains the content description text for the user-generated content 1102, and further obtains text information such as text of the main description for the user-generated content 1102, text of the style for the user-generated content 1102, and text of the theme color for the user-generated content 1102 from the generated content description text.
[0111] The computing device then extracts keywords from the obtained various types of text information, and gives a summary description information suitable for the current scene and the target object appearing in the current scene according to the scene.
[0112] For example, the script 1104 “Ever pet deserves a big stage.” is generated according to the user-generated content 1102.
[0113] In some embodiments, the computing device can further generate multiple candidate scripts by using the visual large model, and score the multiple candidate scripts by using a scoring model. The candidate script whose score reaches or exceeds a threshold score in the multiple candidate scripts is determined as the target script 1104. Additionally, keywords and semantic completeness in the generated candidate scripts can be used as reference parameters for scoring the candidate scripts.
[0114] The above describes, in connection with FIG. 11, an example of a schematic diagram for generating an image according to some embodiments of the present disclosure, for generating a corresponding summary description text of an example of a user-generated content. The following describes, in connection with FIG. 12, an example of a schematic diagram for training of an image generation model for generating an image according to some embodiments of the present disclosure.
[0115] In the example 1200, the image generation model is trained. The computing device can first obtain sample image prompt information 1202. The sample image prompt information 1202 can be sample image prompt information for a sample background image. The sample image prompt information is text information, and the user can adjust the sample image prompt information according to needs.
[0116] The computing device then applies the sample image prompt information 1202 to the image generation model 1206 to generate a predicted image 1208. The predicted image 1208 can be a generated background image. After generating the predicted image 1208, the computing device further compares the predicted image 1208 and the sample image 1204 to determine the difference between the two, so as to further adjust the parameters of the image generation model to complete the training of the image generation model. Additionally, when it is necessary to generate a map using the image generation model, the image generation model can also be trained in the above manner for the map. In some embodiments, the user can train the image generation model according to different sample images and sample image prompt information for the sample images, so that the image generation model improves the generation ability for different types of images.
[0117] The above describes, in connection with FIG. 12, an example of a schematic diagram for training of an image generation model for generating an image according to some embodiments of the present disclosure. The following describes, in connection with FIG. 13, a schematic block diagram of an apparatus 1300 for generating an image according to embodiments of the present disclosure.
[0118] As shown in FIG. 12, the apparatus 1300 includes a content description text generation module 1302 configured to generate a content description text for a user-generated content based on the user-generated content; a set of template element generation module 1304 configured to generate a set of template elements of a template for the user-generated content based on the content description text; and a target combined image generation module 1306 configured to generate a target combined image based on the set of template elements and the user-generated content.
[0119] In some embodiments, the set of template elements includes a background image, and the template element generation module 1304 includes: a first image prompt information generation module configured to generate a first image prompt information for the user-generated content based on the content description text; and a background image generation module configured to generate the background image for the template based on the first image prompt information.
[0120] In some embodiments, the set of template elements further comprises at least one of: a summary description text or a poster, and the template element generation module 1304 further comprises at least one of: a summary description text generation module configured to generate, based on the content description text, the summary description text for the template; or a poster generation module configured to generate, based on the content description text, the poster for the template.
[0121] In some embodiments, the poster generation module comprises: a second image prompt information generation module configured to generate, based on the content description text, second image prompt information for the user-generated content; and a poster generation module configured to generate, based on the second image prompt information, the poster for the template.
[0122] In some embodiments, the first image prompt information comprises at least one of: a content of the background image or a color of the background image.
[0123] In some embodiments, the background image generation module comprises: an image generation model application module configured to generate the background image by applying the first image prompt information to an image generation model, the image generation model being a diffusion model.
[0124] In some embodiments, the training module of the image generation model comprises: a sample image prompt information and sample image obtaining module configured to obtain sample image prompt information and sample images; a predicted image obtaining module configured to obtain predicted images by applying the sample image prompt information to the image generation model; and a parameter adjustment module configured to adjust parameters of the image generation model based on the sample images and the predicted images.
[0125] In some embodiments, the target combined image generation module 1306 comprises: a set of candidate combined image generation modules configured to generate a set of candidate combined images based on the set of template elements and the user-generated content; and a target combined image selection module configured to select the target combined image from the set of candidate combined images.
[0126] In some embodiments, the set of candidate combined image generation modules comprises: a first and second plurality of position determination module configured to determine a first plurality of positions available for placing template elements in the set of template elements and a second plurality of positions available for placing the user-generated content; and a set of candidate combined image generation modules configured to generate the set of candidate combined images by placing the template elements in the first plurality of positions and placing the user-generated content in the second plurality of positions, respectively.
[0127] In some embodiments, the set of candidate combined images includes a plurality of predetermined rule determination modules configured to determine a plurality of predetermined rules for placing a set of template elements and user-generated content, and a set of candidate combined image generation modules configured to generate the set of candidate combined images based on the plurality of predetermined rules.
[0128] In some embodiments, the target combined image selection module includes a set of score determination modules configured to determine a set of scores for the set of candidate combined images, and a target combined image selection module configured to select the target combined image from the set of candidate combined images based on the set of scores, the target combined image having a score that exceeds a threshold score.
[0129] In some embodiments, the content description text module 1302 includes a machine learning model application module configured to obtain the content description text for the user-generated content by applying the user-generated content to a machine learning model.
[0130] In some embodiments, the user-generated content is an image or a video, and the machine learning model is a visual model.
[0131] FIG. 14 shows a schematic block diagram of an example device 1400 that can be used to implement embodiments of the present disclosure. The computing device 102 in FIG. 1 can be implemented with the device 1400. As shown, the device 1400 includes a central processing unit (CPU) 1401, which can perform various suitable actions and processes according to computer program instructions stored in a read-only memory (ROM) 1402 or computer program instructions loaded into a random access memory (RAM) 1403 from the storage unit 808. Various programs and data required by the device 1400 to operate can also be stored in the RAM 1403. The CPU 1401, the ROM 1402, and the RAM 1403 are connected to each other by a bus 1404. An input / output (I / O) interface 1405 is also connected to the bus 1404.
[0132] A plurality of components in the device 1400 are connected to the I / O interface 1405, including an input unit 1406, such as a keyboard, a mouse, etc., an output unit 1407, such as various types of displays, a speaker, etc., a storage unit 1408, such as a magnetic disk, an optical disk, etc., and a communication unit 1409, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1409 allows the device 1400 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.
[0133] The various processes and processes described above, such as the method 200 and examples 300, 1200, can be performed by the processing unit 1401. For example, in some embodiments, the method 200 and examples 300, 1200 can be implemented as a computer software program tangibly embodied in a machine readable medium, such as the storage unit 1408. In some embodiments, portions or all of the computer program can be loaded and / or installed onto the device 1400 via the ROM 1402 and / or the communication unit 1409. When the computer program is loaded onto the RAM 1403 and executed by the CPU 1401, one or more acts of the example methods 200 and examples 300, 1200 described above can be performed.
[0134] The present disclosure can be a method, apparatus, system, and / or computer program product. A computer program product can include a computer readable storage medium (or media) having computer readable program instructions thereon for performing various aspects of the present disclosure.
[0135] The computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer readable storage medium can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer readable storage medium include the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or punched tape, a
[0136] The computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network can comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.
[0137] Computer readable program instructions for carrying out operations of the present disclosure can be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The computer readable program instructions can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate array (FPGA), or programmable logic array (PLA) can execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.
[0138] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0139] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can include random access memory (RAM), read only memory (ROM), electrically erasable programmable read only memory (EEPROM), flash memory or other data storage device. When the computer readable program instructions are loaded into the computer and other programmable data processing apparatus, a series of operational steps are implemented that provide processes such that the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0140] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0141] The flow diagrams and the block diagrams in the drawings are presented to illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flow diagrams and the block diagrams can represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logic functions. In some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks can sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flow diagrams, and combinations thereof, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and
[0142] Embodiments of the present disclosure have been described above, and the description is intended to be illustrative of the embodiments and not restrictive. Many modifications and variations of the described embodiments are possible and are within the scope of the disclosure. The selection of terms is intended to best describe the principles of the embodiments, practical application, or technical improvements over the technology found in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A method for generating an image, comprising: generating, based on user-generated content, a content description text for the user-generated content; generating, based on the content description text, a set of template elements for a template for the user-generated content; and generating, based on the set of template elements and the user-generated content, a target combined image.
2. The method of claim 1, wherein the set of template elements comprises a background image, and wherein generating, based on the content description text, a set of template elements for a template for the user-generated content comprises: generating, based on the content description text, first image prompt information for the user-generated content; and generating, based on the first image prompt information, the background image for the template.
3. The method of claim 2, wherein the set of template elements further comprises at least one of a summary description text or a sticker, and wherein generating, based on the content description text, a set of template elements for a template for the user-generated content further comprises at least one of: generating, based on the content description text, the summary description text for the template; or generating, based on the content description text, the sticker for the template.
4. The method of claim 3, wherein generating, based on the content description text, the sticker for the template comprises: generating, based on the content description text, second image prompt information for the user-generated content; and generating, based on the second image prompt information, the sticker for the template.
5. The method of claim 2, wherein the first image prompt information comprises at least one of a content of a background image or a color of a background image.
6. The method of claim 2, wherein generating, based on the first image prompt information, the background image for the template comprises: generating the background image by applying the first image prompt information to an image generation model, the image generation model being a diffusion model.
7. The method of claim 6, wherein training of the image generation model comprises: obtaining sample image prompt information and sample images; obtaining a predicted image by applying the sample image prompt information to the image generation model; and adjusting parameters of the image generation model based on the sample images and the predicted image.
8. The method of claim 1, wherein generating, based on the set of template elements and the user-generated content, a target combined image comprises: generating, based on the set of template elements and the user-generated content, a set of candidate combined images; and selecting the target combined image from the set of candidate combined images.
9. The method of claim 8, wherein generating, based on the set of template elements and the user-generated content, a set of candidate combined images comprises: determining a first plurality of positions available for placement of a template element in the set of template elements and a second plurality of positions available for placement of the user-generated content; and selecting, from the first plurality of positions and the second plurality of positions, a set of positions for placement of the set of template elements and the user-generated content. The set of candidate combined images is generated by placing the set of template elements at the first plurality of positions and placing the user-generated content at the second plurality of positions, respectively. 10.The method of claim 8, wherein generating a set of candidate combined images based on the set of template elements and the user-generated content comprises: determining a plurality of predetermined rules for placing the set of template elements and the user-generated content; and generating the set of candidate combined images based on the plurality of predetermined rules. 11.The method of claim 8, wherein selecting the target combined image from the set of candidate combined images comprises: determining a set of scores for the set of candidate combined images; and selecting the target combined image from the set of candidate combined images based on the set of scores, the target combined image having a score that exceeds a threshold score. 12.The method of claim 1, wherein generating a content description text for the user-generated content comprises: obtaining a content description text for the user-generated content by applying the user-generated content to a machine learning model. 13.The method of claim 11, wherein the user-generated content is an image or a video, and the machine learning model is a visual model. 14.An apparatus for generating an image, comprising: a content description text generation module configured to generate a content description text for user-generated content based on the user-generated content; a set of template elements generation module configured to generate a set of template elements for a template of the user-generated content based on the content description text; and a target combined image generation module configured to generate a target combined image based on the set of template elements and the user-generated content. 15.An electronic device, comprising: at least one processor; and a memory device for storing at least one program, which, when executed by the at least one processor, causes the at least one processor to implement the method according to any one of claims 1-13. 16.A computer-readable storage medium having stored thereon a computer program, which, when executed by a processor, implements the method according to any one of claims 1-13. 17.A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1-13.
Citation Information
Patent Citations
Method and device for generating image, equipment and medium
CN117593404A
Template generation method and device, equipment and storage medium
CN117744617A
Method, device and equipment for creating works and storage medium
CN118138843A
Video generation method and device, electronic equipment, storage medium and program product
CN118433466A
Method and apparatus for generating video template, and electronic device
WO2024160128A1