Image template generation method and device, electronic equipment and storage medium
By generating a cover template that includes a background layer, a text layer, and a sticker layer, the problem of fixed style of video cover templates is solved, and efficient generation and high-quality effects of personalized video covers are achieved.
Patent Information
- Application Number
- CN202410431134.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-10
- Publication Date
- 2025-10-17
AI Technical Summary
In the prior art, the template style of the video cover is fixed and cannot meet the personalized needs of users, resulting in poor quality of the generated video cover and high acquisition cost.
By generating an initial cover image based on the first image model, parsing the background layer and text layer in the basic template data, and using the second image model to generate a cover template including a sticker layer, the video cover image is generated in combination with the configuration instructions.
It enables rapid construction and personalized control of cover templates, reduces design difficulty and acquisition costs, and generates high-quality video covers that match the title and background.
Smart Images

Figure CN120804362A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present disclosure relate to the technical field of artificial intelligence, and particularly relate to an image template generation method and device, electronic equipment and storage medium. BACKGROUND
[0002] Currently, in a video creation scenario, a video work created by a user is usually displayed in the form of a video cover after being uploaded and shared to a video platform. Therefore, the video cover plays a role in displaying the video content, and a high-quality video cover can better attract platform users to click and watch the video work.
[0003] In the prior art, for the production of a video cover, a video frame in a video is usually intercepted, a title information is inserted in the video frame in combination with a fixed template style, so as to generate a video cover.
[0004] However, the prior art has problems such as fixed template style and high acquisition cost, which leads to that the image template cannot meet the personalized needs of users and affects the quality of the generated video cover. SUMMARY
[0005] Embodiments of the present disclosure provide an image template generation method and device, electronic equipment and storage medium to overcome the problems of fixed template style and high acquisition cost.
[0006] In a first aspect, embodiments of the present disclosure provide an image template generation method, comprising:
[0007] processing a first prompt word based on a first image model to generate an initial cover picture, wherein the first prompt word is used at least to describe title information in the initial cover picture; generating basic template data from the initial cover picture, wherein the basic template data includes a background layer and a text layer, the background layer is used to carry background elements in the initial cover picture, and the text layer is used to carry text elements in the initial cover picture; processing the basic template data based on a second image model to generate a cover template, wherein the cover template includes the background layer, the text layer and a sticker layer, the layer content of the sticker layer is determined based on the basic template data, and the cover template is used to generate a video cover picture in combination with configuration instructions for layers in the cover template, and the configuration instructions are used to configure elements in the layers.
[0008] In a second aspect, embodiments of the present disclosure provide an image template generation device, comprising:
[0009] A first generation module is configured to process a first prompt word based on a first image model to generate an initial cover picture, wherein the first prompt word is used at least to describe title information in the initial cover picture.
[0010] The parsing module is configured to generate basic template data according to the initial cover image, the basic template data including a background layer and a text layer, the background layer being used to carry background elements in the initial cover image, and the text layer being used to carry text elements in the initial cover image;
[0011] The second generation module is configured to process the basic template data based on a second image model to generate a cover template, wherein the cover template includes the background layer, the text layer, and a sticker layer, layer content of the sticker layer being determined based on the basic template data, and the cover template being used to generate a video cover image in combination with configuration instructions for layers in the cover template, the configuration instructions being used to configure elements in the layers.
[0012] In a third aspect, an electronic device is provided, and the electronic device includes a processor and a memory.
[0013] The memory stores computer-executable instructions.
[0014] The processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the image template generation method according to the first aspect and various possible designs of the first aspect.
[0015] In a fourth aspect, a computer-readable storage medium is provided, and the computer-readable storage medium stores computer-executable instructions. When a processor executes the computer-executable instructions, the image template generation method according to the first aspect and various possible designs of the first aspect is implemented.
[0016] In a fifth aspect, a computer program product is provided, and the computer program product includes a computer program. When a processor executes the computer program, the image template generation method according to the first aspect and various possible designs of the first aspect is implemented.
[0017] The image template generation method, device, electronic device and storage medium provided in this embodiment generate an initial cover image by processing a first prompt word based on a first image model, wherein the first prompt word is used to at least describe the title information in the initial cover image; generate basic template data based on the initial cover image, wherein the basic template data includes a background layer and a text layer, the background layer is used to carry the background elements in the initial cover image, and the text layer is used to carry the text elements in the initial cover image; based on a second image model, process the basic template data to generate a cover template, wherein the cover template includes the background layer, the text layer and a sticker layer, the layer content of the sticker layer is determined based on the basic template data, and the cover template is used to generate a video cover image in combination with configuration instructions for the layers in the cover template, and the configuration instructions are used to configure the elements in the layers. By first generating an initial cover image containing title information based on the first prompt word, then generating basic template data including a background layer and a text layer based on the initial cover image, and finally generating a cover template including a background layer, a text layer and a sticker layer based on the basic template data, the rapid construction of a cover template containing title information is achieved. In this process, the style of the title information is controlled to meet the personalized needs of users. At the same time, a sticker layer matching the title and background can be generated, thereby achieving automated, high-quality on-demand generation of cover templates, reducing design difficulty and the cost for users to obtain cover templates. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0019] Figure 1 A diagram of an application scenario of the image template generation method provided in an embodiment of the present disclosure;
[0020] Figure 2 Schematic diagram of the process of the image template generation method provided in the embodiment of the present disclosure Figure 1 ;
[0021] Figure 3 for Figure 2 A flowchart of a specific implementation method of step S102 in the embodiment shown;
[0022] Figure 4 A schematic diagram of a process for generating basic template data provided by an embodiment of the present disclosure;
[0023] Figure 5 A data structure schematic diagram of a cover template provided by an embodiment of the present disclosure is shown in the following table.
[0024] Figure 6 A flow schematic diagram of an image template generation method provided by an embodiment of the present disclosure is shown in the following table. Figure 2
[0025] Figure 7 A flow chart of a specific implementation mode of step S204 in the embodiment shown in the following table. Figure 6
[0026] Figure 8 A flow chart of a specific implementation mode of step S2042 in the embodiment shown in the following table. Figure 7
[0027] Figure 9 A process schematic diagram of generating a sticker layer data provided by an embodiment of the present disclosure is shown in the following table.
[0028] Figure 10 A flow chart of a specific implementation mode of step S205 in the embodiment shown in the following table. Figure 6
[0029] A structure block diagram of an image template generation apparatus provided by an embodiment of the present disclosure is shown in the following table. Figure 11
[0030] A structure schematic diagram of an electronic device provided by an embodiment of the present disclosure is shown in the following table. Figure 12
[0031] A hardware structure schematic diagram of an electronic device provided by an embodiment of the present disclosure is shown in the following table. Figure 13 DETAILED DESCRIPTION
[0032] In order to make the objects, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be described clearly and completely below with reference to the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, rather than all the embodiments of the present disclosure. Based on the embodiments in the present disclosure, all other embodiments obtained by a person of ordinary skill in the art without creative labor fall within the protection scope of the present disclosure.
[0033] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present disclosure are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation portal for user to choose authorization or refusal.
[0034] The application scenarios of the embodiments of the present disclosure are explained as follows:
[0035] Figure 1 An application scenario diagram of the image template generation method provided by the embodiments of the present disclosure is provided. The image template generation method provided by the embodiments of the present disclosure can be applied to an application (APP, Application) with a video cover generation function, such as a video editing application, a short video application, a smart assistant application, and the like. More specifically, it can be applied to an application scenario of generating a video cover or a cover template before a video work is published. The execution subject of the present embodiment can be a terminal device running the application with the video cover generation function, a server deploying a server corresponding to the application, or other electronic devices with similar functions.
[0036] In some embodiments, the terminal device or the server can implement the image template generation method provided by the embodiments of the present disclosure by running various computer executable instructions or computer programs. For example, the computer executable instructions can be program-level commands, machine instructions, or software instructions. The computer program can be a native program or a software module in an operating system; it can be a local application, that is, a program that needs to be installed in an operating system to run, or a small program embedded in any APP, that is, a program running based on a browser environment. In summary, the above computer executable instructions can be any form of instructions, and the above computer programs can be any form of application programs, modules or plug-ins, and the specific implementation form can be configured as needed. Further, in some embodiments, the server can be a standalone physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud services, cloud storage, cloud communication, cloud database, cloud computing, cloud function, network service, middleware service, domain name service, security service, content distribution network (CDN), and big data and artificial intelligence platform, and the like basic cloud computing services, wherein the cloud service can be an interactive processing service for calling by the terminal device.
[0037] Reference Figure 1As shown in the figure, taking a terminal device as an example, after the terminal device runs the application program with the video cover generation function, for the generated video or the video draft loaded in the application program, the function of "creating a cover template" is triggered through user operation, then a cover template is created by using the image template generation scheme provided by the embodiment of the disclosure, for example, as shown in the figure, the cover template includes "cover background" and "cover title", which can be further modified based on needs; then, by triggering the function of "generating a video cover", the cover template is used to generate a corresponding video cover (the video cover contains the title "ABCD") for the above-mentioned video or video draft; then, the video cover and the corresponding video (draft) are published to the video platform, and the publishing process of the video work is completed; at the same time, the cover template is saved in the terminal device locally and / or the server, so as to realize subsequent reuse.
[0038] In the prior art, the generation scheme for the video cover is usually to generate a corresponding video cover by using the fixed style template provided in the application program or platform, so that the style of the cover template is fixed and the number is limited. When a user needs to configure a cover template meeting the personalized needs, the cover background, title and other information can only be manually selected and input, and the layout setting is performed, which leads to a long time-consuming and high difficulty in the configuration process of the template, and there is a problem of poor collocation effect of the template, which affects the visual effect of the generated video cover.
[0039] The embodiment of the disclosure provides an image template generation method to solve the above-mentioned problems.
[0040] Reference Figure 2 , Figure 2 Flowchart of the image template generation method provided by the embodiment of the disclosure Figure 1 The method of the embodiment can be applied in terminal identification, and the image template generation method includes:
[0041] Step S101: processing a first prompt word based on a first image model to generate an initial cover image, wherein the first prompt word is used to describe at least title information in the initial cover image.
[0042] Exemplarily, referring to Figure 1The application scenario diagram shows that after the terminal device runs the application program, the terminal device receives the first prompt word input by the user through the interaction interface set in the application program, and then the terminal device processes the first prompt word by calling the first image model, and generates an initial cover image matched with the first prompt word by the first image model. The first image model is a deep learning model that can be used to generate images, for example, a stable diffusion model (Stable Diffusion). The first prompt word includes a natural language-based description text or a word, which is used to guide the model to output specified content, i.e., title information in the initial cover image. Further, the title information described by the first prompt word includes various contents, such as title content, font color, position, size, font, and the like of the title information, thereby realizing control of the visual performance of the title text in the initial cover image.
[0043] Further, in a possible implementation, the first prompt word is also used to describe the background content of the initial cover image. In this case, through the first prompt word, the control of the background and the title of the initial cover image output by the first image model can be realized at the same time. The specific implementation process of generating a corresponding image through an image processing model and a specified prompt word is prior art known to those skilled in the art, and will not be described in detail here. In another possible implementation, while inputting the first prompt word to the first image model, the initial image or video information is also input to the first image model. The video information is descriptive information for the target video, such as video name, video classification keyword, and video image feature, without limitation. The first image model determines the background of the initial cover image according to the initial image or video information, and then generates the initial cover image in combination with the first prompt word. Further, optionally, the first image model can also extract the video information of the target video by receiving the target video. The specific implementation mode is determined based on the implementation mode of the video information, which will not be described one by one here.
[0044] Step S102: generating basic template data according to the initial cover image, the basic template data including a background layer and a text layer, the background layer being used to carry the background element in the initial cover image, and the text layer being used to carry the text element in the initial cover image.
[0045] Exemplarily, after obtaining the initial cover picture, the terminal device parses the initial cover picture, generates a background layer and a text layer based on the content in the initial cover picture, and further generates basic template data containing the background layer and the text layer based on the background layer and the text layer. First, the layers in the steps of the embodiment are simply explained. A layer is an image used to carry image elements. In the layer, positions other than the positions of the image elements are in a transparent state, so as to avoid occlusion of the background image when superimposed on the upper layer of the background image. In the steps of the embodiment, the background layer and the text layer, i.e., the basic template data, are generated by identifying the image content in the initial cover picture. The background layer is used to carry the background elements in the initial cover picture, such as a background image. The text layer is used to carry the text elements in the initial cover picture, such as title text. Further, when the basic template data contains only one text layer, all text elements are arranged in the text layer. When the basic template data contains multiple text layers, each text layer carries one text element. For example, the text layer layer_1 is used to carry the text element corresponding to the main title text, and the text layer layer_2 is used to carry the text element corresponding to the subtitle text.
[0046] Further, in a possible implementation manner, as shown in Figure 3 the specific implementation manner of step S102 includes:
[0047] Step S1021: parsing the image content of the initial cover picture to obtain a title area.
[0048] Step S1022: performing text element extraction on the initial cover picture based on the title area to generate at least one text layer.
[0049] Step S1023: performing content filling on the title area of the initial cover picture to generate a background layer.
[0050] Step S1024: generating basic template data based on the at least one text layer and the background layer.
[0051] Exemplarily, since the initial cover image generated in the previous step is an AI image generated based on the first prompt word describing the title information and the image generation model (first image model), the content in the initial cover image is controlled by the first prompt word. Therefore, the initial cover image will contain the title information required as the video cover image. By identifying the content of the initial cover image, the text element corresponding to the above-mentioned title information and the position corresponding to the text element, that is, the title area, can be identified. Afterwards, the text elements are extracted from the title area, and the extracted image elements are set to a transparent image to generate a text layer; on the other hand, the content is filled in the title area, and the image generation technology is used to make the title area consistent with the surrounding background, that is, the original text elements are eliminated from the background to generate a background layer. Afterwards, based on the above-mentioned text layer and background layer, basic template data is generated.
[0052] Figure 4 A schematic diagram of a process for generating basic template data provided by an embodiment of the present disclosure, such as Figure 4 As shown, after obtaining the initial cover image, the terminal device parses the initial cover image to generate a background layer (shown as L1 in the figure) and two text layers (shown as L2 and L3 in the figure). Among them, an image is set in the background layer. The image is based on the background image of the initial cover image (such as the "mountain" and "highway" in the initial cover image), and the title area is filled. The image is also the background element. A text element is set in each of the two text layers, such as the "main title" and "subtitle" in the initial cover image. Afterwards, based on the above background layer and two text layers, the basic template data is constructed.
[0053] Step S103: Based on the second image model, the basic template data is processed to generate a cover template, wherein the cover template includes a background layer, a text layer and a sticker layer, and the layer content of the sticker layer is determined based on the basic template data. The cover template is used to generate a video cover image in combination with the configuration instructions for the layers within the cover template, and the configuration instructions are used to configure the elements within the layers.
[0054] Exemplarily, after obtaining the basic template data, the terminal device calls the second image processing to process the basic template data, and generates a sticker layer with content matching the text layer and the background layer in the basic template data, i.e., the layer content of the sticker layer is determined based on the basic template data. Similar to the text layer, the sticker layer is also a layer for carrying image elements, and the sticker elements in the sticker layer are image elements for decorating and beautifying the text layer and the background layer. Then, the cover template is generated by combining the background layer, the text layer and the sticker layer. In the cover template, the three layers (the background layer, the text layer and the sticker layer) can all modify the content in the layers in response to the configuration instruction input by the user, for example, modifying the layer content in the text layer. More specifically, for example, the layer content in the text layer is the text element in the initial cover image by default, and the content (text content) of the text element can be modified to other content by responding to the configuration instruction. Then, the corresponding video cover image can be generated based on the content of each layer in the cover template.
[0055] Further, the cover template has a first configuration interface and / or a second configuration interface.
[0056] The first configuration interface is configured to configure the elements in the target layer in response to a first configuration instruction to generate a video cover image. The second configuration interface is configured to obtain video information of a target video in response to a second configuration instruction, and configure at least one element in the target layer according to the video information to generate a video cover image of the target video. The target layer includes at least one of a background layer, a text layer and a sticker layer.
[0057] Figure 5 The data structure of the cover template provided by the embodiments of the present disclosure is shown in Figure 5As shown, the cover template can include a data part and an interface part, where the data part stores the data corresponding to the plurality of layers in a specific format, and the specific implementation manner can be set as needed, which will not be described in detail here; the interface part includes a first configuration interface and / or a second configuration interface. Taking the case of including both the first configuration interface and the second configuration interface as an example, when generating a video cover based on the cover template, the configuration of image elements in any one or more of the background layer, the text layer, and the sticker layer can be realized by calling the first configuration interface, such as changing the text content of the text element in the text layer, changing the background image in the background layer, and the like. And by calling the second configuration interface, the video information of the target video is obtained, and at least one element in the target layer is configured according to the video information, for example, the text content of the text element in the text layer is changed according to the video information, the background image in the background layer is changed, and the like. The specific meaning of the video information has been introduced in step 101, which will not be described here.
[0058] In the step of the embodiment, the first configuration interface and the second configuration interface provided by the cover template can realize manual configuration and automatic configuration of the elements in the layer. In the implementation process of automatic configuration, the video information of the target video is used to configure the layer content in the cover template, so that the video cover generated based on the cover template can match the content of the target video, and the efficiency and quality of generating the video cover of a specific target video are improved.
[0059] In the embodiment, the initial cover is generated by processing the first prompt based on the first image model, where the first prompt is used to describe at least the title information in the initial cover; the base template data is generated based on the initial cover, where the base template data includes a background layer and a text layer, the background layer is used to carry the background elements in the initial cover, and the text layer is used to carry the text elements in the initial cover; the cover template is generated based on the second image model, where the cover template includes the background layer, the text layer, and a sticker layer, the layer content of the sticker layer is determined based on the base template data, the cover template is used to generate a video cover in combination with the configuration instructions for the layers in the cover template, and the configuration instructions are used to configure the elements in the layers. By generating an initial cover containing title information based on the first prompt, generating base template data including a background layer and a text layer based on the initial cover, and finally generating a cover template including a background layer, a text layer, and a sticker layer based on the base template data, the rapid construction of a cover template containing title information is realized, the control of the title information style is realized in the process to meet the personalized needs of users, and a sticker layer matching the title and the background can be generated, realizing the automatic and high-quality on-demand generation of the cover template, reducing the design difficulty and the cost of users obtaining the cover template.
[0060] refer to Figure 6 , Figure 6 Schematic diagram of the process of the image template generation method provided in the embodiment of the present disclosure Figure 2 In this embodiment Figure 2 Based on the embodiment shown, step S103 is further refined, and the image template generation method includes:
[0061] Step S201: Processing a first prompt word based on a first image model to generate an initial cover image, wherein the first prompt word is at least used to describe title information in the initial cover image.
[0062] Step S202: Generate basic template data based on the initial cover image. The basic template data includes a background layer and a text layer. The background layer is used to carry the background elements in the initial cover image, and the text layer is used to carry the text elements in the initial cover image.
[0063] Step S203: Rendering the background layer and the text layer into a model input image according to the basic template data.
[0064] Step S204: Processing the model input image through the second image model to obtain sticker layer data, where the sticker layer data includes at least one different sticker layer.
[0065] For example, in combination Figure 7 In the solution of the illustrated embodiment, after the basic template data is generated, the background layer and the text layer in the basic template data are merged and rendered into a frame of model input image. Afterwards, the model input image is used as input to call the second image model for processing, and the sticker layer data output by the second image model is obtained. The sticker layer data contains at least one sticker layer. The specific meaning and implementation of the sticker layer have been introduced in detail in the previous embodiment and will not be repeated here. In the above steps, the step of rendering the background layer and the text layer into the model input image and then inputting it into the second image model depends on the specific interface implementation and model training process of the second image model, that is, in the above implementation method, "one picture" is used as the input of the second image model; in other possible implementation methods, a collection of multiple layers (that is, basic template data) can also be used as the input of the second image model, and the paper layer data output by the second image model can be obtained. There is no specific limitation on this.
[0066] In one possible implementation, Figure 8 As shown, the specific implementation of step S204 includes:
[0067] Step S2041: extracting, by the second image model, first image features of the model input image, wherein the first image features represent semantic content features of the background elements and the text elements in the model input image.
[0068] Step S2042: generating sticker layer data according to the first image features of the model input image.
[0069] Exemplarily, in the implementation process, after the model input image is input into the second image model, first, the semantic content features corresponding to the background elements and the text elements in the model input image, i.e., the first image features, are extracted through the network structure of the second image model, such as one or more neural network layers; wherein the semantic content features and the content meanings expressed by the background elements and the text elements, for example, the background elements include background pictures, and the corresponding semantic content features are picture contents, more specifically, for example, representing “forest”, “meteor in the starry sky”, “vehicle driving on the highway” and the like; and the semantic content features of the text elements, i.e., the content of the text, such as the main title and the subtitle of the cover. Of course, the above-mentioned semantic content features can be words representing specific meanings, or feature vectors, feature matrices and the like representing the above-mentioned content, which can be set as needed.
[0070] Further, optionally, before step S2042, it further includes:
[0071] Step S2040: extracting, by the second image model, second image features of the model input image, wherein the second image features represent style content features of the background elements and the text elements in the model input image.
[0072] Correspondingly, when step S2040 is executed, the specific implementation of step S2042 is:
[0073] Step S2042A: generating sticker layer data according to the first image features and the second image features of the model input image.
[0074] In one possible implementation, the second image model extracts the first image features and the second image features of the model input image at the same time, wherein the second image features represent the style content features of the background elements and the text elements, specifically, the style content features of the background elements, such as representing the transparency and filter of the background image; and the style content features of the text elements, such as representing the font style, font color and font size of the text elements (such as the cover title). The implementation of the second image features is similar to that of the first image features, which can be implemented in the form of feature matrix, feature vector and the like, and will not be described here. The ability of the second image model to extract the first image features and the second image features from the model input image is formed in the training process of the second image model, and the specific implementation process will not be introduced here.
[0075] Afterwards, one or more corresponding sticker layers are generated based on the first and second image features extracted by the second image model. This process can be performed by the second image model or by other independent image generation models, without limitation. The sticker elements in the sticker layers of the sticker layer data generated in this way match the layer content of the background layer and the text layer in terms of content, style, and other dimensions, making the generated sticker elements more harmonious and realistic, and improving the visual effect of the video cover image generated by the cover template.
[0076] Furthermore, in a possible implementation, as Figure 9 As shown, the specific implementation of step S2042 includes:
[0077] Step S2042-1: Generate a target sticker layer based on the first image feature of the model input image.
[0078] Step S2042-2: Obtain a rendering effect image based on the model input image and the target sticker layer.
[0079] Step S2042-3: If the rendering effect image does not meet the preset conditions, the model input image is updated to the rendering effect image, and the process returns to step S2042-1.
[0080] Step S2042-4: If the rendering effect image satisfies the preset conditions, sticker layer data is generated based on the target sticker layer.
[0081] Exemplarily, for the specific implementation of step S2042, first, the second image model generates a target sticker layer according to the first image feature of the model input image, and the target sticker layer is a background transparent picture including at least one sticker element, that is, a graph generation process of generating a matching image according to the image feature. Then, the model input image rendered by the background layer and the text layer is merged with the target sticker layer to obtain a rendering effect picture; then, the second image model judges the rendering effect picture, if it meets the preset condition, for example, the second image model judges that the confidence of the rendering effect picture is less than the confidence threshold, or the image feature distance of the rendering effect picture is far away from the specific target image feature, it is considered that the current generated rendering effect picture cannot meet the requirement of the video cover picture, then the rendering effect picture generated in the current loop round is taken as the new model input image, and the above process is repeated. If the rendering effect picture meets the preset condition, for example, the loop number reaches the loop threshold, or the second image model judges that the confidence of the rendering effect picture is greater than the confidence threshold, the target sticker layer is output; if the preset condition is met after multiple loop rounds, the target sticker layer obtained in each loop round is output respectively, so as to obtain the sticker layer data including multiple target sticker layers.
[0082] Figure 9 A process schematic diagram for generating sticker layer data provided by the embodiment of the present disclosure is shown in FIG. 6. Figure 8 Exemplarily, first, based on the background layer and the text layer in the basic template data, the model input image Input is generated, then the feature extraction is performed on the model input image Input, and based on the obtained feature Feature, the target sticker layer U is obtained, and the first obtained target sticker layer is, for example, U(1); then, the model input image Input and the target sticker layer U are rendered into a rendering effect picture P. Then, based on the judgment result of the rendering effect picture P, if the preset condition is not met (NO in the figure), the model input image is updated as the rendering effect picture, that is, Input=P, and the step of performing the feature extraction on the model input image Input is returned to further generate other target sticker layers, for example, the target sticker layers U(2), U(3), …, U(n); if the preset condition is met (YES in the figure), the set of the target sticker layers generated in each loop round in the above step, for example, the target sticker layers U(1), U(2), …, U(n), generates the sticker layer data.
[0083] In this embodiment, the image features of the model input image are extracted in a loop, and the corresponding sticker layers are generated based on the image features, thereby generating sticker layer data containing multiple sticker layers. Each sticker layer (sticker element) in the sticker layer is the best sticker element generated based on the current model input image. Not only does it match the sticker element with the background element and the text element, but it also improves the matching degree between the sticker elements in different sticker layers and the image harmony, ultimately improving the visual effect of the video cover generated by the cover template.
[0084] Further, optionally, the specific implementation of step S204 provided in this embodiment further includes:
[0085] Step S2043: Obtain a second prompt word, which is used to represent positioning information based on semantic content feature description of the background element and / or the text element.
[0086] Correspondingly, in the case of performing step S2043, the specific implementation of step S2042 is:
[0087] Step S2042B: Generate sticker layer data by the second prompt word and the first image feature, wherein the sticker element of the sticker layer in the sticker layer data is located at the target position corresponding to the positioning information.
[0088] Illustratively, further, the second image model can also accept the second prompt word and optimize the position of the sticker element based on the positioning information described by the second prompt word. Wherein the second prompt word is used to represent the positioning information based on the semantic content feature description of the background element and / or the text element, for example, the semantic content feature of the text element represents the cover title, and the positioning information represented by the second prompt word is the information positioned based on the "cover title", for example, the content of the second prompt word is "add a sticker element below the cover title" or "add a halo sticker on the outline of the cover subtitle". For another example, the semantic content feature of the background element (background image) represents "forest, stone, and stream", and the content of the second prompt word is, for example, "add a sticker element on the stone", and so on. In this embodiment, by automatically generating the sticker element and further limiting the positional relationship between the sticker element and the background element and the text element through the second prompt word, the flexibility and effect of the cover template are further improved.
[0089] It should be noted that the above steps S2040 and S2043 are optional execution steps, which can be set as needed. In the case of executing the above two steps respectively, the specific implementation of step S2042 has been introduced in the above embodiment; when the above steps S2040 and S2043 are executed simultaneously, correspondingly, the specific implementation of step S2042 is:
[0090] Step S2042C: generating the sticker layer data by the first image feature, the second image feature and the second prompt word. Thus, the style content feature and the position feature of the sticker element in the layer are automatically configured at the same time, the matching degree of the sticker layer data with the image element in the background layer and the script layer is improved, and the quality of the cover template is improved.
[0091] Correspondingly, when step S2042 is executed in three ways of step S2042A, step S2042B and step S2042C, Figure 10 The specific implementation mode of step S2042 provided by the illustrated embodiment is also correspondingly refined into the corresponding implementation mode, specifically:
[0092] When step S2042 is executed in step S2042A, the implementation mode of step S2042-1 is to generate the target sticker layer according to the first image feature and the second image feature of the model input image.
[0093] When step S2042 is executed in step S2042B, the implementation mode of step S2042-1 is to generate the target sticker layer according to the first image feature and the second prompt word of the model input image.
[0094] When step S2042 is executed in step S2042C, the implementation mode of step S2042-1 is to generate the target sticker layer according to the first image feature, the second image feature and the second prompt word of the model input image.
[0095] The specific implementation mode of the above-mentioned three implementation modes of step S2042-1 has been introduced in the previous steps, and the rest of the steps (step S2042-2 to step S2042-4) are not changed.
[0096] Step S205: combining the corresponding script layer, background layer and sticker layer according to the base template data and sticker layer data to generate a cover template, wherein the cover template is used to generate a video cover picture in combination with a configuration instruction for the layers in the cover template, and the configuration instruction is used to configure the elements in the layers.
[0097] Illustratively, finally, after obtaining the sticker layer data, one or more sticker layers in the above-mentioned sticker layer data are superimposed with the background layer and the script layer to generate a cover template. Then, each layer in the cover template can respond to the configuration instruction to modify the content, such as modifying the script content of the script element in the layer and modifying the sticker element in the sticker layer, so as to obtain the required video cover picture.
[0098] In a possible implementation, after step S204 is performed, the method further includes the following steps:
[0099] Step S205A: processing the model input image by the second image model to obtain layer position data, the layer position data being used to represent the layer relationship between the sticker layer, the background layer and / or the text layer.
[0100] Correspondingly, as shown in Figure 2 the specific implementation of step S205 includes the following steps:
[0101] Step S2051: generating a layer sequence according to the layer position data, the layer sequence including the layer identifiers of the layers constituting the cover template.
[0102] Step S2052: sequentially combining the text layer, the background layer and the sticker layer according to the layer sequence to generate the cover template.
[0103] As shown in the foregoing embodiments, the non-transparent regions of the layers of the cover template will affect each other, and therefore, in order to avoid the problem of effective information being blocked, the display order of the layers needs to be sorted. In this embodiment, after the sticker layer data is generated or while the sticker layer data is being generated, the second image processing model is further used to generate corresponding layer position data by using the model input image, the layer position data representing the layer relationship between the sticker layer, the background layer and / or the text layer, for example, the background layer being under the text layer, the text layer Layer_1 being above the sticker layer Layer_2, and the like. Then, all the layers are sorted based on the layer position relationship to obtain the layer identifiers of the layers, and finally, the layers are combined based on the layer identifiers to form the cover template.
[0104] In this embodiment, the second image model is used to generate the layer position data while or after the sticker layer data is generated, and then the layer identifiers of the layers are determined based on the layer position data, and the layers are sorted based on the layer identifiers to form the cover template, so that the display and blocking order of the layers is controlled, the effective information is prevented from being blocked, and the effect of the video cover image generated by the cover template is improved.
[0105] In this embodiment, the implementation of steps S201-S202 is the same as that of steps S101-S102 in the embodiment of the present disclosure, which will not be repeated here. Figure 11
[0106] Corresponding to the image template generation method of the foregoing embodiments, Figure 11 A structural block diagram of an image template generation apparatus is provided for the embodiments of the present disclosure. The method introduced in the above embodiments can be executed by the image template generation apparatus, which can be implemented in a software and / or hardware manner, and can be integrated in an electronic device with certain data processing functions. The electronic device can include, but is not limited to, a mobile terminal with large data processing capability, and a desktop computer, a supercomputer, and other fixed terminals with large data processing capability.
[0107] For ease of illustration, only parts related to the embodiments of the present disclosure are shown. Referring to Figure 12 , the image template generation apparatus 3 includes:
[0108] The first generation module 31 is configured to process a first prompt word based on a first image model to generate an initial cover image, wherein the first prompt word is used to describe at least title information in the initial cover image.
[0109] The parsing module 32 is configured to generate basic template data from the initial cover image, wherein the basic template data includes a background layer and a text layer, the background layer is used to carry background elements in the initial cover image, and the text layer is used to carry text elements in the initial cover image.
[0110] The second generation module 33 is configured to process the basic template data based on a second image model to generate a cover template, wherein the cover template includes the background layer, the text layer, and a sticker layer, the layer content of the sticker layer is determined based on the basic template data, and the cover template is used to generate a video cover image in combination with configuration instructions for layers in the cover template, and the configuration instructions are used to configure elements in the layers.
[0111] According to one or more embodiments of the present disclosure, the parsing module 32 is specifically configured to: parse image content of the initial cover image to obtain a title area; perform text element extraction on the initial cover image based on the title area to generate at least one text layer; perform content filling on the title area of the initial cover image to generate a background layer; and generate the basic template data based on the at least one text layer and the background layer.
[0112] According to one or more embodiments of the present disclosure, the second generation module 33 is specifically configured to: render the background layer and the text layer into a model input image according to the basic template data; process the model input image through the second image model to obtain sticker layer data, wherein the sticker layer data includes at least one different sticker layer; and combine the corresponding text layer, the background layer, and the sticker layer according to the basic template data and the sticker layer data to generate the cover template.
[0113] According to one or more embodiments of the present disclosure, the second generation module 33 is specifically configured to: extract, by the second image model, first image features of the model input image, where the first image features represent semantic content features of the background elements and the text elements in the model input image; and generate the sticker layer data according to the first image features of the model input image.
[0114] According to one or more embodiments of the present disclosure, the second generation module 33 is further configured to: extract, by the second image model, second image features of the model input image, where the second image features represent style content features of the background elements and the text elements in the model input image; and generate the sticker layer data according to the first image features and the second image features of the model input image.
[0115] According to one or more embodiments of the present disclosure, the second generation module 33 is specifically configured to: generate a target sticker layer according to the first image features of the model input image; obtain a rendering effect image based on the model input image and the target sticker layer; if the rendering effect image does not satisfy a preset condition, update the model input image to the rendering effect image, and return to the step of generating the target sticker layer according to the first image features of the model input image; and if the rendering effect image satisfies the preset condition, generate the sticker layer data based on the target sticker layer.
[0116] According to one or more embodiments of the present disclosure, the second generation module 33 is further configured to: obtain a second prompt word, where the second prompt word is used to represent positioning information described based on the semantic content features of the background elements and / or the text elements; and generate the sticker layer data according to the second prompt word and the first image features, where a sticker element of a sticker layer in the sticker layer data is located at a target position corresponding to the positioning information.
[0117] According to one or more embodiments of the present disclosure, the second generation module 33 is further configured to: obtain layer position data by processing the model input image based on the second image model, where the layer position data is used to represent a layer relationship between the sticker layer and the background layer and / or the text layer; and generate a layer sequence according to the layer position data, where the layer sequence includes layer identifiers of each layer constituting the cover template; and sequentially combine the text layer, the background layer and the sticker layer to generate the cover template according to the layer sequence.
[0118] According to one or more embodiments of the present disclosure, the cover template has a first configuration interface and / or a second configuration interface; the first configuration interface is used to configure elements within the target layer in response to the first configuration instruction to generate a video cover image; the second configuration interface is used to obtain video information of the target video in response to the second configuration instruction, and configure at least one element within the target layer according to the video information to generate a video cover image of the target video; wherein the target layer includes at least one of a background layer, a text layer, and a sticker layer.
[0119] The first generation module 31, the analysis module 32 and the second generation module 33 are connected in sequence. The image template generation device 3 provided in this embodiment can implement the technical solution of the above method embodiment, and its implementation principle and technical effect are similar, which will not be repeated in this embodiment.
[0120] Figure 12 A schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure is shown in FIG. Figures 2-10 As shown, the electronic device 4 includes:
[0121] A processor 41, and a memory 42 communicatively connected to the processor 41;
[0122] Memory 42 stores computer-executable instructions;
[0123] The processor 41 executes the computer execution instructions stored in the memory 42 to implement the following Figures 2-10 The image template generation method in the illustrated embodiment.
[0124] Optionally, the processor 41 and the memory 42 are connected via a bus 43 .
[0125] For related instructions, please refer to Figures 2-10 The relevant descriptions and effects corresponding to the steps in the corresponding embodiments can be understood, and no further details are given here.
[0126] The present invention provides a computer-readable storage medium that stores computer-executable instructions. When the computer-executable instructions are executed by a processor, the computer-executable instructions are used to implement the present invention. Figures 2-10 The image template generation method provided in any one of the corresponding embodiments.
[0127] The present invention provides a computer program product, including a computer program, which implements the present invention when executed by a processor. Figure 13 The image template generation method provided in any one of the corresponding embodiments.
[0128] In order to implement the above embodiment, the embodiment of the present disclosure further provides an electronic device.
[0129] Reference Figure 13 which shows a structural diagram of an electronic device 900 suitable for use in implementing embodiments of the present disclosure, which can be a terminal device or a server. Among them, the terminal device can include but is not limited to mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, personal digital assistants (PDA), tablet computers (PAD), portable multimedia players (PMP), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), and the like, and fixed terminals such as digital TVs, desktop computers, and the like. Figure 13 The electronic device shown is only an example and should not impose any limitation on the functions and use range of the embodiments of the present disclosure.
[0130] As Figure 13 shown, the electronic device 900 can include a processing device (such as a central processor, a graphics processor, etc.) 901, which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 902 or programs loaded into a random access memory (RAM) 903 from a storage device 908. In the RAM 903, various programs and data required for the operation of the electronic device 900 are also stored. The processing device 901, the ROM 902, and the RAM 903 are connected to each other through a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0131] In general, the following devices can be connected to the I / O interface 905: input devices 906 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, and the like; output devices 907 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, and the like; storage devices 908 including, for example, a magnetic tape, a hard disk, and the like; and communication devices 909. The communication devices 909 can allow the electronic device 900 to communicate with other devices wirelessly or by wire to exchange data. Although The electronic device 900 is shown with various devices, but it should be understood that all of the devices shown are not required, and more or less devices can alternatively be implemented.
[0132] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer readable medium, the computer program comprising program code for performing the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via the communication device 909, or installed from the storage device 908, or installed from the ROM 902. When the computer program is executed by the processing device 901, the above-mentioned functions defined in the methods of the embodiments of the present disclosure are performed.
[0133] It should be noted that the computer readable medium described above in the present disclosure can be a computer readable signal medium or a computer readable storage medium or any combination thereof. The computer readable storage medium may, for example, be, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the above. More specific examples of the computer readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device. In the present disclosure, the computer readable signal medium can include a data signal carried in a baseband or as part of a carrier wave, in which the computer readable program code is carried. Such a propagated data signal can take any of a variety of forms, including but not limited to electro-magnetic, optical, or any suitable combination thereof. The computer readable signal medium can also be any computer readable medium that is not a computer readable storage medium and that can communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer readable medium can be transmitted by any suitable medium, including but not limited to wire, cable, RF (radio frequency), or any suitable combination thereof.
[0134] The computer readable medium described above can be included in the electronic device described above; or can exist separately from the electronic device and be not assembled in the electronic device.
[0135] The computer readable medium described above carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the methods shown in the embodiments described above.
[0136] Computer program code for carrying out operations of the present disclosure can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0137] The flow diagrams and the block diagrams in the drawings are illustrations of architectures, functionalities, and operations of possible implementations of systems, methods, and computer program products according to various embodiments of present disclosure. In this regard, each block in the flow diagrams or block diagrams can represent a module, a procedure, or a part of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or in the reverse order, depending on the functionality involved. It is also noted that each block of the block diagrams and / or flow diagrams and combinations of blocks in the block diagrams and / or flow diagrams can be implemented by dedicated hardware-based systems that perform the specified functions or operations, or by combinations of dedicated hardware and computer instructions.
[0138] The units or modules described in the embodiments of the present disclosure can be implemented by software, or by hardware. In some cases, the name of the unit or module does not constitute a limitation on the unit itself.
[0139] The functions described in this specification can be implemented in part or in whole by one or more hardware logic components. For example, and without limitation, illustrative types of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Program-specific Integrated Circuits (ASICs), Program-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.
[0140] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more of: an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0141] In a first aspect, according to one or more embodiments of the present disclosure, an image template generation method is provided, comprising:
[0142] processing a first prompt word based on a first image model to generate an initial cover image, wherein the first prompt word is used at least to describe title information in the initial cover image; generating basic template data according to the initial cover image, wherein the basic template data includes a background layer and a text layer, the background layer is used to carry background elements in the initial cover image, and the text layer is used to carry text elements in the initial cover image; processing the basic template data based on a second image model to generate a cover template, wherein the cover template includes the background layer, the text layer, and a sticker layer, the layer content of the sticker layer is determined based on the basic template data, and the cover template is used to generate a video cover image in combination with configuration instructions for layers in the cover template, and the configuration instructions are used to configure elements in the layers.
[0143] According to one or more embodiments of the present disclosure, the generating basic template data according to the initial cover image comprises: analyzing image content of the initial cover image to obtain a title area; performing text element extraction on the initial cover image based on the title area to generate at least one text layer; performing content filling on the title area of the initial cover image to generate a background layer; and generating the basic template data based on the at least one text layer and the background layer.
[0144] According to one or more embodiments of the present disclosure, the generating the cover template based on the second image model and the base template data comprises: rendering the background layer and the text layer into a model input image according to the base template data; processing the model input image through the second image model to obtain sticker layer data, the sticker layer data comprising at least one different sticker layer; and combining the corresponding text layer, background layer and sticker layer according to the base template data and the sticker layer data to generate the cover template.
[0145] According to one or more embodiments of the present disclosure, the processing the model input image through the second image model to obtain sticker layer data comprises: extracting first image features of the model input image through the second image model, wherein the first image features represent semantic content features of the background elements and text elements in the model input image; and generating the sticker layer data according to the first image features of the model input image.
[0146] According to one or more embodiments of the present disclosure, the method further comprises: extracting second image features of the model input image through the second image model, the second image features representing style content features of the background elements and text elements in the model input image; and generating the sticker layer data according to the first image features and the second image features of the model input image.
[0147] According to one or more embodiments of the present disclosure, the generating the sticker layer data according to the first image features of the model input image comprises: generating a target sticker layer according to the first image features of the model input image; obtaining a rendering effect image based on the model input image and the target sticker layer; if the rendering effect image does not satisfy a preset condition, updating the model input image to the rendering effect image and returning to execute the step of generating the target sticker layer according to the first image features of the model input image; and if the rendering effect image satisfies the preset condition, generating sticker layer data based on the target sticker layer.
[0148] According to one or more embodiments of the present disclosure, the method further comprises: obtaining a second prompt word, the second prompt word being used to represent positioning information described based on the semantic content features of the background elements and / or text elements; and generating the sticker layer data according to the first image features of the model input image comprises: generating the sticker layer data through the second prompt word and the first image features, wherein a sticker element of a sticker layer in the sticker layer data is located at a target position corresponding to the positioning information.
[0149] According to one or more embodiments of the present disclosure, the method further comprises: processing the model input image through the second image model to obtain layer position data, the layer position data being used to represent the layer relationship of the sticker layer with the background layer and / or the text layer; and combining the corresponding text layer, background layer and sticker layer according to the base template data and sticker layer data to generate the cover template, comprising: generating a layer sequence according to the layer position data, the layer sequence comprising layer identifiers of each layer constituting the cover template; and sequentially combining the text layer, background layer and sticker layer according to the layer sequence to generate the cover template.
[0150] According to one or more embodiments of the present disclosure, the cover template has a first configuration interface and / or a second configuration interface; the first configuration interface is used to configure elements in a target layer in response to a first configuration instruction to generate a video cover picture; the second configuration interface is used to obtain video information of a target video and configure at least one element in a target layer according to the video information in response to a second configuration instruction to generate a video cover picture of the target video; and the target layer comprises at least one of a background layer, a text layer and a sticker layer.
[0151] In a second aspect, according to one or more embodiments of the present disclosure, an image template generation device is provided, comprising:
[0152] A first generation module is configured to generate an initial cover picture based on processing a first prompt word based on a first image model, wherein the first prompt word is used to describe at least title information in the initial cover picture;
[0153] A parsing module is configured to generate base template data according to the initial cover picture, wherein the base template data comprises a background layer and a text layer, the background layer is used to carry background elements in the initial cover picture, and the text layer is used to carry text elements in the initial cover picture;
[0154] A second generation module is configured to generate a cover template based on processing the base template data based on a second image model, wherein the cover template comprises the background layer, the text layer and a sticker layer, the layer content of the sticker layer is determined based on the base template data, and the cover template is used to generate a video cover picture in combination with a configuration instruction for layers in the cover template, the configuration instruction being used to configure elements in the layers.
[0155] According to one or more embodiments of the present disclosure, the parsing module is specifically configured to: parse image content of the initial cover picture to obtain a title area; perform text element extraction on the initial cover picture based on the title area to generate at least one script layer; perform content filling on the title area of the initial cover picture to generate a background layer; and generate basic template data based on the at least one script layer and the background layer.
[0156] According to one or more embodiments of the present disclosure, the second generation module is specifically configured to: render the background layer and the script layer into a model input image according to the basic template data; process the model input image through the second image model to obtain sticker layer data, the sticker layer data including at least one different sticker layer; and combine the corresponding script layer, background layer, and sticker layer based on the basic template data and sticker layer data to generate the cover template.
[0157] According to one or more embodiments of the present disclosure, when processing the model input image through the second image model to obtain sticker layer data, the second generation module is specifically configured to: extract first image features of the model input image through the second image model, wherein the first image features represent semantic content features of background elements and text elements in the model input image; and generate the sticker layer data according to the first image features of the model input image.
[0158] According to one or more embodiments of the present disclosure, the second generation module is further configured to: extract second image features of the model input image through the second image model, the second image features representing style content features of the background elements and text elements in the model input image; and when generating the sticker layer data according to the first image features of the model input image, the second generation module is specifically configured to: generate the sticker layer data according to the first image features and the second image features of the model input image.
[0159] According to one or more embodiments of the present disclosure, when generating the sticker layer data according to the first image features of the model input image, the second generation module is specifically configured to: generate a target sticker layer according to the first image features of the model input image; obtain a rendering effect picture based on the model input image and the target sticker layer; if the rendering effect picture does not satisfy a preset condition, update the model input image to the rendering effect picture, and return to execute the step of generating the target sticker layer according to the first image features of the model input image; and if the rendering effect picture satisfies the preset condition, generate sticker layer data based on the target sticker layer.
[0160] According to one or more embodiments of the present disclosure, the second generation module is further configured to: obtain a second prompt word, the second prompt word being used to represent positioning information described based on the semantic content feature of the background element and / or the text element; and when generating the sticker layer data according to the first image feature of the model input image, the second generation module is specifically configured to: generate the sticker layer data by using the second prompt word and the first image feature, wherein a sticker element of a sticker layer in the sticker layer data is located at a target position corresponding to the positioning information.
[0161] According to one or more embodiments of the present disclosure, the second generation module is further configured to: process the model input image by using the second image model to obtain layer position data, the layer position data being used to represent a layer relationship between the sticker layer and the background layer and / or the text layer; and when generating the cover template by combining the corresponding text layer, the background layer and the sticker layer according to the base template data and the sticker layer data, the second generation module is specifically configured to: generate a layer sequence according to the layer position data, the layer sequence including layer identifiers of each layer constituting the cover template; and perform ordered combination on the text layer, the background layer and the sticker layer according to the layer sequence to generate the cover template.
[0162] According to one or more embodiments of the present disclosure, the cover template has a first configuration interface and / or a second configuration interface; the first configuration interface is used to configure elements in a target layer in response to a first configuration instruction to generate a video cover picture; and the second configuration interface is used to obtain video information of a target video and configure at least one element in a target layer according to the video information in response to a second configuration instruction to generate a video cover picture of the target video; wherein the target layer includes at least one of a background layer, a text layer and a sticker layer.
[0163] In a third aspect, according to one or more embodiments of the present disclosure, an electronic device is provided, including: at least one processor and a memory;
[0164] The memory stores computer execution instructions;
[0165] The at least one processor executes the computer execution instructions stored in the memory, so that the at least one processor performs the image template generation method as described in the above first aspect and various possible designs of the first aspect.
[0166] In a fourth aspect, according to one or more embodiments of the present disclosure, a computer readable storage medium is provided, and the computer readable storage medium has stored therein computer executing instructions which, when executed by a processor, implement the image template generation method according to the first aspect and various possible designs of the first aspect.
[0167] In a fifth aspect, according to one or more embodiments of the present disclosure, a computer program product is provided, and the computer program product includes a computer program which, when executed by a processor, implements the image template generation method according to the first aspect and various possible designs of the first aspect.
[0168] The above description merely illustrates the preferred embodiments of the present disclosure and the principles of the applied technologies. It should be understood by those skilled in the art that the disclosed scope of the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and should also cover other technical solutions formed by the combinations of the above technical features or equivalent features, without departing from the above disclosed concepts. For example, the above technical features can be replaced with the technical features disclosed in the present disclosure (but not limited to) having similar functions to form technical solutions.
[0169] In addition, although each operation is described in a particular order, this should not be understood as requiring the operations to be performed in the particular order shown or in a sequential order. In certain circumstances, multitasking and parallel processing can be advantageous. Similarly, although several implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments can also be combined in a single embodiment. Conversely, various features described in the context of a single embodiment can also be separated and implemented in multiple embodiments.
[0170] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.
Claims
1. A method for generating an image template, characterized in that: include: Processing a first prompt word based on a first image model to generate an initial cover image, wherein the first prompt word is used to at least describe title information in the initial cover image; Generate basic template data based on the initial cover image, wherein the basic template data includes a background layer and a text layer, wherein the background layer is used to carry the background elements in the initial cover image, and the text layer is used to carry the text elements in the initial cover image; Based on the second image model, the basic template data is processed to generate a cover template, wherein the cover template includes the background layer, the text layer and the sticker layer, and the layer content of the sticker layer is determined based on the basic template data. The cover template is used to generate a video cover image in combination with the configuration instructions for the layers within the cover template, and the configuration instructions are used to configure the elements within the layers.
2. The method according to claim 1, characterized in that The step of generating basic template data according to the initial cover image includes: Parsing the image content of the initial cover image to obtain a title area; Extracting text elements from the initial cover image based on the title area to generate at least one text layer; Filling the title area of the initial cover image with content to generate a background layer; Based on the at least one text layer and the background layer, basic template data is generated.
3. The method according to claim 1, characterized in that The step of processing the basic template data based on the second image model to generate a cover template includes: Rendering the background layer and the text layer into a model input image according to the basic template data; Processing the model input image through the second image model to obtain sticker layer data, wherein the sticker layer data includes at least one different sticker layer; According to the basic template data and sticker layer data, the corresponding text layer, background layer and sticker layer are combined to generate the cover template.
4. The method according to claim 3, characterized in that The step of processing the model input image by the second image model to obtain sticker layer data includes: extracting a first image feature of the model input image through the second image model, wherein the first image feature represents semantic content features of background elements and text elements in the model input image; The sticker layer data is generated according to the first image feature of the model input image.
5. The method according to claim 4, characterized in that The method further comprises: extracting second image features of the model input image through the second image model, where the second image features represent style and content features of background elements and text elements in the model input image; Generating the sticker layer data according to the first image feature of the model input image includes: The sticker layer data is generated according to the first image feature and the second image feature of the model input image.
6. The method according to claim 4, characterized in that Generating the sticker layer data according to the first image feature of the model input image includes: Generate a target sticker layer based on the first image feature of the model input image; Obtaining a rendering effect image based on the model input image and the target sticker layer; If the rendering effect image does not meet the preset conditions, the model input image is updated to the rendering effect image, and the step of generating the target sticker layer according to the first image feature of the model input image is returned to execution; If the rendering effect image satisfies the preset conditions, sticker layer data is generated based on the target sticker layer.
7. The method according to claim 4, characterized in that The method further comprises: Acquire a second prompt word, where the second prompt word is used to represent positioning information based on the semantic content feature description of the background element and / or text element; Generating the sticker layer data according to the first image feature of the model input image includes: The sticker layer data is generated by using the second prompt word and the first image feature, wherein the sticker elements of the sticker layer in the sticker layer data are located at the target position corresponding to the positioning information.
8. The method according to claim 3, characterized in that The method further comprises: Processing the model input image using the second image model to obtain layer position data, wherein the layer position data is used to represent the layer relationship between the sticker layer and the background layer and / or the text layer; The method of generating the cover template by combining the corresponding text layer, background layer and sticker layer according to the basic template data and sticker layer data includes: generating a layer sequence according to the layer position data, wherein the layer sequence includes layer identifiers of the layers constituting the cover template; According to the layer sequence, the text layer, background layer and sticker layer are orderly combined to generate the cover template.
9. The method according to claim 1, characterized in that The cover template has a first configuration interface and / or a second configuration interface; The first configuration interface is used to configure elements in the target layer in response to the first configuration instruction to generate a video cover image; The second configuration interface is used to obtain video information of the target video in response to the second configuration instruction, and configure at least one element in the target layer according to the video information to generate a video cover image of the target video; The target layer includes at least one of a background layer, a text layer, and a sticker layer.
10. An image template generating device, characterized in that: include: A first generating module, configured to process a first prompt word based on a first image model to generate an initial cover image, wherein the first prompt word is used to at least describe title information in the initial cover image; A parsing module, configured to generate basic template data based on the initial cover image, wherein the basic template data includes a background layer and a text layer, wherein the background layer is configured to carry background elements in the initial cover image, and the text layer is configured to carry text elements in the initial cover image; The second generation module is used to process the basic template data based on the second image model to generate a cover template, wherein the cover template includes the background layer, the text layer and the sticker layer, and the layer content of the sticker layer is determined based on the basic template data. The cover template is used to generate a video cover image in combination with the configuration instructions for the layers in the cover template, and the configuration instructions are used to configure the elements in the layers.
11. An electronic device, characterized in that: include: processor and memory; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor performs the image template generation method according to any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions. When the processor executes the computer-executable instructions, the image template generation method according to any one of claims 1 to 9 is implemented.
13. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the image template generation method according to any one of claims 1 to 9 is implemented.