Picture generation method and related device
By obtaining the video name to filter the target image template and combining it with large model matching, the problem of poor matching between the image template and the video image is solved, and a more effective image is generated to meet the needs of the video to attract users.
Patent Information
- Application Number
- CN202511139059.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-14
- Publication Date
- 2025-10-03
AI Technical Summary
In the prior art, the matching degree between the image template and the video image is not high, resulting in poor presentation effect of the produced image.
By obtaining the video name, determining the target video and filtering out the target image template representing the video content from the image template collection, using the large model for screening and matching, combining frame extraction operations and filling into the target image template, the target image is generated.
The matching degree between the image template and the video image has been improved to ensure that the generated image has a better presentation effect and can attract users to click and watch the video.
Smart Images

Figure CN120751216A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image generation technology, and in particular to an image generation method and related devices. Background Art
[0002] To attract users to click and watch videos, it's necessary to create relevant images based on the videos, such as advertisements and posters, and then place these images on video playback platforms. Currently, this is primarily done manually, requiring the manual selection of video images and image templates. These templates can be pre-designed image frames or layouts, and the creation of finished images by adding the video image to the template. When the image template and video image don't match well, the resulting image can suffer from poor presentation quality. Summary of the Invention
[0003] In view of the above problems, this application provides a method and related device for generating an image, which can improve the matching degree between the template and the video image to ensure the image presentation effect. The specific solution is as follows:
[0004] The first aspect of the present application provides a method for generating an image, comprising:
[0005] Obtaining task information; wherein the task information includes a video name;
[0006] Determine a target video with the video name;
[0007] Filtering a target picture template representing the video content of the target video from a picture template set; wherein the picture template set includes a plurality of picture templates;
[0008] Fill the target video image into the target image template to obtain the target image.
[0009] In a possible implementation, selecting a target picture template representing the video content of the target video from the picture template set includes:
[0010] Acquire data information of the target video; wherein the data information indicates the video content of the target video;
[0011] Filling the target video's data information into a prompt word template to obtain a target prompt word; wherein the prompt word template describes a mapping relationship between an image template and video content;
[0012] Based on the target prompt word, the large model is called to screen the picture templates in the picture template set to obtain a target picture template representing the video content of the target video.
[0013] In a possible implementation, based on the target prompt word, calling the large model to screen the image templates in the image template set to obtain the target image template representing the video content of the target video includes:
[0014] Filtering the picture templates applied to the video content of the target video based on the field information of the picture templates in the picture template set to obtain filtered picture templates;
[0015] Based on the target prompt word, the large model is called to filter the filtered image templates to obtain a target image template representing the video content of the target video.
[0016] In a possible implementation, the task information further includes picture delivery information, and the picture delivery information includes information required to deliver the picture to the video playback platform;
[0017] The step of screening the picture templates applied to the video content of the target video based on the field information of the picture templates in the picture template set to obtain the screened picture templates includes:
[0018] Searching for a picture template in the picture template set, the picture template having the picture placement information included in the field information of the picture template, and obtaining the searched picture template;
[0019] Picture templates containing target information in the field information of the searched picture templates are filtered from the searched picture templates to obtain filtered picture templates; wherein the target information represents the video content of the target video.
[0020] In a possible implementation, the task information further includes a target object, and the target video image includes a frame-drawing image;
[0021] Filling the target video image into the target image template to obtain the target image includes:
[0022] Selecting a frame extraction image with the same aspect ratio as the picture filling area of the target picture template from the frame extraction images of the target video to obtain a first frame extraction image;
[0023] Filtering the frame abstraction image containing the target object from the first frame abstraction image to obtain a second frame abstraction image;
[0024] Fill the second extracted frame image into the image filling area to obtain the target image.
[0025] In a possible implementation, the task information further includes the total number of the target images;
[0026] Filling the target video image into the target image template to obtain the target image includes:
[0027] When the number of target pictures obtained after filling the pictures of the target video into one target picture template is less than the total number, the pictures of the target video are filled into multiple target picture templates respectively so that the number of target pictures obtained reaches the total number.
[0028] In a possible implementation, the method further includes:
[0029] Get frame extraction parameters;
[0030] According to the frame extraction parameters, a frame extraction operation is performed on the target video to obtain the frame extraction image having the target object.
[0031] A second aspect of the present application provides an image generation system, comprising:
[0032] A task information acquisition module is configured to acquire task information; wherein the task information includes a video name;
[0033] a target video determining module, configured to determine a target video having the video name;
[0034] The screening module is configured to screen out a target picture template representing the video content of the target video from the picture template set; wherein the picture template set includes multiple picture templates
[0035] The filling module is configured to fill the target video image into the target image template to obtain the target image.
[0036] A third aspect of the present application provides a computer program product, comprising computer-readable instructions. When the computer-readable instructions are executed on an electronic device, the electronic device implements the image generation method of the first aspect or any implementation of the first aspect.
[0037] A fourth aspect of the present application provides an electronic device, comprising at least one processor and a memory connected to the processor, wherein:
[0038] The memory is used to store computer programs;
[0039] The processor is used to execute the computer program so that the electronic device can implement the picture generation method of the above-mentioned first aspect or any implementation manner of the first aspect.
[0040] In a fifth aspect, the present application provides a computer storage medium, which carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can use the image generation method of the above-mentioned first aspect or any implementation of the first aspect.
[0041] With the help of the above technical solution, the image generation method and related devices provided by this application can obtain the target video with the video name by obtaining the video name, and filter the target image template representing the video content of the target video from the image template set, so as to ensure the matching degree between the filtered target image template and the target video, and fill the image of the target video into the target image template, so that the generated target image has a better presentation effect, which helps to attract users to click and watch the video. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale.
[0043] Figure 1 A flow chart of a method for generating an image provided in this application;
[0044] Figure 2 A schematic diagram of an image template database provided for this application;
[0045] Figure 3 This is an example of a picture template provided for this application;
[0046] Figure 4 Another example of an image template provided for this application;
[0047] Figure 5 Another example of an image template provided for this application;
[0048] Figure 6 A structural diagram of an image generation system provided in this application;
[0049] Figure 7 This is a schematic diagram of the structure of an electronic device provided in this application. DETAILED DESCRIPTION
[0050] The following describes the embodiments of the present application in conjunction with the accompanying drawings. The terms used in the implementation methods of the present application are only used to explain the specific embodiments of the present application and are not intended to limit the present application.
[0051] The embodiments of the present application are described below in conjunction with the accompanying drawings. Those skilled in the art will appreciate that, with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0052] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequential order. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances, and this is merely a way of distinguishing the objects of the same attributes when describing them in the embodiments of the present application. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, so that the process, method, system, product or equipment comprising a series of units need not be limited to those units, but may include other units that are not clearly listed or inherent to these processes, methods, products or equipment.
[0053] Reference Figure 1 , Figure 1 A schematic diagram of a process for generating an image provided in an embodiment of the present application is shown in FIG. Figure 1 As shown, an image generation method provided in an embodiment of the present application may include steps 101 to 104, and these steps are described in detail below.
[0054] Step 101: Obtain task information; wherein the task information includes a video name.
[0055] This task information is the image task information that needs to be produced. For example, in the field of film and television advertising, advertising image materials are needed for delivery. In order to attract users to click and watch videos, relevant advertising images need to be produced according to different videos. Then this task information is the advertising image task information that needs to be produced for the video. This task information includes the video name. The video name is obtained through the film list, and the advertising image of the video can be automatically produced.
[0056] In addition to the video name, the task information may also include but is not limited to image delivery information, target objects, and the total number of images. The image delivery information includes the information required to deliver images to the video playback platform, which may include but is not limited to delivery channels, channel sizes, products, delivery scenarios, and production methods. Among them, delivery channels refer to various video playback platforms, channel sizes refer to the ad space sizes of the channels, products refer to the version types of the video playback platform, including regular versions and express versions, delivery scenarios include information flow and opening screens, and production methods include production only, production and delivery, and scheduled delivery.
[0057] For the automatic production of advertising images, the task information can include user requirements, such as the total number of images to be produced, whether a certain style of advertising image material is required or not, whether images of target objects are required or not, etc. The target objects may include target people, target scenery, target animals, etc. Based on this task information, advertising images that meet the user's requirements can be generated during the image production process.
[0058] Step 102: Determine a target video with a video name.
[0059] By using the video name in the task information, the corresponding video can be searched from the video database to obtain the target video with the video name. The video database can be a relational database such as MySQL.
[0060] After determining the target video, the video content of the target video can be obtained. The video content may include a video theme, which may reflect the subject matter and channel of the target video; wherein the subject matter reflects the content type of the target video, such as suspense, sweet pet, science fiction, historical drama, etc., and the channel represents the video type of the target video, such as TV series, movies, variety shows, etc.
[0061] The video database records the subject matter and channel of the video. After the target video is determined by the video name, the subject matter and channel of the target video can be obtained from the video database. Specifically, the subject matter and channel information of the target video can be obtained according to the video name through an SQL (Structured Query Language) query statement.
[0062] Step 103: Filter out a target picture template representing the video content of the target video from the picture template set; wherein the picture template set includes a plurality of picture templates.
[0063] The picture template set can be part of the picture templates in the picture template database, or it can be all the picture templates in the picture template database. A picture template is a pre-designed picture frame with a fixed layout structure and visual elements. Since the layout structure and visual elements in different picture templates are different, different picture templates can present different template styles and template styles. The template style can be reflected by the visual elements in the picture template, and the template style can be reflected by the layout structure of the picture fill box in the picture template. On the basis of the layout structure, the template style can be further reflected by the visual elements. The layout structure in the picture template is the arrangement framework of the content in the picture template, such as grid layout, symmetrical layout, asymmetrical layout, etc. The visual elements in the picture template are the specific visual components that constitute the picture template. The visual components may include but are not limited to text, images, colors, textures, etc.
[0064] Furthermore, a target image template whose template style and template style can represent the video content of the target video can be screened from the image template set. The video content can be presented by subject matter and channel, so an image template whose template style is suitable for the subject matter of the target video and whose template style is suitable for the channel of the target video can be selected. Of course, an image template whose template style is suitable for the channel of the target video and whose template style is suitable for the subject matter of the target video can also be selected.
[0065] Image template database such as Figure 2 As shown, the picture template example is as follows Figure 3 As shown. The fields of the image template include template style, template type, template style, delivery channel, usage scenario, template size, template ID (Identity document), product, drama title type, adapted channel, and frame extraction type. Among them, delivery channel refers to various video playback platforms, usage scenarios include information flow and opening screen, product refers to the version type of the video playback platform, including standard version and express version, drama title type includes search box and film and television drama logo (logotype), the template contains drama title type, and the video name is reflected in the form of search box and / or film and television drama logo, and frame extraction type includes single-person image, multi-person image, and frame extraction ratio.
[0066] Template styles include primary, secondary, general, and IP styles. Primary styles include, but are not limited to, ancient, modern, contemporary, and general. Secondary styles include, but are not limited to, romance, suspense, war, comedy, fantasy, horror, cartoon, campus, and adult animation. General-style image templates are suitable for all video themes. IP-style image templates are suitable for videos with corresponding video names. IP-style image templates carry the video name logo and / or images of video characters and are only applicable to a specific video.
[0067] Template types include style templates and IP (Intellectual Property) templates. Style templates include universal design elements and layouts. For example, the romance template is suitable for both historical and modern romance films. IP templates are customized for specific TV series and are not universal. They are tailored to the unique style, theme, and plot of a particular TV series. They include fill-in-box styles, stickers, the title, and unique text fonts relevant to the specific TV series. IP templates can further highlight the uniqueness of a TV series and enhance its visual recognition. Figure 3 and Figure 4 For style templates, Figure 3 The template style is modern, adult animation, Figure 4 The template style is Universal. Figure 5It is an IP template. The video name is also filled in the template category, and the template style is IP.
[0068] Template styles include but are not limited to clean templates, single-image templates, puzzle templates, filter templates, copy templates, cutout templates, comic templates, sports templates, skit templates, children's templates, and novel templates. Clean templates are applicable to all channels and are highly versatile, serving as a fallback. Single-image templates are suitable for channels such as movies, TV series, variety shows, and documentaries, focusing on a key image to convey the main message. Puzzle templates are suitable for channels such as movies, TV series, variety shows, and documentaries, combining multiple images to showcase different angles or related plot points. Filter templates are suitable for channels such as movies, TV series, variety shows, and documentaries, using filters to create a suitable atmosphere that matches the content style. Copy templates are suitable for channels such as movies, TV series, variety shows, and documentaries, focusing on text and using design to clearly convey key information. Cutout templates are suitable for channels such as movies, TV series, variety shows, and documentaries, cutting out the main image to enhance its presence or for creative compositing. Comic templates are suitable for channels such as animation, using a comic-style design to showcase the characteristics of the anime. The Sports template is suitable for sports channels, incorporating sports-related elements and using dynamic design to create an exciting atmosphere. The Skit template is suitable for original content channels, featuring a simple design that quickly attracts viewers and reflects the fast-paced and engaging nature of skits. The Children's template is suitable for children's channels, using bright colors and cartoon characters to create a fun atmosphere that is easy for children to understand. The Novel template is suitable for novel channels, using a custom novel style design.
[0069] Template styles are tailored to the video's subject matter, which reflects its content type. For example, if the video's content type is a romantic romance, the template style could be romance. Template styles also apply to the video's channel, which represents the target video's type. For a TV series, template styles could include clean templates, single-image templates, puzzle templates, filter templates, text templates, and cutout templates.
[0070] In practical applications, based on the subject matter and channel of the target video, a target picture template with a target template style and a target template style can be screened out from the picture template collection, so that the target template style is adapted to the subject matter of the target video, the target template style is suitable for the channel of the target video, and the obtained target picture template conforms to the content type and video type of the template video.
[0071] Of course, this image template is user-editable. Editing an image template can be divided into information editing and style editing. Information editing includes editing the template name, applicable channels, applicable scenarios (information flow, channel), applicable channels (TV series, movies, variety shows, etc.), template type (style template, IP template), template first-level style tags (ancient, modern, contemporary, general), and model second-level style tags (romance, suspense, war, comedy, fairy tale, horror, cartoon, campus, etc.). Style editing includes editing template size, TV series and film / TV series name fill-in box, frame extraction, stickers, filters, text and fonts, and other elements.
[0072] In a possible implementation, selecting a target image template representing the video content of a target video from the image template set includes:
[0073] Obtaining data information of a target video; wherein the data information indicates the video content of the target video;
[0074] Fill the target video's data information into the prompt word template to obtain the target prompt word; wherein the prompt word template describes the mapping relationship between the image template and the video content;
[0075] Based on the target prompt word, the large model is called to screen the image templates in the image template set to obtain the target image template that represents the video content of the target video.
[0076] The target video's profile information indicates the target video's video content. The profile information may include content description information of the target video. The content description information may include a video introduction that describes the core content of the video.
[0077] A prompt template is a structured text framework designed to help users efficiently generate high-quality AI (Artificial Intelligence) instructions. A prompt template can include a role definition, task description, input information, and output requirements. It can also include example references. The role definition specifies the AI's role (e.g., expert, assistant, screenwriter, etc.) to guide it in adjusting its response style. The task description includes the specific tasks the AI needs to complete (e.g., generation, analysis, etc.). Input information can include background information, data, and other content for the AI to consider. Output requirements can include constraints such as output format, length, and style.
[0078] To better enable the large model to match an image template suitable for the target video's content, the input information for the prompt word template can include a mapping between the image template and the video content, such as a sports image template for sports-related video content. By inserting the target video's profile information into the prompt word template, the large model can understand the profile information and determine the target video's content. The target image template representing the target video's content can then be retrieved through the mapping between the image template and the video content. The target video's profile information and the mapping between the image template and the video content serve as input for the prompt word template.
[0079] Furthermore, the target video's profile information may include, in addition to the target video's content description, the video title, genre, and channel. The target video's content description, genre, and channel are used to populate the prompt word template to obtain the target prompt word. This prompt word template can guide the macro model in selecting image templates. The video title, target video's content description, genre, and channel populated into the prompt word template can help the macro model find an image template that matches the target video. Based on the target video's content description, genre, and channel, the macro model can select target image templates whose template style and template format are suitable for the target video's video content.
[0080] Optionally, the large model can be a multimodal large model, specifically a GPT4O large model. Before calling the large model to screen the target image template, training data can be input into the large model to train the large model. The training data includes template images generated by the image template and their corresponding annotation information such as subject matter and channel. The template images are feature extracted and learned through a convolutional neural network, so that the large model can understand the visual effects of the template.
[0081] For example, the prompt word template content is as follows:
[0082] You're a master of film and television advertising creative design, familiar with the characteristics of various film and television themes and how to present them. When assigned a task, you need to accurately analyze the input "Film and TV Background Information," provide matching template information from a professional design perspective, and return the matching template style, template style, and template ID.
[0083] Input requirements
[0084]
Background information of film and TV series
[0085] Video Name:
[0086] Channel:
[0087] theme:
[0088] Video Introduction:
[0089] Template requirements
[0090] [Template Style] Carefully select from the following film and television genre categories, closely integrating the core elements of the plot, emotional tone, and target audience:
[0091] [01 - Ancient Style and Puppetry, 02 - Martial Arts, 03 - Mature Emotion, 04 - Sweet Love, 05 - Sad Romance, 06 - Suspense, 07 - Action, 08 - Comedy, 09 - Military War, 10 - Warm Friendship, 11 - Historical Drama, 12 - Nostalgia, 13 - Life Conflict, 14 - European and American / Adult Animation, 15 - Children's Animation, 16 - Variety Show, 17 - Police and Spy, 18 - Inspirational Growth, 19 - Food Record, 20 - Dynamic Competition, 21 - General, 22 - Sports, 30 - Inspirational, 31 - Function Introduction]
[0092] Template style requirements
[0093] [Template style] should be closely based on the plot and film channels, and carefully select multiple from the following reference tags:
[0094] [Clean template, single image template, puzzle template, filter template, copy template, cutout template, comic template, sports template, skit template, children's template, novel template]
[0095] The template styles are described as follows:
[0096] Clean version template: Applicable to all channels, highly versatile, and used as a backup.
[0097] Single image template: Suitable for movies, TV series, variety shows, and documentary channels. It focuses on displaying a key image and uses it to convey the main information.
[0098] Jigsaw Puzzle Template: Suitable for movies, TV series, variety shows, and documentary channels. It can put multiple pictures together to show different angles or related plots.
[0099] Filter template: Suitable for movies, TV series, variety shows, and documentary channels. Use filters to create a suitable atmosphere that fits the content style.
[0100] Copywriting template: Suitable for use in movies, TV series, variety shows, and documentary channels. It mainly displays text and uses design to make the text clearly convey key information.
[0101] Cutout template: Suitable for movies, TV series, variety shows, and documentary channels. It can cut out the main subject of the image to make it more prominent or be used for creative synthesis.
[0102] Comic template: Suitable for anime channels, designed in a comic style to showcase anime features.
[0103] Sports template: Suitable for sports channels, adding sports-related elements and using dynamic design to create a passionate atmosphere.
[0104] Short drama template: Suitable for original channels, with a simple design that quickly attracts viewers and is in line with the fast-paced and interesting characteristics of short dramas.
[0105] Children's template: Suitable for children's channels, using bright colors and cartoon images to create a fun atmosphere and make it easier for children to understand.
[0106] Novel template: suitable for novel channels, designed using a customized novel style.
[0107] Output requirements
[0108] Output must be presented in the following format:
[0109] Template style: [Specific video subject category selected]
[0110] Template style: [specific template style label]
[0111] Template id: [specific template id]
[0112] The above is a prompt word template. By filling the prompt word template with the video title, target video content description (i.e., video introduction), genre, and channel, you can obtain the target prompt word. Then, based on this target prompt word, the large model is invoked to filter image templates in the image template collection to obtain a target image template with the target template style and template pattern. The video introduction can be obtained from the video database, the template style can be obtained from the template style type database, and the template pattern can be obtained from the template database. This template database also stores basic information about the image template, such as template size, delivery scenario, product, and channel.
[0113] For example, the output of the large model is as follows:
[0114] Template styles: 17-Police and Spy, 12-Nostalgic Era, 10-Warm Friendship, 18-Inspirational Growth;
[0115] Template styles: puzzle template, filter template, single image template, copy template;
[0116] Template ID: 1201 (Nostalgic Era + Puzzle Template), 1002 (Warm Friendship + Retro Filter Template), 1205 (Nostalgic Era + Nostalgic Filter Template).
[0117] Then you can get the target image template based on the template id.
[0118] By calling the large model to filter image templates, intelligent template matching can be achieved. In addition, the prompt word template is universal. For different video content, you only need to fill the video name, video content description information, subject matter and channel into the fixed prompt word template to generate prompt words. It is simple, convenient and highly operational.
[0119] In one possible implementation, based on the target prompt word, the large model is called to filter the image templates in the image template set to obtain the target image template representing the video content of the target video, including:
[0120] Filtering the image templates applied to the video content of the target video based on the field information of the image templates in the image template set to obtain the filtered image templates;
[0121] Based on the target prompt word, the large model is called to filter the filtered image templates to obtain the target image template that represents the video content of the target video.
[0122] If there are a large number of image templates in the image template set, you can first filter out the image templates that are applied to the video content of the target video. Specifically, you can filter out the image templates that match the subject matter and channel of the target video, and then call the large model to perform a second screening on the filtered image templates to obtain the target image template that represents the video content of the target video, which can improve the accuracy of selecting image templates. The fields of the image template include template style and template style. The template style has a corresponding relationship with the subject matter of the video, and the template style has a corresponding relationship with the channel of the video. By filtering the fields, you can filter out image templates that match the subject matter and channel of the target video. In actual applications, you can compare the subject matter and channel corresponding to the image template with the channel and subject matter of the video, and use methods such as cosine similarity to calculate the matching degree between the image template and the video, and select multiple templates with high matching degrees.
[0123] Before calling the large model to screen image templates, the image templates are preliminarily screened by the subject matter and channel of the target video to obtain image templates whose template style is compatible with the subject matter of the target video and whose template style is suitable for the target video channel. This improves the accuracy of template matching, reduces the number of image templates that can be screened by the large model, improves the processing speed of the large model, and helps to screen out target image templates with higher matching degrees.
[0124] In a possible implementation, based on the field information of the picture templates in the picture template set, the picture templates applied to the video content of the target video are screened to obtain the screened picture templates, including:
[0125] Searching for an image template whose field information includes image placement information from the image template collection, and obtaining the searched image template;
[0126] Picture templates containing target information in field information of the searched picture templates are filtered from the searched picture templates to obtain filtered picture templates; wherein the target information represents the video content of the target video.
[0127] For the screening of image templates, it is necessary to satisfy that the image template matches the task information. The task information includes the image delivery information, which includes the delivery channel, channel size, product, and delivery scenario. When screening image templates that match the image delivery information, it is necessary to satisfy that the image template size matches the channel size. The delivery channel of the image template must be at least part of the delivery channels in the task information. The product field content of the image template must include at least part of the products in the task information. The usage scenarios of the image template must include at least part of the delivery scenarios in the task information.
[0128] Furthermore, the target information may include a target template style, a target template style, and a video name, wherein the target template style, the target template style, and the video name may represent the video content of the target video. After the image templates are screened, the image templates containing the image placement information are included in the field information of the image templates, and the searched image templates are then searched for image templates whose field information contains the target template style, the target template style, and the video name, thereby obtaining the screened image templates.
[0129] When filtering picture templates from searched picture templates whose field information contains the target template style, you can search for IP style templates and style templates corresponding to the subject matter of the target video. If there are no IP style, first-level style, or second-level style picture templates, you can choose a general style picture template.
[0130] When filtering the image templates whose field information contains the target template style from the searched image templates, you can search for the template style corresponding to the channel of the target video. If the channel is sports, you can choose the sports template. If the channel is original (short drama), you can choose the short drama template. If the channel is animation, you can choose the animation template.
[0131] When filtering image templates whose field information contains the video name from the searched image templates, you can search whether the video has a corresponding drama logo. If there is a drama logo, you can select the template with the drama logo. If there is no drama logo, you can select the template with a drama search box.
[0132] It should be noted that if, based on the field information of the image templates in the image template collection, image templates that match the subject matter and channel of the target video are screened, the screened image templates are then filtered based on the target prompt word and the large model is called to filter the screened image templates. If no image template that matches the subject matter and channel of the video is obtained, a clean version template applicable to all videos can be used. In addition, the screened image templates obtained can be multiple image templates, and the number of target image templates can also be multiple, which facilitates subsequent image production.
[0133] By performing preliminary screening of image templates based on information such as image delivery information, target template style, target template style, and video name, image templates that meet the task requirements can be found, thereby improving the accuracy of image template screening.
[0134] Step 104: Fill the target video image into the target image template to obtain the target image.
[0135] When populating the target video's images into the target image template, you can search for images corresponding to the target video in an image database, such as Redis. If the database doesn't contain the corresponding images, you can extract frames from the target video. This can be done using a video processing tool, such as FFmpeg.
[0136] The resulting image is inserted into the target image template. Specifically, an image processing library can be used to synthesize the image according to the template design to obtain the target image. This image processing library can be Pillow, a Python image processing library. The target image is the generated advertising material image, and advertising material images can be batch-produced based on the number of images required in the task information. The target video image is inserted into the target image template. This can be done with a single image or multiple images.
[0137] Optionally, the target video image may include a frame extraction image.
[0138] In one possible implementation, the image generation method provided in this application further includes:
[0139] Obtaining frame extraction parameters; wherein the frame extraction parameters include at least a frame extraction time interval and a frame extraction image format;
[0140] According to the frame extraction parameters, a frame extraction operation is performed on the target video to obtain a frame extraction image with a target object; wherein the target object includes a target person and / or a target definition.
[0141] The frame extraction parameters can include the frame extraction time interval and image format. The frame extraction time interval and image format can be obtained based on the channel of the target video, so that the target video can be frame extracted according to the time interval and image format. When obtaining the frame extraction image, it can be filtered according to the desired target person, and the frame extraction image of the target object can be extracted through face recognition methods. In addition, when obtaining the frame extraction image, it can also be filtered according to the target clarity. Images with high clarity can be prioritized. If the number of selected frame extraction images exceeds a threshold, frame extraction images with the target clarity are obtained. If the number of selected frame extraction images is less than the threshold, dark images with lower clarity can be selected to ensure that the number of selected frame extraction images exceeds the threshold. In practical applications, this threshold can be 10 or other values, and this application does not specifically limit this. The final frame extraction image can be a single-person image or a multi-person image, and the number of single-person and multi-person images can both exceed the aforementioned threshold. If the number of frame extraction images does not meet the threshold requirement, an error can be reported to notify the task initiator that the number of frame extraction images is insufficient.
[0142] When obtaining the framed images, frame extraction is performed according to the frame extraction time interval, so that an appropriate number of framed images can be obtained to avoid too many or too few framed images. Frame extraction is performed according to the frame extraction format, so that the obtained framed images meet the requirements of the image template for the image format. Framed images can also be obtained according to the characters to meet the task requirements, and framed images with a certain clarity are obtained, ensuring the visual effect of generating the final target image.
[0143] In one possible implementation, the target video image is filled into the target image template to obtain the target image, including:
[0144] Selecting a frame image with the same aspect ratio as the image filling area of the target image template from the frame images of the target video to obtain a first frame image;
[0145] Filter out the frame image containing the target object from the first frame image to obtain a second frame image;
[0146] Fill the second extracted frame image into the image filling area to obtain the target image.
[0147] When selecting a frame from the target video's frame that matches the target image template's image fill area's width-to-height ratio, you can use the template's frame ratio field to select an image from the target video's frame that matches the frame ratio. In practical applications, you can use image recognition methods (such as OpenCV) to detect and crop images to ensure that the fill image's size and position meet the template requirements, thereby obtaining the first frame. Specifically, deep learning algorithms can be used to identify and crop characters and scenes in the frame to better meet template filling requirements.
[0148] Then, based on the target object in the task information, such as a target person, target scenery, or target animal, the first abstract image can be filtered to obtain a second abstract image containing the target object. Of course, if the task information does not contain the target object, the first abstract image can be filtered to obtain an abstract image containing an alternative object, for example, to obtain an abstract image containing the top N candidate objects.
[0149] Optionally, when selecting the extracted frame images, you can prioritize the earliest extracted frame images in the order in which they were stored, so that you can use the latest image resources.
[0150] When populating an abstract image into an image template, you can filter for images that match the template's image fill area's width-to-height ratio, avoiding issues with inappropriately sized abstract images and meeting intelligent filling requirements. You can also filter for abstract images containing the target object required by the task, populating the template with the appropriate size abstract image containing the target object to obtain an image that meets the task requirements. This improves the compatibility between the abstract image and the image template, ensuring that automatically generated images better meet production requirements.
[0151] In one possible implementation, the target video image is filled into the target image template to obtain the target image, including:
[0152] When the number of target pictures obtained after filling the pictures of the target video into a target picture template is less than the total number of target pictures, the pictures of the target video are filled into multiple target picture templates respectively so that the number of target pictures obtained reaches the total number of template pictures.
[0153] If the number of target images obtained based on images that can be matched based on a target image template meets the required number of images, which is the total number of target images, then the images of the target video can be added to a target image template to obtain the target images. If the number of target images obtained based on images that can be matched based on a target image template does not meet the required number of images, that is, when the number of target images obtained after adding the images of the target video to a target image template is less than the required number of images, the images can be added to multiple target image templates separately so that the number of target images obtained reaches the required number of images. In practical applications, the images can be added to the first target image template, the difference between the number of target images obtained and the required number of images is calculated, the images can then be added to the second target image template, and the difference between the number of target images obtained and the required number of images is calculated again. The images can then be added to each target image template in turn until the calculated number of target images is no less than the required number of images. If the number of target images obtained after extracting all images and all target image templates is still less than the required number of images, the task initiator can be notified of the insufficient image production.
[0154] The production task for the target image can be determined by whether it is a complete video. For example, if the video is a TV series, if the TV series is currently airing, a production task can be initiated daily. The daily demand for images is the total demand divided by the preset number of days. If the total demand is for a week, calculated based on a five-day workweek, the preset number of days is 5. If the TV series is finished, a production task can be initiated based on the total demand, that is, the image demand is the total demand. If the number of target images obtained after a round of tasks reaches the image demand, production ends. If the image demand is not reached, multiple rounds of production can be performed, with steps 101 to 104 executed in each round until the image demand is reached.
[0155] When the number of target images produced using one target image template does not meet the total number of images required by the task, multiple target image templates can be used to jointly produce target images to ensure that the final number of target images reaches the total number of images and meets the quantity requirement of production images.
[0156] Optionally, the target image can be delivered to the finished film database according to the production method, which includes production only, production and delivery, and scheduled delivery; among them, production only means manual delivery after the target image is produced, production and delivery means delivery after the target image is produced, and scheduled delivery means producing the target image in advance and delivering it at a specified time.
[0157] After obtaining the target image, you can also add personalized elements to the target image according to the needs and preferences of different users, such as different fonts, special effects, etc.
[0158] The image production method provided by the present application, when producing images, such as advertising material images, the user only needs to enter a playlist to obtain the video name in the playlist, and obtain the subject matter and channel of the target video, and filter the image template that is suitable for the subject matter of the target video from the image template collection, and the image template that is suitable for the channel of the target video, so that the obtained target image template can be made to conform to the content type and video type of the template video, so that after the image of the target video is filled into the target image template, the matching degree between the template and the video image can be guaranteed, so that the generated target image has a better presentation effect, and the fit between the produced image and the video is improved, which helps to attract users to click and watch videos, increase the platform's traffic and user stickiness. Intelligent filling is performed according to the target character information and template size to ensure the content richness and visual effect of the advertising material. The present application realizes a fully automated process from receiving the task to generating advertising material images, without the need for manual design one by one, solving the problems of low efficiency of manual selection of stills, inaccurate template matching, and cumbersome and time-consuming overall production process. It can quickly process a large number of film and television drama playlists, improve the efficiency of image material production, and reduce labor costs and time costs.
[0159] The above describes a method for generating an image according to an embodiment of the present application. The following describes a system for executing the above method for generating an image.
[0160] See also Figure 6 , Figure 6 This is a structural diagram of a picture generation system provided in an embodiment of the present application. Figure 6 As shown, the image generation system includes:
[0161] The task information acquisition module 601 is configured to acquire task information; wherein the task information includes a video name.
[0162] The target video determination module 602 is configured to determine a target video having a video name.
[0163] The screening module 603 is configured to screen out a target picture template representing the video content of the target video from the picture template set; wherein the picture template set includes multiple picture templates.
[0164] The filling module 604 is configured to fill the target video image into the target image template to obtain the target image.
[0165] In a possible implementation, the screening module 603 is specifically configured to:
[0166] Obtaining data information of a target video; wherein the data information indicates the video content of the target video;
[0167] Fill the target video's data information into the prompt word template to obtain the target prompt word; wherein the prompt word template describes the mapping relationship between the image template and the video content;
[0168] Based on the target prompt word, the large model is called to screen the image templates in the image template set to obtain the target image template that represents the video content of the target video.
[0169] In a possible implementation, the screening module 603 is further configured to:
[0170] Filtering the image templates applied to the video content of the target video based on the field information of the image templates in the image template set to obtain the filtered image templates;
[0171] Based on the target prompt word, the large model is called to filter the filtered image templates to obtain the target image template that represents the video content of the target video.
[0172] In a possible implementation, the task information further includes picture delivery information, and the picture delivery information includes information required to deliver the picture to the video playback platform;
[0173] The screening module 603 is further configured to:
[0174] Searching for an image template whose field information includes image placement information from the image template collection, and obtaining the searched image template;
[0175] Picture templates containing target information in field information of the searched picture templates are filtered from the searched picture templates to obtain filtered picture templates; wherein the target information represents the video content of the target video.
[0176] In a possible implementation, the task information further includes a target object, and the target video image includes a frame-by-frame image;
[0177] The filling module 604 is specifically configured to:
[0178] Selecting a frame image with the same aspect ratio as the image filling area of the target image template from the frame images of the target video to obtain a first frame image;
[0179] Filter out the frame image containing the target object from the first frame image to obtain a second frame image;
[0180] Fill the second extracted frame image into the image filling area to obtain the target image.
[0181] In a possible implementation, the task information also includes the total number of target images;
[0182] The filling module 604 is further configured to:
[0183] Fill the target video image into the target image template to obtain the target image, including:
[0184] When the number of target images obtained after filling the target video images into a target image template is less than the total number, the target video images are filled into multiple target image templates respectively so that the number of target images obtained reaches the total number.
[0185] In one possible implementation, the image generation system further includes:
[0186] The frame extraction module is configured to obtain frame extraction parameters;
[0187] According to the frame extraction parameters, a frame extraction operation is performed on the target video to obtain a frame extraction image with the target object.
[0188] An electronic device is also provided in an embodiment of the present application. Figure 7 , which shows a schematic diagram of the structure of an electronic device suitable for implementing the embodiments of the present application. The electronic device in the embodiments of the present application may include, but is not limited to, fixed terminals such as mobile phones, laptops, PDAs (personal digital assistants), PADs (tablet computers), desktop computers, etc. Figure 7 The electronic device shown is merely an example and should not limit the functions and scope of use of the embodiments of the present application.
[0189] like Figure 7 As shown, the electronic device may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 701, which can perform various appropriate actions and processes based on programs stored in a read-only memory (ROM) 702 or programs loaded from a storage device 708 into a random access memory (RAM) 703. When the electronic device is powered on, the RAM 703 also stores various programs and data required for the operation of the electronic device. The processing device 701, ROM 702, and RAM 703 are interconnected via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0190] Typically, the following devices may be connected to the I / O interface 705: an input device 706 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 707 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 708 including, for example, a memory card, a hard disk, etc.; and a communication device 709. The communication device 709 may allow the electronic device to communicate with other devices wirelessly or by wire to exchange data. Figure 7The electronic device is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead.
[0191] The electronic device can implement the above-mentioned picture generation method.
[0192] An embodiment of the present application also provides a computer program product including computer-readable instructions. When the computer-readable instructions are executed on an electronic device, the electronic device implements any one of the image generation methods provided in the embodiments of the present application.
[0193] A computer-readable storage medium is also provided in an embodiment of the present application. The storage medium carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement any one of the image generation methods provided in the embodiment of the present application.
[0194] It should also be noted that the device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the drawings of the device embodiments provided in this application, the connection relationship between the modules indicates that there is a communication connection between them, which can be specifically implemented as one or more communication buses or signal lines.
[0195] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course can also be implemented by special hardware including application-specific integrated circuits, special CPUs, special memories, special components, etc. In general, all functions performed by computer programs can be easily implemented with corresponding hardware, and the specific hardware structures used to implement the same function can also be diverse, such as analog circuits, digital circuits or special circuits, etc. However, for the present application, software program implementation is a better implementation method in most cases. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer's floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk or optical disk, etc., and includes a number of instructions to enable a computer device (which can be a personal computer, training equipment, or network equipment, etc.) to execute the methods described in each embodiment of the present application.
[0196] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.
[0197] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, a computer, a training device or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website, a computer, a training device or a data center. The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a training device, a data center, etc. that includes one or more available media integrations. The available medium can be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).
Claims
1. A method for generating an image, characterized in that: include: Obtaining task information; wherein the task information includes a video name; Determine a target video with the video name; Filtering a target picture template representing the video content of the target video from a picture template set; wherein the picture template set includes a plurality of picture templates; Fill the target video image into the target image template to obtain the target image.
2. The image generation method according to claim 1, wherein: The step of selecting a target picture template representing the video content of the target video from the picture template set includes: Acquire data information of the target video; wherein the data information indicates the video content of the target video; Filling the target video's data information into a prompt word template to obtain a target prompt word; wherein the prompt word template describes a mapping relationship between an image template and video content; Based on the target prompt word, the large model is called to screen the picture templates in the picture template set to obtain a target picture template representing the video content of the target video.
3. The image generation method according to claim 2, characterized in that: The method of calling a large model to screen the image templates in the image template set based on the target prompt word to obtain a target image template representing the video content of the target video includes: Filtering the picture templates applied to the video content of the target video based on the field information of the picture templates in the picture template set to obtain filtered picture templates; Based on the target prompt word, the large model is called to filter the filtered image templates to obtain a target image template representing the video content of the target video.
4. The image generation method according to claim 3, characterized in that: The task information also includes picture delivery information, which includes information required to deliver the picture to the video playback platform; The step of screening the picture templates applied to the video content of the target video based on the field information of the picture templates in the picture template set to obtain the screened picture templates includes: Searching for a picture template in the picture template set, the picture template having the picture placement information included in the field information of the picture template, and obtaining the searched picture template; Picture templates containing target information in the field information of the searched picture templates are filtered from the searched picture templates to obtain filtered picture templates; wherein the target information represents the video content of the target video.
5. The image generation method according to any one of claims 1 to 4, characterized in that: The task information also includes a target object, and the target video image includes a frame extraction image; Filling the target video image into the target image template to obtain the target image includes: Selecting a frame extraction image with the same aspect ratio as the picture filling area of the target picture template from the frame extraction images of the target video to obtain a first frame extraction image; Filtering the frame abstraction image containing the target object from the first frame abstraction image to obtain a second frame abstraction image; Fill the second extracted frame image into the image filling area to obtain the target image.
6. The image generation method according to any one of claims 1 to 4, characterized in that: The task information also includes the total number of target images; Filling the target video image into the target image template to obtain the target image includes: When the number of target pictures obtained after filling the pictures of the target video into one target picture template is less than the total number, the pictures of the target video are filled into multiple target picture templates respectively so that the number of target pictures obtained reaches the total number.
7. The image generation method according to claim 5, characterized in that: Also includes: Get frame extraction parameters; According to the frame extraction parameters, a frame extraction operation is performed on the target video to obtain the frame extraction image having the target object.
8. A picture generation system, characterized in that: include: A task information acquisition module is configured to acquire task information; wherein the task information includes a video name; a target video determining module, configured to determine a target video having the video name; a screening module configured to screen out a target picture template representing the video content of the target video from a picture template set; wherein the picture template set includes a plurality of picture templates; The filling module is configured to fill the target video image into the target image template to obtain the target image.
9. A computer program product, characterized in that The method comprises computer-readable instructions, which, when executed on an electronic device, enable the electronic device to implement the image generation method according to any one of claims 1 to 7.
10. An electronic device, characterized in that: comprising at least one processor and a memory connected to the processor, wherein: The memory is used to store computer programs; The processor is configured to execute the computer program so that the electronic device can implement the image generation method according to any one of claims 1 to 7.
11. A computer storage medium, characterized in that The storage medium carries one or more computer programs, and when the one or more computer programs are executed by an electronic device, the electronic device can implement the image generation method according to any one of claims 1 to 7.