Poster generation method and device, electronic equipment and storage medium

CN117671558BActive Publication Date: 2026-09-18BEIJING QIYI CENTURY SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311559292.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-21
Publication Date
2026-09-18
Estimated Expiration
2043-11-21

AI Technical Summary

Technical Problem

[0005]本申请实施例的目的在于提供一种海报生成方法、装置、电子设备及存储介质,以解决通过设计师人工设计海报的方式,往往效率低下,且,十分耗费人力的问题

Benefits of technology

[0054] This application provides a poster generation method, apparatus, electronic device, and storage medium. In this embodiment, firstly, a target video frame is determined in a target video. Then, at least one human body location region, a text location region, and a target face location region are extracted from the target video frame. Next, multiple frames are determined in the target video frame based on the target face location region. Finally, a target frame is determined from the multiple frames based on the at least one human body location region, text location region, and target face location region. Finally, the target video frame is cropped based on the target frame to obtain the poster corresponding to the target video. In this way, a corresponding poster can be intelligently generated based on the target video frame in the target video, thereby improving poster generation efficiency and saving manpower.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117671558B_ABST
    Figure CN117671558B_ABST
Patent Text Reader

Abstract

The application provides a poster generation method and device, electronic equipment and storage medium. The method comprises the following steps: determining a target video frame in a target video; extracting at least one human body position region, a text position region and a target face position region from the target video frame, wherein the target face position region is used to represent the face position of a target object; determining a plurality of frames based on the target face position region in the target video frame, wherein each frame contains the target face position region; determining a target frame from a plurality of frames based on at least one human body position region, a text position region and a target face position region; and cropping the target video frame based on the target frame to obtain a poster corresponding to the target video. In this way, the corresponding poster can be intelligently generated based on the target video frame in the target video, thereby improving the poster generation efficiency and saving manpower.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to a poster generation method, apparatus, electronic device, and storage medium. Background Technology

[0002] Posters are crucial materials for film and television marketing and promotion. They interpret the themes and connotations of a film or television work, influencing and shaping the audience's first impression. A good poster can effectively promote and sell the product or service. Poster design for film and television works often focuses on the stars, leveraging their celebrity influence to promote and publicize the series.

[0003] Currently, poster design for film and television works generally relies on manual design by designers. In the poster design process, designers need to base their designs on the content of the film or television work itself, and combine aesthetics, plot, and celebrity effects, using techniques such as cropping and image editing to create the final poster image.

[0004] However, this method of manually designing posters by designers is often inefficient and very labor-intensive. Summary of the Invention

[0005] The purpose of this application is to provide a poster generation method, apparatus, electronic device, and storage medium to solve the problem that manually designing posters by designers is often inefficient and extremely labor-intensive. The specific technical solution is as follows:

[0006] Firstly, this application provides a poster generation method, including:

[0007] Identify the target video frame within the target video;

[0008] Extract at least one human body location region, text location region, and target face location region from the target video frame, wherein the target face location region is used to characterize the face location of the target object;

[0009] Based on the target face location region, multiple frames are determined in the target video frame, wherein each frame contains the target face location region;

[0010] A target frame is determined from multiple frames based on at least one of the human body location regions, the text location regions, and the target face location regions;

[0011] The target video frame is cropped based on the target bounding box to obtain the poster corresponding to the target video.

[0012] In one possible implementation, determining multiple bounding boxes in the target video frame based on the target face location region includes:

[0013] A corresponding face detection box is determined based on the target face location region, wherein the face detection box is a rectangular box that includes the target face location region;

[0014] The face detection box is expanded multiple times according to a preset length and a preset aspect ratio to obtain multiple frames. The height of each frame is n preset lengths higher than the height of the face detection box, and the aspect ratio of each frame is the preset aspect ratio.

[0015] In one possible implementation, determining a target bounding box from a plurality of bounding boxes based on at least one of the human body location regions, the text location regions, and the target face location regions includes:

[0016] For each frame, in at least one of the human body position regions, a first human body position region that matches the target human face position region contained in the frame, and a second human body position region that does not match the target human face position region are determined.

[0017] If there is only one corresponding first human body location region, the frame is determined as the first candidate frame;

[0018] Determine whether each first candidate box intersects with the text location area, and determine the first candidate box that does not intersect with the text location area as the second candidate box;

[0019] Based on the target face location region, the first human body location region, and the second human body location region, a target box is determined in the second candidate box.

[0020] In one possible implementation, determining, within at least one of the human body location regions, a first human body location region that matches a target face location region included in the frame, and a second human body location region that does not match the target face location region, includes:

[0021] For each human body location region, a first intersection-union ratio is determined between the target face location region and the human body location region contained in the frame;

[0022] If the first intersection-union ratio is greater than or equal to a preset threshold, the human body location region is determined to be the first human body location region that matches the target human face location region;

[0023] If the first crossover ratio is less than a preset threshold, the human body location region is determined to be a second human body location region that does not match the target face location region.

[0024] In one possible implementation, determining the target box in the second candidate box based on the target face location region, the first human body location region, and the second human body location region includes:

[0025] Obtain a poster template, which contains a standard human face area;

[0026] For each second candidate box, determine whether the target face location region in the second candidate box is inside the standard face region;

[0027] If the target face location region is within the standard face region, the second candidate box is determined as the third candidate box;

[0028] For each third candidate box, determine the area ratio between the target face location region and the standard face region within the third candidate box;

[0029] If the area ratio is within a preset area ratio range, the third candidate box is determined as the fourth candidate box;

[0030] For each fourth candidate box, an evaluation score is calculated based on the target face location region, the standard face region, the first human body location region, and the second human body location region within the fourth candidate box.

[0031] The fourth candidate box with the highest corresponding evaluation score is determined as the target box.

[0032] In one possible implementation, calculating the evaluation score corresponding to the fourth candidate box based on the target face location region, the standard face region, the first human body location region, and the second human body location region in the fourth candidate box includes:

[0033] Determine the distance between the center abscissa of the target face location region and the center abscissa of the standard face region;

[0034] Determine the second intersection-union ratio between the fourth candidate box and the first human body location region, and the third intersection-union ratio between the fourth candidate box and the second human body location region;

[0035] Determine the area ratio between the fourth candidate box and the target video frame;

[0036] Substituting the distance, the second intersection-union ratio, the third intersection-union ratio, and the area ratio into a preset formula, the evaluation score corresponding to the fourth candidate box is obtained;

[0037] The preset formula is as follows:

[0038] Evaluation score = (1 - distance) + (1 - third intersection-union ratio) + second intersection-union ratio + area ratio.

[0039] In one possible implementation, after cropping the target video frame based on the target bounding box to obtain the poster corresponding to the target video, the method further includes:

[0040] Obtain the target rendering strategy;

[0041] The poster is rendered according to the target rendering strategy.

[0042] Secondly, this application provides a poster generating apparatus, comprising:

[0043] The first determining module is used to determine the target video frame in the target video;

[0044] An extraction module is used to extract at least one human body location region, a text location region, and a target face location region from the target video frame, wherein the target face location region is used to characterize the face location of the target object;

[0045] The second determining module is used to determine multiple frames in the target video frame based on the target face location region, wherein each frame contains the target face location region;

[0046] The third determining module is used to determine a target frame from a plurality of frames based on at least one of the human body location regions, the text location regions, and the target face location regions.

[0047] The cropping module is used to crop the target video frame based on the target bounding box to obtain the poster corresponding to the target video.

[0048] Thirdly, an electronic device is provided, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus;

[0049] Memory, used to store computer programs;

[0050] When a processor executes a program stored in memory, it implements any of the steps described in the first aspect.

[0051] Fourthly, a computer-readable storage medium is provided, characterized in that the computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of any of the methods described in the first aspect.

[0052] Fifthly, a computer program product containing instructions is provided, which, when run on a computer, causes the computer to execute any of the poster generation methods described above.

[0053] Beneficial effects of the embodiments in this application:

[0054] This application provides a poster generation method, apparatus, electronic device, and storage medium. In this embodiment, firstly, a target video frame is determined in a target video. Then, at least one human body location region, a text location region, and a target face location region are extracted from the target video frame. Next, multiple frames are determined in the target video frame based on the target face location region. Finally, a target frame is determined from the multiple frames based on the at least one human body location region, text location region, and target face location region. Finally, the target video frame is cropped based on the target frame to obtain the poster corresponding to the target video. In this way, a corresponding poster can be intelligently generated based on the target video frame in the target video, thereby improving poster generation efficiency and saving manpower.

[0055] Of course, implementing any product or method of this application does not necessarily require achieving all of the advantages described above at the same time. Attached Figure Description

[0056] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0057] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0058] One or more embodiments are illustrated by way of example with reference numerals in the accompanying drawings. These illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are denoted as similar elements. Unless otherwise stated, the figures in the drawings are not to be limited by scale.

[0059] Figure 1 A flowchart illustrating a poster generation method provided in this application embodiment;

[0060] Figure 2 A flowchart illustrating another poster generation method provided in this application embodiment;

[0061] Figure 3 This is a schematic diagram of the structure of a poster generation device provided in an embodiment of this application;

[0062] Figure 4This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0063] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0064] The following disclosure provides numerous different embodiments or examples for implementing various structures of the invention. To simplify the disclosure, specific examples of components and arrangements are described below. These are merely examples and are not intended to limit the scope of the invention. Furthermore, reference numerals and / or letters may be repeated in different examples. Such repetition is for simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or arrangements discussed.

[0065] Figure 1 This is a flowchart illustrating a poster generation method provided in an embodiment of this application. This method can be applied to one or more electronic devices such as smartphones, laptops, desktop computers, portable computers, and servers. Furthermore, the execution entity of this method can be hardware or software. When the execution entity is hardware, it can be one or more of the aforementioned electronic devices. For example, a single electronic device can execute this method, or multiple electronic devices can cooperate with each other to execute this method. When the execution entity is software, this method can be implemented as multiple software programs or software modules, or as a single software program or software module. No specific limitations are made here.

[0066] like Figure 1 As shown, the method specifically includes:

[0067] S101, Determine the target video frame in the target video.

[0068] This application provides a poster generation method for generating a corresponding poster for a target video based on target video frames. The number of target video frames can be one or more.

[0069] In one embodiment, a user-specified video frame can be used as the target video frame. This allows the user to flexibly specify the target video frame according to their actual needs.

[0070] In another embodiment, a video frame containing a target object can be identified as the target video frame. The target object can be a user-specified person or a lead actor determined from the cast list of the target video. In this way, the target video frame can be intelligently determined based on the content of each video frame.

[0071] S102, extract at least one human body location region, text location region and target face location region from the target video frame, wherein the target face location region is used to characterize the face location of the target object.

[0072] In applications, a celebrity recognition algorithm for detecting celebrity faces can be pre-trained based on deep learning algorithms, as well as a person detection algorithm for detecting human bodies can be pre-trained based on deep learning algorithms.

[0073] Based on this, in this embodiment, the target video frame can be input into a pre-trained celebrity recognition algorithm, which outputs the object ID of the target object in the target video frame and the target face location region where the target object's face is located. Additionally, the target video frame can be input into a pre-trained person detection algorithm, which outputs the human body location region in the target video frame. Furthermore, to identify whether there is dialogue in the target video frame, the target video frame can also be input into an OCR (Optical Character Recognition) algorithm, which outputs the text location region.

[0074] S103, based on the target face location region, determine multiple frames in the target video frame, wherein each frame contains the target face location region.

[0075] In this embodiment of the application, the specific implementation of determining multiple frames in the target video frame based on the target face location region may include:

[0076] A corresponding face detection box is determined based on the target face location region. The face detection box is a rectangular box that includes the target face location region. The face detection box is expanded multiple times according to a preset length and a preset aspect ratio to obtain multiple frames. The height of each frame is n preset lengths higher than the height of the face detection box, and the aspect ratio of each frame is the preset aspect ratio.

[0077] The face detection bounding box is a rectangle that only contains the location of the target face.

[0078] In this embodiment, using the face detection bounding box as a reference, the height of the face detection bounding box is extended by a preset length each time, and then a determined width is calculated using a preset aspect ratio (e.g., 3:4). The width of the face detection bounding box is then expanded according to this width to obtain the corresponding frame for this expansion. Furthermore, each expanded frame is slid up, down, left, and right, resulting in a corresponding frame each time. It should be noted that each slid must ensure that the target object's face is within the corresponding frame. Thus, multiple frames are obtained.

[0079] S104, a target frame is determined from the plurality of frames based on at least one of the human body location regions, the text location regions, and the target face location regions.

[0080] S105, based on the target bounding box, the target video frame is cropped to obtain the poster corresponding to the target video.

[0081] The following provides a unified explanation of S104 and S105:

[0082] In this embodiment, after obtaining multiple frames, the frames can be filtered based on at least one human body location region, text location region, and target face location region to determine the target frame that meets the expected effect from the multiple frames. Then, the target video frame is cropped based on the position of the target frame to obtain a single-person poster containing the target object.

[0083] The specific method for determining the target frame from multiple frames based on at least one of the human body location regions, the text location regions, and the target face location regions will be explained in detail in the following embodiments, and will not be elaborated here.

[0084] In another embodiment, after S105, the following steps may be included: obtaining a target rendering strategy and rendering the poster according to the target rendering strategy.

[0085] The target rendering strategy can be set by the user according to actual needs. For example, rendering corresponding text at a preset position on the poster, setting corresponding filter effects for the poster, and so on. Thus, the poster can be rendered according to user requirements.

[0086] In this embodiment, firstly, a target video frame is determined in the target video. Then, at least one human body location region, a text location region, and a target face location region are extracted from the target video frame. Next, multiple frames are determined in the target video frame based on the target face location region. Finally, a target frame is determined from these multiple frames based on the at least one human body location region, text location region, and target face location region. Finally, the target video frame is cropped based on the target frame to obtain the poster corresponding to the target video. In this way, a corresponding poster can be intelligently generated based on the target video frame in the target video, thereby improving poster generation efficiency and saving manpower.

[0087] See Figure 2 This is a flowchart illustrating another embodiment of the poster generation method provided in this application. Figure 2 The process shown above Figure 1 Based on the illustrated process, this paper describes how to determine a target bounding box from multiple bounding boxes based on at least one of the human body location regions, the text location regions, and the target face location regions. For example... Figure 2 As shown, the process may include the following steps:

[0088] S201, for each frame, in at least one of the human body position regions, determine a first human body position region that matches the target face position region contained in the frame, and a second human body position region that does not match the target face position region.

[0089] In applications, a video frame often contains one or more people, and correspondingly, one or more human body location regions can be identified.

[0090] Based on this, in the embodiments of this application, in at least one of the human body location regions, a first human body location region matching the target face location region contained in the frame is determined, and a second human body location region not matching the target face location region is determined. This can be implemented by: for each human body location region, determining a first intersection-union ratio (IUU) between the target face location region contained in the frame and the human body location region; if the IUU is greater than or equal to a preset threshold, determining the human body location region as a first human body location region matching the target face location region; if the IUU is less than the preset threshold, determining the human body location region as a second human body location region not matching the target face location region.

[0091] The preset threshold is generally 1, that is, when a certain human body location area completely contains the target human face location area, the human body location area is considered to be the human body of the target object; otherwise, the human body location area is considered not to be the human body of the target object.

[0092] S202, when the number of corresponding first human body position areas is one, the frame is determined as the first candidate frame.

[0093] In this embodiment, when there is only one corresponding first human body position area, it means that the target object's face in the frame contains only one independent human body, and in this case, the frame is determined as the first candidate frame. When there are multiple corresponding first human body position areas, it means that due to misalignment or other reasons, the target object's face in the frame corresponds to multiple human bodies. In this case, the target object's human body in the frame is affected by misaligned human bodies and is not suitable as a poster. When there are zero corresponding first human body position areas, it means that the frame only contains the target object's face and does not contain a corresponding complete human body, similar to a headshot, and is not suitable as a poster.

[0094] S203, determine whether each of the first candidate boxes intersects with the text location area, and determine the first candidate box that does not intersect with the text location area as the second candidate box.

[0095] In this embodiment, when the first candidate box intersects with the text location area, it means that there is text (usually dialogue) in the first candidate box. In this case, the first candidate box is not suitable as a poster. When the first candidate box does not intersect with the text location area, it means that there is no text in the first candidate box. In this case, the first candidate box is used as the second candidate box, and subsequent poster frame selection is continued.

[0096] S204, Based on the target face location region, the first human body location region, and the second human body location region, determine the target box in the second candidate box.

[0097] In this embodiment of the application, determining the target box in the second candidate box based on the target face location region, the first human body location region, and the second human body location region may specifically include the following steps:

[0098] Step A1: Obtain a poster template, which contains a standard face area;

[0099] Step A2: For each second candidate box, determine whether the target face location region in the second candidate box is inside the standard face region;

[0100] Step A3: If the target face location region is within the standard face region, the second candidate box is determined as the third candidate box;

[0101] Step A4: For each third candidate box, determine the area ratio between the target face location region and the standard face region within the third candidate box;

[0102] Step A5: If the area ratio is within the preset area ratio range, the third candidate box is determined as the fourth candidate box;

[0103] Step A6: For each fourth candidate box, calculate the evaluation score corresponding to the fourth candidate box based on the target face location region, the standard face region, the first human body location region, and the second human body location region in the fourth candidate box;

[0104] Step A7: The fourth candidate box with the highest corresponding evaluation score is determined as the target box.

[0105] There can be one or more poster templates. Each poster template sets the standard position of the face (i.e., the standard face area).

[0106] In this embodiment, for each poster template, firstly, based on the standard face region in the template, a second candidate box is selected from the second candidate box, and the second candidate box is selected as the third candidate box, where the target face position region is inside the standard face region. Thus, the target face position region in the selected third candidate box conforms to the allowed area range in the poster template.

[0107] Then, based on the area ratio of the target face location region to the standard face region, the third candidate box with an area ratio within the preset area ratio range is selected as the fourth candidate box. Thus, the area ratio of the target face location region to the standard face region in the selected fourth candidate box conforms to the preset area ratio range allowed by the poster template.

[0108] Next, based on the target face location region, standard face region, first human body location region, and second human body location region in the fourth candidate box, the evaluation score corresponding to the fourth candidate box is calculated, and the fourth candidate box with the highest evaluation score is determined as the target box.

[0109] Specifically, the distance between the center abscissa of the target face location region and the center abscissa of the standard face region is determined; the second intersection-union ratio (IUU) between the fourth candidate box and the first human body location region and the third IUU between the fourth candidate box and the second human body location region are determined; the area ratio between the fourth candidate box and the target video frame is determined; and the distance, the second IUU, the third IUU, and the area ratio are substituted into a preset formula to obtain the evaluation score corresponding to the fourth candidate box. The preset formula is as follows: Evaluation score = (1 - distance) + (1 - third IUU) + second IUU + area ratio.

[0110] Among them, the larger the "1-distance", the smaller the distance between the center x-coordinate of the target face location region and the center x-coordinate of the standard face region, that is, the closer the target face location region is to the center of the standard face region; the larger the "1-third intersection-union ratio", the fewer other human bodies are contained in the frame; the larger the "second intersection-union ratio", the larger the human body of the target object in the frame; the larger the "area ratio", the larger the area of ​​the frame relative to the target video frame.

[0111] In this way, the target face location area can be centered on the standard face area. The frame contains fewer other human figures, the target object's human figure is larger, and the fourth candidate frame with the larger area is determined as the target frame. This allows for the selection of a more aesthetically pleasing frame as the target frame.

[0112] Based on the same technical concept, embodiments of this application also provide a poster generation device, such as... Figure 3 As shown, the device includes:

[0113] The first determining module 301 is used to determine the target video frame in the target video;

[0114] Extraction module 302 is used to extract at least one human body position region, text position region and target face position region from the target video frame, wherein the target face position region is used to characterize the face position of the target object;

[0115] The second determining module 303 is used to determine multiple frames in the target video frame based on the target face location region, wherein each frame contains the target face location region;

[0116] The third determining module 304 is used to determine a target frame from a plurality of frames based on at least one of the human body position regions, the text position regions, and the target face position regions.

[0117] The cropping module 305 is used to crop the target video frame based on the target frame to obtain the poster corresponding to the target video.

[0118] In this embodiment, firstly, a target video frame is determined in the target video. Then, at least one human body location region, a text location region, and a target face location region are extracted from the target video frame. Next, multiple frames are determined in the target video frame based on the target face location region. Finally, a target frame is determined from these multiple frames based on the at least one human body location region, text location region, and target face location region. Finally, the target video frame is cropped based on the target frame to obtain the poster corresponding to the target video. In this way, a corresponding poster can be intelligently generated based on the target video frame in the target video, thereby improving poster generation efficiency and saving manpower.

[0119] Based on the same technical concept, embodiments of this application also provide an electronic device, such as... Figure 4 As shown, it includes a processor 111, a communication interface 112, a memory 113, and a communication bus 114, wherein the processor 111, the communication interface 112, and the memory 113 communicate with each other through the communication bus 114.

[0120] Memory 113 is used to store computer programs;

[0121] When processor 111 executes a program stored in memory 113, it performs the following steps:

[0122] Identify the target video frame within the target video;

[0123] Extract at least one human body location region, text location region, and target face location region from the target video frame, wherein the target face location region is used to characterize the face location of the target object;

[0124] Based on the target face location region, multiple frames are determined in the target video frame, wherein each frame contains the target face location region;

[0125] A target frame is determined from multiple frames based on at least one of the human body location regions, the text location regions, and the target face location regions;

[0126] The target video frame is cropped based on the target bounding box to obtain the poster corresponding to the target video.

[0127] The communication bus mentioned in the above electronic devices can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.

[0128] The communication interface is used for communication between the aforementioned electronic devices and other devices.

[0129] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0130] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0131] In another embodiment provided in this application, a computer-readable storage medium is also provided, which stores a computer program that, when executed by a processor, implements the steps of any of the poster generation methods described above.

[0132] In another embodiment provided in this application, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute any of the poster generation methods described above.

[0133] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0134] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented using software plus a general-purpose hardware platform, or of course, using hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0135] It should be understood that the terminology used herein is for the purpose of describing particular exemplary embodiments only and is not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “described” as used herein may also include the plural forms. The terms “comprising,” “including,” “containing,” and “having” are inclusive and therefore indicate the presence of the stated features, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, elements, components, and / or combinations thereof. The method steps, processes, and operations described herein are not construed as requiring them to be performed in a particular order described or illustrated unless the order of performance is explicitly indicated. It should also be understood that additional or alternative steps may be used.

[0136] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A poster generation method, characterized in that, The method includes: Identify the target video frame within the target video; Extract at least one human body location region, text location region, and target face location region from the target video frame, wherein the target face location region is used to characterize the face location of the target object; Based on the target face location region, multiple frames are determined in the target video frame, wherein each frame contains the target face location region; A target frame is determined from multiple frames based on at least one of the human body location regions, the text location regions, and the target face location regions; The target video frame is cropped based on the target bounding box to obtain the poster corresponding to the target video; The step of determining a target bounding box from multiple bounding boxes based on at least one of the human body location regions, the text location regions, and the target face location regions includes: For each frame, in at least one of the human body position regions, a first human body position region that matches the target human face position region contained in the frame, and a second human body position region that does not match the target human face position region are determined. If there is only one corresponding first human body location region, the frame is determined as the first candidate frame; Determine whether each first candidate box intersects with the text location area, and determine the first candidate box that does not intersect with the text location area as the second candidate box; Based on the target face location region, the first human body location region, and the second human body location region, a target box is determined in the second candidate box.

2. The method according to claim 1, characterized in that, The determination of multiple bounding boxes in the target video frame based on the target face location region includes: A corresponding face detection box is determined based on the target face location region, wherein the face detection box is a rectangular box that includes the target face location region; The face detection box is expanded multiple times according to a preset length and a preset aspect ratio to obtain multiple frames. The height of each frame is n preset lengths higher than the height of the face detection box, and the aspect ratio of each frame is the preset aspect ratio.

3. The method according to claim 1, characterized in that, The step of determining, within at least one of the human body location regions, a first human body location region that matches the target face location region contained in the frame, and a second human body location region that does not match the target face location region, includes: For each human body location region, a first intersection-union ratio is determined between the target face location region and the human body location region contained in the frame; If the first intersection-union ratio is greater than or equal to a preset threshold, the human body location region is determined to be the first human body location region that matches the target human face location region; If the first crossover ratio is less than a preset threshold, the human body location region is determined to be a second human body location region that does not match the target face location region.

4. The method according to claim 1, characterized in that, The step of determining a target box in the second candidate box based on the target face location region, the first human body location region, and the second human body location region includes: Obtain a poster template, which contains a standard human face area; For each second candidate box, determine whether the target face location region in the second candidate box is inside the standard face region; If the target face location region is within the standard face region, the second candidate box is determined as the third candidate box; For each third candidate box, determine the area ratio between the target face location region and the standard face region within the third candidate box; If the area ratio is within a preset area ratio range, the third candidate box is determined as the fourth candidate box; For each fourth candidate box, an evaluation score is calculated based on the target face location region, the standard face region, the first human body location region, and the second human body location region within the fourth candidate box. The fourth candidate box with the highest corresponding evaluation score is determined as the target box.

5. The method according to claim 4, characterized in that, The step of calculating the evaluation score corresponding to the fourth candidate box based on the target face location region, the standard face region, the first human body location region, and the second human body location region in the fourth candidate box includes: Determine the distance between the center abscissa of the target face location region and the center abscissa of the standard face region; Determine the second intersection-union ratio between the fourth candidate box and the first human body location region, and the third intersection-union ratio between the fourth candidate box and the second human body location region; Determine the area ratio between the fourth candidate box and the target video frame; Substituting the distance, the second intersection-union ratio, the third intersection-union ratio, and the area ratio into a preset formula, the evaluation score corresponding to the fourth candidate box is obtained; The preset formula is as follows: Evaluation score = (1 - distance) + (1 - third intersection-union ratio) + second intersection-union ratio + area ratio.

6. The method according to claim 1, characterized in that, After cropping the target video frame based on the target bounding box to obtain the poster corresponding to the target video, the method further includes: Obtain the target rendering strategy; The poster is rendered according to the target rendering strategy.

7. A poster generating device, characterized in that, The device includes: The first determining module is used to determine the target video frame in the target video; An extraction module is used to extract at least one human body location region, a text location region, and a target face location region from the target video frame, wherein the target face location region is used to characterize the face location of the target object; The second determining module is used to determine multiple frames in the target video frame based on the target face location region, wherein each frame contains the target face location region; The third determining module is used to determine a target frame from a plurality of frames based on at least one of the human body location regions, the text location regions, and the target face location regions. The cropping module is used to crop the target video frame based on the target box to obtain the poster corresponding to the target video; The third determining module is specifically used for: For each frame, in at least one of the human body position regions, a first human body position region that matches the target human face position region contained in the frame, and a second human body position region that does not match the target human face position region are determined. If there is only one corresponding first human body location region, the frame is determined as the first candidate frame; Determine whether each first candidate box intersects with the text location area, and determine the first candidate box that does not intersect with the text location area as the second candidate box; Based on the target face location region, the first human body location region, and the second human body location region, a target box is determined in the second candidate box.

8. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, when executing a program stored in memory, implements the steps of the method described in any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the method described in any one of claims 1-6.

Citation Information

Patent Citations

  • Poster generation method and device

    CN106792150A

  • Video abstract generation method and device

    CN108882057A