Video Cover Generation Method, Apparatus and Storage Medium
By predicting the click-through rate of the target video keyframe image and text information, the video cover with the highest click-through rate is generated, which solves the problem of inaccurate video cover generation and improves the video click-through rate.
Patent Information
- Application Number
- CN202011262319.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-11-12
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2040-11-12
AI Technical Summary
In the prior art, video cover generation is inaccurate and cannot accurately reflect the characteristics of the video.
By obtaining the keyframe collection of the target video, text recognition is performed, the video cover generation model is used to predict the click-through rate of the combination of keyframe images and text information, and the cover with the highest click-through rate is generated.
The generated video cover can better reflect the characteristics of the video, take into account the behavioral characteristics of the video viewer, and improve the click rate after the video and cover is placed.
Smart Images

Figure CN114491151B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of video processing, and in particular, to a method, an apparatus, and a storage medium for generating a video cover. Background Art
[0002] With the development of Internet technology, there are more and more videos spread on the Internet. In order to facilitate users to understand the video content, a video cover can be generated for the video, and the video cover is displayed to the users, and the users can select the video to watch according to the video cover. A good video cover can improve the click-through rate and the viewing duration of the video. Therefore, how to generate a video cover is crucial. However, in the related art, the video cover is randomly extracted from the video or specified by the user, and cannot accurately reflect the characteristics of the video. Summary of the Invention
[0003] The present disclosure provides a method, an apparatus, and a storage medium for generating a video cover to at least solve the problem that the generation of the video cover in the related art is inaccurate. The technical solutions of the present disclosure are as follows:
[0004] According to a first aspect of an embodiment of the present disclosure, there is provided a method for generating a video cover, including:
[0005] Obtaining a set of key frames of a target video, where the set of key frames includes at least one key frame image;
[0006] Performing text recognition on each key frame image in the set of key frames to obtain a text set corresponding to the target video, where the text set includes at least one piece of text information;
[0007] Inputting the key frame images in the set of key frames and the text information in the text set into a video cover generation model, predicting the click-through rate of each cover formed by combining each key frame image and each piece of text information, and outputting the identification information of the cover with the highest predicted click-through rate, where the identification information of the cover represents the input order information of the key frame image and the text information forming the cover;
[0008] Taking the key frame image corresponding to the identification information as the target key frame image and taking the text information corresponding to the identification information as the target text information;
[0009] Generating a cover of the target video according to the target key frame image and the target text information.
[0010] In an exemplary embodiment, the obtaining a set of key frames of a target video includes: detecting key frame information of the target video to obtain key frame images of the target video; and determining a set of key frames corresponding to the target video according to the key frame images.
[0011] In an exemplary embodiment, the detection of key frame information for the target video to obtain the key frame images of the target video includes at least one of the following:
[0012] Perform face detection on the target video, and select the frame images containing faces from the target video as key frame images;
[0013] Perform clarity detection on the target video, sort the frame images of the target video in descending order of clarity, and select a preset number of frame images ranked at the front from the target video as key frame images;
[0014] Perform highlight segment detection on the target video, and select the frame images containing highlight segments from the target video as key frame images.
[0015] In an exemplary embodiment, the recognition of text for each key frame image in the key frame set to obtain the text set corresponding to the target video includes: performing text recognition on each key frame image in the key frame set to obtain the recognized text corresponding to the key frame image, and performing denoising processing on the recognized text corresponding to each key frame image; performing semantic recognition on the denoised recognized text, removing text content with similar semantics, and obtaining text information; determining the text set corresponding to the target video according to the text information.
[0016] In an exemplary embodiment, generating the cover of the target video according to the target key frame image and the target text information includes: obtaining the depth information of the target key frame image; determining the text configuration information of the target key frame image according to the depth information; and setting the target text information in the target key frame image based on the text configuration information to obtain the cover of the target video.
[0017] In an exemplary embodiment, before obtaining the key frame set of the target video, it further includes: obtaining a cover set of the sample video, the cover set including at least one cover, and each cover having a unique cover identifier; obtaining the click-through rate corresponding to each cover in the cover set; sorting the covers in the cover set in descending order according to the click-through rate, and selecting the cover identifier of the first cover as the cover label of the cover set; and training the video cover generation model according to the cover label of the cover set and the covers in the cover set.
[0018] In an exemplary embodiment, the obtaining of the cover set of the sample video includes: obtaining the sample video; performing key frame detection on the sample video to obtain a sample key frame set, where the sample key frame set includes at least one sample key frame image; performing text recognition on each sample key frame image in the sample key frame set to obtain a sample text set, where the sample text set includes at least one sample text information; combining each sample key frame image in the sample key frame set with each sample text information in the sample text set to generate at least one cover; and determining the cover set of the sample video according to the generated covers.
[0019] In an exemplary embodiment, the combining each sample key frame image in the sample key frame set with each sample text information in the sample text set to generate at least one cover includes: performing depth information recognition on the sample key frame images in the sample key frame set to obtain the depth information of each sample key frame image; determining the text configuration information of each sample key frame image according to the depth information; and setting the text information in the corresponding sample key frame image based on the text configuration information to obtain at least one cover.
[0020] In an exemplary embodiment, the obtaining of the click-through rate corresponding to each cover in the cover set includes: combining each cover in the cover set with the corresponding sample video to obtain at least one video data; publishing the video data and obtaining the click times of each video data; and determining the click-through rate corresponding to each cover in the cover set according to the click times of the video data and the covers included in each video data.
[0021] In an exemplary embodiment, the training of the video cover generation model according to the cover labels of the cover set and the covers in the cover set includes: extracting the cover images and cover texts of each cover in the cover set; inputting the cover images and the cover texts corresponding to the same cover identifier into the initial video cover generation model, extracting the image feature vector of the cover image and the text feature vector of the cover text, combining the image feature vector and the text feature vector to form a graphic-text feature vector, and outputting the click-through rate prediction result of the graphic-text feature vector, where the prediction result carries the cover identifier of the cover corresponding to the graphic-text feature vector; comparing the cover identifier carried by the prediction result with the cover label to calculate a loss value; adjusting the parameters of the initial video cover generation model according to the loss value, and training the adjusted initial video cover generation model based on the cover labels of the cover set and the covers in the cover set until the adjustment of the parameters of the initial video cover generation model is stopped when a preset training stop condition is satisfied, and obtaining the video cover generation model.
[0022] According to a second aspect of the embodiments of the present disclosure, there is provided a video cover generating device, including:
[0023] A key frame set obtaining unit, configured to obtain a key frame set of a target video, where the key frame set includes at least one key frame image;
[0024] A text set obtaining unit, configured to perform text recognition on each key frame image in the key frame set to obtain a text set corresponding to the target video, where the text set includes at least one piece of text information;
[0025] An identification information obtaining unit, configured to input the key frame images in the key frame set and the text information in the text set into a video cover generation model, perform click-through rate prediction on each cover formed by combining each key frame image and each text information, and output the identification information of the cover with the highest predicted click-through rate, where the identification information of the cover represents the input order information of the key frame image and the text information forming the cover;
[0026] A cover information determining unit, configured to use the key frame image corresponding to the identification information as the target key frame image and the text information corresponding to the identification information as the target text information;
[0027] A cover generating unit, configured to generate a cover of the target video according to the target key frame image and the target text information.
[0028] In an exemplary embodiment, the key frame set obtaining unit includes: a key frame image obtaining module, configured to perform key frame information detection on the target video to obtain key frame images of the target video; a key frame set determining module, configured to determine a key frame set corresponding to the target video according to the key frame images.
[0029] In an exemplary embodiment, the key frame image obtaining module is configured to perform at least one of the following:
[0030] Perform face detection on the target video, and select a frame image including a face from the target video as a key frame image;
[0031] Perform clarity detection on the target video, sort the frame images of the target video in descending order of clarity, and select a preset number of frame images ranked at the front from the target video as key frame images;
[0032] Perform exciting segment detection on the target video, and select a frame image including an exciting segment from the target video as a key frame image.
[0033] In an exemplary embodiment, the text set acquisition unit includes: a text recognition module, which is configured to perform text recognition on each of the key frame images in the key frame set to obtain a recognized text corresponding to the key frame image; a processing module, which is configured to perform denoising on the recognized text corresponding to each of the key frame images; perform semantic recognition on the recognized text after denoising, eliminate text content with similar semantics, and obtain text information; and a text set determination module, which is configured to determine the text set corresponding to the target video based on the text information.
[0034] In an exemplary embodiment, the cover generation unit includes: a depth information acquisition module configured to acquire depth information of the target key frame image; a text configuration information acquisition module configured to determine text configuration information of the target key frame image according to the depth information;
[0035] The first cover generation module is configured to set the target text information in the target key frame image based on the text configuration information to obtain the cover of the target video.
[0036] In an exemplary embodiment, the video cover generation device also includes: a cover set acquisition unit, configured to acquire a cover set of sample videos, the cover set including at least one cover, each cover having a unique cover identifier; a click rate acquisition unit, configured to acquire the click rate corresponding to each cover in the cover set; a cover label determination unit, configured to sort each cover in the cover set in descending order according to the click rate, and select the cover identifier of the first cover as the cover label of the cover set; and a model training unit, configured to train the video cover generation model based on the cover labels of the cover set and the covers in the cover set.
[0037] In an exemplary embodiment, the cover set acquisition unit includes: a sample video acquisition module, configured to acquire a sample video; a sample key frame image acquisition module, configured to perform key frame detection on the sample video to obtain a sample key frame set, wherein the sample key frame set includes at least one sample key frame image; a sample text information acquisition module, configured to perform text recognition on each sample key frame image in the sample key frame set to obtain a sample text set, wherein the sample text set includes at least one sample text information; a second cover generation module, configured to combine each sample key frame image in the sample key frame set with each sample text information in the sample text set to generate at least one cover; and a cover set determination module, configured to determine the cover set of the sample video based on the generated cover.
[0038] In an exemplary embodiment, the second cover generation module includes: a depth information acquisition sub-module configured to perform depth information recognition on the sample key frame images in the sample key frame set to obtain the depth information of each sample key frame image; a text configuration information acquisition sub-module configured to determine the text configuration information of each of the sample key frame images according to the depth information; and a cover generation sub-module configured to set the text information in the corresponding sample key frame image based on the text configuration information to obtain at least one cover.
[0039] In an exemplary embodiment, the click-through rate acquisition unit includes: a video data acquisition module configured to combine each cover in the cover set with the corresponding sample video to obtain at least one piece of video data; a click count acquisition module configured to publish the video data and obtain the click count of each piece of video data; and a click-through rate determination module configured to determine the click-through rate corresponding to each cover in the cover set according to the click count of the video data and the covers included in each piece of video data.
[0040] In an exemplary embodiment, the model training unit is configured to:
[0041] Extract the cover image and cover text of each cover in the cover set;
[0042] Input the cover image and the cover text corresponding to the same cover identifier into the initial video cover generation model, extract the image feature vector of the cover image and the text feature vector of the cover text, combine the image feature vector and the text feature vector to form a graphic-text feature vector, and output the click-through rate prediction result of the graphic-text feature vector based on the initial video cover generation model of the video cover generation model, where the prediction result carries the cover identifier of the cover corresponding to the graphic-text feature vector;
[0043] Compare the cover identifier carried by the prediction result with the cover label, and calculate a loss value;
[0044] Adjust the parameters of the initial video cover generation model according to the loss value, and train the adjusted initial video cover generation model based on the cover labels of the cover set and the covers in the cover set until the preset training stop condition is met, then stop adjusting the parameters of the initial video cover generation model to obtain the video cover generation model.
[0045] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including: a processor; and a memory for storing executable instructions of the processor; wherein, the processor is configured to execute the instructions to implement the video cover generation method as described in the first aspect above.
[0046] According to a fourth aspect of the embodiments of the present disclosure, a storage medium is provided. When instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the video cover generation method as described in the first aspect.
[0047] According to a fifth aspect of the embodiments of the present disclosure, a computer program product is provided. When instructions in the computer program product are executed by a processor of a computer device, the computer device is enabled to execute the video cover generation method as described in the first aspect.
[0048] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects:
[0049] The video cover generation method, apparatus, and storage medium provided by the embodiments of the present disclosure obtain a set of key frames of a target video, perform text recognition on each key frame image in the set of key frames to obtain a text set corresponding to the target video, input the key frame images in the set of key frames and the text information in the text set into a video cover generation model, predict the click-through rate of each cover formed by combining each key frame image and each text information, output the identification information of the cover with the highest predicted click-through rate, use the key frame image corresponding to the identification information as the target key frame image, and use the text information corresponding to the identification information as the target text information; generate a cover for the target video according to the target key frame image and the target text information. This method inputs the key frame pictures and text information of the target video into the video cover generation model, predicts the click-through rate of the covers formed by combining each key frame image and each text information, and uses the key frame image and text information corresponding to the cover with the highest predicted click-through rate to construct the cover of the target video. Among them, the key frame image corresponding to the cover with the highest predicted click-through rate is the frame image in the target video that best represents the video image content, and the text information corresponding to the cover with the highest predicted click-through rate is the text information in the target video that best represents the video text content. The cover generated based on this key frame image and text information can better reflect the characteristics of the target video, solving the problem of inaccurate video cover generation in the related art. At the same time, by predicting the click-through rate to screen out the key frame images and text information for generating the cover from the set of key frames and the text set, the cover not only shows the characteristics of the video itself but also takes into account the behavioral characteristics of video viewers, that is, it has the characteristics that match the preferences of most video viewers, thereby improving the click-through rate after the video and the cover are put on the market.
[0050] In the embodiments of the present disclosure, by obtaining key frame images of a sample video and text information corresponding to each key frame image, generating at least one cover based on the key frame images and the text information, randomly presenting the covers and the corresponding sample videos to users, obtaining the click-through rates of users for each cover, and using all the covers of the sample video and the cover identifier of the cover with the highest click-through rate as training data to train a video cover generation model, the video cover generation model can automatically learn the ability to select the video cover with the highest click-through rate from multiple covers, improving the accuracy of the video cover generation model in predicting click-through rates.
[0051] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure and do not constitute an improper limitation of the present disclosure.
[0053] Figure 1 It is a schematic diagram of an implementation environment of a video cover generation method shown according to an exemplary embodiment.
[0054] Figure 2 It is a schematic flowchart of a video cover generation method shown according to an exemplary embodiment.
[0055] Figure 3 It is a flowchart of a video cover generation method shown according to an exemplary embodiment.
[0056] Figure 4 It is a flowchart of a training method of a video cover generation model shown according to an exemplary embodiment.
[0057] Figure 5 It is a schematic diagram of a video cover generation model shown according to an exemplary embodiment.
[0058] Figure 6 It is a flowchart of a video cover generation method based on a video cover generation model shown according to an exemplary embodiment.
[0059] Figure 7 It is a schematic structural diagram of a video cover generation device shown according to an exemplary embodiment.
[0060] Figure 8 It is a schematic structural diagram of another video cover generation device shown according to an exemplary embodiment.
[0061] Figure 9It is a block diagram of a terminal of a video cover generation method shown according to an exemplary embodiment.
[0062] Figure 10 It is a block diagram of a server of a video cover generation method shown according to an exemplary embodiment. Detailed implementation manners
[0063] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0064] It should be noted that the terms "first", "second", etc. in the specification, claims and above-mentioned drawings of the present disclosure are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are only examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0065] Figure 1 It is a schematic diagram of an implementation environment of a video cover generation method shown according to an exemplary embodiment; please refer to Figure 1 This implementation environment includes: client 01, server 03.
[0066] Client 01 may include: entity devices of types such as smart phones, tablet computers, laptop computers, digital assistants, smart wearable devices, vehicle-mounted terminals, etc., and may also include software running on the entity devices, such as application programs with video key frame extraction functions, etc. The client 01 may communicate with the server 03 based on the browser / server mode (Browser / Server, B / S) or the client / server mode (Client / Server, C / S).
[0067] The client 01 can extract key frame images of the target video, perform optical character recognition (OCR) on each key frame image to obtain text information, input the key frame images and text information of the target video into a pre-trained video cover generation model, use the video cover generation model to predict the click-through rate for each cover formed by combining each key frame image and each text information, output the identification information of the cover with the highest predicted click-through rate, generate the cover of the target video based on the key frame image and text information corresponding to the identification information, and transmit the cover of the target video to the server 03. In a preferred embodiment, the client 01 can also perform machine learning based on the covers corresponding to the sample videos and the click-through rates of the covers to obtain the video cover generation model.
[0068] The server 03 may include an independently operating server, a distributed server, or a server cluster composed of multiple servers.
[0069] Figure 2 is a flowchart showing a video cover generation method according to an exemplary embodiment. As Figure 2 shown, the video cover generation method of the present disclosure includes two parts, namely a model training part and a model application part. The model training part mainly trains an initial video cover generation model based on training samples (cover labels of the cover set corresponding to the sample videos, and cover images and cover texts of each cover in the cover set) to obtain the video cover generation model; the model application part mainly inputs the target video data (key frame images in the key frame set corresponding to the target video and text information in the text set) into the video cover generation model, outputs the identification information of the cover with the highest predicted click-through rate among the covers formed by combining each key frame image and text information, and then can determine the target key frame image from each key frame image, determine the target text information from each text information, and generate the cover of the target video based on the target key frame image and target text information.
[0070] Figure 3 is a flowchart showing a video cover generation method according to an exemplary embodiment. As Figure 3 shown, a video cover generation method of the present disclosure includes the following steps:
[0071] S301, obtain a key frame set of the target video, where the key frame set includes at least one key frame image.
[0072] S303, perform optical character recognition (OCR) on each key frame image in the key frame set to obtain a text set corresponding to the target video, where the text set includes at least one text information.
[0073] S305. Input the key-frame images in the key-frame set and the text information in the text set into the video cover generation model, predict the click-through rate for each cover formed by combining each key-frame image with each text information, and output the identification information of the cover with the highest predicted click-through rate. The identification information of the cover represents the input sequence information of the key-frame image and the text information that form the cover.
[0074] S307. Use the key-frame image corresponding to the identification information as the target key-frame image and the text information corresponding to the identification information as the target text information.
[0075] S309. Generate the cover of the target video based on the target key-frame image and the target text information.
[0076] The video cover generation method provided by the embodiments of the present disclosure obtains a text set corresponding to the target video by performing text recognition on each key-frame image in the key-frame set of the target video, inputs the key-frame images in the key-frame set and the text information in the text set into the video cover generation model, predicts the click-through rate for each cover formed by combining each key-frame image with each text information, outputs the identification information of the cover with the highest predicted click-through rate, and generates the cover of the target video based on the key-frame image and the text information corresponding to the identification information. This method inputs the key-frame pictures and text information of the target video into the video cover generation model, predicts the click-through rate of the covers formed by combining each key-frame image with each text information, and uses the key-frame image and the text information corresponding to the cover with the highest predicted click-through rate to construct the cover of the target video. Among them, the key-frame image corresponding to the cover with the highest predicted click-through rate is the frame image in the target video that best represents the video image content, and the text information corresponding to the cover with the highest predicted click-through rate is the text information in the target video that best represents the video text content. The cover generated based on this key-frame image and text information can better reflect the characteristics of the target video, solving the problem of inaccurate video cover generation in the related art. At the same time, by predicting the click-through rate to screen out the key-frame images and text information for generating the cover from the key-frame set and the text set, the cover not only shows the characteristics of the video itself but also takes into account the behavioral characteristics of video viewers, that is, it has the characteristics that match the preferences of most video viewers, thereby improving the click-through rate after the video and the cover are put on the market.
[0077] In a possible implementation, obtaining the key-frame set of the target video includes:
[0078] Detect the key-frame information of the target video to obtain the key-frame images of the target video;
[0079] Determine the key-frame set corresponding to the target video according to the key-frame images.
[0080] In a possible implementation, performing key frame information detection on a target video to obtain a key frame image of the target video includes at least one of the following:
[0081] Perform face detection on the target video, and select frame images containing faces from the target video as key frame images;
[0082] Performing clarity detection on the target video, sorting the clarity of the frame images of the target video in descending order, and selecting a preset number of frame images that are ranked first in the target video as key frame images;
[0083] Perform highlight detection on the target video, and use the frame images corresponding to the highlight in the target video as key frame images.
[0084] Each key frame image in the key frame set is determined by detecting key frame information of the target video. These key frame images are frame images that contain face images, wonderful pictures or have high definition in the target video. They can better reflect the image content characteristics of the target video, making the video cover generated subsequently based on the key frame images more representative.
[0085] In a possible implementation, performing text recognition on each key frame image in the key frame set to obtain a text set corresponding to the target video includes:
[0086] Performing text recognition on each key frame image in the key frame set to obtain recognized text corresponding to the key frame image, and performing denoising on the recognized text corresponding to each key frame image;
[0087] Perform semantic recognition on the recognized text after denoising, remove text contents with similar semantics, and obtain text information;
[0088] Determine the text set corresponding to the target video according to the text information.
[0089] By extracting the text from the key frame image and then performing deduplication and denoising processing on the extracted text, the meaningless and repeated text content is eliminated to obtain concise text information. This text information matches the image content of the key frame image and can accurately reflect the text content characteristics of the target video.
[0090] In a possible implementation, generating a cover of a target video according to a target key frame image and target text information includes:
[0091] Obtain the depth information of the target key frame image;
[0092] Determine text configuration information of the target key frame image according to the depth information;
[0093] Set the target text information in the target key-frame image based on the text configuration information to obtain the cover of the target video.
[0094] In an embodiment of the present disclosure, key-frame images are extracted from the target video, text information is extracted from the key-frame images, and then a target key-frame image is determined from each key-frame image through a video cover generation model, a target text information is determined from each text information, and based on the analysis result of the depth information of the target key-frame image, the target text information is set at an appropriate position in the target key-frame to obtain the cover of the target video. This cover has a presentation effect with coordinated text and images, can comprehensively display the characteristics of the target video, and enhance the attractiveness for users to click on the target video.
[0095] In a possible implementation manner, before obtaining the set of key-frame images of the target video, it further includes:
[0096] Obtain a set of covers of the sample video, where the set of covers includes at least one cover, and each cover has a unique cover identifier;
[0097] Obtain the click-through rate corresponding to each cover in the set of covers;
[0098] Sort the covers in the set of covers in descending order according to the click-through rate, and select the cover identifier of the first cover as the cover label of the set of covers;
[0099] Train the video cover generation model according to the cover label of the set of covers and the covers in the set of covers.
[0100] In a possible implementation manner, obtaining the set of covers of the sample video includes:
[0101] Obtain the sample video;
[0102] Perform key-frame detection on the sample video to obtain a set of sample key-frames, where the set of sample key-frames includes at least one sample key-frame image;
[0103] Perform text recognition on each sample key-frame image in the set of sample key-frames to obtain a set of sample texts, where the set of sample texts includes at least one sample text information;
[0104] Combine each sample key-frame image in the set of sample key-frames with each sample text information in the set of sample texts to generate at least one cover;
[0105] Determine the set of covers of the sample video according to the generated covers.
[0106] In a possible implementation manner, combining each sample key-frame image in the set of sample key-frames with each sample text information in the set of sample texts to generate at least one cover includes:
[0107] Perform depth information recognition on the sample key frame images in the set of sample key frames, and obtain the depth information of each sample key frame image;
[0108] Determine the text configuration information of each sample key frame image according to the depth information;
[0109] Set the text information in the corresponding sample key frame image based on the text configuration information to obtain at least one cover.
[0110] In a possible implementation, obtaining the click-through rate corresponding to each cover in the cover set includes:
[0111] Combine each cover in the cover set with the corresponding sample video to obtain at least one video data;
[0112] Publish the video data and obtain the click count of each video data;
[0113] Determine the click-through rate corresponding to each cover in the cover set according to the click count of the video data and the covers included in each video data.
[0114] The embodiments of the present disclosure adopt a key frame information detection method to extract the sample key frame images of the sample video, then extract the sample text information from the sample key frame images, combine the sample key frame images and the sample text information into multiple covers, and associate each cover with the sample video for delivery, obtain the true click-through rate of each cover, and generate sample data for training the model according to each cover and the corresponding click-through rate. The sample data is real and reliable, and can ensure that the model trained based on the sample data has a high prediction accuracy.
[0115] In a possible implementation, training a video cover generation model according to the cover labels of the cover set and the covers in the cover set includes:
[0116] Extract the cover image and cover text of each cover in the cover set;
[0117] Input the cover image and cover text corresponding to the same cover identifier into the video cover generation model, extract the image feature vector of the cover image and the text feature vector of the cover text, combine the image feature vector and the text feature vector to form a graphic-text feature vector, and output the click-through rate prediction result of the graphic-text feature vector based on the initial deep learning model of the video cover generation model. The prediction result carries the cover identifier of the cover corresponding to the graphic-text feature vector;
[0118] Compare the cover identifier carried in the prediction result with the cover label, and calculate the loss value;
[0119] Adjust the parameters of the initial deep learning model according to the loss value, and train the adjusted initial deep learning model based on the cover labels of the cover set and the covers in the cover set until the preset training stop condition is met, then stop adjusting the parameters of the initial deep learning model to obtain a video cover generation model.
[0120] In the embodiments of the present disclosure, by obtaining the key frame images of the sample video and the text information corresponding to each key frame image, generating at least one cover according to the key frame images and the text information, randomly delivering the covers and the corresponding sample videos to users, and obtaining the click-through rates of the users for each cover, using all the covers of the sample video and the cover identifier of the cover with the highest click-through rate as training data to train the video cover generation model, the video cover generation model can automatically learn the ability to select the video cover with the highest click-through rate from multiple covers, and improve the accuracy of the video cover generation model in predicting the click-through rate.
[0121] Figure 4 It is a flowchart of a method for training a video cover generation model shown according to an exemplary embodiment. As Figure 4 shown, this method is used in a terminal and includes the following steps.
[0122] S401, obtain a cover set of the sample video, where the cover set includes at least one cover, and each cover has a unique cover identifier.
[0123] The cover of a video refers to the picture presented before the video is played, generally composed of representative frame images and text introductions. In the embodiments of the present disclosure, obtaining the cover set of the sample video may include:
[0124] 1. Obtain the sample video; perform key frame detection on the sample video to obtain a sample key frame set.
[0125] Among them, the sample key frame set includes at least one sample key frame image. The sample key frame image is a frame image in the sample video suitable for use as a cover, generally containing more video content information, and can be detected by key frame information detection methods such as face detection, clarity detection, and wonderful moment detection.
[0126] Specifically, the key frame information of the sample video can be detected by any of the following methods to obtain the key frame images of the sample video.
[0127] First, perform face detection on the sample video, and select the frame images containing faces in the sample video as key frame images.
[0128] In a possible implementation, face detection is performed on each frame image of the sample video, and the frame images containing faces are used as key frame images. For example, a face detection model is used to detect each frame image in the sample video to determine whether each frame image contains a face.
[0129] In another possible implementation, a face filtering method can be adopted to filter out the frame images in the sample video that do not contain faces, and the remaining frame images after the filtering process are used as key frame images. For example, for each frame image in the sample video, when the previous frame image of this image contains a face, the next frame image of this image also contains a face, and this image does not include a face, it can be considered that this image does not include a face or the face is blocked, and then this frame image is filtered out.
[0130] In another possible implementation, expression detection can also be performed on each frame image in the sample video, and the frame images containing smiling faces are selected as key frame images. For example, a face detection model is used to detect each frame image in the sample video to determine whether each frame image contains a face. On this basis, a smiling face detection model is used to perform expression recognition on each frame image containing a face to determine whether the frame image contains a smiling face.
[0131] Second, perform clarity detection on the sample video, sort the frame images of the sample video in descending order of clarity, and select a preset number of frame images ranked at the front from the sample video as key frame images.
[0132] In the quality evaluation without reference images, the clarity of an image is an important indicator to measure the quality of the image. It can correspond well to people's subjective feelings. Graphics with high clarity can bring a better visual experience and are suitable to be used as the cover of a video.
[0133] For each frame image in the sample video, an image quality evaluation model can be used to evaluate the image quality of each frame image and obtain an image quality score. Among them, the content of the image quality evaluation by the image quality evaluation model includes the clarity of the image, whether the image color is pure, image noise, and whether there are a large number of overexposed or underexposed areas in the image, etc.
[0134] In a possible implementation, the image scores of each frame image can be sorted in descending order, and a preset number of frame images ranked at the front are selected as key frame images. The image quality score represents the clarity of the frame image. The larger the score, the higher the clarity. Selecting a preset number of frame images ranked at the front as key frame images can ensure that the video cover is presented with high-quality pictures.
[0135] In addition, a score threshold can be preset, and the image score of each frame image in the sample video is compared with the preset score, and the frame image with an image score greater than the score threshold is used as a key frame image. For sample videos with higher quality, the score threshold can be appropriately increased, and for sample videos with lower quality, the score threshold can be reduced, so as to obtain a key frame image that matches the clarity of the sample video.
[0136] Third, the highlight segments of the sample video are detected, and the frame images corresponding to the highlight segments in the sample video are used as key frame images.
[0137] For a video, not all frames are worth appreciating. Some frames only contain background environment without people or other moving objects. When evaluating whether a video contains highlights, it is often determined based on whether the video contains people or moving objects, that is, a segment consisting of continuous frames containing people or moving objects will be regarded as a highlight of the video. Specifically, a highlight recognition model can be trained to determine the highlights of a sample video through the highlight recognition model, and each frame image corresponding to the highlight is used as a key frame image, wherein the sample data used to train the highlight recognition model can include a video and annotated highlights of the video.
[0138] By detecting key frame information of sample videos, frame images containing faces, actions or with higher definition are extracted as key frame images. Key frame images can better represent the video content and provide better materials for subsequent cover selection.
[0139] 2. Perform text recognition on each sample key frame image in the sample key frame set to obtain a sample text set.
[0140] Among them, the sample text set includes at least one sample text information, and the sample text set can be determined by the following method: performing text recognition on each key frame image in the key frame set to obtain the recognized text corresponding to the key frame image, and denoising the recognized text corresponding to each key frame image; performing semantic recognition on the recognized text after denoising, eliminating text content with similar semantics, and obtaining text information; and determining the text set corresponding to the target video based on the text information.
[0141] In a possible implementation, an OCR (Optical Character Recognition) software can be used to detect and recognize the key-frame images, convert the text content on the key-frame images into editable text, and then splice the recognized text on the same key-frame image according to the coordinates of the text on the key-frame image to obtain the editable text corresponding to the key-frame image. Denoising processing is performed on the editable text to remove special punctuation marks and small text in the text. After that, semantic recognition is performed on the texts corresponding to each key-frame image to remove text content with similar and identical semantics. For example, when there are multiple texts with the same or similar semantics, only the first text that appears for the first time is retained, and other texts are removed.
[0142] 3. Combine each sample key-frame image in the sample key-frame set with each sample text information in the sample text set to generate at least one cover.
[0143] The sample key-frame set includes at least one sample key-frame image, and the sample text set includes at least one sample text information. Combine each sample key-frame image in the sample key-frame set with each sample file information in the sample text set to obtain multiple covers. For example, if the sample key-frame set includes N sample key-frame images and the sample text set includes M sample text information, then the number of generated covers is at least N×M.
[0144] In a possible implementation, depth information recognition can be performed on the sample key-frame images in the sample key-frame set to obtain the depth information of each sample key-frame image; determine the text configuration information of each sample key-frame image according to the depth information; and set the text information in the corresponding sample key-frame image based on the text configuration information to obtain at least one cover.
[0145] The image depth refers to the number of bits used to store each pixel and is also used to measure the color resolution of the image. The image depth determines the number of colors that each pixel of a color image may have, or determines the number of gray levels that each pixel of a grayscale image may have, and determines the maximum number of colors that may appear in a color image, or the maximum gray level in a grayscale image. Based on the characteristics of the image depth, by analyzing the depth information of the sample key-frame image, the position and area suitable for placing text in the sample key-frame image can be found, and further, the number of texts to be adapted, the size and contrast of the texts, etc. can be determined to generate a suitable text and image as the cover of the video. Specifically, the depth information of the sample key-frame image can be analyzed, the area with relatively single color number or relatively low gray level can be used as the candidate area, face detection can be performed on the candidate area, and the area with the largest area among the candidate areas that do not contain faces can be used as the target area, and this target area is the area where text can be filled. After determining the position and area of the target area, obtain the preset text setting rules (the text setting rules include the number of lines of text allowed to be set on the sample key-frame image and the number of texts in each line), and determine each text content that can be set on the sample key-frame image according to the text setting rules and each sample text information. Each text content includes at least one piece of sample text information, and each text content is correspondingly set in the target area on the sample key-frame image in a way that fills the target area, and the color or gray level of the text content is determined according to the color or gray level around the target area, where the color or gray level of the text content is different from the color or gray level around the target area in the sample key-frame image.
[0146] In a possible implementation manner, the following method can be used to obtain the image depth information: 1) Perform Gaussian blur processing on the key-frame image to obtain a blurred image; 2) Detect the texture edges of the key-frame image, and divide the key-frame image into areas with relatively large texture gradients (defined as D areas) and areas with relatively small texture gradients (defined as F areas); 3) For the pixel points in the D area, calculate the proportionality factor of each pixel point according to the fuzzy estimation method; 4) For each pixel point in the F area, perform Kalman filtering to estimate the proportionality factor of each pixel point; 5) According to the focusing information of the key-frame image, convert the proportionality factor of each pixel point into the relative depth value of each pixel point.
[0147] 4. Determine the cover set of the sample video according to the generated cover.
[0148] Form a cover set with all the covers generated in step 3, and establish a corresponding relationship between the cover set and the sample video.
[0149] S403, obtain the click-through rate corresponding to each cover in the cover set.
[0150] In a possible implementation, each cover in the cover set can be combined with the corresponding sample video to obtain at least one video data; the video data is published, and the click count of each video data is obtained; according to the click count of the video data and the covers included in each video data, the click-through rate corresponding to each cover in the cover set is determined.
[0151] All the covers corresponding to the sample video are randomly published to different users, and the click-through rate of each cover is determined according to the click count of each cover. The click-through rate of a cover is the ratio of the click count of this cover to the total click count of all covers in the cover set.
[0152] S405, sort each cover in the cover set in descending order according to the click-through rate, and select the cover identifier of the first cover as the cover label of the cover set.
[0153] The click-through rate can be used to characterize the popularity of each cover in the cover set. Sort each cover in the cover set in descending order according to the click-through rate. The cover ranked first is the cover with the highest click-through rate, that is, the cover most favored by users. Using it as the cover label for subsequent training of the video cover generation model can improve the accuracy of the video cover generation model in predicting the cover with the highest click-through rate.
[0154] S407, train the video cover generation model according to the cover label of the cover set and the covers in the cover set.
[0155] In a possible implementation, training the video cover generation model may include:
[0156] Extract the cover image and cover text of each cover in the cover set; input the cover image and cover text corresponding to the same cover identifier into the video cover generation model, extract the image feature vector of the cover image and the text feature vector of the cover text, combine the image feature vector and the text feature vector to form a graphic-text feature vector, and based on the initial deep learning model of the video cover generation model, output the click-through rate prediction result of the graphic-text feature vector. The prediction result carries the cover identifier of the cover corresponding to the graphic-text feature vector; compare the cover identifier carried by the prediction result with the cover label to calculate the loss value; adjust the parameters of the initial deep learning model according to the loss value, and train the adjusted initial deep learning model based on the cover label of the cover set and the covers in the cover set until the preset training stop condition is met, then stop adjusting the parameters of the initial deep learning model to obtain the video cover generation model. Among them, the training stop condition can be that the loss value reaches a preset value or the number of training times reaches a preset number, or the training of the model can also be stopped when the loss value does not decrease significantly compared with the loss value obtained last time.
[0157] In a possible implementation, the video cover generation model is asFigure 5 As shown, it may include two feature extraction layers, a feature splicing layer, a pooling layer, and a regression layer. Among them, one feature extraction layer uses a convolutional neural network to extract image features from the input cover image, and inputs the extracted image features into the feature splicing layer. The other feature extraction layer uses a convolutional neural network to extract text features from the input cover text, and inputs the extracted text features into the feature splicing layer. The feature splicing layer uses a convolutional neural network to perform an inner product on the image features and text features, and splices the eigenvalues obtained from the inner product to obtain a combined text and image feature. The combined text and image feature is input into the pooling layer, and the pooling layer compresses the combined text and image feature into a fixed dimension and then inputs it into the regression layer. In the regression layer, the softmax function is used to predict the click-through rate, and the cover identifier of the cover with the highest predicted click-through rate is obtained. The cross-entropy loss is calculated between the obtained cover identifier and the cover label, and the parameters of the model are adjusted based on the cross-entropy loss. Among them, text recognition can be performed on the cover, and the recognition result is converted into a computable structured vector through the Word2vec model, and the obtained vector is used as the cover text and input into the feature extraction layer.
[0158] It should be noted that in the embodiments of the present disclosure, only the execution subject is taken as an example of a terminal. In another embodiment, the training method provided by the embodiments of the present disclosure can also be executed by a server, and the embodiments of the present disclosure do not limit the execution subject.
[0159] In the embodiments of the present disclosure, by obtaining the key frame images of the sample video and the text information corresponding to each key frame image, at least one cover is generated according to the key frame images and the text information, the cover and the corresponding sample video are randomly presented to the user, the click-through rate of each cover by the user is obtained, and all the covers of the sample video and the cover identifier of the cover with the highest click-through rate are used as training data to train the video cover generation model, which can enable the video cover generation model to automatically learn the ability to select the video cover with the highest click-through rate from multiple covers, and improve the accuracy of the video cover generation model in predicting the click-through rate.
[0160] Figure 6 is a flowchart of a video cover generation method based on a video cover generation model shown according to an exemplary embodiment, applied to a terminal. Refer to Figure 6 and includes the following steps.
[0161] S601. Obtain a set of key frames of the target video.
[0162] The terminal stores multiple videos. The target video can be any one of the videos stored in the terminal that needs to generate a video cover. Moreover, the videos stored in the terminal can be obtained by the terminal through shooting, or downloaded from the server by the terminal, or obtained by shooting from other terminals. The videos stored in the server can be provided by the publisher to the maintainer, and stored in the server by the maintainer, or sent to the server by the terminal, or sent to the server by other devices.
[0163] The key frame set includes at least one key frame image. In a possible implementation, key frame information detection can be performed on the target video to obtain the key frame images of the target video; the key frame set corresponding to the target video is determined according to the key frame images. Specifically, the key frame images of the target video can be obtained through at least one of the following:
[0164] First, perform face detection on the target video, and select the frame images containing faces from the target video as key frame images;
[0165] Second, perform clarity detection on the target video, sort the frame images of the target video in descending order of clarity, and select the preset number of frame images ranked at the front from the target video as key frame images;
[0166] Third, perform highlight segment detection on the target video, and select the frame images containing highlight segments from the target video as key frame images.
[0167] The specific implementation manner of obtaining the key frame images of the target video is similar to the implementation manner of obtaining the sample key frame images in the sample video in step S401. For details, please refer to step S401 and will not be elaborated here.
[0168] By performing key frame information detection on the target video, the frame images containing faces, actions or with high clarity in the target video are extracted as key frame images. The key frame images can better represent the content of the target video and provide better materials for the subsequent selection of the cover.
[0169] S603. Perform text recognition on each key frame image in the key frame set to obtain a text set corresponding to the target video.
[0170] The text set includes at least one piece of text information. In a possible implementation, the text set corresponding to the target video can be determined in the following way: perform text recognition on each key frame image in the key frame set to obtain the recognized text corresponding to the key frame image, and perform denoising processing on the recognized text corresponding to each key frame image; perform semantic recognition on the denoised recognized text, eliminate the text content with the same and similar semantics, and obtain the text information; determine the text set corresponding to the target video according to the text information.
[0171] By extracting the text in the key frame images, and then performing duplicate and noise removal processing on the extracted text to eliminate the text content without actual meaning and duplicates, a concise text information that matches the image content of the key frame images is obtained, which can accurately reflect the text content characteristics of the target video.
[0172] S605. Obtain a video cover generation model.
[0173] In the embodiments of the present disclosure, the video cover generation model has been trained, and the terminal stores the video cover generation model. When generating a video cover for the target video, the stored video cover generation model can be obtained. Among them, the video cover generation model can be trained through step S401-step S407, or can also be trained by other methods.
[0174] S607. Input the key frame images in the key frame set and the text information in the text set into the video cover generation model, predict the click-through rate of each cover formed by combining each key frame image and each text information, and output the identification information of the cover with the highest predicted click-through rate.
[0175] See Figure 5 the model structure of, input the key frame images in the key frame set and the text information in the text set into the video cover generation model. Each key frame image has an image input order, and each text information has a text input order. Extract features from the key frame images to obtain image features, extract features from the text information to obtain text features, splice the text features and image features to obtain combined features. The combined features represent the features of the cover composed of the text information corresponding to the text features and the key frame images corresponding to the image features. Therefore, the combined features can be regarded as the features of the cover. Input the features of the cover into the pooling layer for dimensionality reduction processing, and input the features after dimensionality reduction processing into the regression layer to predict the click-through rate of the cover corresponding to the features. Based on the predicted click-through rates of each cover, output the identification information of the cover with the highest click-through rate among all covers obtained by combining the key frame images in the key frame set and the text information in the text set. Among them, the identification information of the cover represents the input order information of the key frame image and text information forming the cover, including the input order information of the corresponding key frame image and the input order information of the corresponding text information.
[0176] S609. Use the key frame image corresponding to the identification information as the target key frame image, and use the text information corresponding to the identification information as the target text information.
[0177] S611. Generate a cover for the target video according to the target key frame image and the target text information.
[0178] In a possible implementation, generating the cover of the target video includes: obtaining the depth information of the target key-frame image; determining the text configuration information of the target key-frame image according to the depth information; and setting the target text information in the target key-frame image based on the text configuration information to obtain the cover of the target video.
[0179] Among them, the text configuration information includes the position and area suitable for placing text in the target key-frame image, as well as the color and contrast of the text. Specifically, the depth information of the target key-frame image can be analyzed, and the area with relatively single color number or low gray level can be used as the candidate area. Face detection is performed on the candidate area, and the area with the largest area in the candidate area that does not contain a face is used as the target area. This target area is the area suitable for placing text. The color or gray level of the surrounding area of the target area is detected, and the color or gray level value of the target text information is determined according to the color or gray level value of the surrounding area. The color of the target text information does not coincide with the color of the surrounding area, or the gray level value of the target text information is different from the gray level value of the surrounding area.
[0180] In the embodiments of the present disclosure, after obtaining the cover of the target video, the terminal can display the target video and the corresponding video cover for the user to view. When the user triggers the video cover, the terminal detects the trigger operation and can play the target video. Among them, the trigger operation can be a click operation, a long-press operation, a swipe operation, etc. The target video can be a movie, a TV drama and other film and television works, a food video, a beauty video, a funny video, etc.
[0181] It should be noted that the embodiments of the present disclosure only take the execution entity as the terminal as an example. In another embodiment, the video cover generation method provided by the embodiments of the present disclosure can also be executed by a server. For example, the server receives the target video sent by the terminal, obtains the key-frame set and text set corresponding to the target video, inputs the key-frame image corresponding to the key-frame set and the text information corresponding to the text set into the video cover generation model, outputs the identification information of the cover with the highest predicted click-through rate, generates the cover of the target video according to the key-frame image and text information corresponding to the identification information, sends the target video and the cover to the terminal, and the terminal displays the cover of the target video for the user to view. When detecting the trigger operation of the user on the video cover, play the target video for the user to watch.
[0182] In another embodiment, the method can be applied to a terminal and a server. The terminal obtains a set of key frames and a set of texts corresponding to the target video according to the target video, and sends the target video, the set of key frames and the set of texts corresponding to the target video to the server. The server inputs the key frame images corresponding to the set of key frames and the text information corresponding to the set of texts into a video cover generation model, outputs the identification information of the cover with the highest predicted click-through rate, generates the cover of the target video according to the key frame image and the text information corresponding to the identification information, and sends the target video and the cover to the terminal. The terminal displays the cover of the target video for the user to view. When a trigger operation on the video cover is detected, the target video is played for the user to watch.
[0183] The video cover generation method of the present disclosure inputs the key frame pictures and text information of the target video into a video cover generation model, predicts the click-through rate of the covers formed by combining each key frame image and each text information, and uses the key frame image and the text information corresponding to the cover with the highest predicted click-through rate to construct the cover of the target video, which can take into account the characteristics of the video itself and the behavioral characteristics of video viewers, and improve the click-through rate after the video and the cover are put on the market.
[0184] The method of the present disclosure starts from the characteristics of short video consumers' habits, behaviors, etc., utilizes a large number of user behavioral characteristics of the platform, recommends good short video covers, edits text covers and display methods that users like, etc. Under the condition of generating a small number of video covers for the target video, it covers the preferences of most users and maximally improves the click-through rate and viewing duration of the video.
[0185] Figure 7 It is a schematic structural diagram of a video cover generation device shown according to an exemplary embodiment. Refer to Figure 7 , the video cover generation device includes:
[0186] A key frame set acquisition unit 710, configured to acquire a set of key frames of the target video, and the set of key frames includes at least one key frame image;
[0187] A text set acquisition unit 720, configured to perform text recognition on each key frame image in the set of key frames to obtain a set of texts corresponding to the target video, and the set of texts includes at least one piece of text information;
[0188] An identification information acquisition unit 730, configured to input the key frame images in the set of key frames and the text information in the set of texts into a video cover generation model, predict the click-through rate of each cover formed by combining each key frame image and each text information, and output the identification information of the cover with the highest predicted click-through rate. The identification information of the cover represents the input order information of the key frame image and the text information forming the cover;
[0189] A cover information determination unit 740, configured to use the key-frame image corresponding to the identification information as the target key-frame image and the text information corresponding to the identification information as the target text information;
[0190] A cover generation unit 750, configured to generate a cover for the target video according to the target key-frame image and the target text information.
[0191] In a possible implementation, please refer to Figure 8 , the key-frame set acquisition unit 710 may include: a key-frame image acquisition module 711, configured to detect key-frame information of the target video and acquire the key-frame image of the target video; a key-frame set determination module 712, configured to determine the key-frame set corresponding to the target video according to the key-frame image.
[0192] In another possible implementation, the key-frame image acquisition module 711 is configured to perform at least one of the following: perform face detection on the target video, and select a frame image containing a face from the target video as the key-frame image; perform clarity detection on the target video, sort the clarity of the frame images of the target video in descending order, and select a preset number of frame images sorted at the front from the target video as the key-frame image; perform highlight segment detection on the target video, and select a frame image containing a highlight segment from the target video as the key-frame image.
[0193] In a possible implementation, the text set acquisition unit 720 may include: a text recognition module 721, configured to perform text recognition on each key-frame image in the key-frame set to obtain the recognized text corresponding to the key-frame image; a processing module 722, configured to perform denoising processing on the recognized text corresponding to each key-frame image; and perform semantic recognition on the denoised recognized text, and eliminate text contents with similar semantics to obtain text information; a text set determination module 723, configured to determine the text set corresponding to the target video according to the text information.
[0194] In a possible implementation, the cover generation unit 750 may include: a depth information acquisition module 751, configured to acquire the depth information of the target key-frame image; a text configuration information acquisition module 752, configured to determine the text configuration information of the target key-frame image according to the depth information; a first cover generation module 753, configured to set the target text information in the target key-frame image based on the text configuration information to obtain the cover of the target video.
[0195] The video cover generation device disclosed in the embodiments of the present disclosure inputs the key frame images and text information of the target video into a video cover generation model, predicts the click-through rate of the covers formed by combining each key frame image and each text information, and uses the key frame image and text information corresponding to the cover with the highest predicted click-through rate to construct the cover of the target video, taking into account both the characteristics of the video itself and the behavioral characteristics of video viewers, solving the problem of inaccurate video cover generation in the related art, and being able to improve the click-through rate after the video and the cover are put on the market.
[0196] Please refer to Figure 8 In a possible implementation manner, the video cover generation device may further include: a cover set acquisition unit 810, configured to acquire a cover set of a sample video, the cover set including at least one cover, and each cover having a unique cover identifier; a click-through rate acquisition unit 820, configured to acquire the click-through rate corresponding to each cover in the cover set; a cover label determination unit 830, configured to sort the covers in the cover set in descending order according to the click-through rate, and select the cover identifier of the first cover as the cover label of the cover set; and a model training unit 840, configured to train the video cover generation model according to the cover label of the cover set and the covers in the cover set.
[0197] In a possible implementation manner, the cover set acquisition unit 810 may include: a sample video acquisition module 811, configured to acquire a sample video; a sample key frame image acquisition module 812, configured to perform key frame detection on the sample video to obtain a sample key frame set, the sample key frame set including at least one sample key frame image; a sample text information acquisition module 813, configured to perform text recognition on each sample key frame image in the sample key frame set to obtain a sample text set, the sample text set including at least one sample text information; a second cover generation module 814, configured to combine each sample key frame image in the sample key frame set with each sample text information in the sample text set to generate at least one cover; and a cover set determination module 815, configured to determine the cover set of the sample video according to the generated covers.
[0198] In a possible implementation manner, the second cover generation module 814 may further include: a depth information acquisition sub-module 8141, configured to perform depth information recognition on the sample key frame images in the sample key frame set to obtain the depth information of each sample key frame image; a text configuration information acquisition sub-module 8142, configured to determine the text configuration information of each sample key frame image according to the depth information; and a cover generation sub-module 8143, configured to set the text information in the corresponding sample key frame image based on the text configuration information to obtain at least one cover.
[0199] In a possible implementation, the click-through rate acquisition unit 820 may include: a video data acquisition module 821 configured to combine each cover in the cover set with the corresponding sample video to obtain at least one piece of video data; a click count acquisition module 822 configured to publish the video data and obtain the click counts of each piece of video data; and a click-through rate determination module 823 configured to determine the click-through rate corresponding to each cover in the cover set according to the click counts of the video data and the covers included in each piece of video data.
[0200] In another possible implementation, the model training unit 840 is configured to: extract the cover images and cover texts of each cover in the cover set; input the cover images and cover texts corresponding to the same cover identifier into the video cover generation model, extract the image feature vector of the cover image and the text feature vector of the cover text, combine the image feature vector and the text feature vector to form a graphic-text feature vector, and based on the initial deep learning model of the video cover generation model, output the click-through rate prediction result of the graphic-text feature vector, where the prediction result carries the cover identifier of the cover corresponding to the graphic-text feature vector; compare the cover identifier carried in the prediction result with the cover label, and calculate to obtain a loss value; adjust the parameters of the initial deep learning model according to the loss value, and train the adjusted initial deep learning model based on the cover labels of the cover set and the covers in the cover set until the preset training stop condition is met, then stop adjusting the parameters of the initial deep learning model to obtain the video cover generation model.
[0201] In the embodiments of the present disclosure, by acquiring the key frame images of the sample video and the text information corresponding to each key frame image, generating at least one cover according to the key frame images and the text information, randomly delivering the covers and the corresponding sample videos to users, and acquiring the click-through rates of the users for each cover, using all the covers of the sample video and the cover identifier of the cover with the highest click-through rate as training data to train the video cover generation model, the video cover generation model can automatically learn the ability to select the video cover with the highest click-through rate from multiple covers, improving the accuracy of predicting the click-through rate of the video cover generation model.
[0202] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.
[0203] Figure 9The block diagram of a terminal for a video cover generation method shown according to an exemplary embodiment. The terminal is used to execute the steps performed by the terminal in the above video cover generation method, and may be a portable mobile terminal, such as: a smart phone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a notebook computer or a desktop computer. The terminal may also be referred to by other names such as user equipment, portable terminal, laptop terminal, desktop terminal, etc.
[0204] Referring to Figure 9 , the terminal may include one or more of the following components: a processing component 902, a memory 904, a power supply component 906, a multimedia component 908, an audio component 910, an input / output (I / O) interface 912, a sensor component 914, and a communication component 916.
[0205] The processing component 902 generally controls the overall operation of the terminal, such as operations associated with display, telephone call, data communication, camera operation, and recording operation. The processing component 902 may include one or more processors 920 to execute instructions to complete all or part of the steps of the above method. In addition, the processing component 902 may include one or more modules to facilitate the interaction between the processing component 902 and other components. For example, the processing component 902 may include a multimedia module to facilitate the interaction between the multimedia component 908 and the processing component 902.
[0206] The memory 904 is configured to store various types of data to support the operation of the terminal. Examples of these data include instructions for any application or method operating on the terminal, contact data, phone book data, messages, images, videos, etc. The memory 904 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.
[0207] The power supply component 906 provides power to various components of the terminal. The power supply component 906 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the terminal.
[0208] The multimedia component 908 includes a screen that provides an output interface between the terminal and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of a touch or swipe action but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 908 includes a front camera and / or a rear camera. When the terminal is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.
[0209] The audio component 910 is configured to output and / or input audio signals. For example, the audio component 910 includes a microphone (MIC) that is configured to receive external audio signals when the terminal is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 904 or transmitted via the communication component 916. In some embodiments, the audio component 910 further includes a speaker for outputting audio signals.
[0210] The I / O interface 912 provides an interface between the processing component 902 and a peripheral interface module, and the peripheral interface module can be a keyboard, a click wheel, buttons, etc. These buttons can include but are not limited to: a home button, a volume button, a power button, and a lock button.
[0211] The sensor component 914 includes one or more sensors for providing status assessments of various aspects of the terminal. For example, the sensor component 914 can detect the on / off state of the terminal, the relative positioning of components, such as the display and keypad of the terminal. The sensor component 914 can also detect a change in the position of the terminal or a component of the terminal, the presence or absence of user contact with the terminal, the orientation or acceleration / deceleration of the terminal, and the temperature change of the terminal. The sensor component 914 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor component 914 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 914 can further include an acceleration sensor, a gyro sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0212] The communication component 916 is configured to facilitate communication between the terminal and other devices in a wired or wireless manner. The terminal can access a wireless network based on a communication standard, such as WiFi, a carrier network (such as 2G, 3G, 4G, or 5G), or a combination thereof. In an exemplary embodiment, the communication component 916 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 916 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra-Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0213] In an exemplary embodiment, the terminal can be implemented by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.
[0214] In an exemplary embodiment, a storage medium including instructions is also provided, such as a memory 904 including instructions, and the above instructions can be executed by a processor 920 of the terminal to complete the above method. Optionally, the storage medium can be a non-transitory computer-readable storage medium. For example, the non-transitory computer-readable storage medium can be ROM, Random Access Memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage devices, etc.
[0215] In an exemplary embodiment, a computer program product is also provided, and the computer program product includes readable program code that can be executed by a processor 920 of the terminal to complete the above method. Optionally, the program code can be stored in a storage medium of the terminal, and the storage medium can be a non-transitory computer-readable storage medium. For example, the non-transitory computer-readable storage medium can be ROM, Random Access Memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage devices, etc.
[0216] Figure 10 is a block diagram of a server for a video cover generation method shown according to an exemplary embodiment. The server 1000 can be used to perform the steps executed by the server in the above video cover generation method. Refer to Figure 10, the server 1000 includes a processing component 1010, which further includes one or more processors, and memory resources represented by a memory 1020 for storing instructions executable by the processing component 1010, such as application programs. The application programs stored in the memory 1020 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 1010 is configured to execute instructions to perform the above video cover generation method.
[0217] The server 1000 may further include a power supply group 1030 configured to perform power management of the server 1000, a wired or wireless network interface 1050 configured to connect the server 1000 to a network, and an input / output (I / O) interface 1040. The server 1000 may operate based on an operating system stored in the memory 1020, such as Windows ServerTM, MacOS XTM, UnixTM, LinuxTM, FreeBSD TM or the like.
[0218] In an exemplary embodiment, a non-transitory computer-readable storage medium is also provided. When the instructions in the storage medium are executed by a processor of a computer device, the computer device is enabled to perform the steps executed by the terminal or the server in the above video cover generation method.
[0219] In an exemplary embodiment, a computer program product is also provided. When the instructions in the computer program product are executed by a processor of a computer device, the computer device is enabled to perform the steps executed by the terminal or the server in the above video cover generation method.
[0220] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0221] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A method for generating a video cover, characterized in that, Including: Obtain a set of key frames of the target video, where the set of key frames includes at least one key frame image; Perform text recognition on each key frame image in the set of key frames to obtain a text set corresponding to the target video, where the text set includes at least one piece of text information; Input the key frame images in the set of key frames and the text information in the text set into a video cover generation model, predict the click-through rate for each cover formed by combining each key frame image and each piece of text information, and output the identification information of the cover with the highest predicted click-through rate. The identification information of the cover represents the input order information of the key frame image and the text information that form the cover; Use the key frame image corresponding to the identification information as the target key frame image, and use the text information corresponding to the identification information as the target text information; Generate a cover for the target video based on the target key frame image and the target text information.
2. The method according to claim 1, wherein The obtaining of the set of key frames of the target video includes: Perform key frame information detection on the target video to obtain the key frame images of the target video; Determine the set of key frames corresponding to the target video according to the key frame images.
3. The method according to claim 2, characterized in that, The performing of key frame information detection on the target video to obtain the key frame images of the target video includes at least one of the following: Perform face detection on the target video, and select the frame images containing faces from the target video as key frame images; Perform clarity detection on the target video, sort the frame images of the target video in descending order of clarity, and select the preset number of frame images ranked at the front from the target video as key frame images; Perform highlight segment detection on the target video, and select the frame images containing highlight segments from the target video as key frame images.
4. The method according to claim 1, wherein The performing of text recognition on each key frame image in the set of key frames to obtain a text set corresponding to the target video includes: Perform text recognition on each key frame image in the set of key frames to obtain the recognized text corresponding to the key frame image; Perform denoising processing on the recognized text corresponding to each key frame image; Perform semantic recognition on the denoised recognized text, remove the text content with similar semantics, and obtain text information; Determine the text set corresponding to the target video according to the text information.
5. The method according to claim 1, wherein Generating the cover of the target video based on the target key frame image and the target text information includes: Obtain the depth information of the target key frame image; Determine the text configuration information of the target key frame image according to the depth information; Based on the text configuration information, set the target text information in the target key frame image to obtain the cover of the target video.
6. The method according to claim 1, wherein Before obtaining the set of key frames of the target video, it further includes: Obtain a set of covers of the sample video, where the set of covers includes at least one cover, and each cover has a unique cover identifier; Obtain the click-through rate corresponding to each cover in the set of covers; Sort the covers in the set of covers in descending order according to the click-through rate, and select the cover identifier of the first cover as the cover label of the set of covers; Train the video cover generation model according to the cover labels of the cover set and the covers in the cover set.
7. The method according to claim 6, characterized in that, The obtaining of the cover set of the sample video includes: Obtain a sample video; Perform key frame detection on the sample video to obtain a sample key frame set, where the sample key frame set includes at least one sample key frame image; Perform text recognition on each sample key frame image in the sample key frame set to obtain a sample text set, where the sample text set includes at least one sample text information; Combine each sample key frame image in the sample key frame set with each sample text information in the sample text set to generate at least one cover; Determine the cover set of the sample video according to the generated covers.
8. The method according to claim 7, characterized in that, The combining each sample key frame image in the sample key frame set with each sample text information in the sample text set to generate at least one cover includes: Perform depth information recognition on the sample key frame images in the sample key frame set to obtain the depth information of each sample key frame image; Determine the text configuration information of each sample key frame image according to the depth information; Set the text information in the corresponding sample key frame image based on the text configuration information to obtain at least one cover.
9. The method according to claim 6, wherein The obtaining of the click-through rate corresponding to each cover in the cover set includes: Combine each cover in the cover set with the corresponding sample video to obtain at least one piece of video data; Publish the video data and obtain the number of clicks of each piece of video data; Determine the click-through rate corresponding to each cover in the cover set according to the number of clicks of the video data and the covers included in each piece of video data.
10. The method according to claim 6, characterized in that, The training of the video cover generation model according to the cover labels of the cover set and the covers in the cover set includes: Extract the cover images and cover texts of each cover in the cover set; Input the cover image and the cover text corresponding to the same cover identifier into the initial video cover generation model, extract the image feature vector of the cover image and the text feature vector of the cover text, combine the image feature vector and the text feature vector to form a graphic-text feature vector, and output the click-through rate prediction result of the graphic-text feature vector, where the prediction result carries the cover identifier of the cover corresponding to the graphic-text feature vector; Compare the cover identifier carried by the prediction result with the cover label to calculate a loss value; Adjust the parameters of the initial video cover generation model according to the loss value, and train the adjusted initial video cover generation model based on the cover labels of the cover set and the covers in the cover set until the preset training stop condition is met, then stop adjusting the parameters of the initial video cover generation model to obtain the video cover generation model.
11. A video cover generation device, characterized in that, Includes: A key frame set acquisition unit configured to acquire a key frame set of a target video, where the key frame set includes at least one key frame image; A text set acquisition unit is configured to perform text recognition on each key frame image in the key frame set to obtain a text set corresponding to the target video, wherein the text set includes at least one piece of text information; an identification information acquisition unit, configured to input the key frame images in the key frame set and the text information in the text set into a video cover generation model, perform click rate prediction on each cover formed by combining each key frame image with each text information, and output identification information of the cover with the highest predicted click rate, wherein the identification information of the cover represents input sequence information of the key frame images and text information forming the cover; a cover information determination unit configured to use the key frame image corresponding to the identification information as a target key frame image and the text information corresponding to the identification information as a target text information; A cover generation unit is configured to generate a cover of the target video according to the target key frame image and the target text information.
12. The device according to claim 11, characterized in that, The key frame set acquisition unit comprises: A key frame image acquisition module is configured to perform key frame information detection on the target video and acquire a key frame image of the target video; The key frame set determination module is configured to determine the key frame set corresponding to the target video according to the key frame image.
13. The device according to claim 12, characterized in that, The key frame image acquisition module is configured to perform at least one of the following: Performing face detection on the target video, and selecting frame images containing faces from the target video as key frame images; Performing definition detection on the target video, sorting the definition of frame images of the target video in descending order, and selecting a preset number of frame images ranked first in the target video as key frame images; Performing highlight segment detection on the target video, and selecting frame images containing highlight segments from the target video as key frame images.
14. The device according to claim 11, characterized in that, The text collection acquisition unit comprises: A text recognition module is configured to perform text recognition on each of the key frame images in the key frame set to obtain a recognition text corresponding to the key frame image; The processing module is configured to perform denoising on the recognized text corresponding to each of the key frame images; perform semantic recognition on the recognized text after denoising, remove text contents with similar semantics, and obtain text information; The text set determination module is configured to determine the text set corresponding to the target video according to the text information.
15. The device according to claim 11, characterized in that The cover generation unit comprises: A depth information acquisition module, configured to acquire depth information of the target key frame image; A text configuration information acquisition module, configured to determine the text configuration information of the target key frame image according to the depth information; The first cover generation module is configured to set the target text information in the target key frame image based on the text configuration information to obtain the cover of the target video.
16. The device according to claim 11, wherein The device also includes: A cover set acquisition unit is configured to acquire a cover set of a sample video, wherein the cover set includes at least one cover, and each cover has a unique cover identifier; A click rate acquisition unit, configured to acquire a click rate corresponding to each cover in the cover set; A cover label determination unit, configured to sort each cover in the cover set in descending order according to the click-through rate, and select the cover identifier of the first cover as the cover label of the cover set; A model training unit, configured to train the video cover generation model according to the cover label of the cover set and the covers in the cover set.
17. The device according to claim 16, wherein The cover set acquisition unit includes: A sample video acquisition module, configured to acquire a sample video; A sample key frame image acquisition module, configured to perform key frame detection on the sample video to obtain a sample key frame set, where the sample key frame set includes at least one sample key frame image; A sample text information acquisition module, configured to perform text recognition on each sample key frame image in the sample key frame set to obtain a sample text set, where the sample text set includes at least one sample text information; A second cover generation module, configured to combine each sample key frame image in the sample key frame set with each sample text information in the sample text set to generate at least one cover; A cover set determination module, configured to determine the cover set of the sample video according to the generated covers.
18. The device according to claim 17, characterized in that, The second cover generation module includes: A depth information acquisition sub-module, configured to perform depth information recognition on the sample key frame images in the sample key frame set to obtain the depth information of each sample key frame image; A text configuration information acquisition sub-module, configured to determine the text configuration information of each sample key frame image according to the depth information; A cover generation sub-module, configured to set the text information in the corresponding sample key frame image based on the text configuration information to obtain at least one cover.
19. The device according to claim 16, characterized in that, The click-through rate acquisition unit includes: A video data acquisition module, configured to combine each cover in the cover set with the corresponding sample video to obtain at least one piece of video data; A click count acquisition module, configured to publish the video data and acquire the click count of each piece of video data; A click-through rate determination module, configured to determine the click-through rate corresponding to each cover in the cover set according to the click count of the video data and the covers included in each piece of video data.
20. The device according to claim 16, characterized in that, The model training unit is configured to: Extract the cover image and cover text of each cover in the cover set; Input the cover image and the cover text corresponding to the same cover identifier into the initial video cover generation model, extract the image feature vector of the cover image and the text feature vector of the cover text, combine the image feature vector and the text feature vector to form a graphic-text feature vector, and output the click-through rate prediction result of the graphic-text feature vector, where the prediction result carries the cover identifier of the cover corresponding to the graphic-text feature vector; Compare the cover identifier carried in the prediction result with the cover label, and calculate to obtain a loss value; Adjust the parameters of the initial video cover generation model according to the loss value, and train the adjusted initial video cover generation model based on the cover labels of the cover set and the covers in the cover set until the preset training stop condition is satisfied, then stop adjusting the parameters of the initial video cover generation model to obtain the video cover generation model.
21. An electronic device, characterized in that, Comprising: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to execute the instructions to implement the video cover generation method according to any one of claims 1 to 10.
22. A storage medium, when the instructions in the storage medium are executed by a processor of an electronic device, enabling the electronic device to execute the video cover generation method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Method and device for generating video cover, electronic equipment and computer readable storage medium
CN109996091A
Preview cover generation method and device, electronic equipment and storage medium
CN111880888A