Video cover generation method and device, computing equipment and storage medium

By extracting the keyframes of the video and performing multi-dimensional character feature extraction and rating screening, combined with visual optimization technology, the problem of mismatch and inefficiency of video cover generation in the existing technology is solved, and efficient and automated video cover generation is achieved, which significantly improves the adaptability and visual attractiveness of the video cover.

CN120050490APending Publication Date: 2025-05-27SHANGHAI BILIBILI TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510116230.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

In the prior art, the video cover generation method has problems such as strong randomness, resulting in mismatch between the cover and the video content, or the need for manual participation, resulting in low processing efficiency and strong subjectivity.

Method used

By obtaining the pending video, extracting multiple keyframes, and performing multi-dimensional character features extraction for each keyframe, calculating scores, filtering candidate frames, and finally visually optimizing the frame image of the candidate frame to generate a video cover.

Benefits of technology

It realizes the automated intelligent generation of video covers, improves the efficiency of video cover generation, significantly improves the adaptability and relevance between the video cover and the core content of the video, and improves the expressiveness and visual appeal of the video cover, increasing the video click rate and playback volume.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120050490A_ABST
    Figure CN120050490A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a video cover generation method and device, computing equipment and a storage medium, and the method comprises the steps: obtaining a to-be-processed video, and extracting a plurality of key frames from the to-be-processed video; for each key frame, performing multi-dimensional character feature extraction on the key frame to obtain feature information of a plurality of character feature dimensions corresponding to the key frame, and calculating a score of the key frame according to the feature information of the plurality of character feature dimensions corresponding to the key frame; screening out candidate frames from the plurality of key frames according to the scores; and performing visual optimization on the frame image of the candidate frame to generate a video cover. Through multi-dimensional character feature fusion, automatic intelligent generation of the video cover is realized, the video cover generation efficiency is effectively improved, the adaptation degree and correlation between the video cover and the video core content are remarkably improved, and the expressive force and visual attraction of the video cover are also effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of Internet technologies, and more particularly, to a method, an apparatus, a computing device, and a storage medium for generating a video cover. Background Art

[0002] With the continuous development of Internet technologies, more and more users like to create videos and upload the created videos to video platforms for sharing. Setting a suitable video cover for a video can not only intuitively reflect the video content, but also improve the viewing willingness of viewers and attract them to click on the video cover to watch the video. In the prior art, usually the following two methods are adopted to determine the video cover: one is that the video platform randomly selects a video frame image from the video as the video cover. This method has strong randomness, which may result in the video cover not matching the video content well and having low attractiveness; the other is that the video uploader manually selects a video frame image as the video cover or uploads the video cover made by him / her. This method involves a lot of manual participation, has low processing efficiency, and is highly subjective. Summary of the Invention

[0003] In view of the above problems, the present application is proposed to provide a method, an apparatus, a computing device, and a storage medium for generating a video cover that can overcome or at least partially solve the above problems.

[0004] According to one aspect of the embodiments of the present application, there is provided a method for generating a video cover, including:

[0005] Obtaining a video to be processed, and extracting a plurality of key frames from the video to be processed;

[0006] For each key frame, performing multi-dimensional human feature extraction on the key frame to obtain feature information of multiple human feature dimensions corresponding to the key frame, and calculating a score of the key frame according to the feature information of multiple human feature dimensions corresponding to the key frame;

[0007] Selecting candidate frames from the plurality of key frames according to the scores;

[0008] Performing visual optimization on the frame images of the candidate frames to generate a video cover.

[0009] Further, for each key frame, performing multi-dimensional human feature extraction on the key frame to obtain feature information of multiple human feature dimensions corresponding to the key frame further includes:

[0010] Using deep learning models corresponding to multiple human feature dimensions to identify feature information of multiple human feature dimensions in the frame image of the key frame.

[0011] Further, the multiple human feature dimensions include: expression feature dimension, gesture feature dimension, posture feature dimension, eye contact feature dimension, body symmetry feature dimension, and movement amplitude feature dimension.

[0012] Further, calculating the score of the key frame according to the feature information of the multiple human feature dimensions corresponding to the key frame further includes:

[0013] Dynamically adjust the feature weights of the multiple human feature dimensions according to the video scene type of the video to be processed;

[0014] Perform a weighted operation based on the feature information of the multiple human feature dimensions corresponding to the key frame and the feature weights of the multiple human feature dimensions to obtain the score of the key frame.

[0015] Further, screening out candidate frames from the multiple key frames according to the score further includes:

[0016] Sort the multiple key frames in descending order of the score, and select the key frames with higher rankings from the sorting results;

[0017] Screen out candidate frames from the key frames with higher rankings according to the preset screening criteria.

[0018] Further, visually optimizing the frame image of the candidate frame to generate a video cover further includes:

[0019] Perform image enhancement, composition optimization, and / or add additional information to the frame image of the candidate frame to generate a video cover.

[0020] According to another aspect of the embodiments of the present application, a video cover generation device is provided, including:

[0021] A key frame extraction module, adapted to obtain the video to be processed and extract multiple key frames from the video to be processed;

[0022] A feature extraction and scoring module, adapted to perform multi-dimensional human feature extraction on each key frame to obtain the feature information of the multiple human feature dimensions corresponding to the key frame, and calculate the score of the key frame according to the feature information of the multiple human feature dimensions corresponding to the key frame;

[0023] A screening module, adapted to screen out candidate frames from the multiple key frames according to the score;

[0024] A cover generation module, adapted to visually optimize the frame image of the candidate frame to generate a video cover.

[0025] According to another aspect of the embodiments of the present application, a computing device is provided, including: a processor, a memory, a communication interface, and a communication bus. The processor, the memory, and the communication interface complete communication with each other through the communication bus;

[0026] The memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform the operations corresponding to the above video cover generation method.

[0027] According to still another aspect of the embodiments of the present application, a computer storage medium is provided, in which at least one executable instruction is stored, and the executable instruction causes the processor to perform the operations corresponding to the above video cover generation method.

[0028] According to yet another aspect of the embodiments of the present application, a computer program product is provided, including at least one executable instruction, and the executable instruction causes the processor to perform the operations corresponding to the above video cover generation method.

[0029] According to the technical solution provided by the embodiments of the present application, the human features in the key frames of the video are extracted from multiple human feature dimensions. The feature information of the multiple human feature dimensions extracted can comprehensively and synthetically reflect the dynamic behaviors and emotional expressions of the people in the key frames. By comprehensively analyzing the feature information of the multiple human feature dimensions and combining with a scoring mechanism, the candidate frames for video cover generation can be accurately screened out; the frame images of the candidate frames are visually optimized, and a video cover that not only shows the core content of the video but also has visual attractiveness can be automatically and quickly generated. It not only realizes the automatic and intelligent generation of the video cover, effectively improves the video cover generation efficiency, significantly enhances the adaptability and relevance between the video cover and the core content of the video, but also effectively improves the expressiveness and visual attractiveness of the video cover, which helps to increase the video click-through rate and play volume.

[0030] The above description is only an overview of the technical solutions of the embodiments of the present application. In order to be able to understand the technical means of the embodiments of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features, and advantages of the embodiments of the present application more obvious and understandable, the following specifically illustrates the specific implementation manners of the embodiments of the present application. Description of the Drawings

[0031] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the embodiments of the present application. And throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0032] Figure 1 A flowchart showing the video cover generation method according to an embodiment of the present application is shown;

[0033] Figure 2 shows a schematic flowchart of a video cover generation method according to another embodiment of the present application;

[0034] Figure 3 shows a structural block diagram of a video cover generation apparatus according to an embodiment of the present application;

[0035] Figure 4 shows a schematic structural diagram of a computing device according to an embodiment of the present application. Detailed implementation manners

[0036] Hereinafter, exemplary embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.

[0037] First, the noun terms involved in one or more embodiments of the present application are explained.

[0038] Video cover: refers to a static image displayed before a video file is played, usually used to attract the user's attention and provide preliminary visual cues about the video content.

[0039] Feature weight: refers to the weight value when evaluating the importance of feature information in multiple human feature dimensions, and is used to dynamically adjust the cover recommendation strategy.

[0040] FFmpeg: is an open-source computer program that can be used to record, convert digital audio and video, and convert them into streams. It can decode, encode, transcode, multiplex, demultiplex, stream media transmission and play various audio and video formats. It contains a rich set of command-line tools and libraries, and is widely used in video processing tasks.

[0041] Multimodal analysis: refers to an analysis method that combines various modal information such as human feature information and video scene types in a video for comprehensive evaluation.

[0042] Figure 1 shows a schematic flowchart of a video cover generation method according to an embodiment of the present application, as Figure 1 shown, the method includes the following steps:

[0043] Step S101, obtain the video to be processed, and extract multiple key frames from the video to be processed.

[0044] Among them, the video to be processed refers to the video file for which a video cover needs to be generated. The video to be processed can be obtained from the user or retrieved from a video platform. Video encoders typically use inter-frame coding techniques to reduce the data volume of the video file, storing only the change information relative to the previous or next frame. This coding method can significantly reduce the size of the video file without significantly affecting the visual quality. That is to say, not every frame in the video contains complete image information. After obtaining the video to be processed, multiple key frames of the video to be processed need to be extracted. The key frames contain complete image information for subsequent processing such as screening and visual optimization based on the key frames to generate the video cover.

[0045] Step S102: For each key frame, extract multi-dimensional human feature information from the key frame to obtain the feature information of multiple human feature dimensions corresponding to the key frame, and calculate the score of the key frame based on the feature information of multiple human feature dimensions corresponding to the key frame.

[0046] Identify the frame image of each key frame and extract features from multiple human feature dimensions to obtain the feature information of multiple human feature dimensions. Among them, the multiple human feature dimensions can include: expression feature dimension, gesture feature dimension, posture feature dimension, eye contact feature dimension, body symmetry feature dimension, and movement amplitude feature dimension, etc. Those skilled in the art can also set other human feature dimensions, such as the facial orientation feature dimension, etc., which are not limited here.

[0047] The extracted feature information of multiple human feature dimensions can comprehensively and synthetically reflect the dynamic behavior and emotional expression of the people in the key frame. Quantify and score based on the feature information of multiple human feature dimensions corresponding to the key frame to obtain the score of the key frame.

[0048] Step S103: Filter out candidate frames from multiple key frames based on the scores.

[0049] The high or low score of the key frame can conveniently reflect the adaptability and relevance between the content of the key frame image and the core content of the video to be processed. After calculating the scores of multiple key frames, the key frames with higher scores can be selected from multiple key frames as candidate frames for the final generation of the video cover.

[0050] Step S104: Perform visual optimization on the frame image of the candidate frame to generate the video cover.

[0051] In the embodiments of the present application, in order to enhance the expressiveness and visual appeal of the video cover, the frame image of the candidate frame is used as the basic image of the video cover, and visual optimization processing such as image enhancement, composition optimization, and / or adding additional information is performed on the frame image of the candidate frame, and the finally processed frame image is used as the generated video cover.

[0052] According to the video cover generation method provided by the embodiments of the present application, the human features in the key frames of the video are extracted from multiple human feature dimensions, and the feature information of the multiple human feature dimensions extracted can comprehensively and synthetically reflect the dynamic behaviors and emotional expressions of the people in the key frames. By comprehensively analyzing the feature information of the multiple human feature dimensions and combining with a scoring mechanism, the candidate frames for video cover generation can be accurately screened out; performing visual optimization on the frame images of the candidate frames can automatically and quickly generate a video cover that not only displays the core content of the video but also has visual appeal, which not only realizes the automatic and intelligent generation of the video cover, effectively improves the video cover generation efficiency, significantly enhances the adaptability and relevance between the video cover and the core content of the video, but also effectively enhances the expressiveness and visual appeal of the video cover, and helps to increase the video click-through rate and play volume.

[0053] Figure 2 The flowchart of the video cover generation method according to another embodiment of the present application is shown, as Figure 2 shown, and the method includes the following steps:

[0054] Step S201, obtain the video to be processed, and extract multiple key frames from the video to be processed.

[0055] The embodiments of the present application can be applied before the user uploads a video to a video platform. The video to be processed provided by the user is obtained, and the video cover is generated for the video to be processed by using the solution provided by the embodiments of the present application, and then the user uploads the video with the video cover to the video platform; the embodiments of the present application can also be applied in the background of the video platform. The video to be processed is obtained from the video platform, and the video cover is generated for the video to be processed by using the solution provided by the embodiments of the present application, and then the video with the video cover is published and displayed on the video platform.

[0056] After obtaining the video to be processed, the video to be processed can be parsed into a frame sequence, and at the same time, the meta-information of the video to be processed is extracted, where the meta-information includes information such as the frame rate and duration of the video to be processed; multiple key frames are extracted from the frame sequence. Specifically, tools such as FFmpeg can be used to split the video to be processed into frames, and frames at specified time intervals are extracted to avoid excessive duplicate content in the extracted frames, forming a frame sequence; then the frames in the frame sequence are screened to remove frames with blurred images and incomplete image information, so as to obtain multiple key frames, realizing the automatic extraction of key frames from the video to be processed.

[0057] Step S202: For each key frame, use the deep learning models corresponding to multiple human feature dimensions to identify the feature information of multiple human feature dimensions in the frame image of the key frame.

[0058] Considering that deep learning technology can be well applied in the fields of image, video and other processing, in the embodiments of the present application, deep learning technology is introduced into the field of video cover generation. The deep learning model can be used to extract multi-dimensional human features from the frame image of each key frame to obtain the feature information of multiple human feature dimensions. Among them, multiple human feature dimensions may include: expression feature dimension, gesture feature dimension, posture feature dimension, eye contact feature dimension, body symmetry feature dimension, movement amplitude feature dimension, etc.

[0059] In specific applications, different deep learning models can be set for different human feature dimensions respectively to extract the feature information of the corresponding human feature dimensions in the frame image of the key frame; or, through a series of model trainings on the training set, a deep learning model applicable to multi-dimensional human feature extraction can be obtained. Then, only need to input the frame image of the key frame into the deep learning model, and the feature information of the corresponding multiple human feature dimensions in the frame image of the key frame can be extracted at one time. Taking the example that different human feature dimensions correspond to different deep learning models, the multi-dimensional human feature extraction process is introduced below.

[0060] (1) For the expression feature dimension, the pre-trained ResNet model or MobileNet model can be used as the deep learning model corresponding to this dimension. Based on the above deep learning model, facial expression classification is performed on the frame image of the key frame to identify common emotions and their intensity scores, so as to obtain the feature information of the expression feature dimension. Among them, common emotions may include smile, happy, surprised, angry, crying, etc. Specifically, facial key points are used to analyze the changes in facial features, and by matching with predefined expression categories, the feature information of the expression feature dimension is obtained. In specific implementation, the ResNet50 model and the Softmax classifier can be selected to construct an expression classification model as the deep learning model corresponding to the expression feature dimension. The frame image of the key frame is input into the expression classification model, and the expression classification model identifies the expression in the frame image and outputs the expression category (such as smile, surprise, etc.) to which the expression in the frame image belongs and its confidence score.

[0061] (2) For the gesture feature dimension, the OpenPose model or the MediaPipe model can be used as the deep learning model corresponding to this dimension. Based on the above deep learning model, the frame images of key frames are recognized to locate the hand key points and analyze the integrity of the gesture. For example, the integrity of the gesture is determined according to the degree of gesture occlusion; according to the positions of the hand key points, combined with the semantic classification model, the semantic meaning of the gesture is recognized, such as gestures like thumbs up, pointing, and clapping, to obtain the feature information of the gesture feature dimension. In specific implementation, the OpenPose model and the CNN classifier can be selected to construct a gesture detection model as the deep learning model corresponding to the gesture feature dimension. Only the hand region image in the frame image of the key frame can be input into the gesture detection model, and the gesture detection model recognizes the gesture in the hand region image and outputs the gesture type to which the gesture belongs and the positions of its key points.

[0062] (3) For the pose feature dimension, the PoseNet model or the HRNet model can be used as the deep learning model corresponding to this dimension. Based on the above deep learning model, the frame images of key frames are recognized to extract the human key point coordinates. According to the human key points, the limb symmetry and dynamic sense are calculated to determine the action amplitude and action aesthetics of the person in the picture, etc., to obtain the feature information of the pose feature dimension.

[0063] (4) For the eye contact feature dimension, the pre-trained ResNet model or the MobileNet model can be used as the deep learning model corresponding to this dimension. Based on the above deep learning model, facial key point detection is performed on the frame images of key frames. The eye positions of the person are recognized through the facial key points, and the degree to which the person looks directly at the camera is determined to obtain the feature information of the eye contact feature dimension. Introducing the feature information of the eye contact feature dimension into the scoring calculation of the key frame helps to enhance the sense of interaction between the user and the picture.

[0064] (5) For the body symmetry feature dimension, the PoseNet model or the HRNet model can be used as the deep learning model corresponding to this dimension. Based on the above deep learning model, the frame images of key frames are recognized to extract the human key point coordinates. According to the human key point coordinates, the left-right symmetry degree of the body movement is calculated to obtain the feature information of the body symmetry feature dimension. Introducing the feature information of the body symmetry feature dimension into the scoring calculation of the key frame helps to improve the aesthetic score of the picture.

[0065] (6) For the movement amplitude feature dimension, the deep learning model corresponding to this dimension can be used to compare the changes in the positions of key points in consecutive frames, calculate the limb movement range using the inter-frame displacement, and analyze the action tension and dynamics to obtain the feature information of the movement amplitude feature dimension. Introducing the feature information of the movement amplitude feature dimension into the scoring calculation of the key frame helps to highlight the dynamic sense of the picture.

[0066] Step S203: Dynamically adjust the feature weights of multiple human feature dimensions according to the video scene type of the video to be processed.

[0067] Considering that videos of different video scene types have different requirements for generating video covers. For example, the video cover of an educational video should highlight the indicative actions of the characters (such as pointing to the blackboard, etc.), the entertainment video should highlight the sense of dynamics, and the sports video should highlight the explosive power and aesthetic feeling of the characters' actions. In the embodiment of the present application, according to the video scene type of the video to be processed, the feature weights of multiple human feature dimensions are dynamically adjusted, and multimodal analysis is fully combined with the human feature information and the video scene type, and then more suitable candidate frames are screened out from the key frames for generating the video cover, realizing scene adaptation optimization, adopting a targeted cover recommendation strategy for different video scene types, and effectively improving the adaptability, accuracy and diversity of the video cover. Among them, a higher feature weight can be set for positive emotions (such as smiling, surprise, etc.) in the expression feature dimension, so as to preferentially recommend pictures that convey positive energy for generating the video cover.

[0068] For example, when the video scene type of the video to be processed is an educational video, the feature weights corresponding to the expression feature dimension, gesture feature dimension and eye contact feature dimension can be increased, and the feature weight of the movement amplitude feature dimension can be decreased. Specifically, natural and positive emotions (such as smiling) can have a higher feature weight, indicative gestures can have a higher feature weight, and eye contact looking directly at the camera can have a higher feature weight, so that when generating the video cover, frames with natural expressions and clear indicative gestures of the characters are preferentially selected.

[0069] When the video scene type of the video to be processed is an entertainment video, the feature weights corresponding to the expression feature dimension, movement amplitude feature dimension and posture feature dimension can be increased. Specifically, exaggerated expressions can have a higher feature weight, and postures with a strong sense of dynamics can have a higher feature weight, so that when generating the video cover, frames with exaggerated expressions and strong dynamics of the characters are preferentially selected.

[0070] When the video scene type of the video to be processed is a sports video or a dance video, the feature weights corresponding to the posture feature dimension, movement amplitude feature dimension and body symmetry feature dimension can be increased. Specifically, large-amplitude movements can have a higher feature weight, so as to capture the moments of key actions (such as dunking, jumping, etc.), so that when generating the video cover, frames that highlight the explosive power and aesthetic feeling of the actions are preferentially selected.

[0071] Step S204: Perform weighted operations based on the feature information of multiple human feature dimensions corresponding to the key frame and the feature weights of multiple human feature dimensions to obtain the score of the key frame.

[0072] For each key frame, multi-dimensional scoring is performed based on the feature information of multiple human feature dimensions corresponding to the extracted key frame. Specifically, the feature information of multiple human feature dimensions can be quantified, for example, quantified as dimension scores, and then the score of the key frame is calculated using a scoring formula. The scoring formula can be:

[0073] S = ω 1 .S 表情 + ω 2 .S 手势 + ω 3 .S 姿势 + ω 4 .S 眼神 + ω 5 .S 对称性 + ω 6 .S 运动幅度

[0074] where S represents the score of the key frame; S 表情 , S 手势 , S 姿势 , S 眼神 , S 对称性 and S 运动幅度 represent the dimension scores of the expression feature dimension, gesture feature dimension, posture feature dimension, eye contact feature dimension, body symmetry feature dimension, and movement amplitude feature dimension respectively; ω 1 , ω 2 , ω 3 , ω 4 , ω 5 and ω 6 represent the feature weights of the expression feature dimension, gesture feature dimension, posture feature dimension, eye contact feature dimension, body symmetry feature dimension, and movement amplitude feature dimension respectively.

[0075] Step S205, select candidate frames from multiple key frames according to the scores.

[0076] After calculating the scores of multiple key frames, sort the multiple key frames in descending order of scores, and select the key frames with higher rankings from the sorting results; according to the preset screening criteria, select the more suitable key frames from the key frames with higher rankings as candidate frames for use in the final video cover generation.

[0077] Specifically, a preset number of key frames with higher rankings can be selected from the sorting results, and those skilled in the art can set the preset number according to actual needs. When the preset number is multiple, secondary screening can be further performed based on preset screening criteria. The preset screening criteria can specifically be an expression clarity criterion, an action integrity criterion, etc., so as to preferentially select key frames with clear expressions and complete actions in the picture as candidate frames. The number of candidate frames can be one or more, and no specific limitation is made here.

[0078] Step S206: Perform image enhancement, composition optimization, and / or add additional information to the frame images of the candidate frames to generate a video cover.

[0079] The content of the frame images of the candidate frames screened through step S205 has a high degree of fitness and relevance to the core content of the video to be processed. In order to further improve the expressiveness and visual attractiveness of the video cover, in the embodiments of the present application, visual optimization processing such as image enhancement, composition optimization, and / or adding additional information is also performed on the frame images of the candidate frames, and the finally processed frame images are used as the video cover.

[0080] Specifically, image enhancement may include: increasing the brightness and contrast of the frame images of the candidate frames to make the picture more attractive, and using filters, etc. to enhance the color saturation and overall aesthetic feeling of the picture. Composition optimization may include: according to the position of the human body main body in the frame images of the candidate frames, cropping the frame images to highlight the human body main body, and using the golden ratio or the rule of thirds, etc. to optimize the picture composition. Adding additional information may include: adding text labels (such as titles, timestamps, etc.), adding additional information such as the logo of the video platform and the logo of the video creator to enhance the recognition degree.

[0081] Among them, in the process of generating the video cover, one or more of the visual optimization processing methods of image enhancement, composition optimization, and adding additional information can be selected according to actual needs to perform visual optimization processing on the frame images of the candidate frames, and no limitation is made here.

[0082] According to the video cover generation method provided by the embodiments of the present application, by comprehensively analyzing multi-dimensional human features such as expressions, gestures, postures, and eye contacts in frame images, through the fusion of multi-dimensional human features, combined with the video scene type and dynamic scoring mechanism, for different video scene types, the feature weights of the feature dimensions of human features are dynamically adjusted, achieving scene adaptation optimization, and being able to more accurately screen out candidate frames that both fully reflect the core content of the video and conform to the video scene type; this solution conveniently realizes the automatic intelligent generation of video covers, effectively improves the video cover generation efficiency, significantly enhances the adaptability and relevance of the video cover to the core content of the video, and further improves the expressiveness and visual attractiveness of the video cover through visual optimization of the frame images of the candidate frames, optimizing the video cover generation effect, helping to increase the video click-through rate and play volume, and having good application value and technical application prospects.

[0083] Figure 3 The structural block diagram of a video cover generation device according to an embodiment of the present application is shown, as Figure 3 shown, the device includes: a key frame extraction module 310, a feature extraction and scoring module 320, a screening module 330, and a cover generation module 340.

[0084] The key frame extraction module 310 is adapted to: obtain a video to be processed and extract multiple key frames from the video to be processed.

[0085] The feature extraction and scoring module 320 is adapted to: for each key frame, perform multi-dimensional human feature extraction on the key frame to obtain the feature information of multiple human feature dimensions corresponding to the key frame, and calculate the score of the key frame according to the feature information of multiple human feature dimensions corresponding to the key frame.

[0086] The screening module 330 is adapted to: screen out candidate frames from multiple key frames according to the scores.

[0087] The cover generation module 340 is adapted to: perform visual optimization on the frame images of the candidate frames to generate a video cover.

[0088] Optionally, the feature extraction and scoring module 320 is further adapted to: use deep learning models corresponding to multiple human feature dimensions to identify the feature information of multiple human feature dimensions in the frame image of the key frame.

[0089] Optionally, multiple human feature dimensions include: an expression feature dimension, a gesture feature dimension, a posture feature dimension, an eye contact feature dimension, a body symmetry feature dimension, and a movement amplitude feature dimension.

[0090] Optionally, the feature extraction and scoring module 320 is further adapted to: dynamically adjust the feature weights of multiple human feature dimensions according to the video scene type of the video to be processed; perform a weighted operation based on the feature information of multiple human feature dimensions corresponding to the key frame and the feature weights of multiple human feature dimensions to obtain the score of the key frame.

[0091] Optionally, the screening module 330 is further adapted to: sort multiple key frames in descending order of score, and select the key frames with higher rankings from the sorting results; screen out candidate frames from the key frames with higher rankings according to a preset screening index.

[0092] Optionally, the cover generation module 340 is further adapted to: perform image enhancement, composition optimization, and / or add additional information to the frame image of the candidate frame to generate a video cover.

[0093] The descriptions of the above modules refer to the corresponding descriptions in the method embodiments and will not be elaborated here.

[0094] According to the video cover generation device provided by the embodiments of the present application, by comprehensively analyzing multi-dimensional human features such as expressions, gestures, postures, and eye contacts in the frame image, through multi-dimensional human feature fusion, combined with the video scene type and a dynamic scoring mechanism, for different video scene types, the feature weights of human feature dimensions are dynamically adjusted, realizing scene adaptation optimization, and being able to more accurately screen out candidate frames that fully reflect the core content of the video and conform to the video scene type; this solution conveniently realizes the automated intelligent generation of video covers, effectively improves the video cover generation efficiency, significantly enhances the adaptability and relevance between the video cover and the core content of the video, and further improves the expressiveness and visual attraction of the video cover by visually optimizing the frame image of the candidate frame, optimizing the video cover generation effect, helping to increase the video click-through rate and play volume, and having good application value and technical application prospects.

[0095] The embodiments of the present application provide a non-volatile computer storage medium, and the computer storage medium stores at least one executable instruction or computer program, and the executable instruction or computer program can enable a processor to perform the operations corresponding to the video cover generation method in any of the above method embodiments.

[0096] The embodiments of the present application provide a computer program product, and the computer program product includes at least one executable instruction or computer program, and the executable instruction or computer program can enable a processor to perform the operations corresponding to the video cover generation method in any of the above method embodiments.

[0097] Figure 4 The structural schematic diagram of a computing device according to an embodiment of the present application is shown, and the specific implementation of the computing device is not limited in the specific embodiments of the present application.

[0098] As Figure 4 shown, the computing device may include: a processor 402, a communications interface 404, a memory 406, and a communication bus 408.

[0099] Among them: The processor 402, the communications interface 404, and the memory 406 communicate with each other through the communication bus 408. The communications interface 404 is used to communicate with network elements of other devices such as clients or other servers. The processor 402 is used to execute the program 410, and specifically can execute the relevant steps in the above-mentioned method embodiments for generating video covers of the computing device.

[0100] Specifically, the program 410 may include program code, and the program code includes computer operation instructions.

[0101] The processor 402 may be a central processing unit CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present application. One or more processors included in the computing device may be of the same type of processor, such as one or more CPUs; or may be of different types of processors, such as one or more CPUs and one or more ASICs.

[0102] The memory 406 is used to store the program 410. The memory 406 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk memory.

[0103] The program 410 is specifically used to cause the processor 402 to execute the video cover generation method in any of the above method embodiments. For the specific implementation of each step in the program 410, reference may be made to the corresponding steps and descriptions in the corresponding units in the above video cover generation embodiments, which will not be elaborated here. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described devices and modules can refer to the corresponding process descriptions in the foregoing method embodiments, which will not be elaborated here.

[0104] The algorithms and displays provided herein are not inherently related to any particular computer, virtual system, or other device. A variety of general-purpose systems may also be used in conjunction with the teachings presented herein. The structure required to construct such systems will be apparent from the above description. Additionally, embodiments of the present application are not directed to any particular programming language. It should be understood that the content of the embodiments of the present application described herein can be implemented using a variety of programming languages, and the description of a particular language above is for the purpose of disclosing the best mode of the embodiments of the present application.

[0105] In the specification provided herein, numerous specific details are set forth. However, it can be understood that embodiments of the present application may be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.

[0106] Similarly, it should be understood that, in order to streamline the present disclosure and assist in understanding one or more of the various inventive aspects, in the above description of the exemplary embodiments of the present application, the various features of the embodiments of the present application are sometimes grouped together into a single embodiment, figure, or description thereof. However, the disclosed method should not be construed as reflecting an intention that the claimed embodiments of the present application require more features than are expressly recited in each claim. Rather, as reflected in the following claims, the inventive aspects lie in less than all the features of the single foregoing disclosed embodiment. Thus, the claims following the detailed description are hereby expressly incorporated into the detailed description, with each claim standing on its own as a separate embodiment of the present application.

[0107] Those skilled in the art will appreciate that the modules in the devices in the embodiments can be adaptively changed and disposed in one or more devices different from the embodiments. The modules or units or components in the embodiments can be combined into one module or unit or component, and in addition, they can be divided into multiple sub-modules or sub-units or sub-components. Except for the fact that at least some of such features and / or processes or units are mutually exclusive, any combination can be used to combine all the features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all the processes or units of any method or device so disclosed. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) can be replaced by an alternative feature that provides the same, equivalent, or similar purpose.

[0108] In addition, those skilled in the art can understand that although some of the embodiments described herein include certain features included in other embodiments rather than other features, the combination of features of different embodiments means that it is within the scope of the embodiments of the present application and forms different embodiments. For example, in the following claims, any one of the claimed embodiments can be used in any combination.

[0109] Each component embodiment of the embodiments of the present application can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art should understand that a microprocessor or a digital signal processor (DSP) can be used in practice to implement some or all of the functions of some or all of the components according to the embodiments of the present application. The embodiments of the present application can also be implemented as a device or apparatus program (for example, a computer program and a computer program product) for executing part or all of the methods described herein. Such a program for implementing the embodiments of the present application can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.

[0110] It should be noted that the above embodiments illustrate the embodiments of the present application rather than limit the embodiments of the present application, and those skilled in the art can design alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in the claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The embodiments of the present application can be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In a unit claim listing several devices, several of these devices can be embodied by the same item of hardware. The use of the words first, second, and third, etc. does not denote any order. These words can be interpreted as names.

Claims

1. A method for generating a video cover, comprising: Obtain a video to be processed, and extract multiple key frames from the video to be processed; For each key frame, perform multi-dimensional character feature extraction on the key frame to obtain feature information of multiple character feature dimensions corresponding to the key frame, and calculate the score of the key frame according to the feature information of multiple character feature dimensions corresponding to the key frame; According to the scores, candidate frames are selected from multiple key frames; Visually optimize the frame image of the candidate frame to generate a video cover.

2. According to the method of claim 1, for each key frame, extracting multi-dimensional character features of the key frame to obtain feature information of multiple character feature dimensions corresponding to the key frame further comprises: The deep learning model corresponding to the multiple character feature dimensions is used to identify feature information of the multiple character feature dimensions in the frame image of the key frame.

3. According to the method of claim 1 or 2, the multiple character feature dimensions include: Expression feature dimension, gesture feature dimension, posture feature dimension, eye contact feature dimension, body symmetry feature dimension and movement amplitude feature dimension.

4. The method according to any one of claims 1 to 3, wherein calculating the score of the key frame according to the feature information of the multiple character feature dimensions corresponding to the key frame further comprises: Dynamically adjusting feature weights of multiple character feature dimensions according to the video scene type of the video to be processed; A weighted calculation is performed based on the feature information of multiple character feature dimensions corresponding to the key frame and the feature weights of the multiple character feature dimensions to obtain a score for the key frame.

5. According to the method according to any one of claims 1 to 4, the step of selecting candidate frames from a plurality of key frames based on the scores further comprises: Sort multiple key frames in descending order of scores, and select the key frame with the highest ranking from the sorting results; The candidate frames are selected from the top-ranked key frames according to a preset screening index.

6. According to the method according to any one of claims 1 to 5, the visual optimization of the frame image of the candidate frame to generate the video cover further comprises: The frame image of the candidate frame is enhanced, the composition is optimized, and / or additional information is added to generate the video cover.

7. A video cover generation device, comprising: A key frame extraction module, adapted to obtain a video to be processed and extract a plurality of key frames from the video to be processed; The feature extraction and scoring module is adapted to extract multi-dimensional character features from each key frame, obtain feature information of multiple character feature dimensions corresponding to the key frame, and calculate the score of the key frame based on the feature information of multiple character feature dimensions corresponding to the key frame; A screening module, adapted to screen candidate frames from a plurality of key frames according to scores; The cover generation module is suitable for visually optimizing the frame image of the candidate frame to generate a video cover.

8. A computing device comprising: A processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other via the communication bus; The memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform operations corresponding to the video cover generation method as described in any one of claims 1-6.

9. A computer storage medium, wherein at least one executable instruction is stored in the storage medium, and the executable instruction enables a processor to execute operations corresponding to the video cover generation method as described in any one of claims 1-6.

10. A computer program product, comprising at least one executable instruction, wherein the executable instruction enables a processor to execute operations corresponding to the video cover generation method according to any one of claims 1 to 6.

Citation Information

Cited By

  • Video cover automatic generation method and device, terminal and storage medium

    CN121397298A