Video cover generation method, terminal and storage medium

CN116910305BActive Publication Date: 2026-08-14CHINA MOBILEHANGZHOUINFORMATION TECH CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-11
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0003]在相关的现有技术中,通常是基于对视频内容的理解,提取视频的特征,从而选择合适的视频帧集合组成动态封面;然而,这些动态封面的推荐方式基本都是基于黑箱模式,例如直接将视频分割为若干段,然后对视频片段进行分析并从中选取动态封面;该过程的可解释性差,动态封面生成的策略简单,从而难以获得最优动态封面,造成动态封面的生成效果较差的问题

Benefits of technology

[0015]本申请实施例提供了一种视频封面生成方法、终端及存储介质,终端获取视频信息的静态封面信息和标签信息;基于静态封面信息、标签信息以及目标封面排序模型进行静态封面排序处理,获得排序结果;其中,目标封面排序模型是利用预设损失函数对初始封面排序模型进行训练获得的;预设损失函数是基于点击率信息和余弦相似度信息计算的;基于排序结果生成视频信息的动态封面。由此可见,在本申请中,终端通过获取视频信息的静态封面信息和标签信息,进而利用目标封面排序模型和标签信息对静态封面信息进行排序,获得排序结果,从而令排序结果能够反映不同静态封面的重要性;与此同时,本申请基于点击率信息和余弦相似度信息计算预设损失函数,以利用预设损失函数对初始封面排序模型进行训练,获得目标封面排序模型,能够提升训练效果,从而提升目标封面排序模型的排序能力,令排序结果的准确性更高;进一步地,基于该排序结果,能够生成最优的动态封面,提升动态封面的生成效果。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116910305B_ABST
    Figure CN116910305B_ABST
Patent Text Reader

Abstract

This application discloses a video cover generation method, terminal, and storage medium. The terminal acquires static cover information and tag information of the video information; static cover sorting is performed based on the static cover information, tag information, and a target cover sorting model to obtain a sorting result; wherein, the target cover sorting model is obtained by training an initial cover sorting model using a preset loss function; the preset loss function is calculated based on click-through rate information and cosine similarity information; dynamic cover of the video information is generated based on the sorting result. This method can obtain the optimal dynamic cover and improve the generation effect of the dynamic cover.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a method for generating video cover, a terminal, and a storage medium. Background Technology

[0002] With the rapid development of the Internet, short video mobile applications (APPs) have sprung up rapidly. In order to enable users to learn about video content more quickly and thus increase the attractiveness of the video and the number of clicks, it is usually necessary to set an eye-catching "first impression" for the video, that is, the cover. However, static covers do not display enough information and are not very interesting, so dynamic covers are needed.

[0003] In existing technologies, dynamic cover images are typically generated by extracting video features based on an understanding of the video content, and then selecting a suitable set of video frames. However, these dynamic cover image recommendation methods are mostly based on a black box model, such as directly dividing the video into several segments, then analyzing the video segments and selecting dynamic cover images from them. This process has poor interpretability and the dynamic cover image generation strategy is simple, making it difficult to obtain the optimal dynamic cover image, resulting in poor dynamic cover image generation effects. Summary of the Invention

[0004] This application provides a video cover generation method, terminal, and storage medium that can obtain the optimal dynamic cover and improve the generation effect of the dynamic cover.

[0005] The technical solution of this application embodiment is implemented as follows:

[0006] In a first aspect, embodiments of this application provide a video cover generation method, the method comprising:

[0007] Obtain static cover and tag information from video information;

[0008] Static cover sorting is performed based on the static cover information, the tag information, and the target cover sorting model to obtain the sorting result; wherein, the target cover sorting model is obtained by training the initial cover sorting model using a preset loss function; the preset loss function is calculated based on click-through rate information and cosine similarity information;

[0009] A dynamic cover image for the video information is generated based on the sorting results.

[0010] Secondly, embodiments of this application provide a terminal, which includes an acquisition unit and a generation unit.

[0011] The acquisition unit is used to acquire static cover information and tag information of video information; and to perform static cover sorting processing based on the static cover information, the tag information and the target cover sorting model to obtain a sorting result; wherein, the target cover sorting model is obtained by training an initial cover sorting model using a preset loss function; the preset loss function is calculated based on click-through rate information and cosine similarity information;

[0012] The generation unit is used to generate a dynamic cover for the video information based on the sorting result.

[0013] Thirdly, embodiments of this application provide a terminal, the terminal including a processor and a memory storing processor-executable instructions, wherein when the instructions are executed by the processor, the method described in the first aspect is implemented.

[0014] Fourthly, embodiments of this application provide a computer-readable storage medium having a program stored thereon, which is applied in a terminal, and when the program is executed by a processor, it implements the method as described in the first aspect.

[0015] This application provides a video cover generation method, terminal, and storage medium. The terminal acquires static cover information and tag information of video information; performs static cover sorting processing based on the static cover information, tag information, and a target cover sorting model to obtain a sorting result; wherein, the target cover sorting model is obtained by training an initial cover sorting model using a preset loss function; the preset loss function is calculated based on click-through rate information and cosine similarity information; and a dynamic cover of the video information is generated based on the sorting result. Therefore, in this application, the terminal acquires static cover information and tag information of video information, and then uses the target cover sorting model and tag information to sort the static cover information to obtain a sorting result, thereby enabling the sorting result to reflect the importance of different static covers; simultaneously, this application calculates a preset loss function based on click-through rate information and cosine similarity information, and uses the preset loss function to train the initial cover sorting model to obtain the target cover sorting model, which can improve the training effect, thereby improving the sorting ability of the target cover sorting model and making the sorting result more accurate; furthermore, based on the sorting result, an optimal dynamic cover can be generated, improving the generation effect of the dynamic cover. Attached Figure Description

[0016] Figure 1 This is a schematic diagram of the implementation process of the video cover generation method proposed in this application. Figure 1 ;

[0017] Figure 2 This is a schematic diagram of the implementation process of the video cover generation method proposed in this application. Figure 2 ;

[0018] Figure 3 This is a schematic diagram illustrating the implementation of the video cover generation method proposed in this application. Figure 1 ;

[0019] Figure 4 This is a schematic diagram illustrating the implementation of the video cover generation method proposed in this application. Figure 2 ;

[0020] Figure 5 This is a schematic diagram illustrating the implementation of the video cover generation method proposed in this application. Figure 3 ;

[0021] Figure 6 This is a schematic diagram of the terminal structure proposed in the embodiments of this application. Figure 1 ;

[0022] Figure 7 This is a schematic diagram of the terminal structure proposed in the embodiments of this application. Figure 2 . Detailed Implementation

[0023] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only for explaining the relevant application and not for limiting the application. Furthermore, it should be noted that, for ease of description, only the parts related to the relevant application are shown in the accompanying drawings.

[0024] With the rapid development of the Internet, short video apps have sprung up. To enable users to quickly learn about video content and thus increase the appeal of videos and user clicks, it is usually necessary to set an eye-catching "first impression" for the video, namely the cover. However, static covers do not display enough information and are generally not very interesting, so dynamic covers are needed.

[0025] In existing technologies, dynamic cover images are typically created by understanding the video content, extracting video features, and then selecting a suitable set of video frames. Understanding the video and combining video frames involves various techniques. For example, a video can be segmented into multiple video clips, and the quality of these clips can be evaluated using deep learning methods, including I3D video feature extraction and Vggish audio feature extraction, to obtain a quality score. Video clips that sequentially meet preset conditions are then used to generate dynamic images as recommended cover images. This method recommends and generates dynamic cover images based on quality, integrating video and audio features, with high-quality video clips serving as dynamic cover images. Alternatively, a video can be segmented into multiple video clips, and a video feature extraction network can be trained based on online feedback data such as user click-through rates, selecting high-scoring clips as dynamic cover images. Multiple dynamic cover images can be determined by segmenting the video into multiple frame image sets. These preset dynamic cover images are then input into a pre-trained video click-through rate prediction model to obtain estimated click-through rates, thereby determining the target dynamic cover image for the target video. Furthermore, cover image recommendations can be made based on various image features, such as image quality, facial features, and aesthetics.

[0026] However, the aforementioned existing technologies often suffer from the following problems: First, dynamic cover recommendation is basically based on a black box model, resulting in poor interpretability and difficulty in obtaining the optimal dynamic cover. Simple strategies based on user feedback as supervision, such as directly segmenting the video into several segments, analyzing the video segments and selecting dynamic cover segments to generate segmented video segments, or extracting features from video segments using 3D convolution to obtain the final dynamic cover, suffer from poor interpretability. Although there is supervisory information, the dynamic cover generation strategy is simple, has significant limitations, and is difficult to obtain the optimal dynamic cover. Different scenarios require different segmentation strategies and analysis methods, resulting in poor operability and portability. Second, dynamic cover recommendation algorithms suffer from the deficiency of limited supervisory information. Supervisory information often only includes information such as user click-through rates. While user click-through rates can reflect the effectiveness of the cover, it is difficult to obtain a large amount of data covering different user groups. This limited supervisory information ultimately leads to poor cover recommendation results. Third, in video content-based cover recommendation technologies, the referenced information dimensions are relatively one-sided, resulting in information gaps. Most of these technologies only include factors such as human recognition and image quality, which can easily lead to the loss of important video information. The applicable scenarios for cover recommendation technologies are narrow, and the recommendation results are not accurate enough.

[0027] To address the problems existing in the video cover generation methods described above, this application provides a video cover generation method, terminal, and storage medium. The terminal acquires static cover information and tag information of the video information; performs static cover sorting processing based on the static cover information, tag information, and a target cover sorting model to obtain a sorting result; wherein, the target cover sorting model is obtained by training an initial cover sorting model using a preset loss function; the preset loss function is calculated based on click-through rate information and cosine similarity information; and generates a dynamic cover of the video information based on the sorting result. This method can obtain the optimal dynamic cover and improve the generation effect of the dynamic cover.

[0028] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.

[0029] This application provides a method for generating video cover images. Figure 1 This is a schematic diagram of the implementation process of the video cover generation method proposed in this application. Figure 1 ,like Figure 1 As shown, the method for generating a video cover on a terminal may include the following steps:

[0030] Step 101: Obtain the static cover information and tag information of the video information.

[0031] In the embodiments of this application, the terminal may first obtain the static cover information and tag information of the video information.

[0032] It should be noted that, in the embodiments of this application, the static cover information may include at least one static cover, which is a video frame selected from the video information.

[0033] It is understood that, in the embodiments of this application, the terminal may acquire the video information before acquiring the static cover information and tag information of the video information.

[0034] For example, in an embodiment of this application, the video information may be a short video data with a duration of five minutes.

[0035] In some embodiments of this application, the terminal can perform frame extraction processing on the video information to obtain key frame information of the video information; then perform multi-dimensional feature detection processing on the key frame information to obtain static cover information.

[0036] Furthermore, in the embodiments of this application, multidimensional feature detection processing may include any one or more of the following: facial expression recognition processing, human posture recognition processing, animal face recognition processing, motion optical flow calculation processing, salient object detection processing, quality analysis processing, and sky region detection processing.

[0037] Furthermore, in the embodiments of this application, the terminal may perform facial expression recognition processing on the keyframe information, and use video frames containing target expressions as static cover information; and / or, perform human posture recognition processing on the keyframe information, and use video frames containing target human postures as static cover information; and / or, perform animal face recognition processing on the keyframe information, and use video frames containing animal faces as static cover information; and / or, perform motion optical flow calculation processing on the keyframe information, and use video frames with motion amplitude greater than an amplitude threshold as static cover information; and / or, perform salient object detection processing on the keyframe information, and use video frames containing salient objects as static cover information; and / or, perform quality analysis processing on the keyframe information, and use video frames with quality scores greater than a score threshold as static cover information; and / or, perform sky region detection processing on the keyframe information, and use video frames with a sky region ratio equal to a preset ratio of other regions as static cover information.

[0038] It should be noted that, in the embodiments of this application, the terminal can be any electronic device with communication and storage functions, such as: tablet computer, mobile phone, e-reader, remote control, personal computer (PC), laptop computer, in-vehicle equipment, smart TV, wearable device, personal digital assistant (PDA), portable media player (PMP), navigation device and other electronic devices.

[0039] It is understood that, in the embodiments of this application, the video information can be a video of any length and type; for example, the video information can be a short video data with a duration of two minutes.

[0040] For example, in an embodiment of this application, the static cover information of the video information may include 10 static covers, all of which are video frames in the video information.

[0041] Furthermore, in the embodiments of this application, the tag information is a text tag generated based on video information that can describe the main idea of ​​the video information.

[0042] In some embodiments of this application, the terminal can perform frame extraction processing on the video information to obtain key frame information of the video information; then perform text tag extraction processing on the key frame information to obtain tag information.

[0043] Step 102: Perform static cover sorting based on static cover information, tag information, and target cover sorting model to obtain the sorting result; wherein, the target cover sorting model is obtained by training the initial cover sorting model using a preset loss function; the preset loss function is calculated based on click-through rate information and cosine similarity information.

[0044] In the embodiments of this application, after the terminal obtains the static cover information and tag information of the video information, it can perform static cover sorting processing based on the static cover information, tag information and target cover sorting model to obtain the sorting result; wherein, the target cover sorting model is obtained by training the initial cover sorting model using a preset loss function; the preset loss function is calculated based on click-through rate information and cosine similarity information.

[0045] It should be noted that, in the embodiments of this application, the sorting results can reflect the importance of different static covers in the static cover information.

[0046] In some embodiments of this application, the terminal can use an image feature extraction network to perform image feature extraction processing on static cover information to obtain the feature vector of static cover information; use a text feature extraction network to perform text feature extraction processing on tag information to obtain the feature vector of tag information; use the click rate corresponding to the static cover information as feedback supervision information to calculate the similarity parameter between the feature vector of static cover information and the feature vector of tag information; and obtain the ranking result based on the similarity parameter.

[0047] For example, in an embodiment of this application, the static cover information includes three static covers, A, B, and C. The static cover information is sorted by static cover, and the sorting result can be B, C, A. That is, in the sorting result, the static cover in the first position is B, the static cover in the second position is C, and the static cover in the third position is A.

[0048] Furthermore, in the embodiments of this application, the initial cover ranking model is trained using a preset loss function to obtain the target cover ranking model; wherein, the preset loss function is calculated based on click-through rate information and cosine similarity information.

[0049] The click-through rate (CTR) information represents the CTR weight of each image frame in the video frame data. The CTR information can be used to train the initial cover ranking model. For example, to obtain the CTR information, users can perform click processing on some video frame data in advance on the terminal to obtain the CTR information. Furthermore, the terminal can use this CTR information to train the initial cover ranking model.

[0050] Furthermore, in the embodiments of this application, the cosine similarity information is calculated based on image feature vectors and text feature vectors; specifically, after inputting the training dataset into the initial cover sorting model, image feature vectors and text feature vectors are obtained, and then cosine similarity information is calculated using the image feature vectors and text feature vectors; wherein, the training dataset includes video frame data and video tag data.

[0051] Furthermore, in the embodiments of this application, the initial cover ranking model includes a preset convolutional module; the process of training the initial cover ranking model using a preset loss function to obtain the target cover ranking model can be as follows: the terminal inputs the training dataset into the initial cover ranking model to obtain image feature vectors and text feature vectors; wherein, the training dataset includes video frame data and video tag data; then, cosine similarity information is calculated based on the image feature vectors and text feature vectors; then, a preset loss function is calculated based on the cosine similarity information and click-through rate information; then, the preset convolutional module is optimized using the preset loss function to obtain an optimized convolutional module; thereby, the target cover ranking model is determined based on the optimized convolutional module.

[0052] In other words, in the embodiments of this application, the initial cover sorting model is trained using a preset loss function to obtain the target cover sorting model, which can improve the training effect, thereby improving the sorting ability of the target cover sorting model and making the sorting results more accurate.

[0053] Step 103: Generate a dynamic cover for the video information based on the sorting results.

[0054] In the embodiments of this application, after the terminal performs static cover sorting processing based on static cover information, tag information and target cover sorting model to obtain the sorting result, it can generate dynamic cover of video information based on the sorting result.

[0055] In some embodiments of this application, the process of the terminal generating a dynamic cover of video information based on the sorting result can be as follows: the terminal selects at least one target static cover from the static cover information according to the sorting result; and generates a dynamic cover according to the target static cover and the cover generation strategy; wherein, the cover generation strategy includes frame interval parameters, total number of frames parameters and resolution parameters.

[0056] For example, in an embodiment of this application, the terminal selects the top five static covers (at least one target static cover) from the static cover information according to the sorting result, and then generates a dynamic cover based on these five static covers and the cover generation strategy.

[0057] Furthermore, in the embodiments of this application, the initial cover sorting model includes a preset convolution module; before the terminal performs static cover sorting processing based on static cover information, label information, and the target cover sorting model to obtain the sorting result, i.e., before step 102, the following steps may also be included:

[0058] Step 104: Input the training dataset into the initial cover sorting model to obtain image feature vectors and text feature vectors; the training dataset includes video frame data and video label data.

[0059] In the embodiments of this application, the terminal performs static cover sorting processing based on static cover information, label information and target cover sorting model. Before obtaining the sorting result, the training dataset can be input into the initial cover sorting model to obtain image feature vectors and text feature vectors; wherein, the training dataset includes video frame data and video label data.

[0060] It should be noted that, in the embodiments of this application, the initial cover sorting model includes an encoder module and a preset convolution module; wherein, the encoder module includes a graph encoder and a text encoder, and the preset convolution module includes a first convolution module and a second convolution module. The first convolution module may include two (Convolutional Neural Networks, CNN) convolutional layers, and the second convolution module may include two CNN convolutional layers; the encoder module is connected to the first convolution module, and the text encoder is connected to the second convolution module.

[0061] Furthermore, in the embodiments of this application, the graph encoder and text encoder can employ pre-trained encoder models, so that the terminal trains only the preset convolutional modules based on transfer learning; for example, the graph encoder uses a visual transformer, the text encoder uses a text transformer, and both the graph encoder and the text encoder load the network weights of the Contrastive Language-Image Pre-training (CLIP) model based on contrastive learning image-text learning, trained on 400 million unwashed image-text pair data from the Internet; by freezing the convolutional network layers therein, the weights in the graph encoder and the text encoder are no longer updated, and only the weights of the preset convolutional modules are trained.

[0062] Furthermore, in the embodiments of this application, the preset convolutional module may employ a convolutional neural network (CNN) convolutional layer, and the preset convolutional module may be trained using a training dataset to obtain the weights of the preset convolutional module.

[0063] It should be noted that, in the embodiments of this application, the training dataset includes video frame data and video label data. The video frame data is used as input to the image feature extraction network corresponding to the graph encoder and the first convolutional module, and the video label data is used as input to the text feature extraction network corresponding to the text encoder and the second convolutional module, thereby realizing the training of the first convolutional module and the second convolutional module.

[0064] It is understood that, in the embodiments of this application, video frame data can be video frame data extracted from a video, and video tag data can be tags extracted from a video, which can be used to describe the video content or theme.

[0065] It is also understood that, in the embodiments of this application, after the training dataset is input into the initial cover sorting model, that is, after the video frame data is input into the graph encoder and the video tag data is input into the text encoder, the image feature vector can be obtained after processing by the graph encoder and the first convolution module, and the text feature vector can be obtained after processing by the text encoder and the second convolution module; thus, the image feature vector is the feature vector corresponding to the video frame data, and the text feature vector is the feature vector corresponding to the video tag data.

[0066] Step 105: Calculate cosine similarity information based on image feature vectors and text feature vectors.

[0067] In the embodiments of this application, after the terminal inputs the training dataset into the initial cover sorting model and obtains the image feature vector and text feature vector, it can calculate the cosine similarity information based on the image feature vector and text feature vector.

[0068] It should be noted that, in the embodiments of this application, the similarity between image feature vectors and text feature vectors is calculated using the cosine similarity method; for example, the calculation method of cosine similarity information can be expressed as the following formula:

[0069]

[0070] Among them, SIM cos That is, to represent cosine similarity information, I i T represents the image feature vector. i This represents the text feature vector.

[0071] Step 106: Calculate the preset loss function based on the cosine similarity information and the click-through rate information; whereby the click-through rate information represents the click-through rate weight corresponding to each image frame in the video frame data.

[0072] In the embodiments of this application, after the terminal calculates the cosine similarity information based on the image feature vector and the text feature vector, it can calculate a preset loss function based on the cosine similarity information and the click-through rate information; wherein, the click-through rate information represents the click-through rate weight corresponding to each image frame in the video frame data.

[0073] It should be noted that, in the embodiments of this application, the click-through rate information represents the click-through rate weight corresponding to each image frame in the video frame data, and the click-through rate information can be used to train the initial cover ranking model; for example, to obtain the click-through rate information, the user can perform click processing on some video frame data in advance on the terminal to obtain the click-through rate information; furthermore, the terminal can use this click-through rate information to train the initial cover ranking model.

[0074] For example, in an embodiment of this application, the calculation of the preset loss function can be expressed as the following formula:

[0075]

[0076] in, This represents the preset loss function, where y represents the click-through rate information. The cosine similarity information has been normalized; δ represents the hyperparameter.

[0077] As can be seen, the preset loss function is the Huber Loss of regression. This loss is an optimized loss of mean squared error loss and mean absolute error loss, which is not easily affected by outliers and is relatively stable. Among them, when δ ~ 0, the Huber loss is close to the mean squared error loss MSE, and when δ ~ ∞, the Huber loss is close to the mean absolute error MAE. This parameter needs to be obtained through training.

[0078] Step 107: Optimize the preset convolutional module using a preset loss function to obtain the optimized convolutional module.

[0079] In the embodiments of this application, after the terminal calculates the preset loss function based on the cosine similarity information and the click-through rate information, it can use the preset loss function to optimize the preset convolution module to obtain the optimized convolution module.

[0080] It should be noted that, in the embodiments of this application, the weights in the preset convolutional module can be optimized using a preset loss function and a backpropagation algorithm to obtain an optimized convolutional module and complete the training.

[0081] Therefore, in the embodiments of this application, in the process of optimizing the preset convolution module using the preset loss function, the cosine similarity information is optimized and fitted based on the click rate information, which can improve the training effect and thus improve the ranking ability of the target cover ranking model.

[0082] Step 108: Determine the target cover sorting model based on the optimized convolution module.

[0083] In the embodiments of this application, after the terminal optimizes the preset convolution module using a preset loss function to obtain the optimized convolution module, it can determine the target cover sorting model based on the optimized convolution module.

[0084] It is understood that, in the embodiments of this application, the target cover sorting model includes an encoder module and an optimized convolution module.

[0085] Furthermore, in the embodiments of this application, the method for the terminal to obtain static cover information and tag information of video information, i.e., the method proposed in step 101, may include the following steps:

[0086] Step 101a: Perform frame extraction on the video information to obtain the keyframe information of the video information.

[0087] In the embodiments of this application, the terminal can perform frame extraction processing on the video information to obtain key frame information of the video information.

[0088] For example, in an embodiment of this application, the frame extraction tool ffmpeg is used to perform frame extraction processing on the video information to obtain the keyframe information of the video information.

[0089] It should be noted that, in the embodiments of this application, extracting keyframe information from video information can improve the efficiency of cover recommendation.

[0090] Step 101b: Perform multi-dimensional feature detection processing on the keyframe information to obtain static cover information.

[0091] In the embodiments of this application, after the terminal performs frame extraction processing on the video information to obtain the key frame information, it can perform multi-dimensional feature detection processing on the key frame information to obtain static cover information.

[0092] Furthermore, in the embodiments of this application, when the terminal performs multi-dimensional feature detection processing on the keyframe information to obtain static cover information, it can perform facial expression recognition processing on the keyframe information and use video frames containing target expressions as static cover information; and / or, perform human posture recognition processing on the keyframe information and use video frames containing target human postures as static cover information; and / or, perform animal face recognition processing on the keyframe information and use video frames containing animal faces as static cover information; and / or, perform motion optical flow calculation processing on the keyframe information and use video frames with motion amplitude greater than an amplitude threshold as static cover information; and / or, perform salient object detection processing on the keyframe information and use video frames containing salient objects as static cover information; and / or, perform quality analysis processing on the keyframe information and use video frames with quality scores greater than a score threshold as static cover information; and / or, perform sky region detection processing on the keyframe information and use video frames with a sky region ratio equal to a preset ratio of other regions as static cover information.

[0093] It should be noted that, in the embodiments of this application, the static cover information may include at least one static cover; for example, the static cover information may include 10 static covers.

[0094] It should be noted that, in the embodiments of this application, the specific processing types included in the multidimensional feature detection processing can be arbitrarily combined, added, or deleted; in other words, the above-mentioned facial expression recognition processing, human posture recognition processing, animal face recognition processing, motion optical flow calculation processing, salient object detection processing, quality analysis processing, and sky area detection processing can be arbitrarily combined, added, or deleted to obtain static cover information.

[0095] For example, in the embodiments of this application, the method of obtaining static cover information by facial expression recognition processing can be based on deep learning convolutional neural networks to extract facial features from key frame information to identify various facial expressions such as anger, fear, happiness, sadness, and naturalness. Then, according to actual needs, video frames that reflect the corresponding emotional color expression (target expression) are recommended as static cover information; for example, if the target expression is happiness, then video frames that reflect the emotional color of happiness are used as static cover.

[0096] For example, in the embodiments of this application, the method of obtaining static cover information by human posture recognition processing can be to use the OpenPose human posture recognition algorithm to identify human key points in key frame information and select video frames containing the target human posture as the cover. For example, if the target human posture is a dancing posture, then the video frames containing the dancing posture will be used as the static cover.

[0097] For example, in the embodiments of this application, the method of obtaining static cover information by animal face recognition processing can be based on deep neural networks and publicly available animal face datasets to train an animal face recognition model, thereby using the animal face recognition model to identify animal faces in key frame information, such as using video frames containing cat or dog faces as static covers.

[0098] For example, in the embodiments of this application, for the method of obtaining static cover information by motion optical flow calculation processing, video frames containing human faces, human bodies and animal faces in the key frame information can be subjected to sparse optical flow calculation with their corresponding subsequent video frames to obtain the motion sequence information of human or animal, and video frames with larger motion amplitudes are recommended in the future; for example, based on the motion sequence information, video frames with motion amplitudes greater than the amplitude threshold are used as static covers.

[0099] For example, in the embodiments of this application, the method for obtaining static cover information through salient object detection processing is as follows: when there are no people or animals in the video, the salient object detection algorithm Luminance Contrast or a deep learning-based salient object detection algorithm is used to perform salient object detection processing on the key frame information, thereby prioritizing the recommendation of video frames with salient objects in the center of the screen as static covers.

[0100] For example, in the embodiments of this application, the method of obtaining static cover information through quality analysis processing can be based on indicators such as the golden ratio, image sharpness, and color richness, or based on deep learning algorithms, to evaluate and score the key frame information and recommend video frames with high scores as static covers; for example, video frames with quality evaluation scores (quality scores) greater than the score threshold can be used as static covers.

[0101] For example, in the embodiments of this application, the method of obtaining static cover information by sky area detection processing can be to use edge detection or region detection algorithms such as Sobel operator to perform sky area detection processing on key frame information. When the ratio of sky area to other ground is 2 / 1 (preset ratio), it is the best viewing angle, and such video frames are recommended as static covers.

[0102] Step 101c: Extract text labels from the keyframe information to obtain label information.

[0103] In the embodiments of this application, after the terminal performs frame extraction processing on the video information to obtain the key frame information, it can perform text tag extraction processing on the key frame information to obtain tag information.

[0104] It should be noted that, in the embodiments of this application, the tag information represents text tags used to describe the main idea of ​​the video information.

[0105] Furthermore, in the embodiments of this application, the process of the terminal performing text tag extraction processing on keyframe information to obtain tag information may be as follows: the terminal uses an image description model to perform text description processing on keyframe information to obtain a set of text descriptions of keyframe information; and performs keyword extraction processing on the set of text descriptions to obtain tag information.

[0106] For example, in the embodiments of this application, the image description model can be an image captioning model, which outputs a text description that accurately reflects the content of the image based on the image content; the image description model can be a deep learning encoder-decoder framework with an attention mechanism, wherein a CNN extracts image features of keyframe information, and a recurrent neural network (RNN) / long short-term memory (LSTM) with an attention mechanism generates text descriptions from the image features to obtain a set of text descriptions of keyframe information.

[0107] For example, in the embodiments of this application, Natural Language Processing (NLP) keyword extraction technology, including TF-IDF, Text-Rank, etc., can be used to extract keywords from the text description set. The keyword extraction threshold can be adjusted according to actual needs, and finally, tag information that can represent the main content or theme of the video can be generated.

[0108] For example, in the embodiments of this application, the label information can be "Happy Life", "Food Diary", etc.

[0109] Furthermore, in the embodiments of this application, the target cover sorting model includes an image feature extraction network and a text feature extraction network; Figure 2 This is a schematic diagram of the implementation process of the video cover generation method proposed in this application. Figure 2 ,like Figure 2 As shown, the method for the terminal to perform static cover sorting based on static cover information, label information, and target cover sorting model to obtain the sorting result, i.e., the method proposed in step 102, may include the following steps:

[0110] Step 102a: Use an image feature extraction network to perform image feature extraction processing on the static cover information to obtain the feature vector of the static cover information.

[0111] In some embodiments of this application, the terminal performs static cover sorting processing based on static cover information, label information, and target cover sorting model to obtain sorting results; in some embodiments of this application, the terminal may first use an image feature extraction network to perform image feature extraction processing on the static cover information to obtain the feature vector of the static cover information.

[0112] It is understood that in the embodiments of this application, the image feature extraction network is used to extract image features, so that after the static cover information is input into the image feature extraction network, the feature vector of the static cover information can be obtained.

[0113] Step 102b: Use a text feature extraction network to extract text features from the label information to obtain the feature vector of the label information.

[0114] In some embodiments of this application, the terminal performs static cover sorting based on static cover information, label information, and a target cover sorting model to obtain a sorting result; in some embodiments of this application, the terminal may first use a text feature extraction network to perform text feature extraction processing on the label information to obtain the feature vector of the label information.

[0115] It is understood that in the embodiments of this application, the text feature extraction network is used to extract text features, so that after the label information is input into the text feature extraction network, the feature vector of the label information can be obtained.

[0116] Step 102c: Using the click-through rate corresponding to the static cover information as feedback supervision information, calculate the similarity parameter between the feature vector of the static cover information and the feature vector of the tag information.

[0117] In the embodiments of this application, after the terminal uses an image feature extraction network to perform image feature extraction processing on static cover information to obtain the feature vector of static cover information, and uses a text feature extraction network to perform text feature extraction processing on tag information to obtain the feature vector of tag information, it can use the click rate corresponding to the static cover information as feedback supervision information to calculate the similarity parameter between the feature vector of static cover information and the feature vector of tag information.

[0118] It should be noted that, in the embodiments of this application, the click-through rate corresponding to the static cover information can be obtained by the terminal detecting the click operation corresponding to the static cover information, thereby obtaining the click-through rate corresponding to the static cover information; that is, the user can click on the static cover information of the recommended display according to their preferences on the terminal, thereby causing the terminal to generate the click-through rate corresponding to the static cover information.

[0119] For example, in an embodiment of this application, the static cover information includes 10 static covers. The user clicks on 5 of them on the terminal. After the terminal detects the clicks on these 5 static covers, it generates the click rate corresponding to the static cover information, which includes the click rate of the 5 clicked static covers and the click rate of the remaining 5 static covers that were not clicked, thus having a click rate of 0.

[0120] For example, in the embodiments of this application, the similarity parameter can also be calculated using the cosine similarity method, as shown in the aforementioned formula (1). However, the image feature vector in formula (1) is replaced with the feature vector of static cover information, and the text feature vector is replaced with the feature vector of label information, thereby calculating the similarity parameter.

[0121] It should be noted that, in the embodiments of this application, in the process of using the target cover ranking model to perform static cover ranking, in addition to using the click rate corresponding to the static cover information as feedback supervision information, the tag information can also be understood as a kind of feedback supervision information in the static cover ranking process when calculating the similarity parameter. Thus, the click rate and tag information corresponding to the static cover information can be regarded as feedback and feedforward in the static cover ranking process, which is conducive to achieving effective ranking of static covers.

[0122] Step 102d: Obtain the sorting results based on the similarity parameters.

[0123] In the embodiments of this application, after the terminal uses the click-through rate corresponding to the static cover information as feedback supervision information and calculates the similarity parameter between the feature vector of the static cover information and the feature vector of the tag information, it can obtain the ranking result based on the similarity parameter.

[0124] It is understood that, in the embodiments of this application, the ranking result obtained based on the similarity parameter is obtained by combining the degree of user preference and the degree of closeness to the main theme of the video information content.

[0125] For example, in an embodiment of this application, the static cover information includes five static covers A, B, C, D, and E, and the final sorting result is: C, B, D, E, A.

[0126] Furthermore, in the embodiments of this application, the method for generating a dynamic cover of video information based on the sorting result by the terminal, i.e., the method proposed in step 103, may include the following steps:

[0127] Step 103a: Select at least one target static cover from the static cover information according to the sorting results.

[0128] In some embodiments of this application, the terminal generates a dynamic cover for the video information based on the sorting results; in some embodiments of this application, the terminal may first select at least one target static cover from the static cover information according to the sorting results.

[0129] For example, in an embodiment of this application, the static cover information includes 10 static covers, and the sorting result includes the serial number corresponding to these 10 static covers; according to the sorting result, the top five static covers (at least one target static cover) in the sorting result are selected in order.

[0130] Step 103b: Generate a dynamic cover based on the target static cover and the cover generation strategy.

[0131] In the embodiments of this application, after the terminal selects at least one target static cover from the static cover information according to the sorting result, it can generate a dynamic cover according to the target static cover and the cover generation strategy; wherein, the cover generation strategy includes frame interval parameters, total number of frames parameters and resolution parameters.

[0132] It should be noted that, in the embodiments of this application, the cover generation strategy includes frame interval parameters, total number of frames parameters, and resolution parameters, etc.

[0133] For example, in an embodiment of this application, there are 5 target static covers. A dynamic cover is generated based on these 5 images. First, the frame interval (frame interval parameter), total number of frames (total number of frames parameter), and resolution (resolution parameter) of the dynamic cover are determined. Then, based on the above cover generation strategy, the 5 images are processed to generate a dynamic cover.

[0134] This application provides a video cover generation method. A terminal acquires static cover information and tag information from video information; performs static cover sorting based on the static cover information, tag information, and a target cover sorting model to obtain a sorting result; wherein, the target cover sorting model is obtained by training an initial cover sorting model using a preset loss function; the preset loss function is calculated based on click-through rate information and cosine similarity information; and a dynamic cover of the video information is generated based on the sorting result. Therefore, in this application, the terminal acquires static cover information and tag information from video information, and then uses the target cover sorting model and tag information to sort the static cover information to obtain a sorting result, thereby enabling the sorting result to reflect the importance of different static covers; simultaneously, this application calculates a preset loss function based on click-through rate information and cosine similarity information, and uses the preset loss function to train the initial cover sorting model to obtain the target cover sorting model, which can improve the training effect, thereby improving the sorting ability of the target cover sorting model and making the sorting result more accurate; furthermore, based on the sorting result, an optimal dynamic cover can be generated, improving the generation effect of the dynamic cover.

[0135] Based on the above embodiments, in another embodiment of this application, the terminal may include a static cover recommendation module, a video tag generation module, a static cover sorting module, and a dynamic cover generation module; for example, Figure 3 This is a schematic diagram illustrating the implementation of the video cover generation method proposed in this application. Figure 1 ,like Figure 3 As shown, video information can be input into the static cover recommendation module and the video tag generation module to obtain static cover information and tag information. Then, the static cover information and tag information can be input into the static cover sorting module to sort the static cover information and obtain the sorting result. Finally, the sorting result can be input into the dynamic cover generation module to generate a dynamic cover.

[0136] For example, in the embodiments of this application, based on the above-mentioned static cover recommendation module, video tag generation module, static cover sorting module and dynamic cover generation module, a method for generating dynamic covers with multimodal fine-grained semantics can be realized. First, the terminal can obtain short video data (video information) uploaded by the user. In order to improve the efficiency of cover recommendation, video key frames (key frame information) can be extracted first, and then the key frames are input into the static cover recommendation module for multi-dimensional feature detection processing to obtain multiple recommended static cover frames (static cover information).

[0137] The static cover recommendation module can be deployed with different detection sub-modules to implement multi-dimensional feature detection processing; for example, Figure 4 This is a schematic diagram illustrating the implementation of the video cover generation method proposed in this application. Figure 2,like Figure 4 As shown, the static cover recommendation module may include sub-modules for facial expression recognition, human posture recognition, animal face recognition, motion optical flow calculation, salient object detection, quality analysis, and sky region detection. Specifically, the facial expression recognition sub-module performs facial expression recognition processing, using video frames containing the target expression as static cover information; the human posture recognition sub-module uses video frames containing the target human posture as static cover information; the animal face recognition sub-module performs animal face recognition processing, using video frames containing animal faces as static cover information; and the motion optical flow calculation sub-module performs motion optical flow calculation processing, which can be used to determine the static cover information after a certain period of time. After processing by the facial expression recognition submodule, human pose recognition submodule, or animal face recognition submodule, sparse optical flow calculation is performed on video frames containing human, human, or animal faces and their corresponding subsequent video frames. Video frames with motion amplitude greater than an amplitude threshold are used as static cover information. The salient object detection submodule can be used to perform salient object detection processing, and video frames containing salient objects are used as static cover information. The quality analysis submodule can be used to perform quality analysis processing, and video frames with quality scores greater than a score threshold are used as static cover information. The sky region detection submodule can be used to perform sky region detection processing, and video frames with a sky region ratio equal to a preset ratio with other regions are used as static cover information.

[0138] For example, in the embodiments of this application, the method of obtaining static cover information by facial expression recognition processing can be based on deep learning convolutional neural networks to extract facial features from key frame information to identify various facial expressions such as anger, fear, happiness, sadness, and naturalness. Then, according to actual needs, video frames that reflect the corresponding emotional color expression (target expression) are recommended as static cover information; for example, if the target expression is happiness, then video frames that reflect the emotional color of happiness are used as static cover.

[0139] For example, in the embodiments of this application, the method of obtaining static cover information by human posture recognition processing can be to use the OpenPose human posture recognition algorithm to identify human key points in key frame information and select video frames containing the target human posture as the cover. For example, if the target human posture is a dancing posture, then the video frames containing the dancing posture will be used as the static cover.

[0140] For example, in the embodiments of this application, the method of obtaining static cover information by animal face recognition processing can be based on deep neural networks and publicly available animal face datasets to train an animal face recognition model, thereby using the animal face recognition model to identify animal faces in key frame information, such as using video frames containing cat or dog faces as static covers.

[0141] For example, in the embodiments of this application, for the method of obtaining static cover information by motion optical flow calculation processing, video frames containing human faces, human bodies and animal faces in the key frame information can be subjected to sparse optical flow calculation with their corresponding subsequent video frames to obtain the motion sequence information of human or animal, and video frames with larger motion amplitudes are recommended in the future; for example, based on the motion sequence information, video frames with motion amplitudes greater than the amplitude threshold are used as static covers.

[0142] For example, in the embodiments of this application, the method for obtaining static cover information through salient object detection processing is as follows: when there are no people or animals in the video, the salient object detection algorithm Luminance Contrast or a deep learning-based salient object detection algorithm is used to perform salient object detection processing on the key frame information, thereby prioritizing the recommendation of video frames with salient objects in the center of the screen as static covers.

[0143] For example, in the embodiments of this application, the method of obtaining static cover information through quality analysis processing can be based on indicators such as the golden ratio, image sharpness, and color richness, or based on deep learning algorithms, to evaluate and score the key frame information and recommend video frames with high scores as static covers; for example, video frames with quality evaluation scores (quality scores) greater than the score threshold can be used as static covers.

[0144] For example, in the embodiments of this application, the method of obtaining static cover information by sky area detection processing can be to use edge detection or region detection algorithms such as Sobel operator to perform sky area detection processing on key frame information. When the ratio of sky area to other ground is 2 / 1 (preset ratio), it is the best viewing angle, and such video frames are recommended as static covers.

[0145] Furthermore, in the embodiments of this application, the above detection sub-modules can be combined, added, or removed according to the actual video content or needs; through the static cover recommendation module, highly representative static cover information can be obtained.

[0146] Furthermore, in the embodiments of this application, the video tag generation module can input video keyframes into the Image Captioning model. The function of this model is to output a text description that accurately reflects the content of the image based on the content of the image, thereby completing the text description processing of the video keyframes and obtaining a text description set. Then, NLP keyword extraction techniques such as TF-IDF and Text-Rank are used to extract keywords from the text description set to obtain tag information.

[0147] Furthermore, in the embodiments of this application, it is also necessary to obtain the click-through rate corresponding to the static cover information; for example, 10 recommended static covers (static cover information) can be obtained for users to choose from, thereby obtaining the click-through rate corresponding to the static cover information.

[0148] Furthermore, in the embodiments of this application, the static cover information is sorted by the static cover sorting module to obtain the sorting result. Specifically, the image feature extraction network in the target cover sorting model can be used to perform image feature extraction processing on the static cover information to obtain the feature vector of the static cover information; at the same time, the text feature extraction network can be used to perform text feature extraction processing on the tag information to obtain the feature vector of the tag information; then, the click rate corresponding to the static cover information is used as feedback supervision information to calculate the similarity parameter between the feature vector of the static cover information and the feature vector of the tag information; finally, the sorting result is obtained based on the similarity parameter.

[0149] For example, Figure 5 This is a schematic diagram illustrating the implementation of the video cover generation method proposed in this application. Figure 3 ,like Figure 5 As shown, static cover information is input into the image feature extraction network in the target cover ranking model to obtain the feature vector of static cover information; label information is input into the text feature extraction network to obtain the feature vector of label information. Then, the feature vectors of static cover information and label information are used to calculate the similarity parameter. In the process of calculating the similarity parameter, the click rate corresponding to the static cover information is used as feedback supervision information to participate, thereby obtaining the ranking result.

[0150] It should be noted that, in the embodiments of this application, the target cover ranking model is obtained by training the initial cover ranking model using a preset loss function; the preset loss function is calculated based on click-through rate information and cosine similarity information.

[0151] The initial cover sorting model includes an encoder module and a preset convolution module. The encoder module includes a graph encoder and a text encoder. The preset convolution module includes a first convolution module and a second convolution module. The first convolution module may include two CNN convolutional layers, and the second convolution module may include two CNN convolutional layers. The encoder module is connected to the first convolution module, and the text encoder is connected to the second convolution module.

[0152] Furthermore, in the embodiments of this application, the graph encoder and text encoder can employ pre-trained encoder models, so that the terminal trains only the preset convolutional modules based on transfer learning; for example, the graph encoder uses a visual transformer, the text encoder uses a text transformer, and both the graph encoder and the text encoder load the network weights trained by CLIP based on 400 million unwashed image-text pair data from the Internet; by freezing the convolutional network layers therein, the weights in the graph encoder and the text encoder are no longer updated, and only the weights of the preset convolutional modules are trained.

[0153] Furthermore, in the embodiments of this application, a preset convolutional module is trained using a training dataset to obtain the weights of the preset convolutional module.

[0154] Furthermore, in the embodiments of this application, the dynamic cover generation module selects at least one target static cover from the static cover information according to the sorting result; and generates a dynamic cover according to the target static cover and the cover generation strategy.

[0155] For example, in an embodiment of this application, there are 5 target static covers. A dynamic cover is generated based on these 5 images. First, the frame interval (frame interval parameter), total number of frames (total number of frames parameter), and resolution (resolution parameter) of the dynamic cover are determined. Then, based on the above cover generation strategy, the 5 images are processed to generate a dynamic cover.

[0156] Therefore, based on the aforementioned static cover recommendation module, video tag generation module, static cover sorting module, and dynamic cover generation module, intelligent dynamic cover generation with multi-dimensional semantics can be achieved. This represents a modular dynamic cover generation method; these modules are detachable and can be customized according to requirements. Furthermore, after obtaining multi-dimensional static cover information, this application can also customize different target cover sorting models and different cover generation strategies according to different business scenarios, thereby obtaining the dynamic cover required by the business.

[0157] Furthermore, this application can construct dynamic covers from multiple recommended static covers. The static covers are sorted based on feedback data such as tag information and click-through rates, ensuring the effectiveness of subsequent dynamic cover generation. In addition, this application performs facial expression recognition, human pose recognition, animal face recognition, motion optical flow calculation, salient object detection, quality analysis, and empty region detection on keyframe information, providing fine-grained comprehensive recommendations of static covers with comprehensive reference information and wide coverage of scenarios, resulting in highly representative static cover information.

[0158] Furthermore, the video cover generation method proposed in this application generates a dynamic cover based on a static cover, breaking the black box model and making it more interpretable; it also modularizes the dynamic cover generation process, allowing for selection and replacement to be customized according to different business scenario needs, making it more flexible and convenient; it avoids the inaccuracy of selecting a dynamic cover solely from video clips, improves the flexibility and convenience of dynamic cover recommendations, and expands application scenarios.

[0159] This application provides a video cover generation method. A terminal acquires static cover information and tag information from video information; performs static cover sorting based on the static cover information, tag information, and a target cover sorting model to obtain a sorting result; wherein, the target cover sorting model is obtained by training an initial cover sorting model using a preset loss function; the preset loss function is calculated based on click-through rate information and cosine similarity information; and a dynamic cover of the video information is generated based on the sorting result. Therefore, in this application, the terminal acquires static cover information and tag information from video information, and then uses the target cover sorting model and tag information to sort the static cover information to obtain a sorting result, thereby enabling the sorting result to reflect the importance of different static covers; simultaneously, this application calculates a preset loss function based on click-through rate information and cosine similarity information, and uses the preset loss function to train the initial cover sorting model to obtain the target cover sorting model, which can improve the training effect, thereby improving the sorting ability of the target cover sorting model and making the sorting result more accurate; furthermore, based on the sorting result, an optimal dynamic cover can be generated, improving the generation effect of the dynamic cover.

[0160] Based on the above embodiments, in another embodiment of this application... Figure 6 This is a schematic diagram of the terminal structure proposed in the embodiments of this application. Figure 1 ,like Figure 6 As shown, the terminal 10 proposed in this embodiment may include an acquisition unit 11 and a generation unit 12.

[0161] The acquisition unit 11 is used to acquire static cover information and tag information of video information; and to perform static cover sorting processing based on the static cover information, the tag information and the target cover sorting model to obtain the sorting result; wherein, the target cover sorting model is obtained by training an initial cover sorting model using a preset loss function; the preset loss function is calculated based on click-through rate information and cosine similarity information.

[0162] The generation unit 12 is used to generate a dynamic cover for the video information based on the sorting result.

[0163] Furthermore, the initial cover sorting model includes a preset convolutional module; the acquisition unit 11 is also used to input a training dataset into the initial cover sorting model to obtain image feature vectors and text feature vectors before performing static cover sorting processing based on the static cover information, the tag information, and the target cover sorting model to obtain the sorting result; wherein the training dataset includes video frame data and video tag data; and to calculate the cosine similarity information based on the image feature vectors and the text feature vectors; and to calculate the preset loss function based on the cosine similarity information and the click-through rate information; wherein the click-through rate information represents the click-through rate weight corresponding to each image frame in the video frame data; and to optimize the preset convolutional module using the preset loss function to obtain an optimized convolutional module; and to determine the target cover sorting model based on the optimized convolutional module.

[0164] Furthermore, the acquisition unit 11 is also used to perform frame extraction processing on the video information to obtain key frame information of the video information; and to perform multi-dimensional feature detection processing on the key frame information to obtain the static cover information; and to perform text tag extraction processing on the key frame information to obtain the tag information.

[0165] Furthermore, the acquisition unit 11 is also configured to perform facial expression recognition processing on the keyframe information, and use video frames containing target expressions as the static cover information; and / or, perform human posture recognition processing on the keyframe information, and use video frames containing target human postures as the static cover information; and / or, perform animal face recognition processing on the keyframe information, and use video frames containing animal faces as the static cover information; and / or, perform motion optical flow calculation processing on the keyframe information, and use video frames with motion amplitude greater than an amplitude threshold as the static cover information; and / or, perform salient object detection processing on the keyframe information, and use video frames containing salient objects as the static cover information; and / or, perform quality analysis processing on the keyframe information, and use video frames with quality scores greater than a score threshold as the static cover information; and / or, perform sky region detection processing on the keyframe information, and use video frames with a sky region ratio equal to a preset ratio of other regions as the static cover information.

[0166] Furthermore, the acquisition unit 11 is also used to perform text description processing on the keyframe information using an image description model to obtain a set of text descriptions of the keyframe information; and to perform keyword extraction processing on the set of text descriptions to obtain the tag information.

[0167] Furthermore, the target cover ranking model includes an image feature extraction network and a text feature extraction network; the acquisition unit 11 is also used to perform image feature extraction processing on the static cover information using the image feature extraction network to obtain the feature vector of the static cover information; and to perform text feature extraction processing on the tag information using the text feature extraction network to obtain the feature vector of the tag information; and to use the click-through rate corresponding to the static cover information as feedback supervision information to calculate the similarity parameter between the feature vector of the static cover information and the feature vector of the tag information; and to obtain the ranking result based on the similarity parameter.

[0168] Furthermore, the generation unit 12 is also configured to select at least one target static cover from the static cover information according to the sorting result; and to generate the dynamic cover according to the target static cover and the cover generation strategy.

[0169] Figure 7 This is a schematic diagram of the terminal structure proposed in the embodiments of this application. Figure 2 ,like Figure 7 As shown, the terminal 10 proposed in this application embodiment may further include a processor 13, a memory 14 storing instructions executable by the processor 13, and further, the terminal 10 may also include a communication interface 15 and a bus 16 for connecting the processor 13, the memory 14 and the communication interface 15.

[0170] In the embodiments of this application, the processor 13 can be at least one of the following: Application-Specific Integrated Circuit (ASIC), Digital Signal Processor (DSP), Digital Signal Processing Device (DSPD), Programmable Logic Device (PLD), Field Programmable Gate Array (FPGA), Central Processing Unit (CPU), controller, microcontroller, and microprocessor. It is understood that for different devices, the electronic device used to implement the above-mentioned processor function can also be other types, and this application embodiment does not specifically limit this. The processor 13 may also include a memory 14, which can be connected to the processor 13. The memory 14 is used to store executable program code, which includes computer operation instructions. The memory 14 may include high-speed RAM memory and may also include non-volatile memory, such as at least two disk drives.

[0171] In embodiments of this application, bus 16 is used to connect communication interface 15, processor 13 and memory 14 and the mutual communication between these devices.

[0172] In embodiments of this application, memory 14 is used to store instructions and data.

[0173] Furthermore, in the embodiments of this application, the processor 13 is used to acquire static cover information and tag information of video information; perform static cover sorting processing based on the static cover information, the tag information and the target cover sorting model to obtain a sorting result; wherein, the target cover sorting model is obtained by training an initial cover sorting model using a preset loss function; the preset loss function is calculated based on click-through rate information and cosine similarity information; and generate a dynamic cover of the video information based on the sorting result.

[0174] In practical applications, the aforementioned memory 14 can be volatile memory, such as random-access memory (RAM); or non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid-state drive (SSD); or a combination of the above types of memory, and provide instructions and data to the processor 13.

[0175] Furthermore, in this embodiment, the functional modules can be integrated into one analysis unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional module.

[0176] If the integrated unit is implemented as a software functional module and is not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or terminal, etc.) or processor to execute all or part of the steps of the method of this embodiment. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0177] This application provides a terminal that acquires static cover information and tag information of video information; performs static cover sorting processing based on the static cover information, tag information, and a target cover sorting model to obtain a sorting result; wherein, the target cover sorting model is obtained by training an initial cover sorting model using a preset loss function; the preset loss function is calculated based on click-through rate information and cosine similarity information; and generates a dynamic cover of the video information based on the sorting result. Thus, in this application, the terminal acquires static cover information and tag information of video information, and then uses the target cover sorting model and tag information to sort the static cover information to obtain a sorting result, thereby enabling the sorting result to reflect the importance of different static covers; simultaneously, this application calculates a preset loss function based on click-through rate information and cosine similarity information, and uses the preset loss function to train the initial cover sorting model to obtain the target cover sorting model, which can improve the training effect, thereby improving the sorting ability of the target cover sorting model and making the sorting result more accurate; furthermore, based on the sorting result, an optimal dynamic cover can be generated, improving the generation effect of the dynamic cover.

[0178] Specifically, the program instructions corresponding to a video cover generation method in this embodiment can be stored on storage media such as optical discs, hard disks, and USB flash drives; when the program instructions corresponding to a video cover generation method in the storage medium are read or executed by a processor, the following steps are included:

[0179] Obtain static cover and tag information from video information;

[0180] Static cover sorting is performed based on the static cover information, the tag information, and the target cover sorting model to obtain the sorting result; wherein, the target cover sorting model is obtained by training the initial cover sorting model using a preset loss function; the preset loss function is calculated based on click-through rate information and cosine similarity information;

[0181] A dynamic cover image for the video information is generated based on the sorting results.

[0182] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of hardware embodiments, software embodiments, or embodiments combining software and hardware aspects. Furthermore, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage and optical storage) containing computer-usable program code.

[0183] This application is described with reference to schematic and / or block diagrams of implementations of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block of the schematic and / or block diagrams can be implemented by computer program instructions, and combinations of blocks in the schematic and / or block diagrams can be implemented. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the schematic and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0184] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in the implementation flow diagram. Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0185] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0186] The above description is merely a preferred embodiment of this application and is not intended to limit the scope of protection of this application.

Claims

1. A method for generating video cover images, characterized in that, The method includes: The static cover information and tag information of the video information are obtained; wherein, the static cover information is determined based on multi-dimensional feature detection processing of the key frame information of the video information; the multi-dimensional feature detection processing includes at least one of the following: human posture recognition processing, animal face recognition processing, motion optical flow calculation processing, salient object detection processing, and sky region detection processing; The training dataset is input into the initial cover sorting model to obtain image feature vectors and text feature vectors; wherein, the training dataset includes video frame data and video tag data; the initial cover sorting model includes a preset convolutional module; Calculate cosine similarity information based on the image feature vector and the text feature vector; A preset loss function is calculated based on the cosine similarity information and click-through rate information; wherein, the click-through rate information represents the click-through rate weight corresponding to each image frame in the video frame data; the preset loss function is the Huber Loss regression; the preset loss function is... ,in, This represents the preset loss function. This refers to the click-through rate information. This represents the cosine similarity information. Indicates hyperparameters; The preset convolutional module is optimized using the preset loss function to obtain the optimized convolutional module. The target cover sorting model is determined based on the optimized convolutional module; Static cover sorting is performed based on the static cover information, the tag information, and the target cover sorting model to obtain a sorting result; wherein, the target cover sorting model is obtained by training the initial cover sorting model using the preset loss function; the preset loss function is calculated based on the click-through rate information and the cosine similarity information; A dynamic cover image for the video information is generated based on the sorting results.

2. The method according to claim 1, characterized in that, The static cover information and tag information of the video information obtained include: The video information is subjected to frame extraction processing to obtain the keyframe information of the video information; The keyframe information is subjected to the multi-dimensional feature detection processing to obtain the static cover information; The keyframe information is processed to extract text tags, thereby obtaining the tag information.

3. The method according to claim 2, characterized in that, The step of performing multi-dimensional feature detection processing on the keyframe information to obtain the static cover information includes: Perform facial expression recognition processing on the keyframe information, and use the video frame containing the target expression as the static cover information; and / or, Perform the human pose recognition processing on the keyframe information, and use the video frame containing the target human pose as the static cover information; and / or, Perform the animal face recognition processing on the keyframe information, and use the video frame containing the animal face as the static cover information; and / or, The motion optical flow calculation is performed on the keyframe information, and video frames with motion amplitude greater than the amplitude threshold are used as the static cover information; and / or, Perform the salient object detection processing on the keyframe information, and use the video frame containing the salient object as the static cover information; and / or, The keyframe information is subjected to quality analysis processing, and video frames with quality scores greater than a score threshold are used as the static cover information; and / or, The keyframe information is subjected to the sky region detection process, and the video frame in which the ratio of the sky region to other regions is equal to the preset ratio is used as the static cover information.

4. The method according to claim 2, characterized in that, The step of extracting text tags from the keyframe information to obtain the tag information includes: The keyframe information is processed using an image description model to obtain a set of text descriptions for the keyframe information. The set of text descriptions is processed by keyword extraction to obtain the tag information.

5. The method according to claim 1, characterized in that, The target cover sorting model includes an image feature extraction network and a text feature extraction network; The static cover sorting process based on the static cover information, the label information, and the target cover sorting model to obtain the sorting result includes: The image feature extraction network is used to perform image feature extraction processing on the static cover information to obtain the feature vector of the static cover information; The text feature extraction network is used to perform text feature extraction processing on the label information to obtain the feature vector of the label information; The click-through rate corresponding to the static cover information is used as feedback supervision information to calculate the similarity parameter between the feature vector of the static cover information and the feature vector of the tag information; The ranking result is obtained based on the similarity parameter.

6. The method according to claim 1 or 5, characterized in that, The process of generating a dynamic cover image for the video information based on the sorting result includes: According to the sorting results, at least one target static cover is selected from the static cover information; The dynamic cover is generated based on the target static cover and the cover generation strategy.

7. A terminal, characterized in that, The terminal includes an acquisition unit and a generation unit. The acquisition unit is used to acquire static cover information and tag information of video information; and to perform static cover sorting processing based on the static cover information, the tag information, and a target cover sorting model to obtain a sorting result; wherein, the static cover information is determined based on multi-dimensional feature detection processing of keyframe information of the video information; the multi-dimensional feature detection processing includes at least one of the following: human pose recognition processing, animal face recognition processing, motion optical flow calculation processing, salient object detection processing, and sky region detection processing; the target cover sorting model is obtained by training an initial cover sorting model using a preset loss function; the preset loss function is calculated based on click-through rate information and cosine similarity information; The acquisition unit is further configured to input a training dataset into the initial cover sorting model to obtain image feature vectors and text feature vectors before performing static cover sorting processing based on the static cover information, the tag information, and the target cover sorting model to obtain the sorting result; wherein, the training dataset includes video frame data and video tag data; the initial cover sorting model includes a preset convolution module; calculates the cosine similarity information based on the image feature vectors and the text feature vectors; calculates the preset loss function based on the cosine similarity information and the click-through rate information; wherein, the click-through rate information represents the click-through rate weight corresponding to each image frame in the video frame data; the preset loss function is the Huber Loss of regression; the preset loss function is... ,in, This represents the preset loss function. This refers to the click-through rate information. This represents the cosine similarity information. The hyperparameters are represented; the preset convolutional module is optimized using the preset loss function to obtain the optimized convolutional module; the target cover sorting model is determined based on the optimized convolutional module; The generation unit is used to generate a dynamic cover for the video information based on the sorting result.

8. A terminal, characterized in that, The terminal includes a processor and a memory storing processor-executable instructions, which, when executed by the processor, implement the method as described in any one of claims 1-6.

9. A computer-readable storage medium having a program stored thereon for use in a terminal, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Video cover extraction method and device based on video title

    CN107918656A

  • Video cover generation method, device and electronic equipment

    CN110446063A

  • Video cover generation method and device and storage medium

    CN114491151A