Video cover recommendation method, device, equipment and computer-readable storage medium
By matching video frames with the cover of the video application platform and using interactive data to recommend the video cover, the problem of low click-through rate caused by insufficient experience of video creators is solved, and a higher video publishing click-through rate is achieved.
Patent Information
- Application Number
- CN202110891456.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-04
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2041-08-04
AI Technical Summary
In the prior art, video creators lack experience leads to inaccurate recommendations for video covers, resulting in low click-through rates after video is published.
By obtaining each video frame of the target video, matching it with the video cover in the video application platform, obtaining interactive data, and recommending video covers based on the interactive data.
Improves the accuracy of video cover recommendations and improves the click rate after video is released.
Smart Images

Figure CN115706836B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a method, device, equipment, and computer-readable storage medium for recommending video covers. Background Art
[0002] With the rapid progress of modern information transmission technology and the popularization of video shooting devices such as smartphones, people's enthusiasm for sharing life by creating videos has developed unprecedentedly, and videos have gradually become one of the main carriers for people to receive information daily. As the information that users see first, the video cover largely determines whether the relevant video will be watched by users. Therefore, selecting a high-quality video cover helps to improve the user experience and facilitate video sharing and promotion.
[0003] In the related art, usually, interval frames are extracted from the video for recommending video covers. The creator selects one from the recommended video frames by dragging or clicking as the video cover. In this way, if a higher click-through rate is desired for the video to be published, it depends on the creator's experience accumulation. For creators without much experience, there is a problem of low click-through rate after the video is published for the recommended video covers. Summary of the Invention
[0004] The embodiments of this application provide a method, device, and computer-readable storage medium for recommending video covers, which can accurately recommend video covers, thereby improving the click-through rate after the video is published.
[0005] The technical solution of the embodiments of this application is implemented as follows:
[0006] The embodiments of this application provide a method for recommending video covers, including:
[0007] Obtain each video frame of the target video;
[0008] Match each of the video frames with the video covers in the video application platform to obtain the video covers that match each of the video frames;
[0009] Respectively obtain the interaction data corresponding to the video covers that match each of the video frames;
[0010] Recommend the video cover for the target video according to the interaction data corresponding to each of the video covers.
[0011] The embodiments of this application provide a method for recommending video covers, including:
[0012] During the process of editing the target video, a cover selection instruction for the target video is received; in response to the cover selection instruction, a cover selection interface is presented, and
[0013] In the cover selection interface, recommended video covers and the number of users interested in the video covers are displayed.
[0014] When a selection operation for a recommended video cover is received, the selected video cover is used as the video cover of the target video.
[0015] An embodiment of the present application provides a video cover recommendation device, including:
[0016] A first acquisition module, configured to acquire each video frame of a target video.
[0017] A matching module, configured to match each of the video frames with video covers in a video application platform to obtain video covers matching each of the video frames.
[0018] A second acquisition module, configured to respectively acquire interaction data corresponding to the video covers matching each of the video frames.
[0019] A recommendation module, configured to recommend a video cover for the target video according to the interaction data corresponding to each of the video covers.
[0020] In the above solution, the matching module is further configured to perform image recognition on each of the video frames from at least one dimension to obtain tag information of each of the video frames in at least one of the dimensions.
[0021] Acquire tag information of video covers in the video application platform in at least one of the dimensions.
[0022] Respectively match the tag information of each of the video frames with the tag information of the video covers in the video application platform to obtain video covers matching each of the video frames.
[0023] In the above solution, when the number of dimensions is at least two, the matching module is further configured to perform the following operations for each of the video frames:
[0024] Respectively match the tag information of the video frame in each dimension with the tag information of the video cover in the corresponding dimension to obtain the matching degree of the video frame and the video cover in each dimension.
[0025] Obtain the weight value of each of the dimensions.
[0026] Based on the weight values of each of the dimensions, perform weighted summation on the matching degrees of the video frame and the video cover in each of the dimensions to obtain the matching degree of the video frame and the video cover.
[0027] Based on the matching degree between the video frames and the video covers, filter the video covers in the video application platform to obtain the video covers that match the video frames.
[0028] In the above solution, the matching module is further configured to extract image features from each of the video frames to obtain first image features of each of the video frames;
[0029] Extract image features from the video covers in the video application platform to obtain second image features of the video covers;
[0030] Perform similarity matching between the first image features of each of the video frames and the second image features of the video covers to obtain the matching degree between each of the video frames and the video covers;
[0031] Based on the matching degree between the video frames and the video covers, filter the video covers in the video application platform to obtain the video covers that match the video frames.
[0032] In the above solution, the matching module is further configured to perform clustering processing on each of the video frames to obtain at least two clustering clusters;
[0033] Obtain the stillness degree of each of the video frames, where the stillness degree is used to indicate the numerical value of the motion energy of the video frame;
[0034] Based on the stillness degree of each of the video frames, extract key video frames from at least two of the clustering clusters;
[0035] Match the extracted key video frames with the video covers in the video application platform to obtain the video covers that match each of the key video frames;
[0036] Respectively use the video covers that match each of the key video frames as the video covers that match the video frames in the corresponding clustering clusters.
[0037] In the above solution, the matching module is further configured to obtain the interaction data of each video cover in the video application platform within a target time period;
[0038] Filter out the video covers with interaction data greater than a quantity threshold from the video covers in the video application platform;
[0039] Match each of the video frames with the filtered video covers to obtain the video covers that match each of the video frames.
[0040] In the above solution, the second obtaining module is further configured to respectively obtain first interaction data before a first time point and second interaction data before a second time point for the video covers that match each of the video frames;
[0041] Among them, the first time point is later than the second time point;
[0042] The recommendation module is further configured to determine the difference between the first interaction data and the second interaction data of each of the video covers;
[0043] Take the ratio of the difference to the second interaction data as the growth rate of the interaction data of the corresponding video cover;
[0044] Recommend the video frame corresponding to the video cover whose growth rate meets the growth condition as the video cover of the target video.
[0045] In the above solution, the second acquisition module is further configured to respectively acquire the interaction data of the video covers matching each of the video frames within the target time period;
[0046] The recommendation module is further configured to predict the number of users interested in the corresponding video frame according to the interaction data of each of the video covers within the target time period;
[0047] Recommend the video frame whose number of users interested in the corresponding video frame meets the number of users as the video cover of the target video.
[0048] In the above solution, the recommendation module is further configured to, when the number of video covers matching the video frame is multiple, acquire the average value of the interaction data corresponding to the multiple video covers;
[0049] Based on the average value of the interaction data, select the video frame whose average value meets the average value condition, and perform the recommendation of the video cover for the target video.
[0050] In the above solution, the matching module is further configured to respectively match each of the video frames with the video covers in the video application platform to obtain the matching degrees between each of the video frames and the video covers in the video application platform;
[0051] When the maximum value of the matching degree reaches the matching degree threshold, use the video cover corresponding to the highest value of the matching degree as the video cover matching the video frame;
[0052] When the maximum value of the matching degree does not reach the matching degree threshold, it is determined that there is no video cover matching the video frame.
[0053] In the above solution, the recommendation module is further configured to screen out the video covers whose interaction data meets the interaction condition from the video covers matching each of the video frames according to the interaction data corresponding to each of the video covers;
[0054] Sort the video frames corresponding to the filtered video covers in descending order of the matching degree between the video frames and the video covers to obtain a video frame sequence;
[0055] Select a target number of video frames starting from the first video frame in the video frame sequence for recommending the video cover of the target video.
[0056] An embodiment of the present application provides a video cover recommendation device, including:
[0057] A receiving module, configured to receive a cover selection instruction for a target video during the editing process of the target video;
[0058] A display module, configured to present a cover selection interface in response to the cover selection instruction, and
[0059] In the cover selection interface, display the recommended video covers and the number of users interested in the video covers;
[0060] A selection module, configured to use the selected video cover as the video cover of the target video when a selection operation for a recommended video cover is received.
[0061] In the above solution, the display module is further configured to, when the recommended video covers are video frames in the target video and the number is multiple, display the thumbnails of the recommended video covers in the playing order of the video covers in the target video;
[0062] Display the change curve of the number of users interested in each video cover in the playing order of each video cover in the target video.
[0063] In the above solution, the display module is further configured to obtain the number of users interested in each video frame in the target video within a first time period;
[0064] Determine the first recommended video cover based on the number of users interested in each video frame;
[0065] In the cover selection interface, display the first recommended video cover and the number of users interested in the first video cover within the first time period;
[0066] Present selection function items corresponding to at least two time periods;
[0067] In response to a trigger operation for the selection function item corresponding to a second time period, switch the displayed first video cover to a second video cover, and switch the number of users to the number of users interested in the second video cover within the second time period.
[0068] An embodiment of the present application provides a computer device, including:
[0069] A memory for storing executable instructions;
[0070] A processor, when executing the executable instructions stored in the memory, implements the video cover recommendation method provided by the embodiment of the present application.
[0071] An embodiment of the present application provides a computer-readable storage medium storing executable instructions, which are used to cause a processor to implement the video cover recommendation method provided by the embodiment of the present application when executed.
[0072] The embodiment of the present application has the following beneficial effects:
[0073] Applying the above embodiment, by obtaining each video frame of the target video; matching each of the video frames with the video covers in the video application platform to obtain the video covers matching each of the video frames; respectively obtaining the interaction data corresponding to the video covers matching each of the video frames; and recommending the video covers of the target video according to the interaction data corresponding to each of the video covers. In this way, since the obtained video covers match the video frames, more video frames that users are interested in can be determined according to the interaction data corresponding to the video covers for video cover recommendation, improving the accuracy of video cover recommendation and thus increasing the click-through rate after the video is released. Description of the Drawings
[0074] Figure 1 is a schematic structural diagram of a video cover recommendation system provided by an embodiment of the present application;
[0075] Figure 2 is a schematic structural diagram of a computer device 500 provided by an embodiment of the present application;
[0076] Figure 3 is a schematic flowchart of a video cover recommendation method provided by an embodiment of the present application;
[0077] Figure 4 is a schematic flowchart of a video cover recommendation method provided by an embodiment of the present application;
[0078] Figure 5 is a schematic diagram of a cover selection interface provided by an embodiment of the present application;
[0079] Figure 6 is a schematic diagram of a cover selection interface provided by an embodiment of the present application;
[0080] Figure 7 is a schematic diagram of a cover selection interface provided by an embodiment of the present application;
[0081] Figure 8It is a schematic flowchart of a method for recommending video covers provided by an embodiment of the present application;
[0082] Figure 9 It is a schematic diagram of disassembling screen elements provided by an embodiment of the present application;
[0083] Figure 10 It is a schematic diagram of tag matching provided by an embodiment of the present application. Detailed implementation manners
[0084] In order to make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be construed as limitations on the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.
[0085] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0086] In the following description, the terms "first / second / third" are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when allowed, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0087] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0088] Before further elaborating on the embodiments of the present application, the nouns and terms involved in the embodiments of the present application are described. The nouns and terms involved in the embodiments of the present application are applicable to the following explanations.
[0089] 1) Responsive to, used to indicate the conditions or states on which the executed operations depend. When the dependent conditions or states are met, one or more executed operations can be real-time or can have a set delay; without special instructions, there is no limit on the execution order of the multiple executed operations.
[0090] 2) The client is any application (App, Application) that can run on a terminal, which can be a native application (Native APP), a web application (Web APP), or a hybrid application (Hybrid APP) on the terminal, and can be for various purposes, such as a social network client, a browser client, and a news client, etc.
[0091] See Figure 1 , Figure 1 which is a schematic architecture diagram of a recommended system for video covers provided by an embodiment of the present application. To support an exemplary application, terminals (exemplarily showing terminals 400-1 and 400-2) are connected to the server 200 through the network 300. The network 300 can be a wide area network, a local area network, or a combination of the two.
[0092] In actual implementation, a client of a video application platform is set on the terminal, and the user can perform operations such as video playing, video editing, and video publishing through this client.
[0093] The terminal is configured to receive a cover selection instruction for a target video during the process of the user editing the target video; in response to the cover selection instruction, send each video frame of the target video to the server;
[0094] The server 200 is configured to match each of the video frames with video covers in the video application platform to obtain video covers that match each of the video frames; respectively obtain interaction data corresponding to the video covers that match each of the video frames; predict the number of users interested in the corresponding video frames according to the interaction data corresponding to each of the video covers; determine the recommended video covers and the number of users interested in the corresponding video covers according to the number of users interested in each video frame, and return them to the terminal;
[0095] The terminal is configured to present a cover selection interface, and in the cover selection interface, display the recommended video covers and the number of users interested in the video covers; when receiving a selection operation for a recommended video cover, use the selected video cover as the video cover of the target video.
[0096] In some embodiments, the server 200 may be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery network (CDN), and big data and artificial intelligence platforms. The terminal may be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server may be directly or indirectly connected through wired or wireless communication methods, and there is no limitation in the embodiments of the present application.
[0097] See Figure 2 , Figure 2 FIG. is a schematic structural diagram of a computer device 500 provided by an embodiment of the present application. In practical applications, the computer device 500 may be Figure 1 the terminal or the server 200 in Figure 1 Taking the server 200 shown as an example, the computer device for implementing the video cover recommendation method of the embodiments of the present application will be described. Figure 2 The computer device 500 shown includes: at least one processor 510, a memory 550, at least one network interface 520, and a user interface 530. Each component in the computer device 500 is coupled together through a bus system 540. It can be understood that the bus system 540 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 540 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear description, in Figure 2 all kinds of buses are labeled as the bus system 540.
[0098] The processor 510 may be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or any conventional processor, etc.
[0099] The user interface 530 includes one or more output devices 531 that enable the presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 530 also includes one or more input devices 532, including user interface components that facilitate user input, such as keyboards, mice, microphones, touch screen displays, cameras, and other input buttons and controls.
[0100] The memory 550 can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical disc drives, etc. The memory 550 optionally includes one or more storage devices that are physically remote from the processor 510.
[0101] The memory 550 includes volatile memory or non-volatile memory, and may also include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), and the volatile memory can be random access memory (RAM). The memory 550 described in the embodiments of the present application is intended to include any suitable type of memory.
[0102] In some embodiments, the memory 550 is capable of storing data to support various operations. Examples of such data include programs, modules, and data structures, or subsets or supersets thereof, which are illustrated below.
[0103] The operating system 551 includes system programs for handling various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and handling hardware-based tasks;
[0104] The network communication module 552 is used to reach other computing devices via one or more (wired or wireless) network interfaces 520. Exemplary network interfaces 520 include: Bluetooth, Wi-Fi (Wireless Fidelity), and USB (Universal Serial Bus), etc.;
[0105] The presentation module 553 is used to enable the presentation of information (e.g., a user interface for operating peripheral devices and displaying content and information) via one or more output devices 531 associated with the user interface 530 (e.g., a display screen, a speaker, etc.);
[0106] The input processing module 554 is used to detect and translate one or more user inputs or interactions from one of one or more input devices 532.
[0107] In some embodiments, the video cover recommendation device provided by the embodiments of the present application can be implemented in software. Figure 2 The video cover recommendation device 555 stored in the memory 550 is shown, which can be software in the form of programs and plugins, etc., and includes the following software modules: a first acquisition module 5551, a matching module 5552, a second acquisition module 5553, and a recommendation module 5554. These modules are logical, and thus can be arbitrarily combined or further split according to the functions implemented.
[0108] The functions of each module will be described below.
[0109] In some other embodiments, the video cover recommendation device provided in the embodiments of the present application may be implemented in a hardware manner. As an example, the video cover recommendation device provided in the embodiments of the present application may be a processor in the form of a hardware decoding processor, which is programmed to execute the video cover recommendation method provided in the embodiments of the present application. For example, the processor in the form of a hardware decoding processor may employ one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.
[0110] Based on the above description of the video cover recommendation system in the embodiments of the present application, the video cover recommendation method provided in the embodiments of the present application will be described below. Refer to Figure 3 , Figure 3 is a flowchart of the video cover recommendation method provided in the embodiments of the present application; in some embodiments, this video cover recommendation method may be implemented independently by a terminal or a server, or implemented cooperatively by a terminal and a server. Taking the independent implementation by the server as an example, the video cover recommendation method provided in the embodiments of the present application includes:
[0111] Step 301: The server obtains each video frame of the target video.
[0112] In actual implementation, the server may obtain all the video frames included in the target video, or may obtain some video frames of the target video. For example, it may preprocess all the video frames included in the target video, filter out the dark frames, blurred frames, and / or low-quality frames in the target video to obtain clear, bright, and / or high-quality video frames. In this way, computing resources in the subsequent matching process can be saved, and the matching accuracy can be improved.
[0113] Step 302: Match each video frame with the video covers in the video application platform respectively to obtain the video covers that match each video frame.
[0114] In some embodiments, each video frame can be matched with a video cover in a video application platform in the following manner: perform image recognition on each video frame from at least one dimension to obtain the label information of each video frame in at least one dimension; obtain the label information of the video covers in the video application platform in at least one dimension; respectively match the label information of each video frame with the label information of the video covers in the video application platform to obtain the video covers that match each video frame.
[0115] Here, multiple label information are set for each dimension, and each label information corresponds to a different category. For example, for the "image content" dimension, multiple categories corresponding to the image content can be set, such as people, animals, landscapes, etc. Each category corresponds to a label information. For example, the label information corresponding to people is "NR-renwu", and the label information corresponding to animals is "NR-dongwu".
[0116] In actual implementation, by performing image recognition on each video frame from at least one dimension, the category to which each video frame belongs in this dimension is determined, and then, according to the corresponding relationship between the category and the label information, the label information of this video frame in this dimension is obtained.
[0117] As an example, for the "image content" dimension, perform image recognition on the image content in the video frame to determine the category to which the image content in this video frame belongs. If the image content category is people, determine the label information of this video frame in the "image content" dimension as "NR-renwu"; for the "image text" dimension, perform recognition on the text content in the video frame to determine the category to which the text in this video frame belongs. When the text in the video frame is the Chinese text "I love you", determine the label information of this video frame in the "image text" dimension as "WZ-zhongwen(woaini)".
[0118] In actual applications, the server can pre-determine the label information of all video covers in the video application platform and store it in the database. When matching is required, the label information of each video cover can be directly obtained from the database without repeatedly determining the label information of the video covers in the video application platform to avoid waste of computing resources. Here, the method for determining the label information of the video covers in the video application platform is the same as that for determining the label information of the video frames, both of which are determined through image recognition.
[0119] In some embodiments, each video frame can be matched with a video cover in a video application platform in the following manner: When the number of dimensions is at least two, the following operations are performed for each video frame: The label information of the video frame in each dimension is respectively matched with the label information of the video cover in the corresponding dimension to obtain the matching degrees of the video frame and the video cover in each dimension; the weight values of each dimension are obtained; based on the weight values of each dimension, the matching degrees of the video frame and the video cover in each dimension are weighted and summed to obtain the matching degree of the video frame and the video cover; based on the matching degree of the video frame and the video cover, the video covers in the video application platform are screened to obtain the video covers that match the video frame.
[0120] In actual implementation, a weight value can be set for each dimension, that is, the importance of each dimension to the matching degree can be different, which can be set manually here, or can be determined by a weight analysis algorithm, such as the principal component analysis method to determine the weights of each dimension. Here, when matching the label information of the video frame in each dimension with the label information of the video cover in the corresponding dimension, it can be a similarity match, that is, calculate the similarity between the label information of the video frame in each dimension and the label information of the video cover in the corresponding dimension, and then use the similarity as the matching degree of the video frame and the video cover in this dimension; it can also be a complete match, that is, when the label information of the video frame in a certain dimension is the same as the label information of the video cover in the corresponding dimension, the matching degree is determined to be 100, otherwise, the matching degree is determined to be 0.
[0121] As an example, when the number of dimensions is three, the dimensions include image content, image text, and composition. Among them, the weight of image content is 49%, the weight of image text is 49%, and the weight of composition is 2%. The label information of a certain video frame in these three dimensions is NR-renwu, WZ-zhongwen(woaini), GT-zuoyou respectively, and the label information of a certain video cover in these three dimensions is NR-renwu, WZ-zho ngwen(hahaha), GT-zuoyou respectively. The matching degree of this video frame and the video cover can be calculated to be 51.
[0122] In some embodiments, each video frame can be matched with a video cover in a video application platform in the following manner: Image features of each video frame are extracted to obtain the first image features of each video frame; image features of the video covers in the video application platform are extracted to obtain the second image features of the video covers; the first image features of each video frame are respectively matched with the second image features of the video covers to obtain the matching degrees between each video frame and the video covers; based on the matching degrees of the video frames and the video covers, the video covers in the video application platform are screened to obtain the video covers that match the video frames.
[0123] In actual implementation, the matching can be directly based on the image features of the video frames and video covers. Here, the image features can be simple color and edge histogram features, local binary pattern features, etc. By extracting these simple image features, computing resources can be saved; or the image features can be extracted through deep learning models such as convolutional neural networks. For example, classic convolutional neural network models such as AlexNet, VGGNet, Inception, etc. can be selected, and high-dimensional feature vectors of video frames or video covers are obtained as image features at the feature layer used to output features.
[0124] After obtaining the first image feature and the second image feature, calculate the similarity between the first image feature and the second image feature, such as Pearson correlation coefficient, cosine similarity, etc. Here, the specific calculation method of the similarity is not limited, and the calculated similarity is used as the matching degree between the corresponding video frame and the video cover; furthermore, based on the matching degree between the video frame and the video cover, the video covers in the video application platform are screened. For example, one or more video covers with the highest matching degree are selected as the video covers matching the corresponding video frame; or the video covers with the matching degree reaching the matching degree threshold are selected as the video covers matching the video frame.
[0125] In some embodiments, the following method can be used to match each video frame with the video covers in the video application platform: perform clustering processing on each video frame to obtain at least two clustering clusters; obtain the stillness degree of each video frame, where the stillness degree is used to indicate the numerical value of the motion energy of the video frame; based on the stillness degree of each video frame, extract key video frames from at least two clustering clusters; match the extracted key video frames with the video covers in the video application platform to obtain the video covers matching the key video frames; and respectively use the video covers matching the key video frames as the video covers matching the video frames in the corresponding clustering clusters.
[0126] Here, since many frames in a video are highly similar, to improve the matching efficiency, it is not necessary to match each video frame, and only the key video frames need to be matched.
[0127] In practical applications, the image features of each video frame can be extracted, the similarity between adjacent video frames is calculated based on the image features, and the video frames with the similarity less than the first similarity threshold are divided into the same clustering cluster; for each clustering cluster, obtain the stillness degree of each video frame in the clustering cluster. Here, the stillness degree is used to indicate the numerical value of the motion energy of the video frame. Since the motion compensation used in video compression will cause blurred artifacts, usually the images with high motion energy will also be more blurred, and the motion energy is inversely proportional to the stillness degree. The stillness degree of the video frames with low motion energy is high. Therefore, here the video frame with the lowest stillness degree can be selected as the key video frame in the clustering cluster.
[0128] In some embodiments, each video frame can be matched with the video covers in the video application platform in the following manner: obtain the interaction data of each video cover in the video application platform within the target time period; screen out the video covers with interaction data greater than the quantity threshold from the video covers in the video application platform; match each video frame with the screened video covers respectively to obtain the video covers matched with each video frame.
[0129] In actual implementation, if the interaction data is not greater than the quantity threshold, it indicates that the video cover cannot attract users. Then, even if it is the video cover matched with a certain video frame, this video frame is not the video cover that should be recommended to users. Based on this, this video cover does not need to be considered during the matching process to reduce the number of video covers to be matched and improve the matching efficiency.
[0130] Among them, the quantity threshold can be set by the system, such as it can be set to 0, 100, 1000, etc., or it can be selected by the creator of the target video, that is, the creator triggers the setting operation of the quantity threshold through the terminal to set the quantity threshold. Here, the determination method and specific value of the quantity threshold are not limited.
[0131] In some embodiments, each video frame can be matched with the video covers in the video application platform in the following manner: match each video frame with the video covers in the video application platform respectively to obtain the matching degrees between each video frame and the video covers in the video application platform; when the maximum value of the matching degrees reaches the matching degree threshold, use the video cover corresponding to the highest value of the matching degrees as the video cover matched with the video frame; when the maximum value of the matching degrees does not reach the matching degree threshold, determine that there is no video cover matched with the video frame.
[0132] In actual implementation, for each video frame, match it with the video covers in the video application platform to obtain multiple matching degrees, sort the video covers according to the matching degrees from high to low, and select the video cover with the highest matching degree according to the sorting result. If this matching degree reaches the matching degree threshold, it indicates that the matching is successful, and use the video cover with the highest matching degree as the video cover matched with this video frame; if this matching degree does not reach the matching degree threshold, it indicates that the matching fails, that is, there is no video cover matched with this video frame. Here, the video frames with failed matching can be directly filtered out, that is, this video frame will not be recommended as a video cover.
[0133] Step 303: Obtain the interaction data corresponding to the video covers matched with each video frame respectively.
[0134] The interactive data here refers to the interactive data generated by the interactive operations performed on the video cover or the video corresponding to the video cover. Among them, the interactive operation is used to represent the mutually dependent operation that occurs by spreading information through language or other means, including operations such as clicking, commenting, liking, collecting, forwarding, etc.; correspondingly, the interactive data can be the number of clicks, the number of comments, the number of likes, the number of collections, the number of forwards, etc. The interactive data obtained here can be one or more of the above-mentioned interactive data.
[0135] Step 304: Recommend the video cover of the target video according to the interactive data corresponding to each video cover.
[0136] In actual implementation, it can be directly sorting the video frames corresponding to the video covers according to the level of the interactive data, and then recommending the video cover of the target video according to the sorting result. For example, when the interactive data is the click volume, select the video frames corresponding to the 50 video covers with the highest click volume for recommendation; or, it can be further calculating according to the interactive data. For example, when the interactive data is the click volume, calculate the change trend of the click volume, and then recommend the video cover of the target video based on the change trend.
[0137] In some embodiments, the interactive data corresponding to the video covers matched with each video frame can be obtained in the following way: respectively obtain the first interactive data of the video covers matched with each video frame before the first time point and the second interactive data before the second time point; where the first time point is later than the second time point; correspondingly, the video cover of the target video can be recommended in the following way: determine the difference between the first interactive data and the second interactive data of each video cover; use the ratio of the difference to the second interactive data as the growth rate of the interactive data of the corresponding video cover; recommend the video frame corresponding to the video cover whose growth rate meets the growth condition as the video cover of the target video.
[0138] In actual implementation, that the growth rate meets the growth condition can be sorting the video covers according to the growth rate of the interactive data, starting from the video cover with the highest growth rate according to the sorting result, selecting the target number of video covers, and then recommending the video frames corresponding to the selected video covers as the video cover of the target video.
[0139] As an example, the first time point can be the current time point, the second time point can be 24 hours before the current time point, and the interactive data is the number of clicks on the video cover. Then, according to the formula: Calculate the click growth rate of the video cover within 24 hours, then sort the video covers from high to low according to the calculated click growth rate to obtain a video cover sequence. Starting from the first video cover in the video cover sequence, select 50 video covers, and use the video frames corresponding to these 50 video covers as the video covers of the target video for recommendation.
[0140] In some embodiments, the interaction data corresponding to the video covers matched with each video frame can be obtained respectively in the following manner: obtain the interaction data of the video covers matched with each video frame within the target time period; correspondingly, the video covers of the target video can be recommended in the following manner: predict the number of users interested in the corresponding video frame according to the interaction data of each video cover within the target time period; use the video frames for which the number of users interested in the corresponding video frame meets the user number condition as the video covers of the target video for recommendation.
[0141] In actual implementation, the target time period can be determined first. The target time period can be set by the system or selected by the user through the client and then sent to the server by the client; the server obtains the target time period and then counts the interaction data of the video covers matched with each video frame within the target time period; after the interaction data is counted, the number of users interested in the corresponding video frame can be predicted. For example, when the interaction data is the number of clicks, the number of clicks can be directly used as the number of users interested in the corresponding video frame; when the interaction data is the number of likes and the number of comments, the number of likes and the number of comments can be weighted and summed to obtain the number of users interested in the corresponding video frame.
[0142] In some embodiments, the video covers of the target video can be recommended in the following manner: when the number of video covers matched with the video frame is multiple, obtain the average value of the interaction data corresponding to the multiple video covers; based on the average value of the interaction data, select the video frames whose average value meets the average value condition for recommending the video covers of the target video.
[0143] In actual implementation, when the number of video covers matched with the video frame is multiple, the interaction data corresponding to each video cover can be obtained, and then the average value of these interaction data can be calculated to avoid the problem that the click volume of a single video cover is too high, resulting in inaccurate final video cover recommendation; after obtaining the average value, select the video frames whose average value meets the average value condition, such as selecting the target number of video frames with the highest corresponding average value, or selecting the video frames whose corresponding average value reaches the average value threshold.
[0144] Here, there may be multiple video covers that match some video frames, and only one video cover that matches other video frames. For these two types of video frames, different processing methods can be adopted. For example, when there are multiple video covers that match a video frame, the above method of calculating the average value can be used; when there is only one video cover that matches a video frame, based on the interaction data, select the video frames whose interaction data meets the interaction conditions for recommending the video covers of the target video.
[0145] In some embodiments, the video cover of the target video can be recommended in the following way: according to the interaction data corresponding to each video cover, screen out the video covers whose interaction data meets the interaction conditions from the video covers that match each video frame; sort the video frames corresponding to the screened video covers in descending order of the matching degree between the video frame and the video cover to obtain a video frame sequence; start from the first video frame in the video frame sequence and select a target number of video frames for recommending the video cover of the target video.
[0146] The interaction data meeting the interaction conditions here can be that the interaction data reaches the interaction threshold, or it can be to select a preset number of video covers with the highest interaction data.
[0147] In actual implementation, since relying solely on the interaction data cannot fully and accurately reflect how many people like a certain video frame. For example, if the interaction data of a certain video cover is very high and it matches video frame 1, but its matching degree with video frame 1 is only 50, then the interaction data of this video cover cannot well reflect how many people like video frame 1. Based on this, the matching degree and the interaction data can be combined to recommend the video cover of the target video to further improve the accuracy of video cover recommendation.
[0148] Applying the above embodiments, by obtaining each video frame of the target video; matching each video frame with the video covers in the video application platform to obtain the video covers that match each video frame; respectively obtaining the interaction data corresponding to the video covers that match each video frame; and recommending the video cover of the target video according to the interaction data corresponding to each video cover. In this way, since the obtained video covers match the video frames, it is possible to determine more video frames that users are interested in for recommending the video cover based on the interaction data corresponding to the video covers, improve the accuracy of video cover recommendation, and thus increase the click-through rate after the video is released.
[0149] The following describes the video cover recommendation method provided by the embodiments of the present application. Refer to Figure 4 , Figure 4It is a schematic flowchart of the method for recommending a video cover provided by an embodiment of the present application; in some embodiments, the method for recommending a video cover can be implemented independently by a terminal or jointly by a terminal and a server. Taking the independent implementation by the terminal as an example, the method for recommending a video cover provided by an embodiment of the present application includes:
[0150] Step 401: During the process of editing a target video, the terminal receives a cover selection instruction for the target video.
[0151] In actual implementation, a client corresponding to the video application platform is set on the terminal, and the user can edit the target video through the client corresponding to the video application platform. Here, the user triggers a cover selection instruction for the target video through the client corresponding to the video application platform, and the terminal receives this cover selection instruction.
[0152] Step 402: In response to the cover selection instruction, a cover selection interface is presented, and in the cover selection interface, the recommended video covers and the number of users interested in the video covers are displayed.
[0153] Here, in the cover selection interface, the recommended video covers and the number of users interested in the video covers are displayed, enabling the user to intuitively know the number of users interested in the video covers, and then select a video cover based on this number of users, which can improve the click-through rate after the target video is released.
[0154] In some embodiments, the recommended video covers and the number of users interested in the video covers can be displayed in the following manner: when the recommended video covers are video frames in the target video and the number is multiple, the thumbnails of the recommended video covers are displayed in the playing order of the video covers in the target video; the change curves of the number of users interested in each video cover are displayed in the playing order of the video covers in the target video.
[0155] Here, the recommended video covers can be video frames in the target video. In actual implementation, it can be to match each video frame of the target video with the video covers in the video application platform to obtain the video covers matching each of the video frames; the interaction data corresponding to each of the video covers matching the video frames is obtained respectively; according to the interaction data corresponding to each of the video covers, a certain number of video frames are selected as the video covers of the target video for recommendation.
[0156] In practical applications, when the number of recommended video covers is multiple, the thumbnails of the recommended video covers and the corresponding change curves are displayed in the playing order of the video covers in the target video. Here, the change curves enable the user to perform a global preview of the number of users interested in all the video covers.
[0157] As an example, Figure 5 is a schematic diagram of a cover selection interface provided by an embodiment of the present application. Refer to Figure 5 , the video frame with the largest number of users interested is selected by default, and the video frame 501 is previewed in a large image, and the number of users interested in the video frame 502 is displayed in text form; the change curve 504 of the number of users interested in each video frame is displayed above the thumbnail 503; in this way, users can intuitively see the comparison of the number of users interested in each video frame.
[0158] In some embodiments, the recommended video cover and the number of users interested in the video cover can be displayed in the cover selection interface in the following manner: obtain the number of users interested in each video frame in the target video during the first time period; determine the recommended first video cover based on the number of users interested in each video frame; in the cover selection interface, display the recommended first video cover and the number of users interested in the first video cover during the first time period; the method further includes: presenting selection function items corresponding to at least two time periods; in response to a trigger operation for the selection function item corresponding to the second time period, switch the displayed first video cover to the second video cover, and switch the number of users to the number of users interested in the second video cover during the second time period.
[0159] In actual implementation, the number of users interested in each video frame at different times is provided to switch the number of users interested in each video frame in different time periods. Here, the first time period is the default. First, the first video cover corresponding to the first time period and the number of users interested in the first video cover during the first time period are displayed; upon receiving a trigger operation for the selection function item of a certain time period, the video cover and the corresponding number of users are switched based on the selected time period.
[0160] Here, in addition to switching the number of users in different time periods, the growth rate of the number of users in different time periods can also be switched.
[0161] As an example, Figure 6 is a schematic diagram of a cover selection interface provided by an embodiment of the present application. Refer to Figure 6, first, a global navigation 601 showing the number of users interested in each video frame 24 hours ago is presented. It is the number of users interested in each video frame determined based on the click situation of the video covers 24 hours ago. And selection function items corresponding to different time periods are presented. When the selection function item 602 corresponding to "current" is clicked, a global navigation 603 showing the number of users interested in each video frame within 24 hours is presented. It is the number of users interested in each video frame determined based on the click situation of the video covers within 24 hours. When the selection function item 604 corresponding to "future" is clicked, combining the click situations of the video covers 24 hours ago and within 24 hours, the growth rate of the number of users interested in the video frames in the future is judged, and a global navigation 605 of the growth rate of the number of users interested in each video frame is presented.
[0162] Step 403: When a selection operation for the recommended video cover is received, use the selected video cover as the video cover of the target video.
[0163] In actual implementation, the selection operation can be triggered by a sliding operation or a click operation. When the selection operation for the video cover is triggered by a sliding operation, during the execution of the sliding operation, a large - image preview of the video frame indicated by the sliding operation is performed, and the number of users interested in the corresponding video frame is displayed in text form. When the sliding operation stops, determine the video frame indicated when the sliding operation stops as the selected video frame.
[0164] When the selection operation for the video cover is triggered by a click operation, upon receiving the click operation, a large - image preview of the video frame indicated by the click operation of the large - image preview is performed, and the number of users interested in the video frame is displayed, and determine the video frame as the selected video frame.
[0165] As an example, Figure 7 is a schematic diagram of the cover selection interface provided by the embodiment of the present application. Refer to Figure 7 , the user clicks on the thumbnail 701 of the target video frame, performs a large - image preview of the target video frame 703, and at the same time, the number of users interested in the target video frame 702 is displayed above the thumbnail 701.
[0166] Applying the above - mentioned embodiment, by enabling the user to intuitively understand how many users are interested in the recommended video cover and selecting the cover for the target video based on the recommended video cover, the click - through rate after the release of the target video can be improved.
[0167] Next, an exemplary application of the embodiments of the present application in an actual application scenario will be described. In actual implementation, all video covers in the video application platform are classified, tags are labeled for each video cover according to the classification results, and the click-through rate of each video cover is counted; when the user selects the video cover of the corresponding target video, each video frame of the target video is matched with all video covers in the video application platform to determine the video covers that match each video frame, and then, based on the click-through rate of the video covers, the number of users interested in each video frame is predicted; according to the number of users interested in each video frame, video covers are recommended, such as recommending the video frame with the highest number of interested users to the creator as the video cover, so as to improve the creation efficiency of the creator and increase the click-through rate and view volume of the creator's video.
[0168] As an example, refer to Figure 5 , by default, the video frame with the largest number of interested users is selected, and the video frame 501 is previewed in a large picture, and the number of users interested in the video frame 502 is displayed in text form; above the thumbnail 503, a change curve (global navigation) 504 of the number of users interested in each video frame is displayed; in this way, the user can intuitively see the comparison of the number of users interested in each video frame.
[0169] Here, based on the presented thumbnails, the user can select other video frames by swiping (the user gesture swipes horizontally left and right on the video frame thumbnail) or clicking. After selecting other video frames, the number of users interested in the selected video frame is displayed. For example, refer to Figure 7 , the user clicks on the thumbnail 701 of the target video frame, the target video frame 703 is previewed in a large picture, and at the same time, the number of users interested in the target video frame 702 is displayed above the thumbnail 701.
[0170] In some embodiments, the number of users interested in each video frame at different time periods can be provided, and time trend judgment can also be added to switch the number of users interested in each video frame or the trend at different time periods.
[0171] As an example, Figure 6 is a schematic diagram of the cover selection interface provided by the present application. Refer to Figure 6, first, a global navigation 601 showing the number of users interested in each video frame 24 hours ago is presented. It is the number of users interested in each video frame determined based on the click situation of the video covers 24 hours ago; and selection function items corresponding to different time periods are presented. When the selection function item 602 corresponding to "Current" is clicked, a global navigation 603 showing the number of users interested in each video frame within 24 hours is presented. It is the number of users interested in each video frame determined based on the click situation of the video covers within 24 hours; when the selection function item 604 corresponding to "Future" is clicked, the growth rate of the number of users interested in the video frames in the future is judged by combining the click situations of the video covers 24 hours ago and within 24 hours, and a global navigation 605 of the growth rate of the number of users interested in each video frame is presented.
[0172] Next, the technical implementation principle of the video cover recommendation method provided by the embodiments of the present application will be described. Figure 8 is a schematic flowchart of the video cover recommendation method provided by the embodiments of the present application. Refer to Figure 8 , the video cover recommendation method provided by the embodiments of the present application includes:
[0173] Step 801: Tag all video covers in the video application platform.
[0174] In actual implementation, first, the picture elements in the video cover can be disassembled into multiple dimensions. Here, taking the video cover being disassembled into 3 dimensions as an example, Figure 9 is a schematic diagram of the disassembly of picture elements provided by the embodiments of the present application. Refer to Figure 9 , the video cover is disassembled into 3 dimensions: image content, image text, and composition; the image content is disassembled by element: people, animals, trees, flowers and plants, rivers and lakes, mountains, sky and white clouds, food, tools, etc.; the image text is disassembled by language: Chinese, English, Spanish, etc.; the composition is disassembled by layout: left - right layout, up - down layout, centered, circular, etc.
[0175] Then, according to the corresponding relationship between the elements and the tags, each dimension of the elements is tagged. Table 1 is a table of the corresponding relationship between elements and tags provided by the embodiments of the present application. Refer to Table 1. For example, if the image content is a person, it is tagged as NR - renwu, if the image content is an animal, it is tagged as NR - dongwu; if there is Chinese in the image, it is tagged as: WZ - zhongwen, if the text is English, it is tagged as: WZ - yingwen; if the image composition is a left - right composition: GT - zuoyou, an up - down composition, it is tagged as GT - shangxia.
[0176] Table 1
[0177]
[0178]
[0179] Next, combine the markings of each element in the video cover to obtain the label corresponding to the video cover. For example, if there are two people side by side in the video cover and the Chinese character "I love you" is written on the cover, then the label of this video cover is: NR-renwu-WZ-zhongwen(woaini)-GT-zuoyou.
[0180] Step 802: Count the number of clicks on all video covers in the video application platform.
[0181] Here, count the number of clicks on each video. If the same video cover is clicked by the same person multiple times, it is counted as 1 person. For example, if the cover NR-renwu-WZ-zhongwen(woaini)-GT-zuoyou is clicked by 100 people, append this number of clicks after the label, that is, NR-renwu-WZ-zhongwen(woaini)-GT-zuoyou-(100), and update it in real time.
[0182] Step 803: Obtain the target video selected by the user.
[0183] Here, the user can select a video stored locally as the target video.
[0184] Step 804: Tag all video frames of the target video.
[0185] Here, automatically detect each video frame in the target video and tag each video frame. Among them, the tagging method is the same as the method of tagging all video covers in the video application platform.
[0186] Step 805: Perform label matching to calculate the number of users interested in each video frame in the target video or the growth trend of the number of users.
[0187] Here, different methods are used for matching and calculation for different time periods.
[0188] When the time period is before the target time point, such as 24 hours ago, first, screen the video covers in the video application platform. The primary selection condition is: the number of clicks > 0; then, according to the calculation formula of the matching score: matching score = 100 * 49% (weight of image content) * image content matching result + 100 * 49% (weight of image text) * image text matching result + 100 * 2% (weight of composition) * composition matching result, calculate the matching score; finally, determine the video covers matching each video frame according to the matching score, and determine the number of users interested in the corresponding video frame according to the number of clicks on the video cover.
[0189] Among them, the matching result is 0 or 1. A matching result of 0 indicates non - matching, and a matching result of 1 indicates matching.
[0190] When the time period is the time period between the target time point and the current time point, such as within 24 hours, first, filter the video covers in the video application platform. The primary selection condition is: the number of clicks after the target time point > 0; then, according to the calculation formula of the matching score: matching score = 100 * 49% (weight of image content) * image content matching result + 100 * 49% (weight of image text) * image text matching result + 100 * 2% (weight of composition) * composition matching result, calculate the matching score; finally, determine the video covers matching each video frame according to the matching score, and determine the number of users interested in the corresponding video frame according to the click - through rate of the video cover.
[0191] When the time period is after the current time point, first, filter the video covers in the video application platform. The primary selection condition is: the number of clicks > 0; then, according to the calculation formula of the matching score: matching score = 100 * 49% (weight of image content) * image content matching result + 100 * 49% (weight of image text) * image text matching result + 100 * 2% (weight of composition) * composition matching result, calculate the matching score; then, determine the video covers matching each video frame according to the matching score; finally, calculate the click - through growth rate of the video cover, such as According to the click - through growth rate of the video cover, determine the growth rate of the number of users interested in the corresponding video frame.
[0192] Taking the time period before 24 hours as an example, first, filter out the video covers in the video application platform with the number of clicks greater than zero; then, perform label matching between each video frame of the target video and the filtered video covers. For example, Figure 10 is the schematic diagram of label matching provided by the embodiment of the present application. Refer to Figure 10 For each video frame, match it with all the filtered video covers. For example, the label of video frame 1 is matched with the labels of video covers 1 - n in the video application platform. Here, calculate the matching score according to the calculation formula of the matching score. The higher the matching score, the higher the matching degree. The video cover with the highest matching score is used as the video cover matching this video frame. It should be noted that the matching score must reach at least 49 points to be a successful match.
[0193] For example, in a video application platform, there are covers: Cover 1 - NR - renwu - WZ - zhongwen(woain i) - GT - zuoyou - (100), Cover 2 - NR - renwu - WZ - zhongwen(hahaha) - GT - zuoyou - (150); One frame of the target video: Video Frame X - NR - renwu - WZ - zhongwen(woaini) - GT - zuoyou, after matching with Cover 1, the score is 100 points, and after matching with Cover 2, the score is 51 points. Then the video frame successfully matches with Cover 1, and the number of users interested in Video Frame X is 100.
[0194] Step 806: Screen out the target number of video frames with the highest number of interested users or the highest growth rate.
[0195] The target number here is user - defined.
[0196] Step 807: Visually display the number of interested users for each frame.
[0197] Applying the above - mentioned embodiments, when creating a video cover, the creator can intuitively understand how many users are interested in the selected video cover. Whether for experienced creators or ordinary creators, it can improve their creation efficiency, increase the click - through rate after the video is released, and enhance the user experience.
[0198] Next, continue to describe the exemplary structure of the video cover recommendation device 555 provided by the embodiments of the present application as software modules. In some embodiments, as Figure 2 shown, the software modules in the video cover recommendation device 555 stored in the memory 550 may include:
[0199] The first acquisition module 5551 is used to acquire each video frame of the target video;
[0200] The matching module 5552 is used to match each of the video frames with the video covers in the video application platform respectively to obtain the video covers that match each of the video frames;
[0201] The second acquisition module 5553 is used to acquire the interaction data corresponding to the video covers that match each of the video frames respectively;
[0202] The recommendation module 5554 is used to recommend the video covers for the target video according to the interaction data corresponding to each of the video covers.
[0203] In some embodiments, the matching module 5552 is further used to perform image recognition on each of the video frames from at least one dimension to obtain the label information of each of the video frames in at least one of the dimensions;
[0204] Obtain the label information of the video cover in the video application platform in at least one of the dimensions;
[0205] Match the label information of each of the video frames with the label information of the video cover in the video application platform respectively to obtain the video cover that matches each of the video frames.
[0206] In some embodiments, the matching module 5552 is further configured to, when the number of the dimensions is at least two, perform the following operations for each of the video frames:
[0207] Match the label information of the video frame in each dimension with the label information of the video cover in the corresponding dimension respectively to obtain the matching degrees of the video frame and the video cover in each dimension;
[0208] Obtain the weight values of each of the dimensions;
[0209] Based on the weight values of each of the dimensions, perform weighted summation on the matching degrees of the video frame and the video cover in each of the dimensions to obtain the matching degree of the video frame and the video cover;
[0210] Based on the matching degree of the video frame and the video cover, screen the video covers in the video application platform to obtain the video cover that matches the video frame.
[0211] In some embodiments, the matching module 5552 is further configured to extract image features from each of the video frames to obtain the first image features of each of the video frames;
[0212] Extract image features from the video covers in the video application platform to obtain the second image features of the video covers;
[0213] Perform similarity matching on the first image features of each of the video frames and the second image features of the video covers respectively to obtain the matching degrees between each of the video frames and the video covers;
[0214] Based on the matching degrees of the video frames and the video covers, screen the video covers in the video application platform to obtain the video covers that match the video frames.
[0215] In some embodiments, the matching module 5552 is further configured to perform clustering processing on each of the video frames to obtain at least two clustering clusters;
[0216] Obtain the stillness degrees of each of the video frames, where the stillness degree is used to indicate the numerical value of the motion energy of the video frame;
[0217] Based on the stillness degrees of each of the video frames, extract key video frames from at least two of the clustering clusters;
[0218] Match each of the extracted key video frames with the video covers in the video application platform to obtain video covers that match each of the key video frames;
[0219] Respectively use the video covers that match each of the key video frames as the video covers that match the video frames in the corresponding clustering cluster.
[0220] In some embodiments, the matching module 5552 is further configured to obtain the interaction data of each video cover in the video application platform within a target time period;
[0221] Filter out the video covers from the video covers in the video application platform whose interaction data is greater than a quantity threshold;
[0222] Match each of the video frames with the filtered video covers to obtain video covers that match each of the video frames.
[0223] In some embodiments, the second obtaining module 5553 is further configured to respectively obtain the first interaction data of the video covers that match each of the video frames before a first time point and the second interaction data before a second time point;
[0224] Wherein, the first time point is later than the second time point;
[0225] The recommendation module 5554 is further configured to determine the difference between the first interaction data and the second interaction data of each of the video covers;
[0226] Use the ratio of the difference to the second interaction data as the growth rate of the interaction data of the corresponding video cover;
[0227] Recommend the video frames corresponding to the video covers whose growth rate meets the growth condition as the video covers of the target video.
[0228] In some embodiments, the second obtaining module 5553 is further configured to respectively obtain the interaction data of the video covers that match each of the video frames within a target time period;
[0229] The recommendation module is further configured to predict the number of users interested in the corresponding video frames according to the interaction data of each of the video covers within the target time period;
[0230] Recommend the video frames for which the number of users interested meets the number of users as the video covers of the target video.
[0231] In some embodiments, the recommendation module 5554 is further configured to, when there are multiple video covers matching the video frame, obtain the average value of the interaction data corresponding to the multiple video covers;
[0232] Based on the average value of the interaction data, select video frames whose average value meets the average value condition, and perform recommendation of video covers for the target video.
[0233] In some embodiments, the matching module 5552 is further configured to match each of the video frames with video covers in a video application platform to obtain the matching degree between each of the video frames and the video covers in the video application platform;
[0234] When the maximum value of the matching degree reaches the matching degree threshold, use the video cover corresponding to the highest value of the matching degree as the video cover matching the video frame;
[0235] When the maximum value of the matching degree does not reach the matching degree threshold, it is determined that there is no video cover matching the video frame.
[0236] In some embodiments, the recommendation module 5554 is further configured to, according to the interaction data corresponding to each of the video covers, screen out video covers whose interaction data meets the interaction condition from the video covers matching each of the video frames;
[0237] Sort the video frames corresponding to the screened video covers in descending order of the matching degree between the video frames and the video covers to obtain a video frame sequence;
[0238] Select a target number of video frames starting from the first video frame in the video frame sequence, and perform recommendation of video covers for the target video.
[0239] An embodiment of the present application provides a video cover recommendation device, including:
[0240] A receiving module, configured to receive a cover selection instruction for a target video during the process of editing the target video;
[0241] A display module, configured to respond to the cover selection instruction, present a cover selection interface, and
[0242] In the cover selection interface, display the recommended video covers and the number of users interested in the video covers;
[0243] A selection module, configured to, when a selection operation for a recommended video cover is received, use the selected video cover as the video cover of the target video.
[0244] In some embodiments, the display module is further configured to, when the recommended video covers are video frames in the target video and the number thereof is multiple, display thumbnails of the recommended video covers in the playing order of the video covers in the target video;
[0245] Display a change curve of the number of users interested in each of the video covers in the playing order of the video covers in the target video.
[0246] In some embodiments, the display module is further configured to obtain the number of users interested in each video frame in the target video within a first time period;
[0247] Determine a recommended first video cover based on the number of users interested in each of the video frames;
[0248] In the cover selection interface, display the recommended first video cover and the number of users interested in the first video cover within the first time period;
[0249] Present selection function items corresponding to at least two time periods;
[0250] In response to a trigger operation on the selection function item corresponding to a second time period, switch the displayed first video cover to a second video cover, and switch the number of users to the number of users interested in the second video cover within the second time period.
[0251] An embodiment of the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method for recommending a video cover as described above in the embodiments of the present application.
[0252] An embodiment of the present application provides a computer-readable storage medium storing executable instructions, where the executable instructions are stored therein. When the executable instructions are executed by a processor, the processor will be caused to execute the method provided in the embodiments of the present application, for example, Figure 3 The method shown.
[0253] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disc, or CD-ROM; or may be various devices including one or any combination of the above memories.
[0254] In some embodiments, the executable instructions may be in the form of a program, software, a software module, a script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including being deployed as a stand-alone program or being deployed as a module, a component, a subroutine, or other unit suitable for use in a computing environment.
[0255] As an example, the executable instructions may or may not correspond to a file in a file system, may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, stored in a single file dedicated to the program being discussed, or, stored in multiple cooperating files (such as files that store one or more modules, subroutines, or portions of code).
[0256] As an example, the executable instructions may be deployed to execute on one computing device, or on multiple computing devices located at one site, or, on multiple computing devices distributed across multiple sites and interconnected by a communication network.
[0257] As described above, the above are only embodiments of the present application and are not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and scope of the present application are all included in the protection scope of the present application.
Claims
1. A method for recommending video covers, characterized in that, Including: Performing clustering processing on each video frame of the target video to obtain at least two clustering clusters; Obtaining the stillness degree of each of the video frames, where the stillness degree is used to indicate the value of the motion energy of the video frame; Extracting key video frames from at least two of the clustering clusters based on the stillness degree of each of the video frames; Matching each of the extracted key video frames with video covers in a video application platform to obtain video covers matching the corresponding clustering clusters; Among them, the video covers matching the corresponding clustering clusters are obtained by label information matching, or by image feature matching, or determined based on the matching degree and the matching degree threshold; Respectively obtaining the interaction data corresponding to the video covers matching each of the clustering clusters; Performing recommendation of the video covers for the target video according to the interaction data corresponding to each of the video covers.
2. The method according to claim 1, wherein The step of matching each of the extracted key video frames with video covers in a video application platform to obtain video covers matching the corresponding clustering clusters includes: Performing image recognition on each of the key video frames from at least one dimension to obtain the label information of each of the key video frames in at least one of the dimensions; Obtaining the label information of the video covers in the video application platform in at least one of the dimensions; Respectively matching the label information of each of the key video frames with the label information of the video covers in the video application platform to obtain video covers matching the corresponding clustering clusters.
3. The method according to claim 2, wherein The step of respectively matching the label information of each of the key video frames with the label information of the video covers in the video application platform to obtain video covers matching the corresponding clustering clusters includes: When the number of the dimensions is at least two, the following operations are performed for each of the key video frames: Respectively matching the label information of the key video frame in each dimension with the label information of the video cover in the corresponding dimension to obtain the matching degree of the key video frame and the video cover in each dimension; Obtaining the weight values of each of the dimensions; Based on the weight values of each of the dimensions, performing weighted summation on the matching degrees of the key video frame and the video cover in each of the dimensions to obtain the matching degree of the key video frame and the video cover; Based on the matching degree of the key video frame and the video cover, screening the video covers in the video application platform to obtain video covers matching the corresponding clustering clusters.
4. The method according to claim 1, characterized in that, The step of matching each of the extracted key video frames with video covers in a video application platform to obtain video covers matching the corresponding clustering clusters includes: Performing image feature extraction on each of the key video frames to obtain the first image features of each of the key video frames; Performing image feature extraction on the video covers in the video application platform to obtain the second image features of the video covers; Performing similarity matching on the first image features of each of the key video frames and the second image features of the video covers respectively to obtain the matching degree between each of the key video frames and the video covers; Based on the matching degree of the key video frame and the video cover, screening the video covers in the video application platform to obtain video covers matching the corresponding clustering clusters.
5. The method according to claim 1, wherein Matching each of the extracted key video frames with the video covers in the video application platform to obtain video covers matching the corresponding clustering clusters, includes: Matching each of the extracted key video frames with the video covers in the video application platform to obtain video covers matching each of the key video frames; Respectively using the video covers matching each of the key video frames as the video covers matching the video frames in the corresponding clustering cluster.
6. The method according to claim 1, wherein Matching each of the extracted key video frames with the video covers in the video application platform to obtain video covers matching the corresponding clustering clusters, includes: Obtaining the interaction data of each video cover in the video application platform within the target time period; Filtering out the video covers with interaction data greater than the quantity threshold from the video covers in the video application platform; Matching each of the key video frames with the filtered video covers to obtain video covers matching the corresponding clustering clusters.
7. The method according to claim 1, characterized in that, Respectively obtaining the interaction data corresponding to the video covers matching each of the clustering clusters, includes: Respectively obtaining the first interaction data of the video covers matching each of the key video frames before the first time point and the second interaction data before the second time point; Wherein, the first time point is later than the second time point; Recommending the video covers for the target video according to the interaction data corresponding to each of the video covers, includes: Determining the difference between the first interaction data and the second interaction data of each of the video covers; Taking the ratio of the difference to the second interaction data as the growth rate of the interaction data of the corresponding video cover; Recommending the key video frames corresponding to the video covers with growth rates meeting the growth condition as the video covers of the target video.
8. The method according to claim 1, wherein Respectively obtaining the interaction data corresponding to the video covers matching each of the clustering clusters, includes: Respectively obtaining the interaction data of the video covers matching each of the key video frames within the target time period; Recommending the video covers for the target video according to the interaction data corresponding to each of the video covers, includes: Predicting the number of users interested in the corresponding key video frames according to the interaction data of each of the video covers within the target time period; Recommending the key video frames with the number of users interested in the corresponding key video frames meeting the user number condition as the video covers of the target video.
9. The method according to claim 1, characterized in that, Recommending the video covers for the target video according to the interaction data corresponding to each of the video covers, includes: When the number of video covers matching the key video frame is multiple, obtaining the average value of the interaction data corresponding to the multiple video covers; Based on the average value of the interaction data, selecting the key video frames with the average value meeting the average value condition for recommending the video covers for the target video.
10. The method according to claim 1, characterized in that, Matching each of the extracted key video frames with the video covers in the video application platform to obtain video covers matching the corresponding clustering clusters, includes: Match each of the key video frames with the video covers in the video application platform to obtain the matching degrees between each of the key video frames and the video covers in the video application platform; When the maximum value of the matching degrees reaches the matching degree threshold, use the video cover corresponding to the highest value of the matching degrees as the video cover matching the corresponding clustering cluster; When the maximum value of the matching degrees does not reach the matching degree threshold, determine that there is no video cover matching the corresponding clustering cluster.
11. The method according to claim 1, characterized in that, The recommending the video cover for the target video according to the interaction data corresponding to each of the video covers includes: According to the interaction data corresponding to each of the video covers, screen out the video covers whose interaction data meet the interaction conditions from the video covers matching each of the key video frames; Sort the key video frames corresponding to the screened video covers in descending order of the matching degrees between the key video frames and the video covers to obtain a video frame sequence; Select a target number of key video frames starting from the first key video frame in the video frame sequence to recommend the video cover for the target video.
12. A method for recommending a video cover, characterized in that including: During the process of editing the target video, receive a cover selection instruction for the target video; In response to the cover selection instruction, present a cover selection interface, and In the cover selection interface, display the recommended video covers and the number of users interested in the video covers; Wherein, the recommended video covers are determined according to the interaction data corresponding to each of the video covers matching the corresponding clustering clusters; the video covers matching the corresponding clustering clusters are obtained by matching the key video frames with the video covers in the video application platform, the key video frames are extracted from at least two clustering clusters based on the stillness degrees of the video frames of the target video, the stillness degree is used to indicate the value of the motion energy of the video frame, and the clustering clusters are obtained by clustering the video frames of the target video; When a selection operation for the recommended video cover is received, use the selected video cover as the video cover of the target video.
13. The method according to claim 12, wherein The displaying the recommended video covers and the number of users interested in the video covers in the cover selection interface includes: When the recommended video covers are video frames in the target video and the number is multiple, display the thumbnails of the recommended video covers in the playing order of the video covers in the target video; Display the change curves of the number of users interested in each of the video covers in the playing order of the video covers in the target video.
14. The method according to claim 12, wherein The displaying the recommended video covers and the number of users interested in the video covers in the cover selection interface includes: Obtain the number of users interested in each video frame in the target video within the first time period; Based on the number of users interested in each of the video frames, determine the recommended first video cover; In the cover selection interface, display the recommended first video cover and the number of users interested in the first video cover within the first time period; The method further includes: Present selection function items corresponding to at least two time periods; In response to a trigger operation on a selection function item corresponding to a second time period, the displayed first video cover is switched to a second video cover, and the number of users is switched to the number of users interested in the second video cover during the second time period.
15. A device for recommending video covers, characterized in that, It includes: A matching module for performing clustering processing on each video frame of a target video to obtain at least two clustering clusters; Obtaining the stillness degree of each of the video frames, where the stillness degree is used to indicate the value of the motion energy of the video frame; Based on the stillness degrees of each of the video frames, extracting key video frames from at least two of the clustering clusters; Matching each of the extracted key video frames with video covers in a video application platform to obtain video covers matching the corresponding clustering clusters; Among them, the video cover matching the corresponding clustering cluster is obtained by tag information matching, or by image feature matching, or determined based on a matching degree and a matching degree threshold; A second obtaining module for respectively obtaining the interaction data corresponding to the video covers matching each of the clustering clusters; A recommendation module for recommending video covers of the target video according to the interaction data corresponding to each of the video covers.
16. A video cover recommendation device, characterized in that, It includes: A receiving module for receiving a cover selection instruction for a target video during the process of editing the target video; A presenting module for presenting a cover selection interface in response to the cover selection instruction, and In the cover selection interface, displaying the recommended video covers and the number of users interested in the video covers; Among them, the recommended video covers are determined according to the interaction data corresponding to the video covers matching the corresponding clustering clusters; the video covers matching the corresponding clustering clusters are obtained by matching key video frames with video covers in a video application platform, and the key video frames are extracted from at least two clustering clusters based on the stillness degrees of each video frame of the target video, where the stillness degree is used to indicate the value of the motion energy of the video frame, and the clustering clusters are obtained by performing clustering processing on each video frame of the target video; A selection module for using the selected video cover as the video cover of the target video when a selection operation on the recommended video cover is received.
17. A computer device, characterized in that, It includes: A memory for storing executable instructions; A processor for implementing the video cover recommendation method according to any one of claims 1 to 14 when executing the executable instructions stored in the memory.
18. A computer-readable storage medium, characterized in that, Stored with executable instructions for implementing the video cover recommendation method according to any one of claims 1 to 14 when being executed by a processor.
19. A computer program product, comprising computer instructions, characterized in that, The computer instructions implement the video cover recommendation method according to any one of claims 1 to 14 when being executed by a processor.
Citation Information
Patent Citations
Video recommendation method and device, storage medium, terminal and server
CN111246255A
Video cover determination method and device, electronic equipment and storage medium
CN111918130A