Method, apparatus, computer device, and storage medium for generating a video cover image
By adding picture-in-picture images matching the video content on the initial cover image of the video to generate the target cover image, the problem of low accuracy of video cover image in the prior art is solved, and the accuracy and efficiency of users clicking to enter the video of interest is improved.
Patent Information
- Application Number
- CN202210011031.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-06
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-01-06
AI Technical Summary
The video cover images generated in the prior art are of low accuracy and cannot accurately represent the theme and key content of the video, making it difficult for network objects to accurately click through the cover image to enter the video of interest.
By acquiring the pending video and its initial cover image, the key feature information of the video is extracted based on the information extraction strategy, candidate materials matching these feature information are obtained, and they are synthesized as picture-in-picture images onto the initial cover image to generate the target cover image.
Improve the accuracy of the video cover image, making it more representative of the theme and key content of the video, reduces the number of videos that network objects need to click, and thus improves the user experience.
Smart Images

Figure CN114372172B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a method, apparatus, computer device, and storage medium for generating a video cover image. Background Art
[0002] With the continuous development of self-media technology, more and more online media platforms can provide services for network objects to publish and watch videos. On the online media platform, a network object can upload a pre-produced video to achieve the purpose of publishing the video. The published video can provide a video preview through the title and cover image of the video, so that the network object can click into the video of interest through the video preview and watch the videos published by other network objects.
[0003] The cover image of a video is usually a video frame selected from the video. However, a single video frame in the video can only represent one video scene, cannot accurately represent the theme of the video, and even less can highlight the key content expressed by the video frame. As a result, the network object cannot accurately click into the video of interest through the cover image of the video, so that the network object needs to click into multiple videos to watch before it can watch the video of interest.
[0004] It can be seen that under the related technology, the accuracy of the generated video cover image is relatively low. Summary of the Invention
[0005] Embodiments of this application provide a method, apparatus, computer device, and storage medium for generating a video cover image to solve the problem of relatively low accuracy of generating a video cover image.
[0006] In a first aspect, a method for generating a video cover image is provided, including:
[0007] Obtain a video to be processed and an initial cover image of the video to be processed;
[0008] Generate a target cover image of the video to be processed based on the video to be processed and the initial cover image;
[0009] Wherein, the target cover image is generated by adding a picture-in-picture image to the initial cover image, and the picture-in-picture image matches the content of the video to be processed.
[0010] In a second aspect, an apparatus for generating a video cover image is provided, including:
[0011] An obtaining module: configured to obtain a video to be processed and an initial cover image of the video to be processed;
[0012] Processing module: Based on the video to be processed and the initial cover image, generate the target cover image of the video to be processed;
[0013] Among them, the target cover image is generated by adding a picture-in-picture image to the initial cover image, and the picture-in-picture image matches the content of the video to be processed.
[0014] Optionally, the processing module is specifically used for:
[0015] Extract the key feature information included in the video to be processed based on the information extraction strategy, where the key feature information is used to characterize: the theme of the video to be processed and the key objects included;
[0016] Obtain each candidate material that matches the key feature information, and select the candidate material that meets the presentation conditions from each candidate material as the target material;
[0017] Use the target material as the picture-in-picture image of the initial cover image, synthesize the target material and the initial cover image, and generate the target cover image.
[0018] Optionally, the processing module is specifically used for:
[0019] Obtain each reference image, where the reference image is an image collected from network resources;
[0020] Extract the video frame features of each video frame to be processed included in the video to be processed, and extract the image features of each reference image;
[0021] Based on each video frame feature and each image feature, determine each candidate material that matches the key feature information from each video frame to be processed and each reference image.
[0022] Optionally, the key feature information includes: word feature, face feature and object feature;
[0023] Then the processing module is specifically used for:
[0024] Determine the graphic similarity between the word feature and each video frame feature, and between the word feature and each image feature respectively;
[0025] Determine the face similarity between the face feature and each video frame feature, and between the face feature and each image feature respectively;
[0026] Determine the object similarity between the object feature and each video frame feature, and between the object feature and each image feature respectively;
[0027] Based on the obtained text-image matching degree, each face matching degree, and each object matching degree, determine each candidate material that matches the key feature information from the respective video frames to be processed and the respective reference images.
[0028] Optionally, the processing module obtains the key feature information by the following method:
[0029] Obtain the release information and subtitle file of the video to be processed, and extract the word features of the keywords included in the release information and the subtitle file;
[0030] Obtain the key video frames in the video to be processed, and extract the face features of the face regions included in the key video frames and the object features of the object regions included therein, where the key video frames are the video frames in the video to be processed that represent video scene switching;
[0031] Use the obtained word features, face features, and object features as the key feature information.
[0032] Optionally, the processing module is specifically configured to:
[0033] Based on each video frame feature and each image feature, determine each candidate image that matches the key feature information from the respective video frames to be processed and the respective reference images;
[0034] Respectively determine the image regions of each candidate image that match the key feature information;
[0035] Based on each image region, perform cropping processing on each candidate image to obtain each candidate material.
[0036] Optionally, the processing module is specifically configured to:
[0037] Based on a clarity evaluation strategy, evaluate the clarity of each candidate material to determine the respective clarity evaluation values of each candidate material;
[0038] Based on a content quality evaluation strategy, evaluate the content quality of each candidate material to determine the respective content quality evaluation values of each candidate material;
[0039] Obtain the weighted sum of each clarity evaluation value and each content quality evaluation value, and use the candidate materials whose weighted sum meets the condition of being greater than the presentation threshold as the target materials.
[0040] Optionally, the processing module is specifically configured to:
[0041] Detect the initial cover image and determine the target object included in the initial cover image;
[0042] Based on the position of the target object in the initial cover image, divide the initial cover image into a target area and a non-target area;
[0043] Based on the shape and size of the non-target area, adjust the size of the target material;
[0044] Overlay the adjusted target material over the non-target area in the initial cover image to generate the target cover image of the video to be processed.
[0045] In a third aspect, a computer program product is provided, including a computer program which, when executed by a processor, implements the method described in the first aspect.
[0046] In a fourth aspect, a computer device is provided, including:
[0047] A memory for storing program instructions;
[0048] A processor for calling the program instructions stored in the memory and executing the method described in the first aspect according to the obtained program instructions.
[0049] In a fifth aspect, a computer-readable storage medium is provided, where the storage medium stores computer-executable instructions for causing a computer to execute the method described in the first aspect.
[0050] In the embodiments of the present application, the picture-in-picture image added to the initial cover image matches the content of the video to be processed. For example, it is related to the theme of the video to be processed or matches the key object included in the video to be processed. Then, the obtained target cover image can more accurately represent the content of the video to be processed, improving the accuracy of generating the video cover.
[0051] Further, the picture-in-picture image can, on the basis of the initial cover image, play a role in emphasizing the content of the video to be processed. For example, it plays a role in emphasizing the theme of the video to be processed or the key object included, avoiding the situation where the represented content is unclear when the initial cover image is presented alone, and further improving the accuracy of generating the video cover. Description of the Drawings
[0052] Figure 1a It is a schematic diagram of the principle of a method for generating a video cover image in the related art;
[0053] Figure 1b It is a first schematic diagram of the principle of a method for generating a video cover image provided by the embodiments of the present application;
[0054] Figure 1cAn application scenario of the method for generating a video cover image provided by an embodiment of the present application;
[0055] Figure 2a A first schematic flow diagram of the method for generating a video cover image provided by an embodiment of the present application;
[0056] Figure 2b A second schematic flow diagram of the method for generating a video cover image provided by an embodiment of the present application;
[0057] Figure 3 A second schematic principle diagram of the method for generating a video cover image provided by an embodiment of the present application;
[0058] Figure 4a A schematic principle of the method for generating a video cover image provided by an embodiment of the present application Figure Three ;
[0059] Figure 4b A fourth schematic principle diagram of the method for generating a video cover image provided by an embodiment of the present application;
[0060] Figure 4c A schematic principle of the method for generating a video cover image provided by an embodiment of the present application Figure Five ;
[0061] Figure 5 A sixth schematic principle diagram of the method for generating a video cover image provided by an embodiment of the present application;
[0062] Figure 6a A schematic principle of the method for generating a video cover image provided by an embodiment of the present application Figure Seven ;
[0063] Figure 6b A schematic principle of the method for generating a video cover image provided by an embodiment of the present application Figure Eight ;
[0064] Figure 6c A ninth schematic principle diagram of the method for generating a video cover image provided by an embodiment of the present application;
[0065] Figure 6d A tenth schematic principle diagram of the method for generating a video cover image provided by an embodiment of the present application;
[0066] Figure 6e An eleventh schematic principle diagram of the method for generating a video cover image provided by an embodiment of the present application;
[0067] Figure 7 A first schematic structural diagram of the apparatus for generating a video cover image provided by an embodiment of the present application;
[0068] Figure 8 This is the second structural schematic diagram of the device for generating video cover images provided by the embodiments of the present application. Specific embodiments
[0069] In order to make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application.
[0070] Some terms in the embodiments of the present application will be explained below to facilitate understanding by those skilled in the art.
[0071] (1) Feeds:
[0072] Feeds is the source of messages, which is the abbreviated form of web feed, news feed, and syndicated feed. Feeds is a data format through which a website disseminates the latest information to users, usually arranged in a timeline manner. The prerequisite for a user to subscribe to a website is that the website provides a source of messages.
[0073] (2) Short video:
[0074] A short video is a way of disseminating Internet content, generally referring to video content with a duration of less than 5 minutes disseminated on new Internet media. With the popularization of mobile terminals and the acceleration of network speed, short, flat, and fast large-traffic dissemination content has gradually gained the favor of major platforms and capital.
[0075] The embodiments of the present application relate to the field of artificial intelligence (AI) and are designed based on machine learning (ML) technology, and can be applied to fields such as cloud computing, intelligent transportation, assisted driving, or maps.
[0076] Artificial intelligence uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, and is a theory, method, technology, and application system that can perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It studies the design principles and implementation methods of various machines, attempts to understand the essence of intelligence, and produces a new intelligent machine that can respond in a way similar to human intelligence, enabling the machine to have the functions of perception, reasoning, and decision-making.
[0077] Artificial intelligence is an interdisciplinary subject that covers a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation interaction systems, mechatronics, and other technologies. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, machine learning / deep learning, autonomous driving, and intelligent transportation. With the development and progress of artificial intelligence, it has been able to conduct research and applications in multiple fields. For example, common fields include smart homes, intelligent customer service, virtual assistants, smart speakers, intelligent marketing, smart wearable devices, driverless, autonomous driving, drones, robots, intelligent healthcare, vehicle networking, autonomous driving, and intelligent transportation. It is believed that with the further development of technology in the future, artificial intelligence will be applied in more fields and play an increasingly important role. The solution provided in the embodiments of this application involves technologies such as deep learning and augmented reality in artificial intelligence, which will be further described through the following embodiments.
[0078] Machine learning is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate human learning behaviors to acquire new knowledge or skills, reorganize existing knowledge structures, and continuously improve their own performance.
[0079] Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications cover all fields of artificial intelligence. The core of machine learning is deep learning, which is a technology to achieve machine learning. Machine learning usually includes technologies such as deep learning, reinforcement learning, transfer learning, inductive learning, artificial neural networks, and rote learning. Deep learning includes technologies such as Convolutional Neural Networks (CNN), deep belief networks, recurrent neural networks, autoencoders, and generative adversarial networks.
[0080] It should be noted that in the embodiments of this application, data related to user portraits, user historical operation records, etc. are involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.
[0081] The following briefly introduces the application fields of the method for generating video cover images provided in the embodiments of this application.
[0082] With the continuous development of self-media technology, more and more online media platforms can provide services for online entities to publish and watch videos. The lowered threshold of content production has led to an exponential growth in the volume of various content being published. The content sources come from various content creation institutions, such as self-media and professionally-generated content (PGC) of institutions, and user-generated content (UGC).
[0083] On online media platforms, an online entity can publish a video by uploading a pre-produced video. The published video can provide a video preview through the video's title and cover image, enabling an online entity to click into a video of interest based on the preview and watch the videos published by other online entities.
[0084] The most crucial factors for an online entity to watch content are the title of the content, the cover image, the author, etc. Please refer to Figure 1a , the cover image is the first impression a video gives to an online entity. When consuming video content, the quality of the cover image greatly affects an online entity's desire to watch. The quality of the cover image has two aspects. On the one hand, it is the quality of the cover image, such as whether it is clear and whether it has defects. On the other hand, it is the information conveyed by the cover image, whether the content is effective and whether it fits the video theme. When an online entity uploads a video, to some extent, it is a creative process, delivering content through the video and summarizing it in the video's title. A good cover image should be as relevant as possible to the title, content, or the scenes shown in the video, as it usually most directly reflects the theme, such as a certain scene, a certain star, or the emphasized points and the relationships between these points.
[0085] The cover image of a video is usually a video frame selected from the video. However, a single video frame in the video can only represent one video scene, unable to accurately represent the theme of the video, let alone highlight the key content expressed by the video frame. As a result, an online entity cannot accurately click into a video of interest through the cover image of the video, forcing the online entity to click into multiple videos to watch before finding the video of interest.
[0086] It can be seen that under the related technologies, the accuracy of the generated video cover image is relatively low.
[0087] To solve the problem of relatively low accuracy of generating video cover images, this application proposes a method for generating a video cover image. After obtaining a video to be processed and the initial cover image of the video to be processed, the method can generate a target cover image for the video to be processed based on the video to be processed and the initial cover image. Please refer to Figure 1b, the target cover image is generated by adding a picture-in-picture image to the initial cover image, and the picture-in-picture image matches the content of the video to be processed.
[0088] In the embodiments of the present application, the picture-in-picture image added to the initial cover image matches the content of the video to be processed. For example, it is related to the theme of the video to be processed or matches the key object included in the video to be processed. Then, the obtained target cover image can more accurately represent the content of the video to be processed, improving the accuracy of generating the video cover.
[0089] Furthermore, the picture-in-picture image can, on the basis of the initial cover image, play a role in emphasizing the content of the video to be processed. For example, it plays a role in emphasizing the theme or the key object included in the video to be processed, avoiding the situation where the content represented by presenting only the initial cover image is unclear, and further improving the accuracy of generating the video cover.
[0090] Next, the application scenarios of the method for generating a video cover image provided by the present application will be described.
[0091] Please refer to Figure 1c , which is a schematic diagram of an application scenario of the method for generating a video cover image provided by the present application. This application scenario includes a client 101 and a server 102. Communication can be carried out between the client 101 and the server 102. The communication method can be communication using wired communication technology, such as communicating by connecting an Ethernet cable or a serial cable; it can also be communication using wireless communication technology, such as communicating through technologies such as Bluetooth or wireless fidelity (WIFI), and the specific method is not limited.
[0092] The client 101 generally refers to a device that can provide a video to be processed to the server 102, such as a terminal device, a web page accessible by the terminal device, or a third-party program accessible by the terminal device, etc. The terminal device can be an intelligent transportation device, a camera, a mobile phone, an intelligent voice interaction device, an intelligent household appliance, a vehicle-mounted terminal, etc. The server 102 generally refers to a device that can process the video to be processed, such as a terminal device or a server, etc. The server includes but is not limited to a cloud server, a local server, or an associated third-party server, etc. Both the client 101 and the server 102 can adopt cloud computing to reduce the occupation of local computing resources; similarly, cloud storage can also be adopted to reduce the occupation of local storage resources.
[0093] As an embodiment, the client 101 and the server 102 can be the same device, and the specific situation is not limited. In the embodiments of the present application, the case where the client 101 and the server 102 are different devices respectively will be taken as an example for introduction.
[0094] Next, based on Figure 1c, taking client 101 as the client and server 102 as the server as an example, the method for generating a video cover image provided by the embodiments of the present application will be specifically introduced.
[0095] Please refer to Figure 2a , which is a schematic flowchart of a method for generating a video cover image provided by the embodiments of the present application.
[0096] S21. Obtain the video to be processed and the initial cover image of the video to be processed.
[0097] The client can obtain the video to be processed and the initial cover image of the video to be processed. For example, the client loads the video to be processed and the initial cover image of the video to be processed in response to a control operation of the target object on the display interface. The control operation can be an operation of uploading a video, or an operation of shooting a video, or an operation of entering an interface for browsing video thumbnails, etc., which is not specifically limited. For another example, the client can receive the video to be processed and the initial cover image of the video to be processed sent by other devices, etc.
[0098] For another example, after obtaining the video to be processed, the client can extract the initial cover image of the video to be processed based on the video to be processed. The initial cover image can be a video frame randomly extracted from the video to be processed, or a key video frame with the largest number of video frames between it and the next key video frame, or a specified video frame selected from the video to be processed in response to a selection operation triggered by the target object, etc., which is not specifically limited.
[0099] S22. Generate a target cover image of the video to be processed based on the video to be processed and the initial cover image.
[0100] After obtaining the video to be processed and the initial cover image of the video to be processed, the client can generate a target cover image of the video to be processed based on the video to be processed and the initial cover image. The target cover image can be a picture-in-picture image determined by the client to match the content of the video to be processed, and the picture-in-picture image is added to the initial cover image. The target cover image can also be that the client sends the video to be processed and the initial cover image to the server, and the server determines a picture-in-picture image that matches the content of the video to be processed and adds the picture-in-picture image to the initial cover image. After the client receives the target cover image sent by the server, the client obtains the target cover image, etc., which is not specifically limited. In the embodiments of the present application, the case where the server sends the target cover image to the client is taken as an example for introduction.
[0101] The process of the server determining the target cover image of the video to be processed will be specifically introduced below. The process of the client determining the target cover image of the video to be processed is similar and will not be elaborated here. Please refer to Figure 2b .
[0102] S201, extract the key feature information contained in the video to be processed based on the information extraction strategy.
[0103] Based on the information extraction strategy, extract the key feature information contained in the video to be processed. The key feature information can be used to characterize the theme of the video to be processed, can also be used to characterize the key objects contained in the video to be processed, can also be used to characterize the theme of the video to be processed, and the key objects contained in the video to be processed. The key feature information can include word features, face features, object features, etc.
[0104] There are various methods to extract the key feature information contained in the video to be processed. Three of them will be introduced as examples below.
[0105] Method 1:
[0106] Obtain the release information and subtitle file of the video to be processed, and extract the word features of the keywords contained in the release information and subtitle file.
[0107] After obtaining the video to be processed, the server can obtain the release information and subtitle file of the video to be processed. The release information of the video to be processed can include the publisher, release time, title, tags, participated topics, etc. of the video to be processed. The server can determine the keywords contained in the release information of the video to be processed according to the title, tags and participated topics, etc. of the video to be processed. After the server obtains the subtitle file of the video to be processed, it can determine the keywords contained in the subtitle file of the video to be processed. Thus, the keywords can, to a certain extent, characterize the theme of the video to be processed and the key objects contained therein.
[0108] The keyword can be an entity in the title, or the entity, sentence with the largest number contained in the subtitle file, or the subtitle corresponding to the video frame with the largest number of bullet screens in the video to be processed, or the person name, animal or intellectual property IP contained in the title, tags and participated topics, etc., or the person name, animal or intellectual property IP with the largest number in the subtitle file, etc.
[0109] The intellectual property IP can be the name of a TV drama, movie or anime, etc. After the server obtains the keywords contained in the release information of the video to be processed, it can extract the word features of the keywords. The word features uniquely characterize the keywords by quantifying the features of the keywords.
[0110] After the server obtains the word features of the keywords, it can use the word features as the key feature information contained in the video to be processed.
[0111] Method 2:
[0112] Obtain the key video frames in the video to be processed, and extract the facial features of the facial regions included in the key video frames.
[0113] After obtaining the video to be processed, the video to be processed can be frame-extracted to obtain the key video frames in the video to be processed. The key video frames are the video frames in the video to be processed that characterize the switching of video scenes. After obtaining the key video frames, the server can perform face detection processing on the key video frames to determine whether the facial regions are included. The number of key video frames can be one or multiple. Based on the key video frames including facial regions, the facial features of the facial regions can be obtained, such as facial embeddings.
[0114] After the server obtains the facial features of the facial regions included in the key video frames, the facial features can be used as the key feature information included in the video to be processed.
[0115] Method 3:
[0116] Obtain the key video frames in the video to be processed, and extract the object features of the object regions included in the key video frames.
[0117] After obtaining the key video frames according to the introduction in Method 2, the server can perform object detection processing on the key video frames to determine whether the object regions are included. The object regions are the regions in the key video frames that contain objects, such as regions containing dynamic objects or regions containing static objects. The number of key video frames can be one or multiple. Based on the key video frames including object regions, the object features of the object regions can be obtained.
[0118] After the server obtains the object features of the object regions included in the key video frames, the object features can be used as the key feature information included in the video to be processed.
[0119] As an embodiment, the server can use only one of the above methods to use the obtained word features, facial features, or object features as the key feature information. The server can also use the above methods in combination. For example, use word features and facial features as the key feature information; for another example, use facial features and object features as the key feature information; for another example, use word features and object features as the key feature information; for another example, use word features, facial features, and object features as the key feature information.
[0120] S202. Obtain each candidate material that matches the key feature information, and select a candidate material that meets the presentation conditions from each candidate material as the target material.
[0121] After obtaining the key feature information, the server can obtain various candidate materials that match the key feature information. The server can determine various candidate materials that match the key feature information from each of the to-be-processed video frames included in the to-be-processed video; it can also determine various candidate materials that match the key feature information from each of the reference images collected from network resources; it can also determine various candidate materials that match the key feature information from each of the to-be-processed video frames and each of the reference images, etc., without specific limitations.
[0122] The following takes the server's determination of various candidate materials that match the key feature information from each video frame and each reference image as an example for introduction.
[0123] The server can collect each reference image from the obtained network resources in real time, can also collect each reference image at a preset time interval, can also receive each reference image in the network resources sent by other devices, etc., without specific limitations. The reference image can be an image for which copyright content needs to be purchased, and the server obtains the reference image by purchasing the copyright content.
[0124] The server can extract the video frame features of each of the to-be-processed video frames included in the to-be-processed video, and extract the image features of each of the reference images. There can be various ways to extract video frame features and image features. For example, extraction can be performed using the CLIP model; for another example, extraction can be performed using classic models such as VGG16, Inception series models, and ResNet; for another example, extraction can be performed using the one-stage face detection network Retinaface, combined with Arcface trained with Asian faces that often appear in annotation tasks and star faces in the task scenario, and the Resnet101 model, etc.
[0125] After obtaining the video frame features and image features, the server can, based on the video frame features and image features, determine various candidate materials that match the key feature information from each of the to-be-processed video frames and each of the reference images. The server can determine various candidate materials that match the key feature information based on the to-be-processed video frames for which the similarity between the video frame features and the key feature information is greater than the similarity threshold, and based on the to-be-processed video frames for which the similarity between the image features and the key feature information is greater than the similarity threshold.
[0126] The server can also rank each video frame to be processed in descending order of similarity based on the similarity between the video frame features and the key feature information, and rank each reference image in descending order of similarity based on the similarity between the image features and the key feature information. The server determines each candidate material that matches the key feature information based on the video frames to be processed and the reference images whose arrangement serial numbers are before the specified serial number.
[0127] For example, taking the key feature information including word features, face features, and object features as an example, the server can respectively determine the graphic similarity between the word features and each video frame feature, and the graphic similarity between the word features and each image feature. The graphic similarity can be calculated using the CLIP model to retrieve candidate materials that match the word features from each video frame to be processed and each reference image.
[0128] The server can also respectively determine the face similarity between the face features and each video frame feature, and the face similarity between the face features and each image feature. Since the differences between faces are relatively small, the face similarity can be determined by combining multiple models to improve the retrieval accuracy. For example, the one-stage face detection network Retinaface is used, combined with Arcface trained with Asian faces that often appear in the annotation task and star faces in the task scenario, and the Resnet101 model to determine the face similarity, so as to retrieve candidate materials that match the face features from each video frame to be processed and each reference image.
[0129] The server can also respectively determine the object similarity between the object features and each video frame feature, and the object similarity between the object features and each image feature. Since the differences between objects are relatively large, a traditional neural network model can be used to determine the object similarity, so as to retrieve candidate materials that match the object features from each video frame to be processed and each reference image. Please refer to Figure 3, the server can use the object features as query conditions, and through the neural network model, determine the object similarity between the object features and the features of each video frame, and the object similarity between the object features and the features of each image. If the similarity is greater than the threshold, the output is 1; if the similarity is less than the threshold, the output is 0, obtaining an output vector containing 0 or 1. The output vector contains multiple element positions, and the multiple element positions respectively correspond to each video frame to be processed and each reference image. An element at an element position being 1 indicates that the video frame feature of the corresponding video frame to be processed or the image feature of the reference image has a similarity greater than the threshold with the object features; an element at an element position being 0 indicates that the video frame feature of the corresponding video frame to be processed or the image feature of the reference image has a similarity not greater than the threshold with the object features. Thus, candidate materials can be obtained based on the video frames to be processed or reference images corresponding to 1 in the output vector.
[0130] After the server obtains various text-image similarities, various face similarities, and various object similarities, it can determine each candidate material that matches the key feature information from each video frame to be processed and each reference image based on the obtained various text-image similarities, various face similarities, and various object similarities; or it can also determine each candidate material that matches the key feature information from each video frame to be processed and each reference image based on the weighted sum of the corresponding text-image similarities, face similarities, and object similarities. The server can use the image with the largest weighted sum as the candidate material, or can also use several images with the largest weighted sums as candidate materials, etc. For example, using the Faiss library for clustering and similarity provides efficient similarity search and clustering for dense vectors and supports searches of up to billions of vectors, enabling very efficient retrieval and matching of vectors. Another example is to use 01 vectors for correlation queries, with similarity marked as 1 and dissimilarity marked as 0, so that each candidate material that matches the key feature information can be retrieved or matched.
[0131] As an embodiment, in order to reduce the storage space occupied by each feature, dimensionality reduction processing can be performed on the video frame features and image features. For example, if the video frame features and image features are in vector form, then they can be converted from floating-point vectors to 01 vectors of 01Bit to achieve the purpose of dimensionality reduction and reduce the storage space occupied by each feature.
[0132] As an embodiment, the server can use the candidate images that match the key feature information in each video frame to be processed and each reference image as candidate materials; or it can also crop the candidate images that match the key feature information and use the cropped candidate images as candidate materials, etc.
[0133] After the server determines each candidate image that matches the key feature information from each to-be-processed video frame and each reference image based on each video frame feature and each image feature, it can determine the image region that matches the key feature information for each candidate image. The server can use a target detection algorithm to determine the image region that matches the key feature information using a rectangular frame; the server can also use an edge recognition algorithm to perform edge recognition on the target that matches the key feature information, and determine the image region that matches the key feature information based on the edge of the target.
[0134] After determining each image region that matches the key feature information, the server can crop each candidate image based on each image region to obtain each candidate material. By removing unnecessary elements from the candidate image, the obtained candidate material can more accurately express the theme of the video to be processed and the key objects contained in the video to be processed.
[0135] After obtaining the candidate materials that match the key feature information, the server can select the candidate materials that meet the presentation conditions from the candidate materials as the target materials. There are many ways to select the candidate materials that meet the presentation conditions from the candidate materials. For example, the server can perform clarity evaluation on each candidate material based on the clarity evaluation strategy to determine the clarity evaluation value of each candidate material. The server can use the candidate materials whose clarity evaluation values are greater than the clarity threshold as the target materials; the server can also sort the candidate materials based on the clarity evaluation values and use the candidate materials ranked before the specified sequence number as the target materials, etc.
[0136] For another example, based on the content quality assessment strategy, the content quality of each candidate material is assessed to determine the content quality assessment value of each candidate material. The server can use the candidate material whose content quality assessment value is greater than the content quality threshold as the target material; the server can also sort the candidate materials based on the content quality assessment value, and use the candidate material ranked before the specified sequence number as the target material, etc. The content quality assessment of each candidate material can assess the content involving advertising and promotion, illegal and irregular content, etc. as a lower content quality assessment value, so that the target material determined according to the content quality assessment value does not contain content involving advertising and promotion, illegal and irregular content.
[0137] As an example, after obtaining the respective clarity evaluation values and the respective content quality evaluation values, the server may perform weighted summation on the corresponding clarity evaluation values and content quality evaluation values to obtain respective weighted sums. After obtaining the respective weighted sums, the server may use the candidate materials whose weighted sums are greater than the presentation threshold as the target materials; the server may also sort the respective candidate materials based on the weighted sums and use the candidate materials ranked before the specified serial number as the target materials.
[0138] S203, use the target material as the picture-in-picture image of the initial cover image, synthesize the target material and the initial cover image, and generate the target cover image of the video to be processed.
[0139] After obtaining the target material, the server may use the target material as the picture-in-picture image of the initial cover image, synthesize the target material and the initial cover image, and generate the target cover image of the video to be processed. Please refer to Figure 4a , which is the initial cover image of the video to be processed. This initial cover image only contains a scene of a woman talking and cannot accurately represent the highlights of the video to be processed, etc. Based on the initial cover image, the network object cannot know the theme of the video to be processed or the key objects contained therein, and is very likely not to click to enter the video to be processed for viewing, thus missing the interesting video. It is also possible that the network object clicks to enter the video to be processed for viewing and only learns that the video to be processed is an uninteresting video after watching the video. Therefore, the network object needs to click to enter multiple videos for viewing before it can view the interesting video.
[0140] After using the target material as the picture-in-picture image of the initial cover image, synthesizing the target material and the initial cover image, and generating the target cover image of the video to be processed, please refer to Figure 4b , if the target cover image contains a scene of a woman talking and the target material of star A saying that she doesn't look alike, then the target cover image can represent that in the video to be processed, the woman said something that doesn't match her appearance, and at the same time, it can also represent that star A participated in the video. Thus, the network object can accurately obtain the theme expressed by the video to be processed and the key objects contained in the video to be processed, enabling the network object to more accurately obtain the interesting video for viewing through the target cover image.
[0141] As an example, there are various methods to synthesize the target material and the initial cover image with the target material as the picture-in-picture image of the initial cover image. For example, the target material can be covered on a specified position in the initial cover image to synthesize the target material and the initial cover image; for another example, the target material can be covered on a position in the initial cover image that does not contain the target object to synthesize the target material and the initial cover image; for another example, the target object included in the initial cover image can be obtained, and the target material and the target object can be re-typeset according to a pre-stored cover template to synthesize the target material and the initial cover image, etc., without specific limitations.
[0142] Taking the process of covering the target material on a position in the initial cover image that does not contain the target object to synthesize the target material and the initial cover image as an example, the following is an introduction.
[0143] The server can detect the initial cover image and determine the target object included in the initial cover image. For example, the server takes the initial cover image as the input of a trained target detection model and obtains the target object included in the initial cover image output by the target detection model. The target object included in the initial cover image output by the target detection model can be marked with a rectangular box, or marked with the edge of the target object, etc., without specific limitations.
[0144] The server can divide the initial cover image into a target area and a non-target area based on the position of the target object in the initial cover image. The target area is the area in the initial cover image that contains the target object, which can be a rectangular area or an area surrounded by the edge of the target object. The non-target area is the area in the initial cover image except the target area.
[0145] The server can adjust the size of the target material based on the shape and size of the non-target area. The server can aim to fill the non-target area with the target material and adjust the size of the target material. For example, if the non-target area is a rectangular area and the shape of the target material is a circle, then the long side of the rectangular area can be used as the diameter of the circle to adjust the size of the target material.
[0146] If the content included in the target material is the same object as the target object included in the target area, then the server can aim to enlarge the target material by a specified multiple relative to the target object included in the target area and adjust the size of the target material. Please refer to Figure 4c , if the target material is the face area of a target object included in the target area, then the target material can be enlarged to 2 times the original to highlight the facial expression of the target object. Since this scene is a funny scene, by highlighting the facial expression of the target object, the purpose of enhancing the comedy effect can be achieved, improving the sense of substitution, so that the target cover image can accurately represent the funny theme of the video to be processed.
[0147] The server can also first determine the proportion of the target area in the initial cover image, and then adjust the size of the target material so that the proportion of the target material in the initial cover image is the same as the proportion of the target area in the initial cover image, and so on.
[0148] After adjusting the size of the target material, the adjusted target material is obtained. The server can cover the adjusted target material over the non-target area of the initial cover image to generate the target cover image of the video to be processed.
[0149] As an embodiment, after the server determines the target object included in the initial cover image, the initial cover image can be cropped to obtain the target object material containing the target object. The server can combine the pre-stored cover template, the obtained target material and the target object material to generate the target cover image. In some cases, the background of the initial cover image is relatively complex, which is likely to cause the problem of chaotic content in the obtained target cover image. Therefore, the target object in the initial cover image can be extracted, combined with a clear and concise cover template, and then the target material can be used as a picture-in-picture image and superimposed on the cover template to generate the target cover image.
[0150] In the embodiments of the present application, each service can be called through each distributed system to work together to implement the method for generating a video cover image provided by the embodiments of the present application. Please refer to Figure 5 .
[0151] The network object can upload the video to be processed through the content provider for video publishing. The content provider includes content production forms such as PGC, UGC, Multi-Channel Network (MCN), or Professional Generated Content + User Generated Content (PUGC). The network object uploads the video to be processed through a mobile terminal or by calling the back-end Application Programming Interface (API) system of the client, providing local or captured video content, or self-media articles or picture sets written, etc. The network object can upload the corresponding cover image while uploading the video to be processed, or select the initial cover image from the video to be processed by the server.
[0152] The content provider communicates with the uplink and downlink content interface service to first obtain the upload server interface address, and then uploads the video to be processed. It calls the content storage service to store the video to be processed in the content database. The uplink and downlink content interface service can also obtain the title, publisher, abstract, cover image, release time, etc. of the video to be processed from the content provider. When calling the content storage service, it can also store the meta-information of the video to be processed, such as the video file size, cover image link, bit rate, file format, title, release time, author, mark of whether it is original, and whether it is a first release, etc. in the content database. The uplink and downlink content interface service can submit the data stored in the content database to the scheduling center service for subsequent content processing and circulation by the scheduling center service.
[0153] Thus, when the network object searches for a video, the content consumer can communicate with the content distribution outlet service to obtain the index information corresponding to the searched video. By communicating with the content storage service, it downloads the streaming media file of the searched video corresponding to the index information from the content database. Thus, the streaming media file can be played through the local player, or communicate with the CDN service deployed at the edge to present the graphic data.
[0154] The content consumer can report the behavior data of the network object browsing the video during the upload and download processes, such as reading speed, completion rate, reading time, stuttering, loading time, play click, etc. to the server, so that the server can provide more user-friendly services for the network object based on the reported data.
[0155] The content consumer can browse the video in the form of Feeds stream, and provide an entry for direct reporting and feedback for low-quality content. This entry is directly connected to the manual review system. The operator can confirm and review through the manual review system, which can be used as sample data for machine model quality filtering in the subsequent process of filtering the videos uploaded by the network object.
[0156] The scheduling center service mainly includes machine processing and the aforementioned manual review processing. Among them, machine processing can perform various quality judgments, such as filtering low-quality content; it can also mark content tags, such as content classification, topic information, etc.; it can also perform content deduplication, etc. The processing results obtained by the scheduling center service can be written into the content database.
[0157] Manual review processing can be achieved by the manual review system calling the manual review service. During the manual review process, the manual review system reads the information in the content database, and at the same time, the results and status of the manual review are also transmitted back to the content database. Therefore, the content database also includes the classification of content during the manual review process, including first-level, second-level, and third-level classifications and tag information. For example, for a video explaining a brand mobile phone, the first-level classification is technology, the second-level classification is smart phones, the third-level classification is domestic mobile phones, and the tag information is brand, model, etc. The scheduling center service can specify different tone enhancement template strategies according to the annotation information of the first-level classification of each video in the content database, etc.
[0158] The manual review service can be a WEB system that takes the results filtered by the machine as input on the link, manually confirms and reviews the results filtered by the machine, writes the reviewed results into the content database record, and at the same time can online evaluate the actual effect of the machine filtering model through the results of the manual review here.
[0159] When performing content deduplication, the scheduling center service can call the content deduplication service. The content deduplication service mainly includes title deduplication, picture deduplication of the cover image, content text deduplication, and video fingerprint and audio fingerprint deduplication. The deduplication process usually vectorizes the title and text of the graphic content, using simmhash and BERT text vectors. When deduplicating the picture vectors, for video content, video fingerprints and audio fingerprints are extracted to construct vectors, and then the distance between the vectors, such as the Euclidean distance, is calculated to determine whether there is duplication. Content deduplication can reduce the amount of content review and ensure that the same content exists only once in the recommended distribution pool, guaranteeing the user experience.
[0160] The scheduling center service is mainly responsible for the entire scheduling process of the transfer of video and graphic content. It receives the uploaded video to be processed through the upstream and downstream content interface service, and then obtains the meta-information of the video from the content database. When receiving the uploaded video to be processed, the scheduling center service can call the picture-in-picture service to generate a target cover image for the video to be processed and store it in the content database; it can also call the picture-in-picture service when a certain video to be processed is searched at the content consumption end to generate a target cover image for the video to be processed and present it through the content consumption end, etc.
[0161] The picture-in-picture service can call the picture material extraction service to determine candidate images that match the key feature information based on the key feature information of the video to be processed, and call the service for selecting and cropping the cover image. Among them, the cover image cropping service mainly performs corresponding screenshots on the candidate images through the original image size and the cropped target size, using human detection, object detection, and OCR text recognition of the candidate images to obtain candidate materials.
[0162] After obtaining candidate materials, the service of selecting and cropping the cover image can be called. Among them, the cover image selection service mainly filters and screens the candidate materials according to basic quality characteristics such as clarity, beauty, unsuitable pictures, mosaics, vulgarity and porn, etc., and removes some low-quality candidate materials that are not suitable for being used as the cover.
[0163] After selecting the target materials that meet the presentation conditions based on each candidate material, the obtained target materials can be stored in the enhanced picture material library for use as the cover image of the subsequent video. The enhanced picture material library is used to save the candidate set for picture enhancement, including the content of the video frames extracted from the video content and the content with purchased copyright. After obtaining the target materials, the template database can be called to select the target template. The picture-in-picture service can synthesize the target materials and the initial cover image of the video to be processed based on the target template to generate the target cover image of the video to be processed. The initial cover image can be the cover image uploaded by the content provider, or the cover image selected from each video frame of the video to be processed by calling the service of selecting and cropping the cover image, etc., without specific limitation.
[0164] The core synthesis principle of the template database is that the main elements of the picture should not be blocked. Several non-target areas can be determined by using the main target detection, such as the left, right, upper or lower side strategies, which can be specifically determined according to the actual situation. The template database can communicate with the intelligent enhancement service to provide strategy display. If there is text, it also includes the font and style configuration strategies for synthesizing the text, etc.
[0165] The intelligent picture-in-picture service can communicate with the scheduling center service to complete the method for generating the video cover image provided by the embodiments of the present application. The scheduling center service includes extracting the title keywords and entity words, then communicating with the picture material extraction service to complete the screening and matching of the target materials, and finally generating the target cover image to achieve the output enhancement effect through the target materials.
[0166] The following is an example introduction to the method for generating the video cover image provided by the embodiments of the present application. Please refer to Figure 6a .
[0167] After the client obtains the video to be processed and the initial cover image of the video to be processed, it can extract the key feature information included in the video to be processed based on the video to be processed, as well as the release information and subtitle file of the video to be processed, or the release information and subtitle file. The key feature information can include one or more of word features, face features or object features. For example, the key feature information is the features of entities such as text, names of people, animals, etc.
[0168] After obtaining one or more of word features, face features, or object features, the client can determine each candidate image that matches the key feature information from each reference image and each video frame to be processed, perform object recognition processing on each candidate image, crop the objects in each candidate image, and obtain each candidate material.
[0169] After obtaining each candidate material that matches the key feature information, the client can select the candidate materials that meet the presentation conditions from each candidate material as the target materials. For example, based on the clarity evaluation strategy and the content quality evaluation strategy in sequence, perform clarity evaluation and content quality evaluation on each candidate material, and select the candidate materials that are relatively clear and have a high content quality. After the client obtains the selected candidate materials that are relatively clear and have a high content quality, it can also perform some other post-processing on these candidate materials to obtain the target materials. For example, according to the area in the initial cover image that does not contain the target object, adjust the size of these candidate materials to obtain the target materials, so that when the target materials are used as the picture-in-picture images and added to the initial cover image, it can neither affect the content expressed by the initial cover image nor further express the theme of the video to be processed or the key objects included, and play a role in emphasizing the theme of the video to be processed or the key objects included.
[0170] The number of target materials can be one or multiple. There are various ways to add the target materials as the picture-in-picture images to the initial cover image. It can be randomly added, or added with the goal of filling the non-target area in the initial cover image. It can also add multiple target materials as a whole to the initial cover image, or obtain a pre-stored addition template and add the corresponding target materials to the position specified by the template, etc., which is not specifically limited.
[0171] Next, taking a short video as an example, the method for generating a video cover image provided by the embodiments of the present application will be introduced by way of example.
[0172] For example, a short video introduces its highlights through the title, "The little raccoon eats grapes on the sofa. After finding that there are no grapes in the bowl, its reaction makes me laugh for a year". In order to express the semantics completely and accurately, the title generally has a relatively large number of words and is usually located in an inconspicuous position. Please refer to Figure 6b , through a video frame included in the short video, it can only express the scene of the little raccoon eating grapes on the sofa, and cannot intuitively and accurately convey the theme of the short video to the network object, that is, the comparison of the little raccoon's reaction when eating grapes and after eating grapes.
[0173] After the server obtains the short video and determines the initial cover image of the short video, please refer to Figure 6c, key feature information of the short video can be extracted based on the title of the short video and the video frames included in the short video. The key feature information may include the word features of keywords, that is, the word features of the entity "raccoon" in the title, and may also include the object features of key objects, that is, the object features of the object area of the raccoon in the video frames.
[0174] The server can use the word features of the keyword "raccoon" and the object area of the key object raccoon as the query subject to determine candidate images that match the keyword and candidate images that match the key object in the video frames included in the short video, and obtain each candidate image.
[0175] After obtaining each candidate image, please refer to Figure 6d , the server can perform deduplication processing on each candidate image, and perform cropping processing on each candidate image according to the image area where each candidate image matches the keyword or the image area where each candidate image matches the key object, to obtain each candidate material.
[0176] The server can perform clarity evaluation on each candidate material based on the clarity evaluation strategy to determine the clarity evaluation value of each candidate material, and perform content quality evaluation on each candidate material based on the content quality evaluation strategy to determine the content quality evaluation value of each candidate material. Finally, based on the weighted sum of the corresponding clarity evaluation value and content quality evaluation value, select the candidate material with the largest weighted sum as the target material.
[0177] After the server obtains the initial cover image and the target material, please refer to Figure 6e , it can perform object detection processing on the initial cover image to determine the non-target area that does not contain the target object in the initial cover image, that is, the area outside the area where the raccoon is located in the initial cover image.
[0178] The server can cover the target material on the non-target area, and combine it with the pre-stored cover template to synthesize the target material and the initial cover image to generate the target cover image.
[0179] Based on the same inventive concept, an embodiment of the present application provides a device for generating a video cover image, which can implement the functions corresponding to the foregoing method for generating a video cover image. Please refer to Figure 7 , the device includes an acquisition module 701 and a processing module 702, where:
[0180] The acquisition module 701: is used to acquire the video to be processed and the initial cover image of the video to be processed;
[0181] The processing module 702: based on the video to be processed and the initial cover image, generate the target cover image of the video to be processed;
[0182] Among them, the target cover image is generated by adding a picture-in-picture image to the initial cover image, and the picture-in-picture image matches the content of the video to be processed.
[0183] In a possible embodiment, the processing module 702 is specifically configured to:
[0184] Extract the key feature information included in the video to be processed based on the information extraction strategy, where the key feature information is used to characterize: the theme of the video to be processed and the key objects included;
[0185] Obtain each candidate material that matches the key feature information, and select the candidate material that meets the presentation conditions from each candidate material as the target material;
[0186] Use the target material as the picture-in-picture image of the initial cover image, and synthesize the target material and the initial cover image to generate the target cover image.
[0187] In a possible embodiment, the processing module 702 is specifically configured to:
[0188] Obtain each reference image, where the reference image is an image collected from network resources;
[0189] Extract the video frame features of each video frame to be processed included in the video to be processed, and extract the image features of each reference image;
[0190] Based on each video frame feature and each image feature, determine each candidate material that matches the key feature information from each video frame to be processed and each reference image.
[0191] In a possible embodiment, the key feature information includes: word features, face features, and object features;
[0192] Then the processing module 702 is specifically configured to:
[0193] Determine the graphic similarity between the word features and each video frame feature, and between the word features and each image feature respectively;
[0194] Determine the face similarity between the face features and each video frame feature, and between the face features and each image feature respectively;
[0195] Determine the object similarity between the object features and each video frame feature, and between the object features and each image feature respectively;
[0196] Based on the obtained graphic matching degrees, each face matching degree, and each object matching degree, determine each candidate material that matches the key feature information from each video frame to be processed and each reference image.
[0197] In a possible embodiment, the processing module 702 obtains the key feature information by the following method:
[0198] Obtain the release information and subtitle file of the video to be processed, and extract the word features of the keywords included in the release information and subtitle file;
[0199] Obtain the key video frames in the video to be processed, and extract the face features of the face regions included in the key video frames and the object features of the object regions included therein, where the key video frames are the video frames in the video to be processed that characterize the video scene switching;
[0200] Use the obtained word features, face features and object features as the key feature information.
[0201] In a possible embodiment, the processing module 702 is specifically configured to:
[0202] Based on each video frame feature and each image feature, determine each candidate image that matches the key feature information from each video frame to be processed and each reference image;
[0203] Respectively determine the image regions of each candidate image that match the key feature information;
[0204] Based on each image region, perform cropping processing on each candidate image to obtain each candidate material.
[0205] In a possible embodiment, the processing module 702 is specifically configured to:
[0206] Based on the clarity evaluation strategy, evaluate the clarity of each candidate material to determine the clarity evaluation value of each candidate material;
[0207] Based on the content quality evaluation strategy, evaluate the content quality of each candidate material to determine the content quality evaluation value of each candidate material;
[0208] Obtain the weighted sum of each clarity evaluation value and each content quality evaluation value, and use the candidate materials whose weighted sum is greater than the presentation threshold as the target materials.
[0209] In a possible embodiment, the processing module 702 is specifically configured to:
[0210] Detect the initial cover image and determine the target object included in the initial cover image;
[0211] Based on the position of the target object in the initial cover image, divide the initial cover image into a target region and a non-target region;
[0212] Based on the shape and size of the non-target region, adjust the size of the target material;
[0213] Overlay the adjusted target material over the non-target area of the initial cover image to generate the target cover image of the video to be processed.
[0214] Please refer to Figure 8 , the above device for generating a video cover image can run on a computer device 800. The current version and historical versions of the data storage program and the application software corresponding to the data storage program can be installed on the computer device 800. The computer device 800 includes a processor 880 and a memory 820. In some embodiments, the computer device 800 may include a display unit 840. The display unit 840 includes a display panel 841 for displaying a user interaction operation interface, etc.
[0215] In a possible embodiment, the display panel 841 may be configured in the form of a liquid crystal display (LCD) or an organic light-emitting diode (OLED), etc.
[0216] The processor 880 is used to read the computer program and then execute the method defined by the computer program. For example, the processor 880 reads the data storage program or file, etc., so as to run the data storage program on the computer device 800 and display the corresponding interface on the display unit 840. The processor 880 may include one or more general-purpose processors and may also include one or more DSPs (Digital Signal Processors) for performing related operations to implement the technical solutions provided by the embodiments of the present application.
[0217] The memory 820 generally includes internal memory and external memory. The internal memory may be a random access memory (RAM), a read-only memory (ROM), and a cache (CACHE), etc. The external memory may be a hard disk, an optical disc, a USB flash drive, a floppy disk, or a tape drive, etc. The memory 820 is used to store the computer program and other data. The computer program includes application programs corresponding to each client, etc. Other data may include an operating system or data generated after the application program is run. The data includes system data (such as configuration parameters of the operating system) and user data. In the embodiments of the present application, the program instructions are stored in the memory 820, and the processor 880 executes the program instructions in the memory 820 to implement any of the methods for generating a video cover image discussed in the foregoing figures.
[0218] The above display unit 840 is used to receive input digital information, character information, or contact touch operations / non-contact gestures, and generate signal inputs related to user settings and function controls of the computer device 800, etc. Specifically, in the embodiments of the present application, the display unit 840 may include a display panel 841. The display panel 841 is, for example, a touch screen, which can collect touch operations of users on or near it (such as users using fingers, styli, or any suitable objects or accessories to operate on the display panel 841 or near the display panel 841), and drive corresponding connection devices according to pre-set programs.
[0219] In a possible embodiment, the display panel 841 may include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch orientation of the player, detects the signals brought by the touch operation, and transmits the signals to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact coordinates, and then sends it to the processor 880, and can also receive and execute commands sent by the processor 880.
[0220] Among them, the display panel 841 can be implemented in multiple types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the display unit 840, in some embodiments, the computer device 800 may further include an input unit 830. The input unit 830 may include an image input device 831 and other input devices 832. The other input devices may include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, power on / off keys, etc.), trackballs, mice, joysticks, etc.
[0221] In addition to the above, the computer device 800 may further include a power supply 890 for powering other modules, an audio circuit 860, a near field communication module 870, and an RF circuit 810. The computer device 800 may further include one or more sensors 850, such as an acceleration sensor, a light sensor, a pressure sensor, etc. The audio circuit 860 specifically includes a speaker 861 and a microphone 862, etc. For example, the computer device 800 can collect the user's voice through the microphone 862 and perform corresponding operations, etc.
[0222] As an embodiment, the number of processors 880 may be one or more. The processor 880 and the memory 820 may be coupled or relatively independent.
[0223] As an embodiment, Figure 8 the processor 880 therein may be used to implement the functions of the acquisition module 701 and the processing module 702 as in Figure 7 .
[0224] As an embodiment, Figure 8The processor 880 therein can be used to implement the functions corresponding to the server or terminal device discussed above.
[0225] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the above method embodiments; and the foregoing storage medium includes: removable storage devices, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks and other various media that can store program codes.
[0226] Alternatively, if the above integrated units of the present invention are implemented in the form of software function modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solutions of the embodiments of the present invention, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. For example, it is embodied through a computer program product. The computer program product is stored in a storage medium and includes several instructions for causing a computer device to execute all or part of the methods described in the various embodiments of the present invention. And the foregoing storage medium includes: removable storage devices, ROM, RAM, magnetic disks, or optical disks and other various media that can store program codes.
[0227] Obviously, those skilled in the art can make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to include these changes and modifications.
Claims
1. A method for generating a video cover image, characterized in that, it includes: obtaining a video to be processed and an initial cover image of the video to be processed; extracting key feature information included in the video to be processed based on an information extraction strategy; wherein, the key feature information is used to characterize: the theme of the video to be processed and the key objects included; obtaining each candidate material that matches the key feature information, and selecting a candidate material that meets the presentation condition from each candidate material as the target material; using the target material as a picture-in-picture image of the initial cover image, synthesizing the target material and the initial cover image to generate a target cover image.
2. The method according to claim 1, characterized in that, obtaining each candidate material that matches the key feature information includes: obtaining each reference image, wherein the reference image is an image collected from network resources; extracting the video frame features of each video frame to be processed included in the video to be processed, and extracting the image features of each reference image; based on each video frame feature and each image feature, determining each candidate material that matches the key feature information from each video frame to be processed and each reference image.
3. The method according to claim 2, characterized in that, the key feature information includes: word feature, face feature and object feature; then based on each video frame feature and each image feature, determining each candidate material that matches the key feature information from each video frame to be processed and each reference image includes: respectively determining the graphic similarity between the word feature and each video frame feature, and between the word feature and each image feature; respectively determining the face similarity between the face feature and each video frame feature, and between the face feature and each image feature; respectively determining the object similarity between the object feature and each video frame feature, and between the object feature and each image feature; based on the obtained graphic matching degrees, each face matching degree and each object matching degree, determining each candidate material that matches the key feature information from each video frame to be processed and each reference image.
4. The method according to claim 3, characterized in that, the key feature information is obtained by the following method: obtaining the release information and subtitle file of the video to be processed, and extracting the word features of the keywords included in the release information and the subtitle file; obtaining key video frames in the video to be processed, and extracting the face features of the face regions included in the key video frames and the object features of the object regions included, wherein the key video frames are video frames in the video to be processed that characterize video scene switching; using the obtained word features, face features and object features as the key feature information.
5. The method according to claim 2, characterized in that, Based on each video frame feature and each image feature, determining, from the respective video frames to be processed and the respective reference images, each candidate material that matches the key feature information, including: Based on each video frame feature and each image feature, determining, from the respective video frames to be processed and the respective reference images, each candidate image that matches the key feature information; Respectively determining, for each of the candidate images, the image region that matches the key feature information; Based on each image region, respectively performing a cropping process on each of the candidate images to obtain each candidate material.
6. The method according to claim 1, wherein, selecting, from each of the candidate materials, a candidate material that meets the presentation condition as the target material, including: Based on a clarity evaluation strategy, performing a clarity evaluation on each of the candidate materials to determine the respective clarity evaluation values of each of the candidate materials; Based on a content quality evaluation strategy, performing a content quality evaluation on each of the candidate materials to determine the respective content quality evaluation values of each of the candidate materials; Obtaining the weighted sum of each clarity evaluation value and each content quality evaluation value, and using the candidate material with a weighted sum greater than the presentation threshold as the target material.
7. The method according to claim 1, wherein, using the target material as the picture-in-picture image of the initial cover image, synthesizing the target material and the initial cover image to generate the target cover image, including: Detecting the initial cover image to determine the target object included in the initial cover image; Based on the position of the target object in the initial cover image, dividing the initial cover image into a target region and a non-target region; Based on the shape and size of the non-target region, adjusting the size of the target material; Covering the adjusted target material over the non-target region in the initial cover image to generate the target cover image.
8. An apparatus for generating a video cover image, wherein, comprising: An acquisition module: configured to acquire a video to be processed and the initial cover image of the video to be processed; A processing module: based on an information extraction strategy, extracting the key feature information included in the video to be processed; obtaining each candidate material that matches the key feature information, and selecting, from each of the candidate materials, a candidate material that meets the presentation condition as the target material; Using the target material as the picture-in-picture image of the initial cover image, synthesizing the target material and the initial cover image to generate a target cover image; wherein, the key feature information is used to characterize: the theme of the video to be processed and the key objects included.
9. A computer program product, comprising a computer program, wherein, when the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer device, wherein, comprising: A memory, configured to store program instructions; A processor, configured to call the program instructions stored in the memory and execute the method according to any one of claims 1 to 7 according to the obtained program instructions.
11. A computer-readable storage medium, characterized in that, the storage medium stores computer-executable instructions for causing a computer to execute the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Cover display method and device, electronic equipment and storage medium
CN113518233A