Cover picture generation method and device based on bullet screen content and electronic equipment

By combining barrage content, image clarity and aesthetic scores, video cover images are generated intelligently, which solves the problem that traditional methods are difficult to accurately match user interests, and achieves more efficient and personalized cover images generation, improving user experience and platform click rate.

CN120075491APending Publication Date: 2025-05-30BEIJING QIYI CENTURY SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510225653.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Traditional video cover image generation methods are difficult to accurately reflect user interests, ignore user interaction behavior and viewing preferences, and the generated cover image fails to accurately capture the hot clips of the video, resulting in insufficient attractiveness of the video.

Method used

By calculating the matching degree between the barrage data of the target video and the preset dictionary, the relevant candidate barrage collection is selected, and the candidate image collection is extracted based on the barrage timestamp. Then, the sharpness of the image is calculated and the candidate cover image is filtered out, and the candidate cover image is finally aesthetically scored, and the image with higher scores is selected as the target cover image.

Benefits of technology

The generated cover image is not only in line with the video theme, but also has a high visual appeal, which can better capture user interests, improve user click-through rate and viewing rate, and improve overall user satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120075491A_ABST
    Figure CN120075491A_ABST
Patent Text Reader

Abstract

The invention relates to a bullet screen content-based cover picture generation method and apparatus, and an electronic device. The method comprises the steps of calculating a matching degree between each bullet screen text of bullet screen data of a target video and a preset dictionary, and performing data screening on the bullet screen data according to the matching degree to obtain a candidate bullet screen set; determining a candidate image set of each bullet screen text in the candidate bullet screen set from the target video according to the timestamp; calculating the definition of each image in the candidate image set, and determining a candidate cover image of each bullet screen text in the candidate bullet screen set from the candidate image set according to the definition to obtain a candidate cover image set; and performing aesthetic scoring on each image in the candidate cover image set to obtain a corresponding scoring result, and determining a target cover image set of the target video from the candidate cover image set according to the scoring result. The technical problem that a traditional cover picture generation method is difficult to accurately match user interests is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and particularly to a method, apparatus, and electronic device for generating a cover image based on bullet screen content. Background Art

[0002] With the popularization of video sharing platforms, the method of generating cover images has become an important means to enhance the attractiveness of videos and the click-through rate of users. Traditional methods for generating video cover images usually include fixed-time capture, key frame extraction, and selection based on aesthetic scores. The fixed-time capture method selects a certain moment in the video as the cover, which is simple and easy to implement, but often fails to reflect the exciting segments and user interest points in the video. The key frame extraction method extracts representative frames from the video as cover images. Although it can capture the key content of the video to a certain extent, it still does not fully consider user behavior and viewing preferences. The selection based on aesthetic scores selects cover images by analyzing the visual aesthetic features of the images. Although it has improved in terms of aesthetics, it still cannot accurately match user interests and the actual hot content of the video. These traditional methods generally have several significant problems. First, they cannot accurately reflect user interest points, ignoring the interactive behavior and viewing preferences of users in the video, resulting in the cover image being unable to effectively attract the target audience. Second, the generated cover images often fail to accurately capture the hot segments of the video and cannot maximize the attractiveness of the video. Finally, these methods usually require complex image processing and algorithm calculations, and different parameters need to be adjusted for different video contents, which is cumbersome and time-consuming, wasting a large amount of computing resources. Therefore, there is an urgent need for a more intelligent, efficient, and personalized method for generating cover images to improve the user experience and increase the click-through rate and view volume of the platform. Summary of the Invention

[0003] This application provides a method, apparatus, and electronic device for generating a cover image based on bullet screen content to solve the technical problem that it is difficult for traditional cover image generation methods to accurately match user interests.

[0004] In a first aspect, the present application provides a method for generating a cover image based on bullet screen content, including: calculating the matching degree between each bullet screen text of the bullet screen data of the target video and a preset dictionary, and performing data screening on the bullet screen data according to the matching degree to obtain a candidate bullet screen set, where the bullet screen data includes multiple bullet screen texts and their timestamps; determining, according to the timestamps, a candidate image set for each bullet screen text in the candidate bullet screen set from the target video; calculating the clarity of each image in the candidate image set, and determining, according to the clarity, a candidate cover image for each bullet screen text in the candidate bullet screen set from the candidate image set to obtain a candidate cover image set; performing an aesthetic score on each image in the candidate cover image set to obtain a corresponding score result, and determining, according to the score result, a target cover image set for the target video from the candidate cover image set.

[0005] In a second aspect, the present application provides a device for generating a cover image based on bullet screen content, including: a first calculation module, configured to calculate the matching degree between each bullet screen text of the bullet screen data of the target video and a preset dictionary, and perform data screening on the bullet screen data according to the matching degree to obtain a candidate bullet screen set, where the bullet screen data includes multiple bullet screen texts and their timestamps; a first determination module, configured to determine, according to the timestamps, a candidate image set for each bullet screen text in the candidate bullet screen set from the target video; a second calculation module, configured to calculate the clarity of each image in the candidate image set, and determine, according to the clarity, a candidate cover image for each bullet screen text in the candidate bullet screen set from the candidate image set to obtain a candidate cover image set; a second determination module, configured to perform an aesthetic score on each image in the candidate cover image set to obtain a corresponding score result, and determine, according to the score result, a target cover image set for the target video from the candidate cover image set.

[0006] As an optional example, the device further includes: an acquisition module, configured to acquire the bullet screen data and a noise word set of the target video before calculating the matching degree between each bullet screen text of the bullet screen data of the target video and a preset dictionary; a processing module, configured to use each bullet screen text of the bullet screen data as the current bullet screen text, and perform the following operations on the current bullet screen text: determining whether the current bullet screen text exists in the noise word set; in the case where the current bullet screen text exists in the noise word set, deleting the current bullet screen text from the bullet screen data; in the case where the current bullet screen text does not exist in the noise word set, retaining the current bullet screen text in the bullet screen data.

[0007] As an optional example, the above first calculation module includes: a first processing unit, which takes each piece of barrage text in the above barrage data as the current barrage text and performs the following operations on the current barrage text: calculating the cosine similarity between the current barrage text and each piece of preset text in the above preset dictionary to obtain a plurality of cosine similarities; determining the cosine similarity with the largest value among the above plurality of cosine similarities as the matching degree between the current barrage text and the preset dictionary; in the case where the matching degree between the current barrage text and the preset dictionary is less than the first target threshold, deleting the current barrage text from the above barrage data; in the case where the matching degree between the current barrage text and the preset dictionary is greater than or equal to the above first target threshold, retaining the current barrage text in the above barrage data.

[0008] As an optional example, the above first determination module includes: a second processing unit, which takes each piece of barrage text in the above candidate barrage set as the current barrage text and performs the following operations on the current barrage text: according to the timestamp of the current barrage text, intercepting the candidate video of the current barrage text in the above target video, where the start time point of the candidate video of the current barrage text is before the timestamp of the current barrage text, and the end time point of the candidate video of the current barrage text is after the timestamp of the current barrage text; performing a frame splitting operation on the candidate video of the current barrage text according to a preset frequency to obtain the candidate image sequence of the current barrage text; calculating the cosine similarity between the current barrage text and each image in its candidate image sequence to obtain the matching degree between the current barrage text and each image in its candidate image sequence; deleting the images with a matching degree lower than the second target threshold from the candidate image sequence of the current barrage text to obtain the candidate image set of the current barrage text.

[0009] As an optional example, the above second calculation module includes: a second processing unit, which takes each piece of barrage text in the above candidate barrage set as the current barrage text and performs the following operations on the current barrage text: calculating the sharpness of each image in the candidate image set of the current barrage text using the Laplace transform algorithm; determining the image with the highest sharpness in the candidate image set of the current barrage text as the candidate cover image of the current barrage text.

[0010] As an optional example, the above second determination module includes: a scoring unit, which uses an aesthetic scoring algorithm to perform an aesthetic score on each image in the above candidate cover image set to obtain a corresponding scoring result; a determination unit, which determines the images in the above candidate cover image set with a scoring result greater than the third target threshold as the target cover image set of the above target video.

[0011] As an alternative example, the above device further includes: an optimization module, configured to perform one or more image optimization operations on each image in the target cover image set of the target video after determining the target cover image set of the target video from the candidate cover image set, where the image optimization operations include color adjustment, contrast adjustment, sharpening adjustment, and shape adjustment.

[0012] In a third aspect, the present application provides a storage medium storing a computer program, where the computer program, when run by a processor, executes the above-mentioned cover image generation method based on bullet screen content.

[0013] In a fourth aspect, the present application further provides an electronic device, including a memory and a processor, where a computer program is stored in the memory, and the processor is configured to execute the above-mentioned cover image generation method based on bullet screen content through the computer program.

[0014] The above technical solutions provided by the embodiments of the present application have the following advantages compared with the prior art:

[0015] The present application calculates the matching degree between each bullet screen text in the bullet screen data of the target video and a preset dictionary, and filters the bullet screen data according to the matching degree to obtain a candidate bullet screen set, where the bullet screen data includes multiple bullet screen texts and their timestamps; determines a candidate image set for each bullet screen text in the candidate bullet screen set from the target video according to the timestamp; calculates the clarity of each image in the candidate image set, and determines a candidate cover image for each bullet screen text in the candidate image set according to the clarity to obtain a candidate cover image set; performs an aesthetic score on each image in the candidate cover image set to obtain a corresponding score result, and determines a target cover image set of the target video from the candidate cover image set according to the score result. In the above method, first, the matching degree between the bullet screen text and the preset dictionary is calculated to filter out the relevant candidate bullet screen set, then the candidate image set is extracted according to the bullet screen timestamp, then the clarity of each image is calculated, the clearest image is selected as the candidate cover image, and finally, the candidate cover image is aesthetically scored, and the image with a higher score is selected as the target cover image. By combining the bullet screen content, image clarity, and aesthetic score, it is ensured that the generated cover image not only conforms to the video theme but also has high visual attractiveness, can better capture the user's interest. In addition, it reduces manual intervention, improves automation and processing efficiency, can quickly generate the most attractive cover image for different video contents, enhances the user click-through rate and viewing rate, and improves the overall user satisfaction, thus achieving the purpose of automatically generating video cover images in combination with user interests, and further solving the technical problem that it is difficult for traditional cover image generation methods to accurately match user interests. Brief Description of the Drawings

[0016] The drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with the present application and, together with the specification, are used to explain the principles of the present application.

[0017] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0018] One or more embodiments are exemplarily illustrated by the pictures in the corresponding drawings. These exemplary illustrations do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings represent similar elements, unless otherwise stated, and the drawings in the figures do not constitute a scale limitation.

[0019] Figure 1 is a flowchart of an optional method for generating a cover image based on barrage content according to an embodiment of the present application;

[0020] Figure 2 is a specific implementation flowchart of an optional method for generating a cover image based on barrage content according to an embodiment of the present application;

[0021] Figure 3 is a schematic structural diagram of an optional device for generating a cover image based on barrage content according to an embodiment of the present application;

[0022] Figure 4 is a schematic diagram of an optional electronic device according to an embodiment of the present application. Detailed Description of the Embodiments

[0023] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts fall within the scope of protection of the present application.

[0024] The following disclosure provides many different embodiments or examples for implementing different structures of the present application. To simplify the disclosure of the present application, the components and settings of specific examples are described below. Of course, they are only examples and are not intended to limit the present application. In addition, the present application may repeat reference numerals and / or letters in different examples. This repetition is for the purpose of simplification and clarity and does not itself indicate the relationship between the various embodiments and / or settings discussed.

[0025] According to a first aspect of an embodiment of the present application, a method for generating a cover image based on bullet screen content is provided. Optionally, as Figure 1 shown, the above method includes:

[0026] S102, calculating the matching degree of each bullet screen text in the bullet screen data of the target video with a preset dictionary, and filtering the bullet screen data according to the matching degree to obtain a candidate bullet screen set, where the bullet screen data includes multiple bullet screen texts and their timestamps;

[0027] S104, determining a candidate image set for each bullet screen text in the candidate bullet screen set from the target video according to the timestamp;

[0028] S106, calculating the clarity of each image in the candidate image set, and determining a candidate cover image for each bullet screen text in the candidate bullet screen set from the candidate image set according to the clarity to obtain a candidate cover image set;

[0029] S108, performing an aesthetic score on each image in the candidate cover image set to obtain a corresponding score result, and determining a target cover image set for the target video from the candidate cover image set according to the score result.

[0030] Optionally, in this embodiment, based on the video barrage data, the target cover image of the video is automatically generated through multi-step screening, clarity calculation, and aesthetic scoring. Specifically, first, the barrage data sent by users on the video platform is collected. The barrage data is the comment text and the corresponding timestamps sent by the audience in the video, including some keywords and action descriptions related to the video. By calculating the matching degree of each barrage text with the preset dictionary, the barrage texts highly relevant to the video content are screened out. The preset dictionary contains words or phrases related to the video theme or specific keywords, such as "A famous scene is coming", "High-energy warning ahead", "Such a cool move", "This shot is amazing", "Love this move", "The movement is very smooth", "The picture is so beautiful", etc. According to the matching degree, the candidate barrage set that meets the conditions is screened out to ensure that the screened barrage texts can represent the hotspots of the video or reflect the interests of users. For each screened barrage text, the video frame corresponding to the moment is extracted from the target video according to its timestamp, forming a candidate image set. Each timestamp corresponds to at least one frame of image, and these images represent the moment when the barrage appears, which may be the key pictures that the audience is interested in. The clarity of each image in the candidate image set is calculated, that is, the visual quality of the image is evaluated, which may involve indicators such as the resolution, contrast, or sharpness of the image. Images with high clarity are more suitable as cover images. According to the clarity score, the clearest image is screened out as the candidate cover image of this barrage. The aesthetic score of each candidate cover image is calculated, that is, the visual attractiveness of the image is evaluated. The aesthetic score usually depends on factors such as the color matching, composition, and lighting of the image. The scoring result reflects the visual attractiveness of the image. Images with high scores are more likely to attract the attention of users. According to the aesthetic score, the images with higher scores are selected from the candidate cover image set as the target cover image of the target video.

[0031] Optionally, in this embodiment, by combining the semantic content of the barrage text, the clarity of the image, and the aesthetic score, the video cover image that can most attract users is generated intelligently. Through accurate screening and scoring mechanisms, it can be ensured that the final cover image is not only relevant to the video content but also has strong visual attractiveness, improving the personalization, relevance, and attractiveness of the cover image. At the same time, it reduces manual intervention, improves efficiency, and adapts to the diversity of video content and the diverse interests of the audience.

[0032] As an optional example, before calculating the matching degree of each barrage text of the barrage data of the target video with the preset dictionary, the above method further includes:

[0033] Obtaining the barrage data of the target video and the noise word set;

[0034] Taking each barrage text of the barrage data as the current barrage text, and performing the following operations on the current barrage text:

[0035] Determine whether the current barrage text exists in the noise word set;

[0036] If the current barrage text exists in the noise word set, delete the current barrage text from the barrage data;

[0037] If the current barrage text does not exist in the noise word set, retain the current barrage text in the barrage data.

[0038] Optionally, in this embodiment, before automatically generating the target cover image of the video based on the barrage data of the video, it is necessary to perform data preprocessing on the barrage data to remove noise words. Specifically, obtain the barrage data of the target video and the noise word set. The noise word set is a predefined list of irrelevant words, such as common meaningless words, advertising words, or words unrelated to the video content. These words usually do not help in the generation of the cover image or the analysis of the video content. For each barrage text, determine whether it belongs to the content in the noise word set. If the current barrage text is a noise word (i.e., unrelated to the video content), then delete the text from the barrage data to avoid it affecting subsequent analysis. If the current barrage text does not belong to the noise word set, then retain the text to ensure that only relevant and meaningful text is used for subsequent operations. By eliminating noise words, the quality of the barrage data is improved, making subsequent calculations more accurate, helping to improve the relevance and attractiveness of the video cover image, ensuring that the final cover image better meets the interests of the audience, and enhancing the experience and click-through rate of platform users.

[0039] As an optional example, calculate the matching degree of each barrage text in the barrage data of the target video with a preset dictionary, and perform data screening on the barrage data according to the matching degree. The obtained candidate barrage set includes:

[0040] Take each barrage text in the barrage data as the current barrage text, and perform the following operations on the current barrage text:

[0041] Calculate the cosine similarity between the current barrage text and each preset text in the preset dictionary to obtain multiple cosine similarities;

[0042] Determine the cosine similarity with the largest value among the multiple cosine similarities as the matching degree of the current barrage text with the preset dictionary;

[0043] If the matching degree of the current barrage text with the preset dictionary is less than the first target threshold, delete the current barrage text from the barrage data;

[0044] If the matching degree of the current barrage text with the preset dictionary is greater than or equal to the first target threshold, retain the current barrage text in the barrage data.

[0045] Optionally, in this embodiment, by calculating the similarity between each barrage text and each text in the preset dictionary, high-quality barrage texts related to the video content are screened out. Specifically, the preprocessed barrage data and the preset dictionary are obtained. For each barrage text, the cosine similarity between it and each text in the preset dictionary is calculated. The cosine similarity is the cosine value of the angle between two text vectors. The closer the value is to 1, the more similar the two texts are, and vice versa. For each barrage text, find the maximum value among the similarities with all texts in the preset dictionary, which represents the strongest matching degree between the current barrage text and the preset dictionary. If the matching degree of the current barrage text is lower than the first target threshold (the set lowest matching standard), it is considered that the barrage has nothing to do with the video content and is deleted from the barrage data. If the matching degree of the current barrage text is higher than or equal to the first target threshold, the barrage text is retained, indicating that it has a high correlation with the video content. The finally screened barrage texts form a candidate barrage set, and these texts are highly relevant to the video content. By matching the barrage texts with the preset dictionary through cosine similarity, barrage texts related to the video content are effectively screened out. The quality of the barrage data is improved, irrelevant or low-matching texts are removed, and the retained barrages are ensured to more accurately reflect the theme and hot content of the video. This provides more effective text input for subsequent cover image generation, thereby enhancing the relevance and attractiveness of the cover image, reducing the interference of noise, and improving the user experience.

[0046] As an optional example, according to the timestamp, the candidate image set of each barrage text in the candidate barrage set determined from the target video includes:

[0047] Take each barrage text in the candidate barrage set as the current barrage text, and perform the following operations on the current barrage text:

[0048] According to the timestamp of the current barrage text, intercept the candidate video of the current barrage text in the target video, where the start time point of the candidate video of the current barrage text is before the timestamp of the current barrage text, and the end time point of the candidate video of the current barrage text is after the timestamp of the current barrage text;

[0049] According to the preset frequency, perform frame splitting on the candidate video of the current barrage text to obtain the candidate image sequence of the current barrage text;

[0050] Calculate the cosine similarity between the current barrage text and each image in its candidate image sequence to obtain the matching degree between the current barrage text and each image in its candidate image sequence;

[0051] Delete the images with a matching degree lower than the second target threshold from the candidate image sequence of the current barrage text to obtain the candidate image set of the current barrage text.

[0052] Optionally, in this embodiment, by analyzing the timestamps of the barrage texts and matching them with the images in the target video, a set of candidate images is generated. Specifically, each barrage has a timestamp indicating the appearance time of the barrage in the video. This timestamp is used to locate a specific period in the video, and then relevant images are extracted from the video. According to the timestamp of the barrage text, a candidate video segment is determined. The start time of this segment is before the timestamp of the current barrage, and the end time is after the timestamp of the current barrage, ensuring that the video segment contains the content before and after the appearance of the barrage, which helps to generate images related to the barrage. For example, a video segment within 10 seconds before and after the timestamp is intercepted, thus avoiding the problem that the image calibrated by the timestamp may not be related to the barrage content due to the certain lag of the barrage. For the candidate video segment, frame splitting operations are performed according to a preset frequency (e.g., intercepting one frame per second) to obtain a series of images. In this way, multiple image frames can be extracted from the video segment to form a candidate image sequence. For each image in the candidate image sequence and the current barrage text, the cosine similarity is calculated. The cosine similarity is an index to measure the relevance between the text and the image content. The higher the value, the more similar the two are. Through this method, the correlation degree between each image and the barrage text can be judged. If the cosine similarity of a certain image is lower than the second target threshold, it means that the image is not relevant or has a low correlation with the barrage text and should be deleted from the candidate image sequence. The remaining images are those with a high similarity to the barrage text, forming a set of candidate images for the current barrage text. By generating a set of candidate images based on the similarity between the barrage text and the video images, the relevance of the images is effectively improved. By calculating the cosine similarity and performing threshold screening, irrelevant images are removed, improving the quality of the set of candidate images and providing more accurate and attractive images for the subsequent generation of cover images.

[0053] As an optional example, calculate the clarity of each image in the set of candidate images, and based on the clarity, determine the candidate cover image for each barrage text in the set of candidate barrages from the set of candidate images. The obtained set of candidate cover images includes:

[0054] Take each barrage text in the set of candidate barrages as the current barrage text, and perform the following operations on the current barrage text:

[0055] Use the Laplace transform algorithm to calculate the clarity of each image in the set of candidate images of the current barrage text;

[0056] Determine the image with the highest clarity in the set of candidate images of the current barrage text as the candidate cover image of the current barrage text.

[0057] Optionally, in this embodiment, the optimal cover image is selected by calculating the sharpness of candidate images to ensure that the final cover image has the best visual effect. Specifically, the Laplace transform algorithm is a method commonly used for image sharpness detection. It reflects the details and sharpness of an image by calculating the second derivative of the gray-scale change of the image. Images with higher sharpness usually have more details and more obvious edges. For each bullet screen text, the algorithm performs the Laplace transform on each image in its candidate image set to obtain the sharpness value of the image. The higher the sharpness value, the clearer the details and edges in the image. Among the candidate image set, the image with the highest sharpness is selected as the candidate cover image for the bullet screen text. Images with high sharpness usually can provide a clearer visual effect and are suitable as video covers. For each bullet screen text in the candidate bullet screen set, its corresponding candidate cover image is selected through the above steps, and finally a candidate cover image set is formed. By calculating the sharpness of the image using the Laplace transform algorithm, the clearest image in the candidate images is effectively selected, ensuring the visual effect of the final cover image, improving the quality of the cover image, and ensuring that the cover image is rich in details, clear, and attractive.

[0058] As an optional example, aesthetic scores are given to each image in the candidate cover image set to obtain corresponding scoring results, and based on the scoring results, the target cover image set of the target video is determined from the candidate cover image set, including:

[0059] Using an aesthetic scoring algorithm, aesthetic scores are given to each image in the candidate cover image set to obtain corresponding scoring results;

[0060] The images in the candidate cover image set with scoring results greater than the third target threshold are determined as the target cover image set of the target video.

[0061] Optionally, in this embodiment, the candidate cover images are evaluated by an aesthetic scoring algorithm, and the most visually attractive image is selected as the target cover image. Specifically, the aesthetic scoring algorithm is an algorithm that analyzes and gives a score based on image features (such as color, contrast, composition, brightness, sharpness, etc.). The algorithm takes into account visual beauty and audience acceptance and provides an aesthetic score for each candidate cover image. The higher the aesthetic score, the more the image meets the audience's aesthetic standards in terms of visual effects, artistry and attractiveness. The aesthetic scoring algorithm is applied to each image in the candidate cover image set to obtain a scoring result for each image. According to the scoring result, the images with a score greater than the third target threshold are screened out and determined as the target cover image set of the target video. The third target threshold is a preset score standard. Only when the image score is higher than this threshold, it is considered to have sufficient aesthetic appeal and can be used as the cover image of the video. Or sort from large to small according to the score, and screen out the top few images to determine the target cover image set of the target video. By applying an aesthetic scoring algorithm, we can accurately score and screen candidate cover images, effectively improving the quality of video covers and ensuring that the final selected cover image has high visual appeal.

[0062] As an optional example, after determining a target cover image set of the target video from the candidate cover image set, the method further includes:

[0063] One or more image optimization operations are performed on each image in the target cover image set, wherein the image optimization operations include color adjustment, contrast adjustment, sharpening adjustment, and shape adjustment.

[0064] Optionally, in this embodiment, the visual effect of the target cover image is further enhanced through image optimization operations to ensure that the target cover image can attract the user's attention. Specifically, a series of optimization operations are performed on each image determined as the target cover image to improve its visual quality and attractiveness. The optimization operations include but are not limited to the following:

[0065] (1) Color adjustment: Aims to enhance the color saturation of the image, making the image's tones more vivid and attracting the audience's attention. This operation can improve the overall color balance of the image and prevent the image from looking too monotonous or dull. Color adjustment can include changing the image's color temperature, color saturation or tones to make the image more visually impactful.

[0066] (2) Contrast adjustment: By adjusting the contrast of an image, the brightness difference of the image can be made more obvious, thereby enhancing the visibility and layering of details.

[0067] (3) Sharpening adjustment: By enhancing the edges and details of the image, the image looks clearer and sharper. This operation can highlight the details of the image, especially applicable to parts with important details or parts that need to be emphasized in the image content.

[0068] (4) Shape adjustment: By changing operations such as the scale, cropping, or rotation of the image to improve the composition of the image, making it more balanced and conforming to visual aesthetics.

[0069] Combined with an example for illustration, this application relates to a method for generating a cover image based on bullet screen content. The process of automatically generating a cover image includes screening valid text from video bullet screen data, calculating the matching degree between the text and a dictionary, selecting relevant video segments and generating candidate images. Then, calculating the image clarity and screening the best cover images, performing aesthetic scoring and image optimization, and finally generating high-quality cover images. The specific implementation process is as Figure 2 shown:

[0070] (1) Bullet screen data screening: Extract bullet screen data from the target video, and obtain a preset set of noise words. Judge each bullet screen text. If the text is in the noise word set, delete it; otherwise, retain the valid bullet screen data;

[0071] (2) Matching degree calculation and bullet screen screening: Calculate the cosine similarity between each valid bullet screen text and each preset text in the preset dictionary to obtain multiple similarity values. According to the cosine similarity, retain the bullet screen texts with a matching degree greater than or equal to the set threshold to obtain a candidate bullet screen set.

[0072] (3) Candidate image generation: According to the timestamp of each candidate bullet screen, intercept relevant video segments containing the bullet screen in the video. Split the video segments by frame extraction according to the preset frequency to generate a series of candidate images. Calculate the cosine similarity between each bullet screen text and each candidate image, and screen out the images with high matching degrees to obtain a candidate image set.

[0073] (4) Image clarity evaluation and selection: Use the Laplace transform algorithm to calculate the clarity of each candidate image. According to the clarity value, select the image with the highest clarity corresponding to each bullet screen text as the candidate cover image to form a candidate cover image set.

[0074] (5) Aesthetic scoring and optimization: Perform aesthetic scoring on the candidate cover images to evaluate their visual attractiveness. According to the aesthetic scoring results, screen out the images with scores higher than the set threshold to determine the final target cover image set.

[0075] (6) Image optimization: Perform a series of optimization operations on the target cover images, such as color adjustment, contrast adjustment, sharpening adjustment, and shape adjustment, to improve the image effect.

[0076] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0077] According to another aspect of the embodiments of the present application, there is also provided a cover image generation device based on barrage content, as Figure 3 shown, including:

[0078] A first calculation module 302, configured to calculate the matching degree between each barrage text of the barrage data of the target video and a preset dictionary, and perform data screening on the barrage data according to the matching degree to obtain a candidate barrage set, where the barrage data includes multiple barrage texts and their timestamps;

[0079] A first determination module 304, configured to determine a candidate image set for each barrage text in the candidate barrage set from the target video according to the timestamp;

[0080] A second calculation module 306, configured to calculate the clarity of each image in the candidate image set, and determine a candidate cover image for each barrage text in the candidate barrage set from the candidate image set according to the clarity to obtain a candidate cover image set;

[0081] A second determination module 308, configured to perform an aesthetic score on each image in the candidate cover image set to obtain a corresponding score result, and determine a target cover image set for the target video from the candidate cover image set according to the score result.

[0082] It should be noted that the first calculation module 302 in this embodiment can be used to execute step S102 in the embodiments of the present application, the first determination module 304 in this embodiment can be used to execute step S104 in the embodiments of the present application, the second calculation module 306 in this embodiment can be used to execute step S106 in the embodiments of the present application, and the second determination module 308 in this embodiment can be used to execute step S108 in the embodiments of the present application.

[0083] As an optional example, the above device further includes:

[0084] An acquisition module, configured to acquire the barrage data of the target video and a noise word set before calculating the matching degree between each barrage text of the barrage data of the target video and the preset dictionary;

[0085] A processing module, which takes each barrage text in the barrage data as the current barrage text and performs the following operations on the current barrage text:

[0086] Determine whether the current barrage text exists in the noise word set;

[0087] If the current barrage text exists in the noise word set, delete the current barrage text from the barrage data;

[0088] If the current barrage text does not exist in the noise word set, retain the current barrage text in the barrage data.

[0089] As an optional example, the first calculation module includes:

[0090] A first processing unit, which takes each barrage text in the barrage data as the current barrage text and performs the following operations on the current barrage text:

[0091] Calculate the cosine similarity between the current barrage text and each preset text in the preset dictionary to obtain multiple cosine similarities;

[0092] Determine the cosine similarity with the largest value among the multiple cosine similarities as the matching degree between the current barrage text and the preset dictionary;

[0093] If the matching degree between the current barrage text and the preset dictionary is less than the first target threshold, delete the current barrage text from the barrage data;

[0094] If the matching degree between the current barrage text and the preset dictionary is greater than or equal to the first target threshold, retain the current barrage text in the barrage data.

[0095] As an optional example, the first determination module includes:

[0096] A second processing unit, which takes each barrage text in the candidate barrage set as the current barrage text and performs the following operations on the current barrage text:

[0097] According to the timestamp of the current barrage text, intercept the candidate video of the current barrage text in the target video, where the start time point of the candidate video of the current barrage text is before the timestamp of the current barrage text, and the end time point of the candidate video of the current barrage text is after the timestamp of the current barrage text;

[0098] According to the preset frequency, perform a frame splitting operation on the candidate video of the current barrage text to obtain the candidate image sequence of the current barrage text;

[0099] Calculate the cosine similarity between the current barrage text and each image in its candidate image sequence to obtain the matching degree between the current barrage text and each image in its candidate image sequence;

[0100] Delete the images with a matching degree lower than the second target threshold from the candidate image sequence of the current bullet screen text to obtain the candidate image set of the current bullet screen text.

[0101] As an optional example, the second calculation module includes:

[0102] A second processing unit for using each bullet screen text in the candidate bullet screen set as the current bullet screen text and performing the following operations on the current bullet screen text:

[0103] Use the Laplace transform algorithm to calculate the sharpness of each image in the candidate image set of the current bullet screen text;

[0104] Determine the image with the highest sharpness in the candidate image set of the current bullet screen text as the candidate cover image of the current bullet screen text.

[0105] As an optional example, the second determination module includes:

[0106] A scoring unit for using an aesthetic scoring algorithm to perform aesthetic scoring on each image in the candidate cover image set to obtain the corresponding scoring result;

[0107] A determination unit for determining the images in the candidate cover image set with a scoring result greater than the third target threshold as the target cover image set of the target video.

[0108] As an optional example, the above device further includes:

[0109] An optimization module for performing one or more image optimization operations on each image in the target cover image set after determining the target cover image set of the target video from the candidate cover image set, where the image optimization operations include color adjustment, contrast adjustment, sharpening adjustment, and shape adjustment.

[0110] For other examples of this embodiment, please refer to the above examples and will not be elaborated here.

[0111] Figure 4 is a schematic diagram of an optional electronic device according to an embodiment of the present application, as Figure 4 shown, including a processor 402, a communication interface 404, a memory 406, and a communication bus 408, where the processor 402, the communication interface 404, and the memory 406 complete communication with each other through the communication bus 408, where,

[0112] The memory 406 is used to store a computer program;

[0113] The processor 402, when executing the computer program stored on the memory 406, implements the following steps:

[0114] Calculate the matching degree between each barrage text in the barrage data of the target video and a preset dictionary, and perform data screening on the barrage data according to the matching degree to obtain a candidate barrage set, where the barrage data includes multiple barrage texts and their timestamps;

[0115] Determine a candidate image set for each barrage text in the candidate barrage set from the target video according to the timestamp;

[0116] Calculate the clarity of each image in the candidate image set, and determine a candidate cover image for each barrage text in the candidate barrage set from the candidate image set according to the clarity to obtain a candidate cover image set;

[0117] Perform an aesthetic score on each image in the candidate cover image set to obtain a corresponding score result, and determine a target cover image set for the target video from the candidate cover image set according to the score result.

[0118] Optionally, in this embodiment, the above communication bus may be a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, etc. The communication bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 4 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus. The communication interface is used for communication between the above electronic device and other devices.

[0119] The memory may include a RAM, and may also include a non-volatile memory, for example, at least one disk memory. Optionally, the memory may also be at least one storage device located far from the foregoing processor.

[0120] As an example, the above memory 406 may but is not limited to include the first calculation module 302, the first determination module 304, the second calculation module 306, and the second determination module 308 in the above cover image generation device based on barrage content. In addition, it may also include but is not limited to other module units in the above cover image generation device based on barrage content, which will not be elaborated in this example.

[0121] The above-mentioned processor may be a general-purpose processor, including but not limited to: CPU (Central Processing Unit, central processing unit), NP (Network Processor, network processor), etc.; it may also be a DSP (Digital Signal Processing, digital signal processor), ASIC (Application Specific Integrated Circuit, application-specific integrated circuit), FPGA (Field-Programmable Gate Array, field-programmable gate array) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0122] Optionally, specific examples in this embodiment may refer to the examples described in the above-mentioned embodiment, and will not be elaborated herein.

[0123] Those of ordinary skill in the art can understand that Figure 4 The structure shown is only schematic. The device for implementing the above-mentioned cover image generation method based on bullet screen content may be a terminal device, which may be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a handheld computer, and a mobile Internet device (Mobile Internet Devices, MID), a PAD and other terminal devices. Figure 4 It does not limit the structure of the above-mentioned electronic device. For example, the electronic device may further include more or fewer components (such as a network interface, a display device, etc.) than those shown in Figure 4 or have a different configuration from that shown in Figure 4 shown.

[0124] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above-mentioned embodiments can be completed by instructing the relevant hardware of the terminal device through a program, and the program can be stored in a computer-readable storage medium. The storage medium may include: a flash drive, a ROM, a RAM, a magnetic disk or an optical disc, etc.

[0125] According to another aspect of the embodiments of the present application, there is also provided a computer-readable storage medium, in which a computer program is stored, and when the computer program is run by a processor, it executes the steps in the above-mentioned cover image generation method based on bullet screen content.

[0126] Optionally, in this embodiment, those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the relevant hardware of the terminal device through a program. This program can be stored in a computer-readable storage medium, and the storage medium can include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, an optical disk, etc.

[0127] The serial numbers of the embodiments of the present application above are only for description and do not represent the advantages or disadvantages of the embodiments.

[0128] If the integrated unit in the above embodiments is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in the above computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing one or more computer devices (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application.

[0129] In the above embodiments of the present application, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0130] In the several embodiments provided by the present application, it should be understood that the disclosed client can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some interfaces. The indirect coupling or communication connection of units or modules can be in an electrical or other form.

[0131] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0132] In addition, in each embodiment of the present application, each functional unit may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.

[0133] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A method for generating a cover image based on barrage content, characterized in that: include: Calculate the matching degree between each bullet-screen text of the bullet-screen data of the target video and a preset dictionary, and filter the bullet-screen data according to the matching degree to obtain a candidate bullet-screen set, wherein the bullet-screen data includes multiple bullet-screen texts and their timestamps; Determine, from the target video, a set of candidate images for each barrage text in the candidate barrage set according to the timestamp; Calculating the clarity of each image in the candidate image set, and determining a candidate cover image of each barrage text in the candidate barrage set from the candidate image set according to the clarity, to obtain a candidate cover image set; An aesthetic score is performed on each image in the candidate cover image set to obtain a corresponding score result, and based on the score result, a target cover image set for the target video is determined from the candidate cover image set.

2. The method according to claim 1, characterized in that: Before calculating the matching degree between each bullet-screen text of the bullet-screen data of the target video and the preset dictionary, the method further includes: Obtaining the barrage data and noise word set of the target video; Take each bullet text of the bullet text data as the current bullet text, and perform the following operations on the current bullet text: Determine whether the current bullet text exists in the noise word set; In the case where the current barrage text exists in the noise word set, deleting the current barrage text from the barrage data; When the current barrage text does not exist in the noise word set, the current barrage text is retained in the barrage data.

3. The method according to claim 1, characterized in that The calculation of the matching degree between each bullet screen text of the target video bullet screen data and a preset dictionary, and the bullet screen data is screened according to the matching degree to obtain a candidate bullet screen set including: Take each bullet text of the bullet text data as the current bullet text, and perform the following operations on the current bullet text: Calculating the cosine similarity between the current bullet comment text and each preset text in the preset dictionary to obtain multiple cosine similarities; Determine the cosine similarity with the largest value among the multiple cosine similarities as the matching degree between the current bullet comment text and the preset dictionary; When the matching degree between the current bullet-screen text and the preset dictionary is less than a first target threshold, deleting the current bullet-screen text from the bullet-screen data; When the matching degree between the current bullet-screen text and the preset dictionary is greater than or equal to the first target threshold, the current bullet-screen text is retained in the bullet-screen data.

4. The method according to claim 1, characterized in that: Determining a candidate image set of each barrage text of the candidate barrage set from the target video according to the timestamp includes: Take each bullet text in the candidate bullet text set as the current bullet text, and perform the following operations on the current bullet text: According to the timestamp of the current barrage text, intercepting a candidate video of the current barrage text in the target video, wherein the starting time point of the candidate video of the current barrage text is before the timestamp of the current barrage text, and the ending time point of the candidate video of the current barrage text is after the timestamp of the current barrage text; According to a preset frequency, a frame splitting operation is performed on the candidate video of the current bullet text to obtain a candidate image sequence of the current bullet text; Calculate the cosine similarity between the current bullet text and each image in its candidate image sequence to obtain the matching degree between the current bullet text and each image in its candidate image sequence; The images whose matching degree is lower than the second target threshold are deleted from the candidate image sequence of the current barrage text to obtain a candidate image set of the current barrage text.

5. The method according to claim 1, characterized in that: The step of calculating the definition of each image in the candidate image set, and determining a candidate cover image of each barrage text in the candidate barrage set from the candidate image set according to the definition, and obtaining the candidate cover image set includes: Take each bullet text in the candidate bullet text set as the current bullet text, and perform the following operations on the current bullet text: Using the Laplace transform algorithm to calculate the clarity of each image in the candidate image set of the current bullet text; The image with the highest definition in the candidate image set of the current barrage text is determined as the candidate cover image of the current barrage text.

6. The method according to claim 1, characterized in that The step of performing aesthetic scoring on each image in the candidate cover image set to obtain a corresponding scoring result, and determining a target cover image set for the target video from the candidate cover image set according to the scoring result comprises: Using an aesthetic scoring algorithm, performing an aesthetic scoring on each image in the candidate cover image set to obtain a corresponding scoring result; The images in the candidate cover image set whose scoring results are greater than a third target threshold are determined as the target cover image set for the target video.

7. The method according to claim 1, characterized in that After determining a target cover image set of the target video from the candidate cover image set, the method further includes: One or more image optimization operations are performed on each image in the target cover image set, wherein the image optimization operations include color adjustment, contrast adjustment, sharpening adjustment, and shape adjustment.

8. A device for generating a cover image based on bullet comment content, characterized in that: include: A first calculation module is used to calculate the matching degree between each bullet text of the bullet data of the target video and a preset dictionary, and to screen the bullet data according to the matching degree to obtain a candidate bullet set, wherein the bullet data includes multiple bullet texts and their timestamps; A first determination module is used to determine, from the target video according to the timestamp, a candidate image set of each barrage text in the candidate barrage set; A second calculation module is used to calculate the clarity of each image in the candidate image set, and determine a candidate cover image of each barrage text in the candidate barrage set from the candidate image set according to the clarity, to obtain a candidate cover image set; The second determination module is used to perform aesthetic scoring on each image in the candidate cover image set to obtain a corresponding scoring result, and determine a target cover image set for the target video from the candidate cover image set based on the scoring result.

9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is executed.

10. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to execute the method according to any one of claims 1 to 7 through the computer program.