Video cover picture processing method and related apparatus

By using a method for determining video cover images based on high-quality comments and user interest similarity, the problem of cover images not matching user interests in existing technologies is solved, enabling personalized cover image display and reducing invalid video playback and waste of network resources.

CN115618056BActive Publication Date: 2026-05-19TENCENT TECH (BEIJING) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TENCENT TECH (BEIJING) CO LTD
Filing Date
2021-07-15
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing methods for determining video cover images fail to reflect the personalized interests of different users, resulting in invalid video playback and wasting network resources.

Method used

Based on the relevance of high-quality comments to video frames and the similarity of interests between the commenting user and the target user, a target cover image is determined from the candidate cover images corresponding to high-quality comments, thereby enabling the display of personalized cover images.

Benefits of technology

By displaying personalized cover images, users can quickly find videos that interest them, avoiding unnecessary playback and saving network resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115618056B_ABST
    Figure CN115618056B_ABST
Patent Text Reader

Abstract

The application discloses a video cover picture processing method and related device, and belongs to the technical field of image processing. The video cover picture processing method comprises the following steps: obtaining a high-quality comment published for a to-be-processed video; determining a video frame, which reaches a specified correlation condition in terms of correlation degree between the high-quality comment and the video frame, as a candidate cover picture corresponding to the high-quality comment based on the correlation degree; and determining a target cover picture to be displayed to a target user from the candidate cover picture corresponding to the high-quality comment according to the interest similarity between a comment publishing user publishing the high-quality comment and the target user to be recommended. The method realizes the determination of a personalized cover picture of a video based on user interaction behaviors, makes the personalized cover picture more in line with the interests of the user, and thus the user can quickly find the video that the user wants to watch through the cover picture, invalid playing of the video is avoided, and a large amount of network resources are saved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to video cover image processing methods and related apparatus. Background Technology

[0002] A video cover image is the image displayed to the user before the video plays. Video cover images help users quickly determine if a video is of interest to them. With the widespread use of computers and the rapid development of the internet, the amount of video content circulating online is increasingly abundant. Faced with a massive amount of videos on various video platforms, setting appropriate cover images can help users find videos of interest more quickly, avoid unnecessary playback, and thus save significant network resources.

[0003] Currently, the commonly used methods for determining video cover images fail to reflect whether the video matches the viewing preferences of different users, leading to invalid video playback and wasting significant network resources. Summary of the Invention

[0004] This application provides a video cover image processing method and related apparatus. When distributing a video to a target user, the method determines the target cover image to be displayed to the target user from the candidate cover images corresponding to the high-quality comments based on the similarity of interests between the commenting user who posted the high-quality comment and the target user. This realizes the determination of personalized cover images for videos based on user interaction behavior, making the personalized cover images more in line with the user's interests. As a result, users can quickly find the videos they want to watch through the cover images, avoiding invalid video playback and saving a lot of network resources.

[0005] In a first aspect, embodiments of this application provide a video cover image processing method, the method comprising:

[0006] Get high-quality comments posted for the video to be processed;

[0007] Based on the correlation between the high-quality comments and the video frames in the video to be processed, the video frames whose correlation reaches the specified correlation conditions are determined as the candidate cover images corresponding to the high-quality comments;

[0008] Based on the similarity of interests between the commenting user who published the high-quality comment and the target user to be recommended, a target cover image is determined from the candidate cover images corresponding to the high-quality comment to be displayed to the target user.

[0009] This application provides a method for determining video cover images. Specifically, based on the relevance between high-quality comments on a video to be processed and video frames in the video, video frames whose relevance reaches a specified relevance condition are determined as candidate cover images corresponding to the high-quality comments. Then, based on the interest similarity between the user who posted the high-quality comment and the target user to be recommended, a target cover image to be displayed to the target user is determined from the candidate cover images corresponding to the high-quality comments. The specified relevance condition refers to the condition that the relevance between the high-quality comment and the video frames in the video to be processed reaches a certain level. It can be understood that any relevance between the high-quality comment and the video frames in the video, as long as it is not the lowest relevance, can be considered as reaching a certain level of relevance. It can also be understood that any relevance higher than a certain threshold, such as higher than the average, can be considered a high relevance. This threshold is not a fixed value and can vary depending on different application scenarios. The target cover image determined by this application embodiment can make the cover image displayed to the target user when the video is not played more in line with the target user's interests, thereby enabling the target user to quickly find the video they want to watch through the cover image, avoiding invalid video playback, and saving a significant amount of network resources.

[0010] In one possible implementation, determining the target cover image to be displayed to the target user from the candidate cover images corresponding to the high-quality comment based on the interest similarity between the commenting user who posted the high-quality comment and the target user to be recommended includes:

[0011] Based on the interest similarity between the commenters who published high-quality comments and the target users, commenters whose interest similarity reaches the specified similarity conditions are identified as candidate users;

[0012] The candidate cover image corresponding to the high-quality comments posted by the candidate users is determined as the target cover image.

[0013] This application provides a possible implementation for determining a target cover image. Specifically, firstly, based on the interest similarity between the commenter who posted a high-quality comment and the target user, commenters whose interest similarity reaches a specified similarity condition are identified as candidate users. Then, the candidate cover images corresponding to the high-quality comments posted by these candidate users are determined as the target cover images. The specified similarity condition refers to a certain level of interest similarity between the commenter who posted a high-quality comment and the target user. This means that any interest similarity between the commenter who posted a high-quality comment and the target user, as long as it is not the minimum interest similarity, can be considered to have reached a certain level of interest similarity. Alternatively, it can be understood that any interest similarity above a certain threshold, such as above the average, can be considered a high level of interest similarity. This threshold is not a fixed value and can vary depending on the application scenario. This application embodiment, based on the interest similarity between the commenter who posted a high-quality comment and the target user, can make the cover image displayed to the target user before the video is played more in line with the target user's interests. This allows the target user to quickly find the video they want to watch through the cover image, avoiding invalid video playback and saving significant network resources.

[0014] In one possible implementation, identifying comment posting users whose interest similarity meets a specified similarity condition as candidate users includes:

[0015] Users who post comments and whose interests are more similar to those of the target user than a first target threshold are identified as candidate users.

[0016] This application provides a possible specific implementation for determining candidate users. Specifically, users who post comments with an interest similarity greater than a first target threshold are identified as candidate users. The first target threshold is not a fixed value and can vary depending on the application scenario. The interest similarity between the target user and the comment posting user can be obtained based on their respective interest tags, which are derived from the user's browsing history, comment data, etc. In this application embodiment, candidate users determined based on the interest similarity between comment posting users who published high-quality comments and the target user have a higher relevance to the target user. Therefore, the target cover image determined from the candidate cover image corresponding to the high-quality comments posted by the candidate users is more consistent with the target user's interests.

[0017] In one possible implementation, determining the target cover image to be displayed to the target user from the candidate cover images corresponding to the high-quality comment based on the interest similarity between the commenting user who posted the high-quality comment and the target user to be recommended includes:

[0018] The target cover image is determined from the candidate cover images corresponding to the high-quality comments based on one or more of the following: the quality rate of the high-quality comments, the relevance of the high-quality comments to the video frames, and the interest similarity between the commenting user who published the high-quality comments and the target user.

[0019] This application provides a possible specific implementation for determining a target cover image. Specifically, the target cover image to be displayed to the target user is determined from the candidate cover images corresponding to the high-quality comments based on one or more of the following: the quality rate of high-quality comments, the relevance of high-quality comments to video frames, and the interest similarity between the comment publisher and the target user. Specifically, the target cover image can be determined from the candidate cover images corresponding to the high-quality comments based on the quality rate of high-quality comments and the interest similarity between the comment publisher and the target user; or, it can be determined from the candidate cover images corresponding to the high-quality comments based on the relevance of high-quality comments to video frames and the interest similarity between the comment publisher and the target user; or, it can be determined from the candidate cover images corresponding to the high-quality comments based on the quality rate of high-quality comments, the relevance of high-quality comments to video frames, and the interest similarity between the comment publisher and the target user. In this embodiment of the application, based on the correlation between high-quality comments and candidate cover images, the cover image displayed to the target user when the video is not played can be more in line with the target user's interests. This allows the target user to quickly find the video they want to watch through the cover image, avoids invalid video playback, and saves a lot of network resources.

[0020] In one possible implementation, determining the target cover image from the candidate cover images corresponding to the high-quality comments based on one or more of the high-quality rate of the high-quality comments, the relevance of the high-quality comments to the video frames, and the interest similarity between the commenting user who posted the high-quality comments and the target user includes:

[0021] If the weighted sum of the quality rate of the first high-quality comment, the relevance of the first high-quality comment to the video frame, and the interest similarity between the commenting user who published the first high-quality comment and the target user is greater than the second target threshold, the candidate cover image corresponding to the first high-quality comment is determined as the target cover image.

[0022] This application provides a possible specific implementation for determining a target cover image. Specifically, it first calculates a weighted sum of the quality rate of a first high-quality comment, the relevance of the first high-quality comment to the video frame, and the interest similarity between the commenter who posted the first high-quality comment and the target user. If this weighted sum is greater than a second target threshold, the candidate cover image corresponding to the first high-quality comment is determined as the target cover image. The second target threshold is not a fixed value and can vary depending on different application scenarios. In this application embodiment, the target cover image determined based on the correlation between high-quality comments and candidate cover images is more aligned with the interests of the target user.

[0023] In one possible implementation, determining video frames whose relevance meets a specified relevance condition as candidate cover images corresponding to the high-quality comment includes:

[0024] Video frames with a relevance greater than a first threshold to the high-quality comment are identified as candidate cover images corresponding to the high-quality comment; the relevance is obtained by inputting the video frame and the high-quality comment into a first neural network model.

[0025] This application provides a possible implementation for determining candidate cover images corresponding to high-quality comments. Specifically, video frames and high-quality comments are input into a first neural network model to obtain the correlation between the video frame and the high-quality comment. If the correlation is greater than a first threshold, the video frame is determined as a candidate cover image corresponding to the high-quality comment. The first threshold is not a fixed value and can vary depending on different application scenarios. This application embodiment, by constructing a correspondence between high-quality comments and candidate cover images, can determine the target cover image from the candidate cover images corresponding to high-quality comments based on this correspondence.

[0026] In one possible implementation, obtaining high-quality comments posted for the video to be processed includes:

[0027] Among the comments posted on the video to be processed, comments that meet the specified quality criteria are identified as high-quality comments. The quality rate is obtained by inputting the comments posted on the video to be processed and the comment interaction information into a second neural network model. The comment interaction information includes one or more of the following: the like rate of the comment, the reply rate of the comment, and the richness of the comment.

[0028] This application provides a possible specific implementation for obtaining high-quality comments on a video to be processed. Specifically, comments posted on the video to be processed and comment interaction information are input into a second neural network model to obtain the quality rate of the comments. The comment interaction information includes one or more of the following: the comment's like rate, the comment's reply rate, and the comment's richness. If the quality rate of a comment meets a specified quality condition, the comment is determined to be a high-quality comment. The specified quality condition refers to a condition where the quality rate of the comment reaches a certain level. It can be understood that a quality rate that is not the lowest is considered to meet a certain level of quality. Alternatively, it can be understood that a quality rate higher than a certain threshold, such as above the average, is considered a high-quality rate. This threshold is not a fixed value and can vary depending on the application scenario. The high-quality comments obtained through this application embodiment can reflect the different interests of different users towards videos to the greatest extent, which is beneficial for subsequently building similar interest-based interactive behaviors between high-quality comments and target users, thereby improving the accuracy of determining the target cover image for the target user.

[0029] Secondly, embodiments of this application provide a video cover image processing method, the method comprising:

[0030] The system receives a target cover image of the video to be processed. The target cover image is determined from the candidate cover image corresponding to the high-quality comments based on the interest similarity between the commenting user and the target user to be recommended for the video to be processed. The candidate cover image is determined based on the video frames in the video to be processed where the relevance between the high-quality comments and the video frames in the video to be processed reaches a specified relevance condition.

[0031] In response to the cover display command corresponding to the video to be processed, the target cover image is displayed.

[0032] This application provides a method for displaying video cover images. Specifically, it first receives a target cover image of a video to be processed, and then displays the target cover image to a target user. The target cover image can be determined based on the relevance between high-quality comments on the video to be processed and video frames in the video. Video frames with relevance meeting specified relevance conditions are identified as candidate cover images corresponding to the high-quality comments. Then, based on the similarity of interests between the user who posted the high-quality comment and the target user to be recommended, the target cover image to be displayed to the target user is determined from the candidate cover images corresponding to the high-quality comments. Through this application embodiment, the target cover image displayed to the target user can better match the target user's interests, enabling the target user to quickly find the video they want to watch through the cover image, avoiding invalid video playback, and saving a significant amount of network resources.

[0033] Thirdly, embodiments of this application provide a video cover image processing apparatus, the apparatus comprising:

[0034] The acquisition unit is used to acquire high-quality comments posted for the video to be processed.

[0035] The determining unit is used to determine the video frames whose relevance reaches a specified relevance condition as candidate cover images corresponding to the high-quality comments based on the relevance between the high-quality comments and the video frames in the video to be processed;

[0036] The determining unit is further configured to determine, from the candidate cover images corresponding to the high-quality comments, a target cover image to be displayed to the target user based on the interest similarity between the commenting user who published the high-quality comment and the target user to be recommended.

[0037] This application provides a method for determining video cover images. Specifically, based on the relevance between high-quality comments on a video to be processed and video frames in the video, video frames whose relevance reaches a specified relevance condition are determined as candidate cover images corresponding to the high-quality comments. Then, based on the interest similarity between the user who posted the high-quality comment and the target user to be recommended, a target cover image to be displayed to the target user is determined from the candidate cover images corresponding to the high-quality comments. The specified relevance condition refers to the condition that the relevance between the high-quality comment and the video frames in the video to be processed reaches a certain level. It can be understood that any relevance between the high-quality comment and the video frames in the video, as long as it is not the lowest relevance, can be considered as reaching a certain level of relevance. It can also be understood that any relevance higher than a certain threshold, such as higher than the average, can be considered a high relevance. This threshold is not a fixed value and can vary depending on different application scenarios. The target cover image determined by this application embodiment can make the cover image displayed to the target user when the video is not played more in line with the target user's interests, thereby enabling the target user to quickly find the video they want to watch through the cover image, avoiding invalid video playback, and saving a significant amount of network resources.

[0038] In one possible implementation, the determining unit is specifically configured to determine comment users whose interest similarity reaches a specified similarity condition as candidate users based on the interest similarity between the comment user who posted the high-quality comment and the target user;

[0039] The determining unit is further configured to determine the candidate cover image corresponding to the high-quality comment posted by the candidate user as the target cover image.

[0040] This application provides a possible implementation for determining a target cover image. Specifically, firstly, based on the interest similarity between the commenter who posted a high-quality comment and the target user, commenters whose interest similarity reaches a specified similarity condition are identified as candidate users. Then, the candidate cover images corresponding to the high-quality comments posted by these candidate users are determined as the target cover images. The specified similarity condition refers to a certain level of interest similarity between the commenter who posted a high-quality comment and the target user. This means that any interest similarity between the commenter who posted a high-quality comment and the target user, as long as it is not the minimum interest similarity, can be considered to have reached a certain level of interest similarity. Alternatively, it can be understood that any interest similarity above a certain threshold, such as above the average, can be considered a high level of interest similarity. This threshold is not a fixed value and can vary depending on the application scenario. This application embodiment, based on the interest similarity between the commenter who posted a high-quality comment and the target user, can make the cover image displayed to the target user before the video is played more in line with the target user's interests. This allows the target user to quickly find the video they want to watch through the cover image, avoiding invalid video playback and saving significant network resources.

[0041] In one possible implementation, the determining unit is further configured to identify comment posting users whose interest similarity to the target user is greater than a first target threshold as candidate users.

[0042] This application provides a possible specific implementation for determining candidate users. Specifically, users who post comments with an interest similarity greater than a first target threshold are identified as candidate users. The first target threshold is not a fixed value and can vary depending on the application scenario. The interest similarity between the target user and the comment posting user can be obtained based on their respective interest tags, which are derived from the user's browsing history, comment data, etc. In this application embodiment, candidate users determined based on the interest similarity between comment posting users who published high-quality comments and the target user have a higher relevance to the target user. Therefore, the target cover image determined from the candidate cover image corresponding to the high-quality comments posted by the candidate users is more consistent with the target user's interests.

[0043] In one possible implementation, the determining unit is further configured to determine the target cover image from the candidate cover image corresponding to the high-quality comment based on one or more of the high-quality rate of the high-quality comment, the relevance of the high-quality comment to the video frame, and the interest similarity between the commenting user who published the high-quality comment and the target user.

[0044] This application provides a possible specific implementation for determining a target cover image. Specifically, the target cover image to be displayed to the target user is determined from the candidate cover images corresponding to the high-quality comments based on one or more of the following: the quality rate of high-quality comments, the relevance of high-quality comments to video frames, and the interest similarity between the comment publisher and the target user. Specifically, the target cover image can be determined from the candidate cover images corresponding to the high-quality comments based on the quality rate of high-quality comments and the interest similarity between the comment publisher and the target user; or, it can be determined from the candidate cover images corresponding to the high-quality comments based on the relevance of high-quality comments to video frames and the interest similarity between the comment publisher and the target user; or, it can be determined from the candidate cover images corresponding to the high-quality comments based on the quality rate of high-quality comments, the relevance of high-quality comments to video frames, and the interest similarity between the comment publisher and the target user. In this embodiment of the application, based on the correlation between high-quality comments and candidate cover images, the cover image displayed to the target user when the video is not played can be more in line with the target user's interests. This allows the target user to quickly find the video they want to watch through the cover image, avoids invalid video playback, and saves a lot of network resources.

[0045] In one possible implementation, the determining unit is further configured to determine the candidate cover image corresponding to the first high-quality comment as the target cover image when the weighted sum of the first high-quality comment's quality rate, the first high-quality comment's relevance to the video frame, and the interest similarity between the commenting user who published the first high-quality comment and the target user is greater than a second target threshold.

[0046] This application provides a possible specific implementation for determining a target cover image. Specifically, it first calculates a weighted sum of the quality rate of a first high-quality comment, the relevance of the first high-quality comment to the video frame, and the interest similarity between the commenter who posted the first high-quality comment and the target user. If this weighted sum is greater than a second target threshold, the candidate cover image corresponding to the first high-quality comment is determined as the target cover image. The second target threshold is not a fixed value and can vary depending on different application scenarios. In this application embodiment, the target cover image determined based on the correlation between high-quality comments and candidate cover images is more aligned with the interests of the target user.

[0047] In one possible implementation, the determining unit is further configured to determine video frames whose relevance to the high-quality comment is greater than a first threshold as candidate cover images corresponding to the high-quality comment; the relevance is obtained by inputting the video frame and the high-quality comment into a first neural network model.

[0048] This application provides a possible implementation for determining candidate cover images corresponding to high-quality comments. Specifically, video frames and high-quality comments are input into a first neural network model to obtain the correlation between the video frame and the high-quality comment. If the correlation is greater than a first threshold, the video frame is determined as a candidate cover image corresponding to the high-quality comment. The first threshold is not a fixed value and can vary depending on different application scenarios. This application embodiment, by constructing a correspondence between high-quality comments and candidate cover images, can determine the target cover image from the candidate cover images corresponding to high-quality comments based on this correspondence.

[0049] In one possible implementation, the determining unit is further configured to identify comments with a quality rate that meet a specified quality condition from the comments posted on the video to be processed as high-quality comments; the quality rate is obtained by inputting the comments posted on the video to be processed and comment interaction information into a second neural network model, and the comment interaction information includes one or more of the following: the like rate of the comment, the reply rate of the comment, and the richness of the comment.

[0050] This application provides a possible specific implementation for obtaining high-quality comments on a video to be processed. Specifically, comments posted on the video to be processed and comment interaction information are input into a second neural network model to obtain the quality rate of the comments. The comment interaction information includes one or more of the following: the comment's like rate, the comment's reply rate, and the comment's richness. If the quality rate of a comment meets a specified quality condition, the comment is determined to be a high-quality comment. The specified quality condition refers to a condition where the quality rate of the comment reaches a certain level. It can be understood that a quality rate that is not the lowest is considered to meet a certain level of quality. Alternatively, it can be understood that a quality rate higher than a certain threshold, such as above the average, is considered a high-quality rate. This threshold is not a fixed value and can vary depending on the application scenario. The high-quality comments obtained through this application embodiment can reflect the different interests of different users towards videos to the greatest extent, which is beneficial for subsequently building similar interest-based interactive behaviors between high-quality comments and target users, thereby improving the accuracy of determining the target cover image for the target user.

[0051] Fourthly, embodiments of this application provide a video cover image processing apparatus, the apparatus comprising:

[0052] A receiving unit is configured to receive a target cover image of a video to be processed. The target cover image is determined from a candidate cover image corresponding to a high-quality comment based on the interest similarity between the commenting user and the target user to be recommended for the video to be processed. The candidate cover image is determined based on video frames in the video to be processed where the relevance between the high-quality comment and the video frames in the video to be processed reaches a specified relevance condition.

[0053] The display unit is configured to display the target cover image in response to a cover display command corresponding to the video to be processed.

[0054] This application provides a method for displaying video cover images. Specifically, it first receives a target cover image of a video to be processed, and then displays the target cover image to a target user. The target cover image can be determined based on the relevance between high-quality comments on the video to be processed and video frames in the video. Video frames with relevance meeting specified relevance conditions are identified as candidate cover images corresponding to the high-quality comments. Then, based on the similarity of interests between the user who posted the high-quality comment and the target user to be recommended, the target cover image to be displayed to the target user is determined from the candidate cover images corresponding to the high-quality comments. Through this application embodiment, the target cover image displayed to the target user can better match the target user's interests, enabling the target user to quickly find the video they want to watch through the cover image, avoiding invalid video playback, and saving a significant amount of network resources.

[0055] Fifthly, embodiments of this application provide a video cover image processing apparatus, which includes a processor and a memory; the memory is used to store computer execution instructions; the processor is used to execute the computer execution instructions stored in the memory to cause the video cover image processing apparatus to perform the method as described in the first aspect and any possible implementation thereof. Optionally, the video cover image processing apparatus further includes a transceiver for receiving or transmitting signals.

[0056] Sixthly, embodiments of this application provide a video cover image processing apparatus, the video cover image processing apparatus including a processor and a memory; the memory is used to store computer execution instructions; the processor is used to execute the computer execution instructions stored in the memory to cause the video cover image processing apparatus to perform the method as described in the second aspect above and any possible implementation. Optionally, the video cover image processing apparatus further includes a transceiver, the transceiver being used to receive signals or transmit signals.

[0057] In a seventh aspect, embodiments of this application provide a computer-readable storage medium for storing instructions or a computer program; when the instructions or the computer program are executed, the method described in the first aspect and any possible implementation is implemented; or, when the instructions or the computer program are executed, the method described in the second aspect and any possible implementation is implemented.

[0058] Eighthly, embodiments of this application provide a computer program product comprising instructions or a computer program; when the instructions or the computer program are executed, the method described in the first aspect and any possible implementation is implemented; or, when the instructions or the computer program are executed, the method described in the second aspect and any possible implementation is implemented.

[0059] Ninthly, embodiments of this application provide a chip including a processor configured to execute instructions, wherein when the processor executes the instructions, the chip performs the method as described in the first aspect and any possible implementation; or, when the processor executes the instructions, the chip performs the method as described in the second aspect and any possible implementation. Optionally, the chip further includes a communication interface configured to receive or transmit signals.

[0060] In a tenth aspect, embodiments of this application provide a system comprising at least one video cover image processing device as described in the third or fifth aspect, or a video cover image processing device as described in the fourth or sixth aspect, or a chip as described in the ninth aspect.

[0061] Furthermore, in the process of performing the method described in the first aspect and any possible implementation above, the processes related to sending and / or receiving information in the above methods can be understood as the process of the processor outputting information, and / or the process of the processor receiving input information. When outputting information, the processor can output the information to a transceiver (or communication interface, or transmitting module) so that the transceiver can transmit it. After the information is output by the processor, it may need to undergo other processing before reaching the transceiver. Similarly, when the processor receives input information, the transceiver (or communication interface, or transmitting module) receives the information and inputs it to the processor. Furthermore, after the transceiver receives the information, the information may need to undergo other processing before being input to the processor.

[0062] Based on the above principles, for example, the information sent mentioned in the aforementioned method can be understood as information output by the processor. Similarly, the information received can be understood as information received by the processor from input.

[0063] Optionally, unless otherwise specified, or unless they contradict their actual function or internal logic in the relevant description, the operations of the processor, such as transmitting, sending, and receiving, can be more generally understood as processor output and receiving, input, and other operations.

[0064] Optionally, in performing the methods described in the first aspect and any possible implementation above, the processor may be a processor specifically designed to perform these methods, or it may be a processor that performs these methods by executing computer instructions stored in memory, such as a general-purpose processor. The memory may be a non-transitory memory, such as read-only memory (ROM), which may be integrated with the processor on the same chip or disposed on different chips. This application does not limit the type of memory or the arrangement of the memory and processor.

[0065] In one possible implementation, at least one of the aforementioned memories is located outside the device.

[0066] In yet another possible implementation, at least one of the aforementioned memories is located within the device.

[0067] In another possible implementation, a portion of the memory of the at least one memory is located inside the device, while another portion is located outside the device.

[0068] In this application, the processor and memory may also be integrated into a single device, that is, the processor and memory can be integrated together.

[0069] In this embodiment, when distributing a video to a target user, the target cover image to be displayed to the target user is determined from the candidate cover images corresponding to the high-quality comments based on the similarity of interests between the commenting user who posted the high-quality comment and the target user. This realizes the determination of a personalized cover image for the video based on user interaction behavior, making the personalized cover image more in line with the user's interests. As a result, users can quickly find the videos they want to watch through the cover image, avoiding invalid video playback and saving a lot of network resources. Attached Figure Description

[0070] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0071] Figure 1This application provides a schematic diagram of a scenario for determining a video cover image.

[0072] Figure 2 A schematic diagram of an architecture for determining a video cover image is provided in an embodiment of this application;

[0073] Figure 3 A flowchart illustrating a method for determining a video cover image provided in an embodiment of this application;

[0074] Figure 4 A schematic diagram of a network model provided in an embodiment of this application;

[0075] Figure 5 This is a schematic diagram of another network model provided in an embodiment of this application;

[0076] Figure 6 This is a schematic diagram of the structure of another network model provided in an embodiment of this application;

[0077] Figure 7a A flowchart illustrating a video cover image display method provided in an embodiment of this application;

[0078] Figure 7b This application provides a schematic diagram illustrating the effect of a video cover image display interface.

[0079] Figure 7c This is a schematic diagram illustrating the effect of another video cover image display interface provided in an embodiment of this application;

[0080] Figure 8 A schematic diagram of a video cover image processing device provided in this application embodiment;

[0081] Figure 9 A schematic diagram of a video cover image processing device provided in this application embodiment;

[0082] Figure 10 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0083] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described below with reference to the accompanying drawings.

[0084] The terms "first" and "second," etc., used in the specification, claims, and drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.

[0085] The term "embodiment" as used herein means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0086] It should be understood that in this application, "at least one (item)" means one or more, "more than one" means two or more, "at least two (items)" means two or three or more, and "and / or" is used to describe the relationship between related objects, indicating that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the related objects before and after are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0087] This application provides a method for processing video cover images. To more clearly describe the solution of this application, some knowledge related to video cover images will be introduced below.

[0088] Personalized cover images: When a video is shown to a user before it is played, the cover image is dynamically displayed based on the user's interests. Video cover images can help users quickly understand whether a video is of interest to them.

[0089] Automatic speech recognition (ASR) text: The text content obtained by converting video speech using ASR.

[0090] Video optical character recognition (OCR) text: The text content converted from video images using OCR.

[0091] With the widespread use of computers and the rapid development of the internet, the amount of video content disseminated online is becoming increasingly abundant. Faced with the massive amount of videos on various video platforms, setting appropriate cover images can help users find videos of interest more quickly, avoid unnecessary playback, and thus save a significant amount of network resources.

[0092] Currently, common methods for determining video cover images include manually creating multiple candidate cover images, randomly selecting one or more frames from the video, or directly fixing the most popular image frame in the video as the cover image. However, the video cover images obtained through these methods cannot reflect whether the video matches the viewing preferences of different users, leading to invalid video playback and wasting significant network resources.

[0093] To address the problem of wasted network resources caused by inappropriate video cover image selection in commonly used methods for determining video cover images, this application provides a new architecture for determining video cover images and a new method for doing so. By implementing the architecture and method provided in this application, the target cover image to be displayed to the target user can be determined from the candidate cover images corresponding to high-quality comments based on the similarity of interests between the commenting user who posted the high-quality comment and the target user. This achieves personalized cover images for videos based on user interaction behavior, making the personalized cover images more in line with user interests. Users can then quickly find the videos they want to watch through the cover image, avoiding invalid video playback and saving significant network resources.

[0094] The embodiments of this application are described below with reference to the accompanying drawings.

[0095] Please see Figure 1 , Figure 1 This is a schematic diagram of a scenario for determining a video cover image, provided as an embodiment of this application.

[0096] In this scenario, the electronic device used to determine the video cover image is a device equipped with a processor that can execute computer instructions. Specifically, the electronic device can be a computer, a server, etc.

[0097] like Figure 1As shown, when distributing a video to the current user Uy, a target cover image needs to be selected as the interface displayed to Uy when the video is not playing. To ensure that the selected target cover image better reflects the highlights of the video and better aligns with the target user's interests, this embodiment constructs candidate cover images related to high-quality comments posted by other users on the video based on a large amount of interactive data. Then, candidate cover images related to high-quality comments posted by users with similar interests to the target user are used as the target cover image displayed to the current user Uy.

[0098] For example, based on a large amount of high-quality comment interaction data from other users on the video, high-quality comments C1 posted by user U1, C2 posted by user U2, ..., and Cx posted by user Ux are selected. Then, candidate cover images corresponding to these high-quality comments are constructed. For instance, based on the relevance of high-quality comment C1 to each video frame, video frames whose relevance meets specified relevance conditions are determined as candidate cover image P1 (e.g., the first one or several video frames with the highest relevance can be selected as candidate cover image P1). This candidate cover image P1 is the cover image displayed to user U1 when the video is not played. Similarly, based on the relevance of high-quality comment C2 to each video frame, video frames whose relevance meets specified relevance conditions are selected as candidate cover images P2 (e.g., the first one or several video frames with the highest relevance can be selected as candidate cover images P2). This candidate cover image P2 is the cover image displayed to user U2 when the video is not played. Similarly, based on the relevance of high-quality comment Cx to each video frame, video frames whose relevance meets specified relevance conditions are selected as candidate cover images Px (e.g., the first one or several video frames with the highest relevance can be selected as candidate cover images Px). This candidate cover image Px is the cover image displayed to user Ux when the video is not played. Finally, based on the interest similarity between the current user and the aforementioned users, the cover image corresponding to the high-quality comment posted by the user with the highest similarity is selected as the target cover image displayed to the current user.

[0099] In this scenario, the target cover image is more aligned with the current user's interests, achieving personalized cover image selection based on user interaction. Because platform users comment on videos after watching them, video frames related to popular comments generally have compelling content. Using these as cover images before playback is more likely to attract other users. When displaying the video to the current user, identifying commenters with similar interests from the video's comment section and showcasing video frames related to their high-quality comments as cover images enhances the video's appeal to the current user, allowing them to quickly find videos of interest and avoiding unnecessary playback, thus conserving significant network resources.

[0100] Please see Figure 2 , Figure 2 This is a schematic diagram of an architecture for determining a video cover image, provided in an embodiment of this application.

[0101] like Figure 2 As shown, the architecture for determining the video cover image in this embodiment mainly includes a high-quality comment filtering module, a high-quality comment-related candidate cover image construction module, and a current user target cover image determination module.

[0102] The high-quality comment filtering module is primarily used to select high-quality comments from a large number of user comments on videos, aiming to construct candidate cover images based on these high-quality comments. Specifically, it identifies comments with a quality rate meeting specified quality criteria as high-quality comments (e.g., selecting the top one or a few comments with the highest quality rate). The quality rate represents the probability that a comment is of high quality; the probability of a comment being considered high-quality is understood as the higher the number of likes and replies a comment receives. The quality rate can be obtained by inputting comments and comment interaction information for the video into a second neural network model. This comment interaction information includes one or more of the following: like rate, reply rate, and comment richness. The high-quality comments obtained through this module can reflect the different interest tendencies of users who posted high-quality comments, which is beneficial for building similar interest interactions between these users and the current user, thereby improving the accuracy of determining the target cover image for the current user.

[0103] The module for constructing candidate cover images related to high-quality comments is primarily used to build candidate cover images associated with high-quality comments selected by the high-quality comment filtering module. The purpose is to determine the target cover image for the current user based on these candidate cover images. Specifically, the video can be divided into several video segments. Video segments with a relevance greater than a second threshold to the high-quality comments posted by the user who posted the video are selected as candidate video segments. This second threshold is not a fixed value and can vary depending on the application scenario. Then, frames are extracted from these candidate video segments. Image frames with a relevance greater than a first threshold to the high-quality comments are selected as candidate cover images for those high-quality comments. This first threshold is also not a fixed value and can vary depending on the application scenario. Alternatively, video frames with a relevance greater than the first threshold to the high-quality comments can be directly selected as candidate cover images. The candidate cover images obtained through this module are more aligned with the interests of the users who posted the high-quality comments, which is beneficial for them to engage in similar interactive behaviors with the current user, thereby improving the accuracy of determining the target cover image for the current user.

[0104] The current user's target cover image determination module is primarily used to determine the target cover image to be displayed to the current user based on the aforementioned candidate cover images. Specifically, this can be achieved by finding high-quality commenters whose interests are similar to the current user's, meeting specified similarity criteria, as candidate users (e.g., high-quality commenters whose interest similarity is greater than a first target threshold). The candidate cover image corresponding to the high-quality comment posted by these candidate users is then used as the current user's target cover image. The target cover image is the image displayed to the current user before the video is played. The aforementioned first target threshold is not a fixed value and can vary depending on the application scenario. The interest similarity between the current user and the high-quality commenters can be obtained based on the current user's interest tags and the high-quality commenters' interest tags. The user's interest tags are obtained from their browsing history, historical comment data, etc. The target cover image determined by this module, based on the interaction behavior between the current user and high-quality commenters with similar interests, better matches the current user's interests, enabling the current user to quickly find the video they want to watch, avoiding invalid video playback, and thus saving significant network resources.

[0105] Based on the above-described framework for determining video cover images, this application also provides a new method for determining video cover images, which will be described below in conjunction with... Figure 3 The method for determining the cover image of this video is explained.

[0106] Please see Figure 3 , Figure 3This application provides a flowchart illustrating a method for determining a video cover image, which includes, but is not limited to, the following steps:

[0107] Step 301: Obtain high-quality comments posted for the video to be processed.

[0108] Electronic devices acquire high-quality comments posted on videos to be processed.

[0109] In this application embodiment, the electronic device is a device equipped with a processor that can execute computer execution instructions. Specifically, the electronic device may be a computer, a server, etc.

[0110] In this step, it is necessary to filter through the massive number of comments posted by users on the video to be processed, and obtain the high-quality comments posted on the video.

[0111] Specifically, among the comments posted on the video to be processed, comments that meet the specified quality criteria are identified as high-quality comments. These specified quality criteria refer to the condition that the comment's quality rate reaches a certain level. This means that any quality rate that is not the lowest can be considered to meet the certain quality criteria. Alternatively, a quality rate above a certain threshold can be considered a high quality rate. For example, the quality rates of various comments can be sorted from lowest to highest, and the top one or a few comments with the highest quality rate can be selected as high-quality comments. Another example is that comments with a quality rate higher than a third threshold can be identified as high-quality comments. This third threshold is not a fixed value and can vary depending on the application scenario. The quality rate represents the probability of a comment being of high quality. The probability of a comment being of high quality is the likelihood that it has reached a certain level of quality. This means that the more likes and replies a comment receives, the higher its probability of being high-quality. In one implementation, the quality rate can be obtained using a second neural network model based on the comments and comment interaction information posted on the video to be processed.

[0112] For more details, please refer to [link / reference]. Figure 4 , Figure 4 This is a schematic diagram of a network model provided in an embodiment of this application.

[0113] like Figure 4As shown, user comments on platform videos are input into a second neural network model for selecting high-quality comments. The left side of this model models the interaction information and literal richness of comments using a fully connected network, while the right side models the comment text content using a Bidirectional Encoder Representation from Transformers (BERT) model for deep learning representation. The interaction information and literal richness of comments include, but are not limited to, the following:

[0114] Comment like rate = Number of likes for this comment / Number of views for the current video;

[0115] Comment reply rate = Number of times this comment has been replied to / Number of times the current video has been played;

[0116] Comment richness = min(1.0, comment word count / ML) * comment word complexity, where ML is the average length of comments on the platform, such as ML = 20;

[0117] Comment word complexity = number of non-repeating characters in the comment / total number of characters in the comment.

[0118] It can be seen that the above Figure 4 The second neural network model in the model is trained on high-quality and low-quality video comment data. This enables the model to input the aforementioned statistical features and text content features, and return the comment quality rate PY_c. Therefore, by modeling the interaction information and literal richness of the comment on the left side of the model, and deeply modeling the text content of the comment on the right side, the model can simultaneously combine the interaction and text content of the comment to obtain the comment quality rate, and the accuracy of the quality rate obtained by this model is relatively high. When the comment quality rate PY_c meets a certain threshold, the comment is considered high-quality. For example, comments with a quality rate greater than the third threshold are considered high-quality comments.

[0119] Step 302: Based on the relevance between high-quality comments and video frames in the video to be processed, determine the video frames whose relevance reaches the specified relevance conditions as candidate cover images corresponding to high-quality comments.

[0120] In this step, based on the high-quality comments obtained through the above process and the relevance between these high-quality comments and video frames in the video to be processed, video frames whose relevance meets a specified relevance condition are identified as candidate cover images corresponding to the high-quality comments. The specified relevance condition refers to the condition that the relevance between the high-quality comments and video frames reaches a certain level. This can be understood as any relevance between a high-quality comment and a video frame that is not at the lowest possible level being considered to meet a certain level of relevance. Alternatively, a relevance level above a certain threshold can be considered a high relevance. For example, the relevance between high-quality comments and each video frame can be sorted from lowest to highest, and the top one or several video frames with the highest relevance can be selected as candidate cover images corresponding to that high-quality comment. Another example is that video frames with a relevance higher than a first threshold can also be identified as candidate cover images corresponding to that high-quality comment.

[0121] The process of determining the candidate cover image corresponding to a high-quality comment specifically involves identifying video frames with a relevance greater than a first threshold to the high-quality comment as candidate cover images. The relevance between the video frame and the high-quality comment can be obtained by inputting both the video frame and the high-quality comment into a first neural network model. The first threshold is not a fixed value and can vary depending on the application scenario. In this embodiment, by constructing a correspondence between high-quality comments and candidate cover images, a target cover image can be determined from the candidate cover images corresponding to high-quality comments based on this correspondence.

[0122] In one implementation, the process of determining the candidate cover image corresponding to the high-quality comment can be further divided into two stages. The first stage involves dividing the video slice into several video segments, and selecting video segments from these segments whose second relevance to the high-quality comment is greater than a second threshold as candidate video segments. Then, in the second stage, frames are extracted from these candidate video segments, and image frames from these candidate video segments whose first relevance to the high-quality comment is greater than a first threshold are selected as candidate cover images. The second and first thresholds are not fixed values ​​and can vary depending on the application scenario. The second relevance characterizes the degree of correlation between the video segments in the video and the high-quality comment, while the first relevance characterizes the degree of correlation between the image frames in the video segments and the high-quality comment.

[0123] Furthermore, a second correlation between the video clip and the high-quality comments can be obtained by inputting the text information of the video clip and the high-quality comments into a third neural network model. The text information of the video clip includes the dialogue text or the subtitle text of the video clip.

[0124] For more details, please refer to [link / reference]. Figure 5 , Figure 5This is a schematic diagram of another network model provided in an embodiment of this application.

[0125] like Figure 5 As shown, the current video is sliced ​​into S segments of length dt. The relevance of each video segment to the aforementioned high-quality comments is then determined. The text information of the video segments and the high-quality comments (text of high-quality interactive comments) are input into... Figure 5 The model in this paper determines the relevance between video clips and high-quality comments. The video clips use dialogue text recognized by Automatic Speech Recognition (ASR) and subtitle text recognized by Optical Character Recognition (OCR) as input features. Keyword information is extracted from the video text and used as the input features for the video clips. The textual keyword information of the video clips and the high-quality comments are respectively processed through a Transformer-Encoder multi-layer self-attention model to construct deep representations. Then, the two deep representations are fused to output the relevance probability of the video clips and the high-quality interactive comments (i.e., the second relevance).

[0126] It can be seen that the above Figure 5 The third neural network model in the model is trained on a constructed dataset of positive and negative related / unrelated video clips and comments. This enables the model to take video clip features and high-quality interactive comment features as input and output the correlation probability between the two. Therefore, by constructing deep representations of the text keyword information of the input video clip and high-quality comments respectively, and then fusing the two deep representations to output the correlation probability of the video clip and high-quality interactive comments, the accuracy of the obtained second relevance score can be relatively high. When the second relevance score of a video clip is greater than the aforementioned second threshold, the video clip is determined to be a candidate video clip for further identification.

[0127] Furthermore, by extracting frames from the candidate video segments obtained in the first stage, the image frames and high-quality comments in the candidate video segments can be input into the first neural network model to obtain the first correlation between the image frames in the candidate video segments and the high-quality comments.

[0128] For more details, please refer to [link / reference]. Figure 6 , Figure 6 This is a schematic diagram of another network model provided in an embodiment of this application.

[0129] like Figure 6 As shown, for the video segment with the highest relevance identified in the first stage, F frames are uniformly extracted (e.g., F=20), and the relevance between each image frame and the high-quality comment is calculated. A specific image frame from this most relevant video segment and the high-quality comment (high-quality interactive comment text) are then input into... Figure 6The model in the video segment determines the correlation between a specific image frame in the most relevant video segment and a high-quality comment. The image frame in the most relevant video segment and the high-quality comment are each processed through a Transformer-Encoder multi-layer self-attention model to construct deep representations. Then, the two deep representations are fused to output the correlation probability between the video frame and the high-quality interactive comment (i.e., the aforementioned first correlation score).

[0130] It can be seen that the above Figure 6 The first neural network model in the model is trained on a constructed dataset of positive and negative images and comments that are related or unrelated. This enables the model to input image frame features and high-quality interactive comment features, and output the correlation probability between the two. Therefore, by constructing deep representations for the input image frame information and high-quality comments separately, and then fusing the two deep representations to output the correlation probability between the video frame and the high-quality interactive comment, the accuracy of the first correlation score can be relatively high. When the first correlation score of an image frame in a video segment is greater than the aforementioned first threshold, the image frame is determined to be a candidate cover image. At this time, the first correlation score between the candidate cover image and the high-quality interactive comment is denoted as Pr_c.

[0131] Based on the first and second stages described above, a video can have multiple high-quality interactive comments. Therefore, the methods described in the first and second stages can be used to construct the most relevant personalized candidate cover image for each high-quality interactive comment, as shown in the table below:

[0132] Video v1 High-quality comment c1 Candidate cover image p1 First relevance Pr_c1 High quality rate PY_c1 Candidate user u1 Video v1 High-quality comment c2 Candidate cover image p2 First relevance Pr_c2 High quality rate PY_c2 Candidate user u2 Video v1 High-quality comment c3 Candidate cover image p3 First relevance Pr_c3 High quality rate PY_c3 Candidate user u3 Video v2 High-quality comment c4 Candidate cover image p4 First relevance Pr_c4 High quality rate PY_c4 Candidate user u4 Video v2 High-quality comment c5 Candidate cover image p5 First relevance Pr_c5 High quality rate PY_c5 Candidate User u5 …… …… …… …… …… …… Video vn High-quality comments cn Candidate cover image pn First relevance Pr_cn High quality rate PY_cn Candidate user un

[0133] Step 303: Based on the interest similarity between the commenting user who posted the high-quality comment and the target user to be recommended, determine the target cover image to be displayed to the target user from the candidate cover images corresponding to the high-quality comment.

[0134] In this step, based on the high-quality comments obtained from the above screening and the candidate cover images corresponding to the high-quality comments, the target cover image to be displayed to the target users without playing the video is determined from the candidate cover images corresponding to the high-quality comments, according to the interest similarity between the commenting user who posted the high-quality comment and the target user to be recommended.

[0135] In one implementation, specifically, based on the interest similarity between the commenting user who posted the high-quality comment and the target user, commenting users whose interest similarity reaches a specified similarity condition are identified as candidate users, and then the candidate cover image corresponding to the high-quality comment posted by the candidate user is identified as the target cover image.

[0136] The aforementioned similarity criteria refer to the condition that the interest similarity between the user who posted the high-quality comment and the target user reaches a certain level. This can be understood as meaning that any interest similarity between the user who posted the high-quality comment and the target user, as long as it is not the lowest possible level, can be considered to meet the certain level of interest similarity. Alternatively, it can be understood that any interest similarity exceeding a certain threshold can be considered a high level of interest similarity. For example, the interest similarity between each high-quality comment user and the target user can be sorted from lowest to highest, and the candidate cover image corresponding to the high-quality comment posted by the user with the highest interest similarity can be selected as the target cover image. Another example is that candidate cover images corresponding to high-quality comments posted by users whose interest similarity to the target user is higher than the first target threshold can also be selected as the target cover image. The target cover image is the cover image displayed to the target user before the video is played. The aforementioned first target threshold is not a fixed value and can vary depending on the application scenario. The interest similarity between the target user and the user who posted the high-quality comment can be obtained based on the target user's interest tags and the interest tags of the user who posted the high-quality comment. The user's interest tags are obtained from the user's historical browsing history, historical comment data, etc.

[0137] For example, the interest similarity between the current target user and the user of high-quality interactive comments is denoted as Ps_c. This is calculated by using the Jaccard similarity coefficient (i.e., the number of intersections of interest tags / the number of unions of interest tags) on the two users' interest tag sets, and then using this coefficient as Ps_c. The user's interest tags are then transferred from the tags of the videos that have been played to the user's interest tags based on the user's effective playback history statistics.

[0138] Alternatively, in another implementation, the target cover image is determined from the candidate cover images corresponding to the high-quality comments based on one or more of the high-quality rate of the high-quality comments, the relevance of the high-quality comments to the video frames, and the interest similarity between the commenting user who posted the high-quality comment and the target user.

[0139] Specifically, the target cover image can be determined from the candidate cover images corresponding to high-quality comments based on the quality rate of high-quality comments and the interest similarity between the commenter who posted the high-quality comment and the target user; alternatively, it can be determined based on the relevance of high-quality comments to video frames and the interest similarity between the commenter who posted the high-quality comment and the target user; or, it can be determined based on the quality rate of high-quality comments, the relevance of high-quality comments to video frames, and the interest similarity between the commenter who posted the high-quality comment and the target user. For example, first calculate the weighted sum of the quality rate of the first high-quality comment, the relevance of the first high-quality comment to the video frame, and the interest similarity between the commenter who posted the first high-quality comment and the target user. If this weighted sum is greater than a second target threshold, the candidate cover image corresponding to the first high-quality comment is determined as the target cover image. The second target threshold is not a fixed value and can vary depending on the application scenario.

[0140] For example, the expected score Pi of each personalized candidate cover image for the current target user is calculated as follows: w1 * quality rate PY_c[Pi] + w2 * first relevance Pr_c[Pi] + w3 * interest similarity between the current target user and the user who posted the high-quality comment Ps_c[Pi]. Here, w1, w2, and w3 are weights, and w1 + w2 + w3 = 1.0. The interactive candidate cover image with the highest expected score is selected and displayed to the current target user. This fully utilizes the vast amount of user interaction data on the video platform and combines it with user interests to achieve personalized cover images. This allows the target user to quickly find the video they want to watch, avoiding invalid video playback and thus saving significant network resources.

[0141] The most common methods for determining video cover images are to manually create multiple cover images as candidates, or to arbitrarily select one or more frames from the video as candidates, or to directly fix the most popular image frame in the video as the video cover image.

[0142] Compared with the commonly used methods for determining video cover images, the video cover image determination method in this embodiment can make the cover image displayed to the target user more in line with the target user's interests when the video is not played. It realizes the determination of personalized cover images of videos based on user interaction behavior, enabling target users to quickly find the videos they want to watch, avoiding invalid video playback, and thus saving a lot of network resources.

[0143] Please see Figure 7a , Figure 7a The following is a flowchart illustrating a video cover image display method provided in an embodiment of this application. The method includes, but is not limited to, the following steps:

[0144] Step 701: Receive the target cover image of the video to be processed.

[0145] The electronic device receives a target cover image of the video to be processed. The target cover image can be determined based on the relevance between high-quality comments on the video and video frames within the video. Video frames with relevance meeting specified relevance conditions are identified as candidate cover images corresponding to the high-quality comments. Then, based on the similarity of interests between the user who posted the high-quality comment and the target user to be recommended the image, the target cover image to be displayed to the target user is determined from the candidate cover images corresponding to the high-quality comments. For details on the implementation method of determining the target cover image, please refer to the above. Figure 3 The methods described above will not be elaborated here.

[0146] Furthermore, the electronic device in this application embodiment is a device equipped with a processor that can execute computer execution instructions. Specifically, the electronic device may be a mobile phone, tablet computer, etc.

[0147] Step 702: In response to the cover display instruction for the corresponding video to be processed, display the target cover image.

[0148] The electronic device responds to a cover image display command for a video to be processed and displays the target cover image to the target user when the video to be processed is not being played.

[0149] The cover display instruction can be determined when the electronic device receives a page rendering request corresponding to the video to be processed. This page rendering request may carry data indicating the target cover image, such as the target location where the target cover image is stored, so that the electronic device can obtain the target cover image from the target location based on the page rendering request and render and display it. Alternatively, it may be data directly used to render the target cover image, so that the target cover image is rendered and displayed directly based on the page rendering request. It should be noted that the foregoing is merely illustrative, and other methods can be used to obtain the cover image as needed, which does not constitute a limitation on this application. In some embodiments, the cover display instruction can be generated when the electronic device receives a page rendering request corresponding to the video to be processed (in which case the cover display instruction may carry data indicating the target cover image); in other embodiments, the cover display instruction can also be directly the page rendering request corresponding to the video to be processed received by the electronic device, which is not limited in this embodiment.

[0150] Furthermore, the process of an electronic device receiving page rendering requests corresponding to the aforementioned videos to be processed will be further explained below. In some embodiments, when a target user clicks on a video software or app running on an electronic device, the video software, upon receiving the user's launch command, needs to display a homepage containing several videos to be processed. At this time, the electronic device receives a page rendering request corresponding to the homepage. This page rendering request may carry data indicating the target cover image corresponding to each video to be processed, or data directly used to render the target cover image. In other embodiments, after the video software is launched, the target user clicks on a sub-option on the homepage of the video software, switching to the subpage corresponding to that sub-option. For example, clicking the "Short Video" sub-option on the homepage switches to the "Short Video" subpage. After receiving the user's page switching command, the video software needs to display a subpage containing several short videos. At this time, the electronic device receives a page rendering request corresponding to that subpage. This page rendering request may carry data indicating the target cover image corresponding to the short video contained in the subpage, or data directly used to render the target cover image.

[0151] As can be seen from step 701 above, the electronic device in this application embodiment can be a mobile phone, tablet computer or other device with display function. Therefore, in this step, when the user is not playing the video to be processed, the target cover image of the video to be processed can be displayed on the display page of the mobile phone, tablet computer or other electronic device.

[0152] Specifically, the following will take mobile phones as an example, combined with... Figure 7b and Figure 7c Further explanation of the display effect of the target cover image.

[0153] Please see Figure 7b , Figure 7b This is a schematic diagram illustrating the effect of a video cover image display interface provided in an embodiment of this application.

[0154] like Figure 7b The image shows the display page of a video application running on a mobile phone. Its top navigation bar includes basic information such as network operator, network signal strength, time, and battery level. Icon 7201 represents the target user, i.e., the currently logged-in user. The target cover image displayed in this step is used to show the video to the target user indicated by 7201. Icon 7202 is the search bar, where the target user can enter keywords to search and find the video they want to watch in the search results displayed below the search bar. Below the search bar, brief information about several short videos is displayed, including video v1 indicated by icon 7203, video v2 indicated by icon 7204, video v3 indicated by icon 7205, video v4 indicated by icon 7206, video v5 indicated by icon 7207, and video v6 indicated by icon 7208.

[0155] When the video software starts, it needs to display a page containing the aforementioned videos. Taking video v1 as an example, when the electronic device receives a page rendering request for this page, it can obtain the cover image display instruction for video v1 and respond to the instruction by displaying the target cover image for video v1 to the target user. Since the target cover image is determined based on the similarity of interests between the user who posted a high-quality comment on video v1 and the target user, it maximizes the reflection of whether video v1 matches the target user's interests. This allows the target user to quickly determine whether video v1 is the video they want to watch, avoiding invalid video playback and saving significant network resources. Similarly, the display process, principle, and effect of the cover images for other videos (video v2, video v3, video v4, video v5, and video v6) are similar to those for video v1 and will not be elaborated upon here.

[0156] The above Figure 7b To demonstrate the page display effect of displaying multiple video cover images, the following example will show the page display effect of displaying a single video v1 cover image.

[0157] Please see Figure 7c , Figure 7c This is a schematic diagram illustrating the effect of another video cover image display interface provided in an embodiment of this application.

[0158] like Figure 7c The image shows the display page of a video application running on a mobile phone. Its top navigation bar includes basic information such as network operator, network signal strength, time, and battery level. The area indicated by icon 7301 is the display area for the aforementioned video v1, and the currently logged-in user is still the target user indicated by icon 7201. Below the display area for video v1 are the video description and comments posted by online users. In the display area for comments on video v1, high-quality comments from various users are displayed in descending order of their quality rate. Icon 7302 represents candidate user u1, who posted a high-quality comment c1 on video v1, which corresponds to the aforementioned candidate cover image p1; icon 7303 represents candidate user u2, who posted a high-quality comment c2 on video v1, which corresponds to the aforementioned candidate cover image p2.

[0159] If the target user is not playing video v1, the electronic device responds to the cover display command for video v1 by displaying the target cover image in the display area of ​​video v1 and showing it to the target user. This target cover image is determined based on the similarity of interests between the target user and other candidate users such as candidate user u1 (who posted a high-quality comment c1) and candidate user u2 (who posted a high-quality comment c2) related to video v1. The target cover image is selected from the candidate cover images corresponding to other high-quality comments, such as candidate cover image p1 for high-quality comment c1 and candidate cover image p2 for high-quality comment c2.

[0160] Through the embodiments of this application, the target cover image displayed to the target user can better match the target user's interests, thereby enabling the target user to quickly find the video they want to watch through the cover image, avoiding invalid video playback, and saving a lot of network resources.

[0161] The methods of the embodiments of this application have been described in detail above. The apparatus of the embodiments of this application is provided below.

[0162] Please see Figure 8 , Figure 8 This is a schematic diagram of a video cover image processing device provided in an embodiment of this application. The video cover image processing device 80 may include an acquisition unit 801 and a determination unit 802, wherein the descriptions of each unit are as follows:

[0163] Unit 801 is used to acquire high-quality comments posted on the video to be processed.

[0164] The determining unit 802 is used to determine the video frames whose relevance reaches a specified relevance condition as candidate cover images corresponding to the high-quality comments based on the relevance between the high-quality comments and the video frames in the video to be processed;

[0165] The determining unit 802 is further configured to determine, from the candidate cover images corresponding to the high-quality comments, a target cover image to be displayed to the target user based on the interest similarity between the commenting user who published the high-quality comment and the target user to be recommended.

[0166] This application provides a method for determining video cover images. Specifically, based on the relevance between high-quality comments on a video to be processed and video frames in the video, video frames whose relevance reaches a specified relevance condition are determined as candidate cover images corresponding to the high-quality comments. Then, based on the interest similarity between the user who posted the high-quality comment and the target user to be recommended, a target cover image to be displayed to the target user is determined from the candidate cover images corresponding to the high-quality comments. The specified relevance condition refers to the condition that the relevance between the high-quality comment and the video frames in the video to be processed reaches a certain level. It can be understood that any relevance between the high-quality comment and the video frames in the video, as long as it is not the lowest relevance, can be considered as reaching a certain level of relevance. It can also be understood that any relevance higher than a certain threshold, such as higher than the average, can be considered a high relevance. This threshold is not a fixed value and can vary depending on different application scenarios. The target cover image determined by this application embodiment can make the cover image displayed to the target user when the video is not played more in line with the target user's interests, thereby enabling the target user to quickly find the video they want to watch through the cover image, avoiding invalid video playback, and saving a significant amount of network resources.

[0167] In one possible implementation, the determining unit 802 is specifically used to determine the comment posting user whose interest similarity reaches a specified similarity condition as a candidate user based on the interest similarity between the comment posting user who posted the high-quality comment and the target user;

[0168] The determining unit 802 is further configured to determine the candidate cover image corresponding to the high-quality comment posted by the candidate user as the target cover image.

[0169] This application provides a possible implementation for determining a target cover image. Specifically, firstly, based on the interest similarity between the commenter who posted a high-quality comment and the target user, commenters whose interest similarity reaches a specified similarity condition are identified as candidate users. Then, the candidate cover images corresponding to the high-quality comments posted by these candidate users are determined as the target cover images. The specified similarity condition refers to a certain level of interest similarity between the commenter who posted a high-quality comment and the target user. This means that any interest similarity between the commenter who posted a high-quality comment and the target user, as long as it is not the minimum interest similarity, can be considered to have reached a certain level of interest similarity. Alternatively, it can be understood that any interest similarity above a certain threshold, such as above the average, can be considered a high level of interest similarity. This threshold is not a fixed value and can vary depending on the application scenario. This application embodiment, based on the interest similarity between the commenter who posted a high-quality comment and the target user, can make the cover image displayed to the target user before the video is played more in line with the target user's interests. This allows the target user to quickly find the video they want to watch through the cover image, avoiding invalid video playback and saving significant network resources.

[0170] In one possible implementation, the determining unit 802 is further configured to identify comment posting users whose interest similarity to the target user is greater than a first target threshold as candidate users.

[0171] This application provides a possible specific implementation for determining candidate users. Specifically, users who post comments with an interest similarity greater than a first target threshold are identified as candidate users. The first target threshold is not a fixed value and can vary depending on the application scenario. The interest similarity between the target user and the comment posting user can be obtained based on their respective interest tags, which are derived from the user's browsing history, comment data, etc. In this application embodiment, candidate users determined based on the interest similarity between comment posting users who published high-quality comments and the target user have a higher relevance to the target user. Therefore, the target cover image determined from the candidate cover image corresponding to the high-quality comments posted by the candidate users is more consistent with the target user's interests.

[0172] In one possible implementation, the determining unit 802 is further configured to determine the target cover image from the candidate cover image corresponding to the high-quality comment based on one or more of the high-quality rate of the high-quality comment, the relevance of the high-quality comment to the video frame, and the interest similarity between the commenting user who published the high-quality comment and the target user.

[0173] This application provides a possible specific implementation for determining a target cover image. Specifically, the target cover image to be displayed to the target user is determined from the candidate cover images corresponding to the high-quality comments based on one or more of the following: the quality rate of high-quality comments, the relevance of high-quality comments to video frames, and the interest similarity between the comment publisher and the target user. Specifically, the target cover image can be determined from the candidate cover images corresponding to the high-quality comments based on the quality rate of high-quality comments and the interest similarity between the comment publisher and the target user; or, it can be determined from the candidate cover images corresponding to the high-quality comments based on the relevance of high-quality comments to video frames and the interest similarity between the comment publisher and the target user; or, it can be determined from the candidate cover images corresponding to the high-quality comments based on the quality rate of high-quality comments, the relevance of high-quality comments to video frames, and the interest similarity between the comment publisher and the target user. In this embodiment of the application, based on the correlation between high-quality comments and candidate cover images, the cover image displayed to the target user when the video is not played can be more in line with the target user's interests. This allows the target user to quickly find the video they want to watch through the cover image, avoids invalid video playback, and saves a lot of network resources.

[0174] In one possible implementation, the determining unit 802 is further configured to determine the candidate cover image corresponding to the first high-quality comment as the target cover image when the weighted sum of the first high-quality comment's quality rate, the first high-quality comment's relevance to the video frame, and the interest similarity between the commenting user who published the first high-quality comment and the target user is greater than a second target threshold.

[0175] This application provides a possible specific implementation for determining a target cover image. Specifically, it first calculates a weighted sum of the quality rate of a first high-quality comment, the relevance of the first high-quality comment to the video frame, and the interest similarity between the commenter who posted the first high-quality comment and the target user. If this weighted sum is greater than a second target threshold, the candidate cover image corresponding to the first high-quality comment is determined as the target cover image. The second target threshold is not a fixed value and can vary depending on different application scenarios. In this application embodiment, the target cover image determined based on the correlation between high-quality comments and candidate cover images is more aligned with the interests of the target user.

[0176] In one possible implementation, the determining unit 802 is further configured to determine video frames whose relevance to the high-quality comment is greater than a first threshold as candidate cover images corresponding to the high-quality comment; the relevance is obtained by inputting the video frame and the high-quality comment into a first neural network model.

[0177] This application provides a possible implementation for determining candidate cover images corresponding to high-quality comments. Specifically, video frames and high-quality comments are input into a first neural network model to obtain the correlation between the video frame and the high-quality comment. If the correlation is greater than a first threshold, the video frame is determined as a candidate cover image corresponding to the high-quality comment. The first threshold is not a fixed value and can vary depending on different application scenarios. This application embodiment, by constructing a correspondence between high-quality comments and candidate cover images, can determine the target cover image from the candidate cover images corresponding to high-quality comments based on this correspondence.

[0178] In one possible implementation, the determining unit 802 is further configured to identify comments with a quality rate that meet a specified quality condition as high-quality comments among the comments posted on the video to be processed; the quality rate is obtained by inputting the comments posted on the video to be processed and comment interaction information into a second neural network model, and the comment interaction information includes one or more of the following: the like rate of the comment, the reply rate of the comment, and the richness of the comment.

[0179] This application provides a possible specific implementation for obtaining high-quality comments on a video to be processed. Specifically, comments posted on the video to be processed and comment interaction information are input into a second neural network model to obtain the quality rate of the comments. The comment interaction information includes one or more of the following: the comment's like rate, the comment's reply rate, and the comment's richness. If the quality rate of a comment meets a specified quality condition, the comment is determined to be a high-quality comment. The specified quality condition refers to a condition where the quality rate of the comment reaches a certain level. It can be understood that a quality rate that is not the lowest is considered to meet a certain level of quality. Alternatively, it can be understood that a quality rate higher than a certain threshold, such as above the average, is considered a high-quality rate. This threshold is not a fixed value and can vary depending on the application scenario. The high-quality comments obtained through this application embodiment can reflect the different interests of different users towards videos to the greatest extent, which is beneficial for subsequently building similar interest-based interactive behaviors between high-quality comments and target users, thereby improving the accuracy of determining the target cover image for the target user.

[0180] According to the embodiments of this application, Figure 8The various units in the illustrated device can be individually or entirely combined into one or more other units, or some of the units can be further divided into multiple functionally smaller units. This achieves the same operation without affecting the technical effects of the embodiments of this application. The above-mentioned units are based on logical function division. In practical applications, the function of one unit can also be implemented by multiple units, or the function of multiple units can be implemented by one unit. In other embodiments of this application, the network device may also include other units. In practical applications, these functions can also be implemented with the assistance of other units, and can be implemented collaboratively by multiple units.

[0181] It should be noted that the implementation of each unit can also refer to the above. Figure 3 The corresponding description of the method embodiments shown.

[0182] exist Figure 8 In the described video cover image processing device 80, when distributing a video to a target user, the target cover image to be displayed to the target user is determined from the candidate cover images corresponding to the high-quality comments based on the similarity of interests between the commenting user who posted the high-quality comment and the target user. This realizes the determination of personalized cover images for videos based on user interaction behavior, making the personalized cover images more in line with the user's interests. As a result, users can quickly find the videos they want to watch through the cover images, avoiding invalid video playback and saving a lot of network resources.

[0183] Please see Figure 9 , Figure 9 This is a schematic diagram of a video cover image processing device provided in an embodiment of this application. The video cover image processing device 90 may include a receiving unit 901 and a display unit 902, wherein the descriptions of each unit are as follows:

[0184] The receiving unit 901 is used to receive a target cover image of the video to be processed. The target cover image is determined from the candidate cover image corresponding to the high-quality comment based on the interest similarity between the commenting user and the target user to be recommended for the high-quality comment on the video to be processed. The candidate cover image is determined based on the video frames in the video to be processed whose relevance to the high-quality comment reaches a specified relevance condition.

[0185] Display unit 902 is configured to display the target cover image in response to a cover display command corresponding to the video to be processed.

[0186] This application provides a method for displaying video cover images. Specifically, it first receives a target cover image of a video to be processed, and then displays the target cover image to a target user. The target cover image can be determined based on the relevance between high-quality comments on the video to be processed and video frames in the video. Video frames with relevance meeting specified relevance conditions are identified as candidate cover images corresponding to the high-quality comments. Then, based on the similarity of interests between the user who posted the high-quality comment and the target user to be recommended, the target cover image to be displayed to the target user is determined from the candidate cover images corresponding to the high-quality comments. Through this application embodiment, the target cover image displayed to the target user can better match the target user's interests, enabling the target user to quickly find the video they want to watch through the cover image, avoiding invalid video playback, and saving a significant amount of network resources.

[0187] According to the embodiments of this application, Figure 9 The various units in the illustrated device can be individually or entirely combined into one or more other units, or some of the units can be further divided into multiple functionally smaller units. This achieves the same operation without affecting the technical effects of the embodiments of this application. The above-mentioned units are based on logical function division. In practical applications, the function of one unit can also be implemented by multiple units, or the function of multiple units can be implemented by one unit. In other embodiments of this application, the network device may also include other units. In practical applications, these functions can also be implemented with the assistance of other units, and can be implemented collaboratively by multiple units.

[0188] It should be noted that the implementation of each unit can also refer to the above. Figure 7a The corresponding description of the method embodiments shown.

[0189] exist Figure 9 In the described video cover image processing device 90, when distributing a video to a target user, the target cover image to be displayed to the target user is determined from the candidate cover images corresponding to the high-quality comments based on the similarity of interests between the commenting user who posted the high-quality comment and the target user. This realizes the determination of personalized cover images for videos based on user interaction behavior, making the personalized cover images more in line with the user's interests. As a result, users can quickly find the videos they want to watch through the cover images, avoiding invalid video playback and saving a lot of network resources.

[0190] Please see Figure 10 , Figure 10This is a schematic diagram of the structure of an electronic device 100 provided in an embodiment of this application. The electronic device 100 may include a memory 1001 and a processor 1002. Optionally, it may also include a communication interface 1003 and a bus 1004, wherein the memory 1001, processor 1002, and communication interface 1003 are interconnected via the bus 1004. The communication interface 1003 is used for data interaction with the aforementioned video cover image processing device 80 or video cover image processing device 90.

[0191] The memory 1001 provides storage space, which can store data such as the operating system and computer programs. The memory 1001 includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), or compact disc read-only memory (CD-ROM).

[0192] Processor 1002 is a module that performs arithmetic and logical operations. It can be one or a combination of processing modules such as a central processing unit (CPU), a graphics processing unit (GPU), or a microprocessor unit (MPU).

[0193] The memory 1001 stores a computer program, and the processor 1002 calls the computer program stored in the memory 1001 to execute the above-mentioned... Figure 3 The method for determining the video cover image shown is as follows:

[0194] Get high-quality comments posted for the video to be processed;

[0195] Based on the correlation between the high-quality comments and the video frames in the video to be processed, the video frames whose correlation reaches the specified correlation conditions are determined as the candidate cover images corresponding to the high-quality comments;

[0196] Based on the similarity of interests between the commenting user who published the high-quality comment and the target user to be recommended, a target cover image is determined from the candidate cover images corresponding to the high-quality comment to be displayed to the target user.

[0197] For details regarding the execution method of the processor 1002, please refer to the above. Figure 3 This will not be elaborated upon here.

[0198] Correspondingly, the processor 1002 can also call the computer program stored in the memory 1001 to execute the above-mentioned functions. Figure 8 The specific details of the method steps performed by the acquisition unit 801 and the determination unit 802 in the video cover image processing apparatus 80 shown can be found in the above description. Figure 8 This will not be elaborated upon here.

[0199] On the other hand, the memory 1001 stores a computer program, and the processor 1002 calls the computer program stored in the memory 1001 to execute the above-mentioned... Figure 7a The method for displaying video cover images is as follows:

[0200] The system receives a target cover image of the video to be processed. The target cover image is determined from the candidate cover image corresponding to the high-quality comments based on the interest similarity between the commenting user and the target user to be recommended for the video to be processed. The candidate cover image is determined based on the video frames in the video to be processed where the relevance between the high-quality comments and the video frames in the video to be processed reaches a specified relevance condition.

[0201] In response to the cover display command corresponding to the video to be processed, the target cover image is displayed.

[0202] For details regarding the execution method of the processor 1002, please refer to the above. Figure 7a This will not be elaborated upon here.

[0203] Correspondingly, the processor 1002 can also call the computer program stored in the memory 1001 to execute the above-mentioned functions. Figure 9 The specific details of the method steps performed by the receiving unit 901 and the display unit 902 in the video cover image processing apparatus 90 shown can be found in the above-described method steps. Figure 9 This will not be elaborated upon here.

[0204] exist Figure 10 In the described electronic device 100, when distributing videos to target users, the target cover image to be displayed to the target user is determined from the candidate cover images corresponding to the high-quality comments based on the similarity of interests between the commenting user who posted the high-quality comments and the target user. This realizes the determination of personalized cover images for videos based on user interaction behavior, making the personalized cover images more in line with the user's interests. As a result, users can quickly find the videos they want to watch through the cover images, avoiding invalid video playback and saving a lot of network resources.

[0205] This application also provides a computer-readable storage medium storing a computer program. When the computer program is run on one or more processors, it can perform the above-mentioned tasks. Figure 3 , Figure 7aThe method shown.

[0206] This application also provides a computer program product, which includes a computer program. When the computer program product runs on a processor, it can achieve the above-mentioned... Figure 3 , Figure 7a The method shown.

[0207] This application also provides a chip, which includes a processor for executing instructions. When the processor executes the instructions, it can achieve the above-mentioned... Figure 3 , Figure 7a The method shown. Optionally, the chip also includes a communication interface for inputting or outputting signals.

[0208] This application also provides a system that includes at least one video cover image processing device 80, video cover image processing device 90, electronic device 100, or chip as described above.

[0209] In summary, when distributing videos to target users, the target cover image is determined from the candidate cover images corresponding to the high-quality comments based on the similarity of interests between the commenters who posted the high-quality comments and the target users. This achieves personalized cover images for videos based on user interaction behavior, making the personalized cover images more in line with users' interests. As a result, users can quickly find the videos they want to watch through the cover images, avoiding invalid video playback and saving a lot of network resources.

[0210] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by hardware related to a computer program. The computer program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The aforementioned storage medium includes various media capable of storing computer program code, such as read-only memory (ROM) or random access memory (RAM), magnetic disks, or optical disks.

Claims

1. A method for processing video cover images, characterized in that, include: Among the comments posted on the video to be processed, comments that meet the specified quality criteria are identified as high-quality comments; the quality rate is obtained by inputting the comments posted on the video to be processed and the comment interaction information into a second neural network model, and the comment interaction information includes one or more of the following: the like rate of the comment, the reply rate of the comment, and the richness of the comment; Based on the correlation between the high-quality comments and the video frames in the video to be processed, the video frames whose correlation reaches the specified correlation conditions are determined as the candidate cover images corresponding to the high-quality comments; Based on the similarity of interests between the commenting user who published the high-quality comment and the target user to be recommended, a target cover image is determined from the candidate cover images corresponding to the high-quality comment to be displayed to the target user.

2. The method according to claim 1, characterized in that, The step of determining the target cover image to be displayed to the target user from the candidate cover images corresponding to the high-quality comments based on the interest similarity between the commenting user who published the high-quality comment and the target user to be recommended includes: Based on the interest similarity between the commenters who published high-quality comments and the target users, commenters whose interest similarity reaches the specified similarity conditions are identified as candidate users; The candidate cover image corresponding to the high-quality comments posted by the candidate users is determined as the target cover image.

3. The method according to claim 2, characterized in that, The step of identifying comment posting users whose interest similarity meets specified similarity criteria as candidate users includes: Users who post comments and whose interests are more similar to those of the target user than a first target threshold are identified as candidate users.

4. The method according to claim 1, characterized in that, The step of determining the target cover image to be displayed to the target user from the candidate cover images corresponding to the high-quality comments based on the interest similarity between the commenting user who published the high-quality comment and the target user to be recommended includes: The target cover image is determined from the candidate cover images corresponding to the high-quality comments based on one or more of the following: the quality rate of the high-quality comments, the relevance of the high-quality comments to the video frames, and the interest similarity between the commenting user who published the high-quality comments and the target user.

5. The method according to claim 4, characterized in that, The step of determining the target cover image from the candidate cover images corresponding to the high-quality comments based on one or more of the following: the quality rate of the high-quality comments, the relevance of the high-quality comments to the video frames, and the interest similarity between the commenting user who posted the high-quality comments and the target user, includes: If the weighted sum of the quality rate of the first high-quality comment, the relevance of the first high-quality comment to the video frame, and the interest similarity between the commenting user who published the first high-quality comment and the target user is greater than the second target threshold, the candidate cover image corresponding to the first high-quality comment is determined as the target cover image.

6. The method according to any one of claims 1 to 5, characterized in that, The step of determining video frames whose relevance meets the specified relevance conditions as candidate cover images corresponding to the high-quality comments includes: Video frames with a relevance greater than a first threshold to the high-quality comment are identified as candidate cover images corresponding to the high-quality comment; the relevance is obtained by inputting the video frame and the high-quality comment into a first neural network model.

7. A method for processing video cover images, characterized in that, include: The system receives a target cover image of the video to be processed. The target cover image is determined from the candidate cover image corresponding to the high-quality comments based on the interest similarity between the commenting user and the target user to be recommended for the video to be processed. The candidate cover image is determined based on the video frames in the video to be processed where the relevance between the high-quality comments and the video frames in the video to be processed reaches a specified relevance condition. In response to the cover display command corresponding to the video to be processed, the target cover image is displayed; Specifically, among the comments posted on the video to be processed, comments that meet the specified quality criteria are identified as high-quality comments. The quality rate is obtained by inputting the comments posted on the video to be processed and the comment interaction information into a second neural network model. The comment interaction information includes one or more of the following: the like rate of the comment, the reply rate of the comment, and the richness of the comment.

8. A video cover image processing device, characterized in that, include: The acquisition unit is used to identify comments that meet a specified quality criteria from the comments posted on the video to be processed as high-quality comments; the quality rate is obtained by inputting the comments posted on the video to be processed and the comment interaction information into a second neural network model, and the comment interaction information includes one or more of the following: the like rate of the comment, the reply rate of the comment, and the richness of the comment; The determining unit is used to determine the video frames whose relevance reaches a specified relevance condition as candidate cover images corresponding to the high-quality comments based on the relevance between the high-quality comments and the video frames in the video to be processed; The determining unit is further configured to determine, from the candidate cover images corresponding to the high-quality comments, a target cover image to be displayed to the target user based on the interest similarity between the commenting user who published the high-quality comment and the target user to be recommended.

9. A video cover image processing device, characterized in that, include: A receiving unit is configured to receive a target cover image of a video to be processed. The target cover image is determined from a candidate cover image corresponding to a high-quality comment based on the interest similarity between the commenting user and the target user to be recommended for the video to be processed. The candidate cover image is determined based on video frames in the video to be processed where the relevance between the high-quality comment and the video frames in the video to be processed reaches a specified relevance condition. The display unit is configured to display the target cover image in response to a cover display command corresponding to the video to be processed; Specifically, among the comments posted on the video to be processed, comments that meet the specified quality criteria are identified as high-quality comments. The quality rate is obtained by inputting the comments posted on the video to be processed and the comment interaction information into a second neural network model. The comment interaction information includes one or more of the following: the like rate of the comment, the reply rate of the comment, and the richness of the comment.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store instructions or computer programs; when the instructions or computer programs are executed, they implement the method as described in any one of claims 1-7.

11. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed, they implement the method as described in any one of claims 1-7.