Video cover generation method, device, equipment, storage medium and program product
By automatically generating video covers, using neural network models to extract image and text features, and combining them with target object labels, we solved the problems of low efficiency and poor click-through rate in barrage-style video cover generation, achieving more efficient cover generation and improved click-through rate.
Patent Information
- Application Number
- CN202111546389.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-16
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2041-12-16
AI Technical Summary
In the existing technology, the generation efficiency of barrage-style video covers is low and needs to be manually produced. It cannot fit the preferences of different types of users, resulting in poor click-through rate improvement effect.
By obtaining the original cover image and associated text of the video, the cover text is automatically determined based on the correlation, and fusion processing is performed to generate the target video cover. The neural network model is used to extract image and text features, and the cover is personalized in combination with the target object label.
The efficiency of video cover generation has been improved. The generated cover can more accurately reflect the key content of the video, increase the click-through rate of the video, and avoid the problem of manually selecting related text that cannot suit user preferences.
Smart Images

Figure CN116266193B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of multimedia technology, and in particular to a method, apparatus, device, storage medium, and program product for generating a video cover. Background Art
[0002] With the popularity of barrage (bullet-screen) culture, video publishers and multimedia platforms are gradually experimenting with creating video covers in the form of barrages. Barrage-style covers can intuitively reflect the video's engagement and key content, and compared to ordinary covers, they can increase video click-through rates and user engagement.
[0003] In related technologies, barrage-style covers need to be manually constructed by video creators, that is, manually selecting barrage content and then placing each barrage content in the appropriate position on the cover.
[0004] However, the generation efficiency of the above-mentioned barrage-style cover is low, and the video publisher needs to manually create the cover, which requires a high degree of professionalism from the video publisher. In addition, the barrage in the cover is selected based on personal preferences, which cannot well fit the preferences of different types of users in the multimedia platform, and has a poor effect on improving the click-through rate of the video. Summary of the Invention
[0005] The embodiments of the present application provide a method, apparatus, device, storage medium, and program product for generating a video cover, which can improve the efficiency of generating video covers, optimize the expressive effect of video covers, and increase the click-through rate of videos. The technical solution is as follows:
[0006] In one aspect, an embodiment of the present application provides a method for generating a video cover, the method comprising:
[0007] Obtaining an original cover image and associated text of the video, wherein the associated text includes at least one of a video comment and a video screencap;
[0008] Determining a cover text based on the correlation between the original cover image and each of the associated texts, wherein the correlation between the cover text and the original cover image is higher than the correlation between the other associated texts and the original cover image;
[0009] The original cover image and the cover text are fused to generate a target video cover.
[0010] On the other hand, an embodiment of the present application provides a device for generating a video cover, the device comprising:
[0011] An acquisition module, configured to acquire an original cover image and associated text of a video, wherein the associated text includes at least one of a video comment and a video screencap;
[0012] A first determining module is configured to determine a cover text based on the correlation between the original cover image and each of the associated texts, wherein the correlation between the cover text and the original cover image is higher than the correlation between the other associated texts and the original cover image;
[0013] The image processing module is used to fuse the original cover image and the cover text to generate a target video cover.
[0014] On the other hand, an embodiment of the present application provides a computer device, which includes a processor and a memory; the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set are loaded and executed by the processor to implement the method for generating a video cover as described in the above aspects.
[0015] On the other hand, an embodiment of the present application provides a computer-readable storage medium, in which at least one computer program is stored. The computer program is loaded and executed by a processor to implement the method for generating a video cover as described in the above aspects.
[0016] According to one aspect of the present application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the video cover generation method provided in various optional implementations of the above aspects.
[0017] The technical solutions provided by the embodiments of the present application include at least the following beneficial effects:
[0018] In the embodiment of the present application, the cover text is determined based on the correlation between the original cover image of the video and the associated text of the video, so as to automatically generate a video cover that incorporates the associated text. This eliminates the need for the video publisher to manually select the associated text, thereby improving the efficiency of cover generation. Furthermore, by mining the associated text, the associated text with a high correlation with the original cover image is selected as the cover text, so that the target video cover can more accurately reflect the key content of the video. This can prevent the associated text manually selected by the video publisher based on personal preferences from not being consistent with the preferences of other video recommendation targets, further optimizing the expressive effect of the video cover, and thus increasing the video's click-through rate. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 is a schematic diagram of an implementation environment provided by an exemplary embodiment of the present application;
[0020] Figure 2 is a flowchart of a method for generating a video cover provided by an exemplary embodiment of the present application;
[0021] Figure 3 is a schematic diagram of two video covers provided by an exemplary embodiment of the present application;
[0022] Figure 4 is a flowchart of a method for generating a video cover provided by another exemplary embodiment of the present application;
[0023] Figure 5 is a schematic diagram of an image evaluation model provided by an exemplary embodiment of the present application;
[0024] Figure 6 is a flowchart of a method for generating a video cover provided by another exemplary embodiment of the present application;
[0025] Figure 7 is a schematic diagram of a first text evaluation model provided by an exemplary embodiment of the present application;
[0026] Figure 8 is a schematic diagram of a second text evaluation model provided by an exemplary embodiment of the present application;
[0027] Figure 9 is a schematic diagram of a process for determining the vertical position of a cover text provided by an exemplary embodiment of the present application;
[0028] Figure 10 is a flowchart of a method for generating a video cover provided by another exemplary embodiment of the present application;
[0029] Figure 11 This is a logical framework diagram of a method for generating a video cover provided by an exemplary embodiment of the present application;
[0030] Figure 12 This is a structural block diagram of a device for generating a video cover provided by an exemplary embodiment of the present application;
[0031] Figure 13 It is a structural block diagram of a computer device provided by an exemplary embodiment of the present application. DETAILED DESCRIPTION
[0032] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0033] In this document, "plurality" refers to two or more. "And / or" describes a relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone. The character " / " generally indicates an "or" relationship between the associated objects.
[0034] Figure 1 A schematic diagram of an implementation environment provided by an embodiment of the present application is shown. This implementation environment is described by taking the method for generating a video cover as applied to a backend server of a video application as an example. The implementation environment may include: a first terminal 110, a backend server 120, and a second terminal 130.
[0035] A video application is installed and running in the first terminal 110. When the first terminal 110 receives a text posting operation (e.g., a bullet screen posting operation, a comment posting operation, etc.) from a user on a target video, the first terminal 110 sends text posting data to the backend server 120. The text posting data includes the associated text, the video identifier, and the time the text was generated. Figure 1 Only one first terminal 110 is shown in the figure, but in different embodiments, there are multiple other terminals that can access the server 120 and send text publishing data.
[0036] A video application is installed and running in the second terminal 130. When the second terminal 130 receives a display instruction for a video push page, the second terminal 130 sends a video push request to the backend server 120. Figure 1 Only one second terminal 130 is shown in the figure, but in different embodiments, there are multiple other terminals that can access the server 120 and send video push requests.
[0037] Optionally, the applications installed on the first terminal 110 and the second terminal 130 are the same, or the applications installed on the two terminals are the same type of applications on different operating system platforms, or the applications installed on the two terminals are different. The first terminal 110 can generally refer to one of multiple terminals, and the second terminal 130 can generally refer to another of the multiple terminals. Alternatively, the first terminal 110 and the second terminal 130 can be the same device. This embodiment uses the first terminal 110 and the second terminal 130 as an example. The first terminal 110 and the second terminal 130 can be the same or different device types, including smartphones, tablet computers, e-book readers, MP3 players, MP4 players, smart TVs, in-vehicle terminals, laptop computers, and desktop computers. Optionally, the first terminal 110 can also be used to send video push requests to the backend server 120, and the second terminal 130 can also be used to send text publishing data to the backend server 120.
[0038] When the backend server 120 receives a video push request from the second terminal 130 and determines that the video recommendation page sent to the second terminal 130 contains a video cover of the target video, the backend server 120 obtains the original cover image of the target video and the associated text sent by the first terminal 110, determines the cover text based on the correlation between each associated text and the original cover image, merges the cover text with the original cover image, obtains the target video cover, and sends the target video cover to the second terminal 130.
[0039] The backend server 120 includes a single server, a server cluster consisting of multiple servers, a cloud computing platform, a virtualization center, etc. The backend server 120 is used to provide backend services for video applications.
[0040] The method for generating a video cover in the present application can be executed separately by a terminal installed with a video application. The terminal receives the original cover image and associated text sent by the background server, generates a target video cover, and displays the target video cover through the video recommendation page in the video application. The method can also be executed separately by the background server of the video application. The background server generates a target video cover based on the original cover image and associated text, and sends the target video cover or the video recommendation page containing the target video cover to the terminal for display by the terminal. In addition, the method can also be executed collaboratively by the above-mentioned terminal and the background server. For example, the background server determines the cover text based on the correlation between the original video image and the associated text, and sends the original video image and the cover text to the terminal. The terminal fuses the original cover image and the cover text to generate a target video cover, and displays the target video cover through the video recommendation page in the video application. The following method embodiments are described by taking the method executed by the background server alone as an example.
[0041] Figure 2 A flowchart of a method for generating a video cover provided by an exemplary embodiment of the present application is shown. This embodiment is described by taking the method executed by a backend server alone as an example, and the method includes the following steps.
[0042] Step 201: Obtain the original cover image and associated text of the video.
[0043] The associated text includes at least one of video barrage and video comments.
[0044] The original cover image refers to the cover image that has not been processed by the associated text fusion of the background server, that is, when the background server receives the video file, the cover image contained in the video file and produced by the video publisher through the terminal, or the cover image automatically generated by the background server based on the video file sent by the terminal (for example, the video screen corresponding to a certain video frame in the video file).
[0045] In one possible implementation, a backend server receives a text publishing request from a terminal, including a terminal of a video publisher and a terminal of a viewer. The text publishing request includes associated text, a video identifier, and the time the text was generated. Based on the text publishing request, the backend server forwards the associated text to other terminals, enabling other viewers to view the associated text corresponding to the video while watching the video. Simultaneously, the backend server generates a video cover incorporating the associated text based on the associated text and sends it to the corresponding push terminal. This allows the push recipient to view the video cover incorporating the associated text on the video recommendation page, thereby attracting viewers to the video and increasing the video's click-through rate.
[0046] Step 202: Determine the cover text based on the correlation between the original cover image and each associated text.
[0047] The correlation between the cover text and the original cover image is higher than the correlation between other associated texts and the original cover image.
[0048] Associated texts such as barrages and comments are generated based on subjective information such as the audience's personal preferences and interpretation of the video content. Therefore, the associated texts usually include associated texts that are highly relevant to the video theme, as well as associated texts that are less relevant to the video theme. If the video cover is generated directly based on all the associated texts, on the one hand, if there are a large number of associated texts, the video cover cannot carry them. On the other hand, associated texts that are less relevant to the video theme have little effect on optimizing the cover expression effect, but may reduce the target object's interest in the video. In one possible implementation, since the original cover image is usually information that is more consistent with the video theme or hot spot, the backend server determines the associated text with a higher correlation as the cover text based on the correlation between the associated text and the original cover image.
[0049] For example, the background server performs image recognition and text recognition, and determines the comment content and / or barrage content used to describe the content related to the original cover image as the cover text, or the background server determines the video playback time corresponding to the video content related to the original cover image, and then determines the cover text based on the barrage corresponding to the video playback time.
[0050] Step 203: fuse the original cover image and the cover text to generate a target video cover.
[0051] After the backend server determines the cover text, it places each cover text in the appropriate position on the original cover image, merges the original cover image and cover text, and generates the target video cover. The target video cover or a video recommendation page containing the target video cover is then sent to the target user's corresponding terminal.
[0052] Indicative, such as Figure 3 The figure shows a comparison diagram of a common video cover and a target video cover. Common video cover 301 includes the video image at a certain moment, the number of video plays, the number of barrage comments, and the video title. In addition to the above, target video cover 302 also includes associated text. As can be seen from the figure, target video cover 302 can more directly express the video content and reflect other viewers' feelings and evaluations of the video, thereby achieving the effect of attracting the target audience to watch the video.
[0053] In summary, in the embodiments of the present application, the cover text is determined based on the correlation between the original cover image of the video and the associated text of the video, so as to automatically generate a video cover that incorporates the associated text. This eliminates the need for the video publisher to manually select the associated text, thereby improving the efficiency of cover generation. Furthermore, by mining the associated text, the associated text with a high correlation with the original cover image is selected as the cover text, so that the target video cover can more accurately reflect the key content of the video. This can prevent the associated text manually selected by the video publisher based on personal preferences from not being consistent with the preferences of other video recommendation targets, further optimizing the expressive effect of the video cover, and thus increasing the video's click-through rate.
[0054] In a possible implementation, the backend server divides the video into multiple video segments, calculates the relevance between the original cover image and each video segment, and uses the associated text corresponding to the video segment with high relevance as the selection range of the cover text. Figure 4 A flowchart of a method for generating a video cover provided by another exemplary embodiment of the present application is shown. This embodiment is described by taking the method executed by a backend server alone as an example, and the method includes the following steps.
[0055] Step 401: Obtain the original cover image and associated text of the video.
[0056] The specific implementation of step 401 can refer to the above step 201, and will not be repeated here in this embodiment of the present application.
[0057] Step 402: Determine a target video segment from the video based on the original cover image.
[0058] The correlation between the video screen corresponding to the target video segment and the original cover image is higher than the correlation between the video screen corresponding to other video segments and the original cover image.
[0059] The video cover image is usually a video screen captured by the video publisher based on the theme or highlights of the video, or a video screen corresponding to a video frame extracted by the background server based on video playback data (such as peak interaction time). Therefore, the original cover image can accurately reflect the theme or highlights of the video. The background server divides the video into multiple video segments, and determines the video segment with a high correlation with the original cover image, that is, the target video segment. The target video segment has a high correlation with the original cover image, so the associated text corresponding to the target video segment has a higher correlation with the original cover image than the associated text corresponding to other video segments, that is, it is more in line with the theme and highlights of the video. Selecting the cover text from the associated text corresponding to the target video segment with a high correlation with the original cover image can not only improve the quality of the cover text, but also quickly screen out a large amount of invalid associated text, thereby improving the generation efficiency of the target video cover.
[0060] In one possible implementation, step 402 includes the following steps:
[0061] Step 402a: Segment the video according to the target number of segments or target duration to obtain at least two candidate video segments.
[0062] Optionally, the backend server segments the video based on a target number of segments or a target duration, dividing the video into at least two candidate video segments, which is not limited in the embodiments of the present application. For example, for a video whose duration is less than a first duration threshold, the backend server segments the video based on the target number of segments; for a video whose duration is greater than a second duration threshold, the backend server segments the video based on the target duration, where the second duration threshold is greater than or equal to the first duration threshold.
[0063] Step 402b: Input the original cover image and the video frames corresponding to the candidate video clips into the image evaluation model to obtain the image relevance score corresponding to the candidate video clips.
[0064] The image evaluation model is trained based on positive and negative sample pairs, where the positive sample pair consists of a sample cover image and a related video clip corresponding to the sample cover image, and the negative sample pair consists of a sample cover image and an irrelevant video clip corresponding to the sample cover image.
[0065] In one possible implementation, before the actual application phase, that is, before obtaining the original video cover image and associated text, the backend server needs to train the image evaluation model. The backend server performs model training based on positive and negative sample pairs. The sample data includes sample cover images and video clips. The sample cover image and the related video clip constitute the positive sample pair, and the sample cover image and the unrelated video clip constitute the negative sample pair. Furthermore, to improve the model's learning ability, the unrelated video clip and the related video clip of the same sample cover image belong to the same video, and the time difference between the unrelated and related video clips is small.
[0066] After model training is completed, in the application phase, the backend server inputs the original cover image of the video and the video frames of the candidate video clips into the image evaluation model to obtain the image relevance score. Specifically, step 402b includes the following steps:
[0067] Step 1: Input the original cover image into the first feature extraction network in the image evaluation model to obtain the cover feature vector corresponding to the original cover image.
[0068] The first feature extraction network is a neural network used to identify and extract image features, such as a convolutional neural network (CNN), a recurrent neural network (RNN), and a generative adversarial network (GAN). The backend server inputs the original cover image into the first feature extraction network in the image evaluation model to obtain a deep representation of the original cover image, namely the cover feature vector.
[0069] Step 2: Input the video frames corresponding to the candidate video clips into the second feature extraction network in the image evaluation model to obtain the video frame feature vectors corresponding to each video frame.
[0070] The backend server then feeds the video frames corresponding to the candidate video clips into a second feature extraction network within the image evaluation model to obtain a depth representation of each video frame, namely, a video frame feature vector. Optionally, the second feature extraction network is the same as the first feature extraction network, or different from the first feature extraction network.
[0071] Typically, a candidate video segment contains a large number of video frames. If all video frames were fed into the model for feature extraction, the computational complexity would be high and inefficient. Furthermore, consecutive video frames are relatively correlated and have similar image content. Therefore, in one possible implementation, the backend server extracts a certain number of video frames from the candidate video segment according to a target number of frames or a target frame interval duration and feeds them into the model.
[0072] Step 3: Through the self-attention mechanism in the image evaluation model, the video frame feature vectors corresponding to each video frame are fused to obtain the segment feature vectors of the candidate video segment.
[0073] Since the video frame feature vectors are used to represent the features of the corresponding single video frames, and the background server needs to judge the correlation between the video clips and the original cover images, the background server also needs to fuse the feature vectors of each video frame through the self-attention mechanism to obtain the segment feature vectors of the candidate video clips.
[0074] Step 4: Input the cover feature vector and the segment feature vector into the fully connected layer of the image evaluation model to obtain the image relevance score.
[0075] The backend server generates an interactive representation of the cover image and the video clip through the fully connected layer of the model based on the matrix composed of the model parameters, the cover feature vector and the clip feature vector, and then obtains the probability of correlation between the original plane image and the candidate video clip based on the interactive representation, that is, the image correlation score.
[0076] Figure 5 A schematic diagram of the model architecture of an image evaluation model is shown. The first and second feature extraction networks in the model both use an efficient network (EfficientNet). The backend server inputs the original cover image into the first EfficientNet to obtain a deep representation of the cover image (cover feature vector), inputs the video frames in the candidate video clip into the second EfficientNet and Self-Attention to obtain a deep representation of the video clip (segment feature vector), and then inputs the deep representation of the cover image and the deep representation of the video clip into the fully connected layer to obtain the probability that the candidate video clip is related to the cover image (image relevance score).
[0077] Step 402c: Determine at least one candidate video segment with the highest image relevance score as the target video segment.
[0078] The backend server sorts the candidate video segments in descending order of relevance scores, and determines one or more candidate video segments with the highest image relevance scores as target video segments.
[0079] Step 403: Based on the correlation between the associated text and the original cover image, determine the cover text from the associated text corresponding to the target video clip.
[0080] After the backend server determines the target video segment, it obtains the video time period corresponding to the target video segment and then determines the associated text corresponding to the video time period. For example, for a non-live video, if the backend server determines that the video segment from 1 minute 15 seconds to 1 minute 45 seconds is the target video segment, the backend server will determine the cover text from the video comments played between 1 minute 15 and 1 minute 45 seconds; for a live video, if the backend server determines that the video segment from 1 minute 15 seconds to 1 minute 45 seconds is the target video segment, the backend server will determine the cover text from the video comments received between 1 minute 15 seconds and 1 minute 45 seconds during the live broadcast.
[0081] In one possible implementation, step 403 includes the following steps:
[0082] Step 403a: In response to the number of texts corresponding to the target video segment being lower than or equal to a text number threshold, all associated texts corresponding to the target video segment are determined as cover texts.
[0083] When the text quantity of all associated texts corresponding to the target video clip is less than or equal to the text quantity threshold (for example, the text quantity threshold is 10, and the text quantity of all associated texts is ≤10), the background server directly determines all associated texts corresponding to the target video clip as cover texts.
[0084] Step 403b: In response to the text quantity corresponding to the target video segment being higher than the text quantity threshold, a cover text is determined from the associated text corresponding to the target video segment based on the correlation between the associated text and the original cover image.
[0085] When the number of all associated texts corresponding to the target video clip is higher than the text number threshold (for example, 10), since the video cover cannot carry all the associated texts, the background server further refines the text with a higher correlation with the original cover image from all the associated texts corresponding to the target video clip as the cover text.
[0086] Step 404 : Based on the correlation between the associated text and the target object label, determine the cover text from the associated text corresponding to the target video clip.
[0087] Among them, the target object tag is used to indicate the target object's orientation to the video type, wherein the target object is the video object to be pushed by the background server. After the background server generates the target video cover, it sends the data containing the target video cover to the terminal of the target object.
[0088] Each object has its own object tag. This tag is selected by the object through the tag setting operation and sent by the terminal to the backend server, or is derived by the backend server based on the object's historical video viewing data, with the object's permission. For example, when the total time object A has watched a certain type of video reaches a duration threshold, or the total number of times reaches a count threshold, the tag corresponding to that type of video is added to object A's object tag.
[0089] In one possible implementation, the video application in the embodiment of the present application provides a video recommendation page. When the terminal receives a display instruction for the video recommendation page, the video recommendation page (or the elements used to constitute the video recommendation page) is obtained from the background server. The video recommendation page contains the video cover of each recommended video. Since there are a large number of different types of users in the network platform, the associated text that is attractive to different types of users is also different. Therefore, the background server determines different cover texts based on the object tags of each object and generates personalized target video covers to achieve the effect of attracting various types of objects to watch the video.
[0090] Similarly, in response to the number of texts corresponding to the target video clip being lower than the text number threshold, the background server determines all associated texts corresponding to the target video clip as cover texts; in response to the number of texts corresponding to the target video clip being higher than the text number threshold, the background server determines the cover texts from the associated texts corresponding to the target video clip based on the correlation between the associated texts and the target object label.
[0091] Step 405: fuse the original cover image and the cover text to generate a target video cover.
[0092] The specific implementation of step 405 can refer to the above step 203, and will not be repeated here in this embodiment of the present application.
[0093] In an embodiment of the present application, the characteristic that the original cover image has a strong correlation with the theme and highlights of the video is utilized, a target video clip that is relatively relevant to the original cover image is selected from the video, and the cover text is determined from the associated text corresponding to the target video clip. This can not only improve the accuracy of the cover text, but also quickly eliminate a large amount of irrelevant text, thereby improving the efficiency of cover text determination and target video cover generation.
[0094] The backend server uses a neural network model to determine the correlation between the associated text and the original cover image and the target object label, and pre-trains the model using network big data to improve the accuracy of the cover text. Figure 6 A flowchart of a method for generating a video cover provided by another exemplary embodiment of the present application is shown. This embodiment is described by taking the method executed by a backend server alone as an example, and the method includes the following steps.
[0095] Step 601: Obtain the original cover image and associated text of the video.
[0096] Step 602: Determine a target video segment from the video based on the original cover image.
[0097] The specific implementation of steps 601 to 602 can refer to the above steps 401 to 402, and will not be repeated here in this embodiment of the present application.
[0098] Step 603: Determine the associated text corresponding to the target video clip as a candidate associated text.
[0099] The backend server obtains the target time period corresponding to the target video clip and identifies the associated text played during the target time period as candidate associated text. Since there are many associated texts corresponding to the video clip, placing all of the candidate associated texts on the original cover image would result in a large amount of overlapping text, which would negatively impact the video cover's presentation. Therefore, it is necessary to select from the candidate associated texts those that are relevant to the main content of the video and the user's interests.
[0100] Step 604 : Determine a first text score of the candidate associated text based on the relevance between the candidate associated text and the original cover image.
[0101] In a possible implementation, the backend server determines the cover text by comprehensively analyzing the relevance between each candidate associated text and the original cover image, and the relevance between the candidate associated text and the target object label.
[0102] Step 604 includes the following steps:
[0103] The original cover image and the candidate associated text are input into a first text evaluation model to obtain a first text score.
[0104] The first text evaluation model is trained based on positive and negative sample pairs. A positive sample pair consists of a sample video frame and a positive sample text, and the sample video frame and the positive sample text are played at the same time. A negative sample pair consists of a sample video frame and a negative sample text, and the sample video frame and the negative sample text are played at different times.
[0105] In one possible implementation, before the application phase, that is, before obtaining the original cover image and associated text of the video, the backend server performs model training on the first text evaluation model based on the sample data. Because associated text such as barrage comments typically has a strong correlation with the video frame corresponding to its release time, sample text in the same sample video whose release time coincides with the playback time of the sample video frame is determined as positive sample text, and sample text whose release time does not coincide with the playback time of the sample video frame is determined as negative sample text.
[0106] The sample text is obtained after preliminary manual screening and elimination of related texts that are not highly relevant to the corresponding video content. Schematically, the sample video contains 100 sample video frames. First, the positive sample text corresponding to each sample video frame is obtained through manual screening, and then the negative sample text of each sample video frame is determined based on the time correspondence. For example, for the 50th sample video frame, in order to improve the model capability, the backend server will use the positive sample text corresponding to the 35th to 45th frames and the 65th to 75th frames, which are relatively close to it, as the negative sample text of the 50th sample video frame.
[0107] Specifically, the process of determining the first text score using the first text evaluation model includes the following steps:
[0108] Step 604a: input the original cover image into the third feature extraction network in the first text evaluation model to obtain a cover feature vector corresponding to the original cover image.
[0109] The third feature extraction network is a neural network used to identify and extract image features, such as a CNN, RNN, or GAN. The backend server inputs the original cover image into the third feature extraction network in the image evaluation model to obtain a deep representation of the original cover image, namely, a cover feature vector. Optionally, the first, second, and third feature extraction networks may be the same or different.
[0110] Step 604b: input the candidate associated text into the text feature extraction network in the first text evaluation model to obtain a text feature vector corresponding to the candidate associated text.
[0111] The text feature extraction network is a neural network used to identify and extract text features. Examples include pre-trained language representation models (Bidirectional Encoder Representation from Transformers, BERT), CNNs, and RNNs. The backend server inputs the candidate associated text into the first text evaluation model to obtain a deep representation of the candidate associated text, namely, the text feature vector corresponding to the candidate associated text.
[0112] Step 604c: Input the cover feature vector and the text feature vector into the fully connected layer of the first text evaluation model to obtain a first text score.
[0113] The backend server generates a fusion representation of the cover image and the candidate associated text through the fully connected layer of the model based on the matrix composed of the model parameters, the cover feature vector and the text feature vector, and then obtains the probability of the correlation between the original plane image and the candidate associated text based on the fusion representation, that is, the first text score.
[0114] Figure 7 A schematic diagram of the model architecture of a first text evaluation model is shown. The model's third feature extraction network uses EfficientNet, and the text feature extraction network uses BERT. The backend server inputs the original cover image into EfficientNet to obtain a deep representation of the cover image (cover feature vector). The candidate associated text is input into BERT to obtain a deep representation of the candidate associated text (text feature vector). The deep representation of the cover image and the deep representation of the candidate associated text are then input into a fully connected layer to obtain the probability that the candidate associated text is related to the cover image (the first text score).
[0115] Step 605 : Determine a second text score of the candidate associated text based on the correlation between the candidate associated text and the target object label.
[0116] The target object label is used to indicate the orientation of the target object to the video type.
[0117] The object tag is information that can reflect the object's interest in the video, so the cover text can be determined based on the correlation between the object tag and the associated text. The backend server generates a personalized target video cover for the object to be pushed to the video. For example, the backend server determines that the target video needs to be pushed to the video recommendation page of object A and the video recommendation page of object B. The object tag corresponding to object A is "home, food, movie", and the object tag corresponding to object B is "office, music, pets, books". The target video covers generated by the backend server for object A and object B may be different.
[0118] In one possible implementation, step 605 includes the following steps:
[0119] The target object label and the candidate associated text are input into the second text evaluation model to obtain a second text score.
[0120] The second text evaluation model is obtained based on training of positive and negative sample pairs, wherein the positive sample pair consists of a sample label corresponding to the sample object and a positive sample text, and the positive sample text is the associated text that receives positive feedback interaction operations from the sample object; the negative sample consists of a sample label and a negative sample text, and the negative sample text is the associated text that receives negative feedback interaction operations from the sample object.
[0121] In one possible implementation, before the application stage, that is, before obtaining the original cover image and associated text of the video, the background server performs model training on the second text evaluation model based on the sample data. Users in the platform can interact with the associated text, such as liking the barrage or comments that they are interested in or agree with, that is, performing a positive feedback interaction operation, and blocking the barrage or comments that they are not interested in or disagree with, that is, performing a negative feedback interaction operation. The background server can construct sample data based on historical feedback operations. Sample texts that receive feedback interaction operations are collected, wherein the sample texts that receive positive feedback interaction operations are used as positive sample texts of the sample labels corresponding to the positive feedback interaction operations, and the sample texts that receive negative feedback interaction operations are used as negative sample texts of the sample labels corresponding to the negative feedback interaction operations. The sample label corresponding to the feedback interaction operation refers to the object label corresponding to the object that triggers the feedback interaction operation.
[0122] Specifically, the process of determining the second text score using the second text evaluation model includes the following steps:
[0123] Step 605a: input the target object label into the first text feature extraction network in the second text evaluation model to obtain a label feature vector corresponding to the target object label.
[0124] The first text feature extraction network is a neural network used to identify and extract text features, such as BERT, CNN, or RNN. The backend server inputs the target object label into the first text feature extraction network to obtain a deep representation of the target object label, namely, a label feature vector corresponding to the target object label.
[0125] In one possible implementation, the target object may correspond to multiple target object tags. The backend server generates a target tag sequence corresponding to the target object based on the tag weights corresponding to each target object tag (for example, tags with larger weights are positioned at the front), and inputs the target tag sequence into the first text feature extraction network to obtain a tag feature vector.
[0126] Step 605b: input the candidate associated text into the second text feature extraction network in the second text evaluation model to obtain a text feature vector corresponding to the candidate associated text.
[0127] Similarly, the second text feature extraction network is a neural network for identifying and extracting text features, such as BERT, CNN, RNN, etc. The backend server inputs the candidate associated text into the second text feature extraction network to obtain a deep representation of the candidate associated text, that is, a text feature vector corresponding to the candidate associated text.
[0128] Optionally, the first text feature extraction network is the same as the second text feature extraction network, or the first text feature extraction network is different from the second text feature extraction network.
[0129] Step 605c: input the label feature vector and the text feature vector into the fully connected layer in the second text evaluation model to obtain a second text score.
[0130] The backend server generates a fused representation of the target object label and the candidate associated text through the fully connected layer of the model, based on the matrix composed of the model parameters, the label feature vector, and the text feature vector. Then, based on the fused representation, it obtains the probability of correlation between the target object label and the candidate associated text, that is, the second text score.
[0131] Figure 8 A schematic diagram of the model architecture of a second text evaluation model is shown. The first text feature extraction network in this model uses BERT, and the second text feature extraction network also uses BERT. The backend server inputs the target object label into the first BERT to obtain a deep representation of the target object label (label feature vector), inputs the candidate associated text into the second BERT to obtain a deep representation of the candidate associated text (text feature vector), and then inputs the deep representation of the target object label and the deep representation of the candidate associated text into the fully connected layer to obtain the probability that the candidate associated text is associated with the target object label (the second text score).
[0132] Step 606 : Determine a text relevance score based on the first text score, the first weight corresponding to the first text score, the second text score, and the second weight corresponding to the second text score.
[0133] The backend server combines the first text score and the second text score to obtain a text relevance score, and then determines the cover content that can reflect the relevant information of the original cover image and fit the personal interests of the video recommendation object.
[0134] Schematically, the calculation formula of the text relevance score gra[i] of the candidate associated text i is as follows:
[0135] gra[i]=x1*rel[i]+x2*int[i]
[0136] Where x1 is the first weight, rel[i] is the first text score, x2 is the second weight, and int[i] is the second text score. The first and second weights can be set by the developer based on their needs.
[0137] Step 607: Determine the n candidate associated texts with the highest text relevance scores as cover texts.
[0138] n is a positive integer.
[0139] In one possible implementation, the backend server sorts the candidate associated texts in descending order of text relevance scores, and determines the n candidate associated texts with the highest text relevance scores as cover texts, which are subsequently fused with the original cover image to construct the target video cover.
[0140] Step 608: fuse the original cover image and the cover text to generate a target video cover.
[0141] The backend server will fuse the cover comments determined in the above steps into the original cover image to construct the target video cover. The text relevance scores of each cover text may be different, that is, the degree of relevance between different cover texts and the target subject's personal interests and video content may vary. Since the target subject usually pays attention to the middle part of each video cover when browsing the video recommendation page, the backend server places the cover text with high scores in a position that the target subject can pay attention to first based on the text relevance score, thereby further optimizing the target video cover.
[0142] In one possible implementation, step 608 includes the following steps:
[0143] Step 608a: Determine the horizontal position of each cover text in the original cover image based on the text generation time.
[0144] Since the user's visual focus is usually affected by the vertical position and the horizontal position has little effect, the selection of the horizontal position respects the order of appearance of the cover texts and determines the position of each cover text in the same line according to the order of text generation time.
[0145] For example, the length of the original cover image is L. Among the text generation times corresponding to the cover bullet comments in the same row (i.e., the same vertical position), the time difference between the earliest text generation time da and the latest text generation time db is dt = db - da. Then the horizontal starting position of each cover bullet comment in this row is: L*(text generation time - da) / dt.
[0146] Step 608b: Determine the vertical position of each cover text in the original cover image based on the text relevance score.
[0147] Among them, the vertical distance between the cover text with a high text relevance score and the horizontal center line is smaller than the vertical distance between the cover text with a low text relevance score and the horizontal center line.
[0148] In the vertical direction, the backend server places the cover texts in descending order based on the text relevance scores of the cover texts, from the middle to the top and bottom. Figure 9As shown, the cover text with a high text relevance score is closer to the center line, and the barrage with a low text relevance score is farther away from the center line.
[0149] Schematically, the overall height of the original cover image is h, and the formula for determining the vertical position is as follows:
[0150] h / 2 - h / 2 * (text relevance score - lowest text relevance score) / (highest text relevance score - lowest text relevance score)
[0151] Each Cover Text may be randomly selected to be displayed above or below the center line.
[0152] Step 608c: Draw the cover text on the original cover image according to the horizontal position and vertical position to obtain the target video cover.
[0153] The backend server draws the cover text on the original cover image according to the calculated horizontal and vertical positions to obtain the target video cover. If there is overlap in the cover text, the backend server randomly retains one of the overlapping cover texts, or the front-end performs special processing on the overlapping part (for example, replacing it with an ellipsis).
[0154] Correspondingly, if the backend server determines the cover text only based on the correlation between the candidate associated text and the original cover image, the backend server determines the vertical position based on the first text score; if the backend server determines the cover text only based on the correlation between the candidate associated text and the object tag, the backend server determines the vertical position based on the second text score.
[0155] In an embodiment of the present application, the backend server uses a neural network model to determine the comprehensive correlation between the candidate associated texts and the original cover image and the target object label, and pre-trains the model based on big data, thereby improving the accuracy of cover text selection; on the other hand, the backend server performs fusion processing based on the scores of each cover text, and places the cover text with a high score in a position where the target object is easy to focus, thereby further optimizing the video cover.
[0156] The above embodiments illustrate the process of a computer device generating a target video cover based on the correlation between associated text and an original cover image. In one possible implementation, before generating the target video cover, the computer device first determines whether a cover incorporating associated text is necessary for the video. After generating the target video cover, the computer device updates the video cover as the text information and the interests of the target subject change. Figure 10 A flowchart of a method for generating a video cover provided by another exemplary embodiment of the present application is shown. This embodiment is described by taking the method executed by a backend server alone as an example, and the method includes the following steps.
[0157] Step 1001: Obtain the original cover image and associated text of the video.
[0158] The specific implementation of step 1001 can refer to the above-mentioned step 601, and will not be repeated here in this embodiment of the present application.
[0159] Step 1002: Perform optical character recognition on the original cover image to determine the cover text in the original cover image.
[0160] The original cover image is the cover image produced by the video publisher through the terminal, or it is the cover image automatically generated by the background server based on the video file sent by the terminal. The original cover image may contain important text content. For example, the video publisher adds text to the cover image in the later stage of video production, or the video screen automatically captured by the terminal or background server as the cover image contains important text content. In this case, if a cover with associated text is generated, the cover text will block the text in the original video screen, and there is little point in constructing the target video cover. Therefore, after the background server obtains the original cover image of the video, it performs optical character recognition (OCR) on the original cover image to determine the cover text in the original cover image and determine whether the target cover image needs to be generated.
[0161] Step 1003: In response to the number of words in the cover text being less than a word count threshold, and / or the ratio of the display area of the cover text to the area of the original cover image being less than a ratio threshold, the cover text is determined based on the correlation between the original cover image and each associated text.
[0162] The backend server uses the number of characters or screen size of the original cover image as a basis for determining whether to construct a target cover image. If the original cover image already contains a large amount of text (for example, more than 32 characters) or a large screen size, the cover text will not be determined and the associated text will not be integrated into the original cover image.
[0163] The process of determining the cover text can refer to the above steps 602 to 607, and will not be repeated here in this embodiment of the present application.
[0164] Step 1004: fuse the original cover image and the cover text to generate a target video cover.
[0165] The specific implementation of step 1004 can refer to the above-mentioned step 608, and will not be repeated here in this embodiment of the present application.
[0166] Step 1005: In response to the video meeting the cover update condition, the video cover is updated based on the original cover image and associated text.
[0167] The cover update condition includes at least one of the following: the increment of the associated text reaches a text increment threshold, and / or the increment of the target object tag reaches a tag increment threshold.
[0168] The two main variables that affect video covers are changes in associated text and changes in the interests of the target object. Therefore, as the associated text and the interests of the target object change, the video cover needs to be dynamically updated to generate a better video cover. When the increment of the associated text corresponding to the video reaches the text increment threshold (for example, 500), the backend server updates the video cover based on the original cover image and associated text, and / or, when the increment of the target object tag reaches the tag increment threshold (for example, 3), the backend server updates the video cover based on the original cover image and associated text. Therefore, the video cover of the same video viewed by the same target object at different times may be different.
[0169] It is worth mentioning that when updating the video cover, the target video clip does not need to be re-determined because the original cover image and video content remain unchanged.
[0170] In this embodiment of the present application, when the original video image contains a large amount of text, the backend server does not perform the steps of determining the cover text and generating the target video cover, thereby avoiding the generation of a meaningless video cover that would adversely affect the original cover image. Furthermore, as the associated text and the interests of the target audience change, the backend server dynamically updates the video cover, further enhancing its appeal to the target audience.
[0171] In combination with the above embodiments, Figure 11 A framework diagram of the method for generating a video cover provided by this application is shown.
[0172] When receiving the associated text publishing operation for the target video, the first terminal 110 sends the associated text corresponding to the associated text publishing operation to the backend server 120 .
[0173] The backend server 120 associates and stores the associated text, the target video, and the time the text was published in a database 121. The backend server 120 retrieves the target video from the database 121 and, through the video processing module 122, segments the target video to obtain multiple candidate video segments. The backend server 120 sends the candidate video segments and the original cover image of the target obtained from the database 121 to the image evaluation model 123 to obtain an image relevance score for each candidate video segment and, based on the image relevance score, identifies the target video segment. The backend server 120 retrieves the associated text corresponding to the target video segment from the associated text of the target video as the candidate associated text. The backend server 120 inputs the candidate associated text and the original cover image of the target video into the first text evaluation model 124 to obtain a first text score indicating the relevance of the candidate associated text to the original cover image. The backend server 120 also inputs the candidate associated text and the target object label into the second text evaluation model 125 to obtain a second text score indicating the relevance of the candidate associated text to the target object label. The backend server 120 inputs the first and second text scores into the text scoring module 126, which combines the two scores to determine the cover text from the candidate associated texts. The backend server 120 inputs the cover text and the original cover image into the cover generation module 127 to obtain the target video cover, and sends the target video cover to the second terminal 130 corresponding to the target object.
[0174] When the second terminal 130 displays the video recommendation page and the recommended videos include the target video, the target video cover generated in the above steps is displayed on the video recommendation page.
[0175] Figure 12 This is a structural block diagram of a device for generating a video cover provided by an exemplary embodiment of the present application. The device is composed of the following components:
[0176] An acquisition module 1201 is configured to acquire an original cover image and associated text of a video, wherein the associated text includes at least one of a video comment and a video screencast;
[0177] A first determining module 1202 is configured to determine a cover text based on the correlation between the original cover image and each of the associated texts, wherein the correlation between the cover text and the original cover image is higher than the correlation between the other associated texts and the original cover image;
[0178] The image processing module 1203 is used to fuse the original cover image and the cover text to generate a target video cover.
[0179] Optionally, the first determining module 1202 includes:
[0180] a first determining unit, configured to determine a target video segment from the video based on the original cover image, wherein a correlation between a video frame corresponding to the target video segment and the original cover image is higher than a correlation between video frames corresponding to other video segments and the original cover image;
[0181] The second determining unit is configured to determine the cover text from the associated text corresponding to the target video segment based on the correlation between the associated text and the original cover image.
[0182] Optionally, the first determining unit is further configured to:
[0183] Segmenting the video according to a target number of segments or a target duration to obtain at least two candidate video segments;
[0184] Inputting the original cover image and the video frames corresponding to the candidate video clip into an image evaluation model to obtain an image relevance score corresponding to the candidate video clip, wherein the image evaluation model is trained based on positive and negative sample pairs, wherein the positive sample pair consists of a sample cover image and a related video clip corresponding to the sample cover image, and the negative sample pair consists of the sample cover image and an unrelated video clip corresponding to the sample cover image;
[0185] At least one of the candidate video segments with the highest image relevance score is determined as the target video segment.
[0186] Optionally, the first determining unit is further configured to:
[0187] Inputting the original cover image into a first feature extraction network in the image evaluation model to obtain a cover feature vector corresponding to the original cover image;
[0188] Inputting the video frames corresponding to the candidate video clips into the second feature extraction network in the image evaluation model to obtain video frame feature vectors corresponding to each video frame;
[0189] Performing feature fusion on the video frame feature vectors corresponding to each video frame through the self-attention mechanism in the image evaluation model to obtain the segment feature vectors of the candidate video segment;
[0190] The cover feature vector and the segment feature vector are input into a fully connected layer in the image evaluation model to obtain the image relevance score.
[0191] Optionally, the second determining unit is further configured to:
[0192] In response to the number of texts corresponding to the target video segment being lower than or equal to a text number threshold, determining all associated texts corresponding to the target video segment as the cover texts;
[0193] In response to the text quantity corresponding to the target video segment being higher than the text quantity threshold, the cover text is determined from the associated text corresponding to the target video segment based on the correlation between the associated text and the original cover image.
[0194] Optionally, the first determining module 1202 further includes:
[0195] a third determining unit, configured to determine a target video segment from the video based on the original cover image, wherein a correlation between a video frame corresponding to the target video segment and the original cover image is higher than a correlation between video frames corresponding to other video segments and the original cover image;
[0196] A fourth determining unit is configured to determine the cover text from the associated text corresponding to the target video segment based on a correlation between the associated text and a target object label, wherein the target object label is used to indicate the target object's orientation to the video type.
[0197] Optionally, the second determining unit is further configured to:
[0198] Determining the associated text corresponding to the target video clip as a candidate associated text;
[0199] determining a first text score of the candidate associated text based on the relevance between the candidate associated text and the original cover image;
[0200] determining a second text score for the candidate associated text based on a correlation between the candidate associated text and a target object label, wherein the target object label is used to indicate the target object's orientation to the video type;
[0201] determining a text relevance score based on the first text score, a first weight corresponding to the first text score, the second text score, and a second weight corresponding to the second text score;
[0202] The n candidate associated texts with the highest text relevance scores are determined as the cover texts, where n is a positive integer.
[0203] Optionally, the second determining unit is further configured to:
[0204] The original cover image and the candidate associated text are input into a first text evaluation model to obtain the first text score. The first text evaluation model is trained based on positive and negative sample pairs, wherein the positive sample pair consists of a sample video frame and a positive sample text, and the playback time of the sample video frame is consistent with that of the positive sample text; the negative sample pair consists of the sample video frame and the negative sample text, and the playback time of the sample video frame is inconsistent with that of the negative sample text.
[0205] Optionally, the second determining unit is further configured to:
[0206] Inputting the original cover image into a third feature extraction network in the first text evaluation model to obtain a cover feature vector corresponding to the original cover image;
[0207] Inputting the candidate associated text into a text feature extraction network in the first text evaluation model to obtain a text feature vector corresponding to the candidate associated text;
[0208] The cover feature vector and the text feature vector are input into a fully connected layer in the first text evaluation model to obtain the first text score.
[0209] Optionally, the second determining unit is further configured to:
[0210] The target object label and the candidate associated text are input into a second text evaluation model to obtain the second text score. The second text evaluation model is obtained based on training of positive and negative sample pairs, wherein the positive sample pair consists of a sample label corresponding to a sample object and a positive sample text, and the positive sample text is the associated text that has received positive feedback interaction operations from the sample object; the negative sample consists of the sample label and negative sample text, and the negative sample text is the associated text that has received negative feedback interaction operations from the sample object.
[0211] Optionally, the second determining unit is further configured to:
[0212] Inputting the target object label into the first text feature extraction network in the second text evaluation model to obtain a label feature vector corresponding to the target object label;
[0213] Inputting the candidate associated text into a second text feature extraction network in the second text evaluation model to obtain a text feature vector corresponding to the candidate associated text;
[0214] The label feature vector and the text feature vector are input into a fully connected layer in the second text evaluation model to obtain the second text score.
[0215] Optionally, the image processing module 1203 includes:
[0216] A fifth determining unit, configured to determine a horizontal position of each cover text in the original cover image based on a text generation time;
[0217] a sixth determining unit, configured to determine a vertical position of each cover text in the original cover image based on the text relevance score, wherein a vertical distance between a cover text having a high text relevance score and a horizontal center line is smaller than a vertical distance between a cover text having a low text relevance score and the horizontal center line;
[0218] An image processing unit is used to draw the cover text on the original cover image according to the horizontal position and the vertical position to obtain the target video cover.
[0219] Optionally, the device further includes:
[0220] A cover update model is used to update the video cover based on the original cover image and the associated text in response to the video meeting the cover update condition, wherein the cover update condition includes at least one of the increment of the associated text reaching a text increment threshold and / or the increment of the target object label reaching a label increment threshold.
[0221] Optionally, the device further includes:
[0222] a text recognition module, configured to perform optical character recognition on the original cover image to determine the cover text in the original cover image;
[0223] The first determining module 1202 includes:
[0224] The seventh determination module is used to determine the cover text based on the correlation between the original cover image and each of the associated texts in response to the number of words in the cover text being less than a word count threshold and / or the proportion of the display area of the cover text to the area of the original cover image being less than a proportion threshold.
[0225] In summary, in the embodiments of the present application, the cover text is determined based on the correlation between the original cover image of the video and the associated text of the video, so as to automatically generate a video cover that incorporates the associated text. This eliminates the need for the video publisher to manually select the associated text, thereby improving the efficiency of cover generation. Furthermore, by mining the associated text, the associated text with a high correlation with the original cover image is selected as the cover text, so that the target video cover can more accurately reflect the key content of the video. This can prevent the associated text manually selected by the video publisher based on personal preferences from not being consistent with the preferences of other video recommendation targets, further optimizing the expressive effect of the video cover, and thus increasing the video's click-through rate.
[0226] Please refer to Figure 13 , which shows a schematic diagram of the structure of a computer device provided by an embodiment of the present application. The computer device can be a terminal with a video application installed, or a background server of a video application. Specifically:
[0227] The computer device 1300 includes a central processing unit (CPU) 1301, a system memory 1304 including a random access memory (RAM) 1302 and a read-only memory (ROM) 1303, and a system bus 1305 connecting the system memory 1304 and the CPU 1301. The computer device 1300 may also include a basic input / output (I / O) controller 1306 for facilitating information transfer between various components within the computer, and a mass storage device 1307 for storing an operating system 1313, application programs 1314, and other program modules 1315.
[0228] In some embodiments, the basic input / output system 1306 includes a display 1308 for displaying information and an input device 1309, such as a mouse or keyboard, for user input. Both the display 1308 and the input device 1309 are connected to the central processing unit 1301 via an input / output controller 1310 connected to the system bus 1305. The basic input / output system 1306 may also include an input / output controller 1310 for receiving and processing input from a variety of other devices, such as a keyboard, mouse, or electronic stylus. Similarly, the input / output controller 1310 also provides output to a display screen, printer, or other types of output devices.
[0229] The mass storage device 1307 is connected to the central processing unit 1301 via a mass storage controller (not shown) connected to the system bus 1305. The mass storage device 1307 and its associated computer-readable media provide non-volatile storage for the computer device 1300. In other words, the mass storage device 1307 may include a computer-readable medium (not shown) such as a hard disk or a compact disc read-only memory (CD-ROM) drive.
[0230] Without loss of generality, the computer-readable medium may include computer storage media and communication media. Computer storage media include volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules or other data. Computer storage media include RAM, ROM, Erasable Programmable Read Only Memory (EPROM), flash memory or other solid-state storage technologies, CD-ROM, Digital Video Disc (DVD) or other optical storage, tape cassettes, magnetic tape, disk storage or other magnetic storage devices. Of course, those skilled in the art will appreciate that the computer storage media are not limited to the above-mentioned ones. The above-mentioned system memory 1304 and mass storage device 1307 can be collectively referred to as memory.
[0231] According to various embodiments of the present application, the computer device 1300 may also be connected to a remote computer on a network such as the Internet for operation. That is, the computer device 1300 may be connected to a network 1312 via a network interface unit 1311 connected to the system bus 1305. Alternatively, the network interface unit 1311 may be used to connect to other types of networks or remote computer systems (not shown).
[0232] The memory also includes at least one instruction, at least one program, code set or instruction set, which is stored in the memory and configured to be executed by one or more processors to implement the above-mentioned method for generating a video cover.
[0233] An embodiment of the present application also provides a computer-readable storage medium, which stores at least one instruction, and the at least one instruction is loaded and executed by a processor to implement the method for generating a video cover as described in the above embodiments.
[0234] According to one aspect of the present application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the video cover generation method provided in various optional implementations of the above aspects.
[0235] Those skilled in the art will appreciate that in one or more of the above examples, the functions described in the embodiments of the present application can be implemented using hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable storage medium or transmitted as one or more instructions or codes on a computer-readable storage medium. Computer-readable storage media include computer storage media and communication media, wherein communication media include any media that facilitates the transmission of computer programs from one place to another. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0236] It is understandable that in the specific implementation of this application, it involves the user's personal data, namely user tags, feedback interaction operations, etc. When the above embodiments of this application are applied to specific products or technologies, it is necessary to obtain user permission or consent, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions.
[0237] The above description is merely an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
Claims
1. A method for generating a video cover, characterized in that: The method comprises: Obtaining an original cover image and associated text of the video, wherein the associated text includes at least one of a video comment and a video screencap; Determining a target video segment from the video based on the original cover image, wherein the correlation between the video frame corresponding to the target video segment and the original cover image is higher than the correlation between the video frames corresponding to other video segments and the original cover image; Determining a cover text from the associated text corresponding to the target video clip based on the correlation between the associated text corresponding to the target video clip and the original cover image, wherein the correlation between the cover text and the original cover image is higher than the correlation between other associated texts and the original cover image; The original cover image and the cover text are fused to generate a target video cover.
2. The method according to claim 1, characterized in that The determining a target video segment from the video based on the original cover image includes: Segmenting the video according to a target number of segments or a target duration to obtain at least two candidate video segments; Inputting the original cover image and the video frames corresponding to the candidate video clip into an image evaluation model to obtain an image relevance score corresponding to the candidate video clip, wherein the image evaluation model is trained based on positive and negative sample pairs, wherein the positive sample pair consists of a sample cover image and a related video clip corresponding to the sample cover image, and the negative sample pair consists of the sample cover image and an unrelated video clip corresponding to the sample cover image; At least one of the candidate video segments with the highest image relevance score is determined as the target video segment.
3. The method according to claim 2, characterized in that The step of inputting the original cover image and the video frames corresponding to the candidate video clip into an image evaluation model to obtain an image relevance score corresponding to the candidate video clip includes: Inputting the original cover image into a first feature extraction network in the image evaluation model to obtain a cover feature vector corresponding to the original cover image; Inputting the video frames corresponding to the candidate video clips into the second feature extraction network in the image evaluation model to obtain video frame feature vectors corresponding to each video frame; Performing feature fusion on the video frame feature vectors corresponding to each video frame through the self-attention mechanism in the image evaluation model to obtain the segment feature vectors of the candidate video segment; The cover feature vector and the segment feature vector are input into a fully connected layer in the image evaluation model to obtain the image relevance score.
4. The method according to claim 1, wherein The determining the cover text from the associated text corresponding to the target video segment based on the correlation between the associated text corresponding to the target video segment and the original cover image includes: In response to the number of texts corresponding to the target video segment being lower than or equal to a text number threshold, determining all associated texts corresponding to the target video segment as the cover texts; In response to the text quantity corresponding to the target video segment being higher than the text quantity threshold, the cover text is determined from the associated text corresponding to the target video segment based on the correlation between the associated text corresponding to the target video segment and the original cover image.
5. The method according to claim 1, wherein The method further comprises: The cover text is determined from the associated text corresponding to the target video segment based on the correlation between the associated text corresponding to the target video segment and the target object label, where the target object label is used to indicate the target object's orientation to the video type.
6. The method according to claim 1, wherein The determining the cover text from the associated text corresponding to the target video segment based on the correlation between the associated text corresponding to the target video segment and the original cover image includes: Determining the associated text corresponding to the target video clip as a candidate associated text; determining a first text score of the candidate associated text based on the relevance between the candidate associated text and the original cover image; determining a second text score for the candidate associated text based on a correlation between the candidate associated text and a target object label, wherein the target object label is used to indicate the target object's orientation to the video type; determining a text relevance score based on the first text score, a first weight corresponding to the first text score, the second text score, and a second weight corresponding to the second text score; The n candidate associated texts with the highest text relevance scores are determined as the cover texts, where n is a positive integer.
7. The method according to claim 6, characterized in that The determining, based on the relevance between the candidate associated text and the original cover image, a first text score of the candidate associated text includes: The original cover image and the candidate associated text are input into a first text evaluation model to obtain the first text score. The first text evaluation model is trained based on positive and negative sample pairs, wherein the positive sample pair consists of a sample video frame and a positive sample text, and the playback time of the sample video frame is consistent with that of the positive sample text; the negative sample pair consists of the sample video frame and the negative sample text, and the playback time of the sample video frame is inconsistent with that of the negative sample text.
8. The method according to claim 7, characterized in that The step of inputting the original cover image and the candidate associated text into a first text evaluation model to obtain the first text score includes: Inputting the original cover image into a third feature extraction network in the first text evaluation model to obtain a cover feature vector corresponding to the original cover image; Inputting the candidate associated text into a text feature extraction network in the first text evaluation model to obtain a text feature vector corresponding to the candidate associated text; The cover feature vector and the text feature vector are input into a fully connected layer in the first text evaluation model to obtain the first text score.
9. The method according to claim 6, characterized in that The determining, based on the relevance between the candidate associated text and the target object label, a second text score of the candidate associated text includes: The target object label and the candidate associated text are input into a second text evaluation model to obtain the second text score. The second text evaluation model is obtained based on training of positive and negative sample pairs, wherein the positive sample pair consists of a sample label corresponding to a sample object and a positive sample text, and the positive sample text is the associated text that has received positive feedback interaction operations from the sample object; the negative sample consists of the sample label and negative sample text, and the negative sample text is the associated text that has received negative feedback interaction operations from the sample object.
10. The method according to claim 9, characterized in that The step of inputting the target object label and the candidate associated text into a second text evaluation model to obtain the second text score includes: Inputting the target object label into the first text feature extraction network in the second text evaluation model to obtain a label feature vector corresponding to the target object label; Inputting the candidate associated text into a second text feature extraction network in the second text evaluation model to obtain a text feature vector corresponding to the candidate associated text; The label feature vector and the text feature vector are input into a fully connected layer in the second text evaluation model to obtain the second text score.
11. The method according to any one of claims 6 to 10, characterized in that: The fusing the original cover image and the cover text to generate a target video cover includes: Determining the horizontal position of each cover text in the original cover image based on the text generation time; Determining a vertical position of each cover text in the original cover image based on the text relevance score, wherein a vertical distance between a cover text with a high text relevance score and a horizontal center line is smaller than a vertical distance between a cover text with a low text relevance score and the horizontal center line; The cover text is drawn on the original cover image according to the horizontal position and the vertical position to obtain the target video cover.
12. The method according to any one of claims 6 to 10, characterized in that: After fusing the original cover image and the cover text to generate a target video cover, the method further includes: In response to the video meeting the cover update condition, the video cover is updated based on the original cover image and the associated text, and the cover update condition includes at least one of the increment of the associated text reaching a text increment threshold and the increment of the target object label reaching a label increment threshold.
13. The method according to any one of claims 1 to 10, characterized in that: The method further comprises: Performing optical character recognition on the original cover image to determine the cover text in the original cover image; In response to the number of words in the cover text being less than a word count threshold, and / or the proportion of the display area of the cover text to the area of the original cover image being less than a proportion threshold, the cover text is determined from the associated text corresponding to the target video clip based on the correlation between the associated text corresponding to the target video clip and the original cover image.
14. A device for generating a video cover, characterized in that: The device comprises: An acquisition module, configured to acquire an original cover image and associated text of a video, wherein the associated text includes at least one of a video comment and a video screencap; a first determining module, configured to determine a target video segment from the video based on the original cover image, wherein the correlation between the video frame corresponding to the target video segment and the original cover image is higher than the correlation between the video frames corresponding to other video segments and the original cover image; Determining a cover text from the associated text corresponding to the target video clip based on the correlation between the associated text corresponding to the target video clip and the original cover image, wherein the correlation between the cover text and the original cover image is higher than the correlation between other associated texts and the original cover image; The image processing module is used to fuse the original cover image and the cover text to generate a target video cover.
15. A computer device, characterized in that: The computer device includes a processor and a memory; the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the method for generating a video cover as described in any one of claims 1 to 13.
16. A computer-readable storage medium, characterized in that The computer-readable storage medium stores at least one computer program, which is loaded and executed by a processor to implement the method for generating a video cover as described in any one of claims 1 to 13.
17. A computer program product, characterized in that The computer program product includes computer instructions, which are stored in a computer-readable storage medium; the processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the terminal executes the method for generating a video cover as described in any one of claims 1 to 13.
Citation Information
Patent Citations
Video cover generation method and device
CN112752121A
Cover image generation method and device, equipment, storage medium and program product
CN113656642A