Video barrage generation method, device, electronic device and storage medium

By analyzing and understanding the video plot text within the specified playback period of the video file, a specific barrage of barrage related to the video plot is generated, which solves the problem of lack of personalization of barrage input methods in the existing technology, and improves the user's barrage interactive experience.

CN119172609BActive Publication Date: 2025-09-02BEIJING QIYI CENTURY SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411201180.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-29
Publication Date
2025-09-02
Estimated Expiration
2044-08-29

AI Technical Summary

Technical Problem

The existing video barrage input methods lack understanding and analysis of video content, and cannot provide a personalized barrage experience closely related to the video plot.

Method used

By obtaining the video plot text within the specified playback period of the video file, extracting the semantic feature information of the video keywords, and encoding it, determining the target video keywords whose similarity index is greater than the threshold value, generating barrage text related to the video plot, and storing the video file identifier, playback period and barrage text.

Benefits of technology

It enables users to express their viewing experience more accurately when entering barrage, and improves the barrage interaction experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119172609B_ABST
    Figure CN119172609B_ABST
Patent Text Reader

Abstract

The present application provides a method, device, electronic device and storage medium for generating video barrage. The method includes: obtaining the video plot text of a video file within a specified video playback period; obtaining the semantic feature information of video keywords in the video plot text; encoding the video plot text and the semantic feature information respectively to obtain a text encoding vector of the video plot text and a semantic feature vector corresponding to the semantic feature information; determining the target video keyword in the video keyword whose similarity index with the video plot text is greater than a threshold value based on the semantic feature vector and the text encoding vector; generating a video barrage text based on the target video keyword; associating and storing the video file identifier, the specified video playback period and the video barrage text of the video file. The present application can generate a barrage vocabulary library related to the video plot, so that users can express their viewing experience more accurately when entering barrage, thereby improving the experience of barrage interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of video barrage processing, and in particular to a method, device, electronic device and storage medium for generating video barrage. Background Art

[0002] With the rapid development of Internet technology, barrage as a unique form of interaction has become increasingly popular on video sharing platforms and social media.

[0003] During video playback, interactive barrage comments on video sharing platforms have become an important way for users to watch videos. However, existing barrage comment input methods mostly rely on preset templates or user-defined input, lacking understanding and analysis of video content, and are unable to provide a personalized barrage comment experience that is closely related to the video plot. Summary of the Invention

[0004] The purpose of the embodiments of this application is to provide a method, device, electronic device, and storage medium for generating video bullet comments, so as to generate a bullet comment vocabulary library related to the video plot, so that users can more accurately express their viewing experience when entering bullet comments, and improve the interactive experience of bullet comments. The specific technical solution is as follows:

[0005] In a first aspect of the present application, a method for generating a video barrage is provided, comprising:

[0006] Get the video plot text of the video file within the specified video playback period;

[0007] Obtaining semantic feature information of video keywords in the video plot text;

[0008] Encoding the video plot text and the semantic feature information respectively to obtain a text encoding vector of the video plot text and a semantic feature vector corresponding to the semantic feature information;

[0009] Determining, based on the semantic feature vector and the text encoding vector, target video keywords whose similarity index with the video plot text is greater than a threshold among the video keywords;

[0010] Generate video barrage text for the video file during the specified video playback period based on the target video keywords;

[0011] The video file identifier of the video file, the designated video playback time period and the video barrage text are stored in association.

[0012] In a second aspect of the present application, a video barrage generation device is provided, comprising:

[0013] A video plot text acquisition module is used to obtain the video plot text of a video file within a specified video playback period;

[0014] A semantic feature information acquisition module, used to acquire semantic feature information of video keywords in the video plot text;

[0015] A text semantic vector acquisition module is used to encode the video plot text and the semantic feature information respectively to obtain a text encoding vector of the video plot text and a semantic feature vector corresponding to the semantic feature information;

[0016] a target keyword determination module, configured to determine, based on the semantic feature vector and the text encoding vector, a target video keyword whose similarity index with the video plot text is greater than a threshold value among the video keywords;

[0017] A video barrage text generation module, configured to generate a video barrage text for the video file during the specified video playback period according to the target video keywords;

[0018] The video barrage text storage module is used to associate and store the video file identifier of the video file, the specified video playback time period and the video barrage text.

[0019] In another aspect of the present application, an electronic device is provided, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other via the communication bus;

[0020] Memory for storing computer programs;

[0021] The processor is used to implement any of the above-mentioned video barrage generation methods when executing the program stored in the memory.

[0022] In another aspect of the implementation of the present application, a computer-readable storage medium is provided, in which instructions are stored. When the computer-readable storage medium is run on a computer, the computer executes any of the above-mentioned video barrage generation methods.

[0023] In another aspect of the implementation of the present application, a computer program product containing instructions is also provided, which, when run on a computer, enables the computer to execute any of the above-mentioned video barrage generation methods.

[0024] The solution provided by the embodiment of the present application is to obtain the video plot text of the video file within the specified video playback period. The semantic feature information of the video keywords in the video plot text is obtained. The video plot text and the semantic feature information are respectively encoded to obtain the text encoding vector of the video plot text and the semantic feature vector corresponding to the semantic feature information. According to the semantic feature vector and the text encoding vector, the target video keyword in the video keyword whose similarity index with the video plot text is greater than the threshold is determined. According to the target video keyword, the video barrage text of the video file in the specified video playback period is generated. The video file identifier, the specified video playback period and the video barrage text of the video file are stored in association. The embodiment of the present application analyzes and understands the video plot text of the video file within the specified video playback period to generate a specific video barrage vocabulary related to the video plot, so that users can express their viewing experience more accurately when inputting barrage, thereby improving the experience of barrage interaction. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art.

[0026] Figure 1 A flowchart of a method for generating video bullet comments provided in an embodiment of the present application;

[0027] Figure 2 A flowchart of a method for obtaining video plot text provided in an embodiment of the present application;

[0028] Figure 3 A flowchart of a method for obtaining semantic feature information provided in an embodiment of the present application;

[0029] Figure 4 A flowchart of a method for screening target video keywords provided in an embodiment of the present application;

[0030] Figure 5 A flowchart of a method for obtaining video barrage text provided in an embodiment of the present application;

[0031] Figure 6 A flowchart of another method for obtaining video barrage text provided in an embodiment of the present application;

[0032] Figure 7 A flowchart of a method for recommending target video barrage text provided in an embodiment of the present application;

[0033] Figure 8 A schematic diagram of a video barrage generation and recommendation process provided in an embodiment of the present application;

[0034] Figure 9 A schematic diagram of the structure of a video barrage generation device provided in an embodiment of the present application;

[0035] Figure 10 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0036] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.

[0037] Figure 1 A flowchart of a method for generating video bullet screen provided in an embodiment of the present application is shown as follows: Figure 1 As shown, the video barrage generation method may include: step 101, step 102, step 103, step 104, step 105 and step 106.

[0038] Step 101: Obtain the video plot text of the video file within a specified video playback period.

[0039] The embodiments of the present application can be applied to scenarios where video plot text is analyzed to generate video comments related to the video plot.

[0040] In this example, the video file may be a variety show video file, or a film and television video file, etc. The type of the video file may be determined according to actual conditions, and this embodiment does not impose any limitation on this.

[0041] In a specific implementation, the playback duration of a video file can be divided into several time periods for plot analysis. The designated video playback period is the video playback period of the video file used to generate video barrage. In this example, the designated video playback period can be the period from the 30th minute to the 35th minute of the video file. For example, taking a film or television video as an example, the designated video playback period can be the 1-minute playback period at the beginning of the film or television video, or the 1-minute playback period at the end.

[0042] The video plot text may be used to indicate text related to the video plot during a specified video playback period of a video file. In this embodiment, the video plot text may include at least one of script text, video character dialogue text, and video screen narration text.

[0043] When generating video comments related to the video plot within the specified video playback period of a video file, the video plot text of the video file within the specified video playback period can be obtained. In a specific implementation, when performing plot analysis on the specified video playback period of a video file, the initial plot text of the video file within the specified video playback period can be collected, and the initial plot text can be preprocessed, such as word segmentation, part-of-speech tagging, syntactic analysis, etc., to obtain the video plot text of the video file within the specified video playback period. For this implementation process, the following implementation will be combined with Figure 2 The present embodiment will be described in detail and will not be described in detail here.

[0044] After obtaining the video plot text of the video file within the specified video playback period, step 102 is executed.

[0045] Step 102: Acquire semantic feature information of video keywords in the video plot text.

[0046] After obtaining the video plot text of the video file within the specified video playback period, the video keywords in the video plot text can be extracted, and the extracted video keywords can be semantically analyzed to obtain the semantic feature information of the video keywords, such as the meaning, part of speech, emotion, etc. of the video keywords. Specifically, the plot text can be processed based on the pre-trained semantic analysis model to obtain candidate semantic feature information of the video keywords of the video plot text, and the candidate semantic feature information can be post-processed (such as deduplication, part of speech restoration, etc.) to obtain semantic feature information. This implementation process will be combined with the following embodiments. Figure 3 The present embodiment will be described in detail and will not be described in detail here.

[0047] After obtaining the semantic feature information of the video keywords in the video plot text, step 103 is executed.

[0048] Step 103: Encoding the video plot text and the semantic feature information respectively to obtain a text encoding vector of the video plot text and a semantic feature vector corresponding to the semantic feature information.

[0049] After obtaining the semantic feature information of the video keywords in the video plot text, the video plot text and the semantic feature information can be encoded separately to obtain the text encoding vector of the video plot text and the semantic feature vector corresponding to the semantic feature information. That is, the semantic feature information of the video plot text and the video keywords are vectorized respectively to obtain the corresponding text encoding vector and semantic feature vector.

[0050] In a specific implementation, the video plot text and video keywords can be vectorized based on a preset model (such as a BERT-based large language model (LLM)) to obtain a vectorized text encoding vector and semantic feature vector. Of course, this is not limited to this. In a specific implementation, other methods can also be used to obtain the text encoding vector and semantic feature vector. This embodiment does not limit the method for obtaining the text encoding vector and semantic feature vector.

[0051] After encoding the video plot text and the semantic feature information to obtain a text encoding vector of the video plot text and a semantic feature vector corresponding to the semantic feature information, step 104 is executed.

[0052] Step 104: Determine, based on the semantic feature vector and the text encoding vector, target video keywords whose similarity index with the video plot text is greater than a threshold among the video keywords.

[0053] The threshold refers to a preset threshold of a similarity index for screening video keywords with high similarity. In this example, the specific value of the threshold can be determined according to business requirements, and this embodiment does not impose any limitation on this.

[0054] The target video keyword refers to a video keyword selected from the video keywords and having a similarity index with the video plot text greater than a threshold.

[0055] After encoding the video plot text and semantic feature information to obtain the text encoding vector of the video plot text and the semantic feature vector corresponding to the semantic feature information, the target video keyword in the video keyword whose similarity index with the video plot text is greater than a threshold value can be determined based on the semantic feature vector and the text encoding vector. Specifically, the semantic feature vector and the text encoding vector can be subjected to correlation analysis to obtain the similarity index between the video keyword and the video plot text, and the target video keyword can be screened out from the video keywords based on the similarity index and the threshold value. This implementation process will be described in conjunction with the following embodiments. Figure 4 The present embodiment will be described in detail and will not be described in detail here.

[0056] After determining the target video keywords whose similarity index with the video plot text is greater than the threshold value among the video keywords based on the semantic feature vector and the text encoding vector, step 105 is executed.

[0057] Step 105: Generate video barrage text for the video file during the specified video playback period based on the target video keywords.

[0058] After obtaining the target video keywords, the video barrage text of the video file during the specified video playback period can be generated based on the target video keywords. Specifically, the target video keywords can be directly used as the video barrage text. The target video keywords can also be optimized to generate the video barrage text.

[0059] In this embodiment, after obtaining the target video keywords, the barrage text generated by the target video keywords can also be updated in combination with the user's historical barrage data or the user's barrage requirements to obtain a video barrage text that meets the user's needs. Figure 5 and Figure 6 The present embodiment will be described in detail and will not be described in detail here.

[0060] After the video barrage text of the video file in the specified video playback period is generated according to the target video keywords, step 106 is executed.

[0061] Step 106: Associate and store the video file identifier of the video file, the designated video playback time period, and the video barrage text.

[0062] After generating the video barrage text of the video file in the specified video playback period based on the target video keywords, the video file identifier of the video file, the specified video playback period and the video barrage text can be associated and stored, that is, the storage format is <video id, playback time, vocabulary (containing one or more video barrage texts)>. When the user enters the barrage, a specific vocabulary can be found based on the video id and the current playback time, and the corresponding words can be recommended from the vocabulary, so that the user can quickly select and complete the barrage related to the video plot for sending.

[0063] The embodiment of the present application analyzes and understands the video plot text of the video file within a specified video playback period to generate a specific video barrage vocabulary library related to the video plot, so that users can express their viewing experience more accurately when inputting barrage.

[0064] Next, combine Figure 2 The process of obtaining video plot text is described in detail.

[0065] Reference Figure 2 , shows a flowchart of the steps of a method for obtaining video plot text provided by an embodiment of the present application. Figure 2 As shown, the method for obtaining video plot text may include: step 201 and step 202.

[0066] Step 201: Obtain the initial plot text of the video file during the specified video playback period.

[0067] In this embodiment, when generating video barrage related to the video plot within the specified video playback period of the video file, the initial plot text of the video file in the specified video playback period can be obtained, such as script text, dialogue text, narration text, etc.

[0068] After the initial plot text of the video file in the specified video playback period is obtained, step 202 is executed.

[0069] Step 202: Process the initial plot text based on a preprocessing method to obtain the video plot text.

[0070] In this embodiment, the preprocessing method may include: at least one of word segmentation processing, part-of-speech tagging processing, and syntactic analysis processing.

[0071] After obtaining the initial plot text of the video file in the specified video playback period, the initial plot text may be processed based on a preprocessing method to obtain the video plot text.

[0072] The embodiment of the present application pre-processes the plot text (such as word segmentation, part-of-speech tagging, syntactic analysis, etc.) in advance, which can facilitate subsequent plot analysis.

[0073] Next, combine Figure 3 The implementation process of obtaining the semantic feature information of video keywords is described in detail.

[0074] Reference Figure 3 , shows a flowchart of the steps of a method for obtaining semantic feature information provided by an embodiment of the present application. Figure 3 As shown, the method for acquiring semantic feature information may include: step 301 and step 302.

[0075] Step 301: Process the video plot text based on a pre-trained semantic analysis model to obtain candidate semantic feature information of video keywords in the video plot text.

[0076] In this embodiment, the pre-trained semantic analysis model refers to a model for performing keyword extraction and keyword semantic analysis on the video plot text. In this example, the semantic analysis model can be, but is not limited to, a BERT-based LLM (Large Language Model 1) model.

[0077] After obtaining the video plot text of the video file, the video plot text can be processed based on a pre-trained semantic analysis model to obtain candidate semantic feature information for video keywords in the video plot text. For example, the keyword semantic information extraction function of the LLM model can be used to extract video keywords from the video plot text and analyze candidate semantic feature information for each video keyword. The candidate semantic feature information may include the meaning, part of speech, and sentiment of the video keyword.

[0078] Specifically, the model processing process can be as follows: 1. Input the video plot text into a semantic analysis model (such as an LLM model). 2. Utilize the context information extraction function in the semantic analysis model to obtain the context information of each word in the video plot text, where the context information may include words and phrases in the preceding and following contexts. 3. Utilize the keyword semantic information extraction function in the semantic analysis model to obtain the semantic information of each keyword in the video plot text, where the semantic information may include the keyword's meaning, part of speech, and sentiment. 4. Feature fusion: Fusion of context information and keyword semantic information to obtain a feature representation of the video plot text, where the feature representation may include the theme, sentiment, and key information of the video plot text. 5. Feature selection: Based on task requirements, semantic feature information related to the video plot text can be selected (e.g., by setting a threshold, sorting, etc.) according to the fused feature representation of the video plot text as candidate semantic feature information.

[0079] After the video plot text is processed based on the pre-trained semantic analysis model to obtain candidate semantic feature information of video keywords in the video plot text, step 302 is executed.

[0080] Step 302: Process the candidate semantic feature information based on a preset post-processing method to obtain the semantic feature information.

[0081] After processing the video plot text based on the pre-trained semantic analysis model to obtain candidate semantic feature information of the video keywords of the video plot text, the candidate semantic feature information can be processed using a preset post-processing method to obtain semantic feature information of the video keywords. In this example, the preset post-processing method may include at least one of: repeated feature removal, feature part-of-speech restoration processing, etc. Among them, repeated feature removal refers to eliminating features that appear repeatedly in the candidate semantic feature information to reduce redundancy and improve processing efficiency. Feature part-of-speech restoration processing refers to restoring words with different forms but the same meaning (such as different tenses of verbs, singular and plural forms of nouns, etc.) to a unified basic form for subsequent semantic analysis and processing.

[0082] The present embodiment improves the accuracy of plot analysis by using a pre-trained semantic analysis model to extract keywords and perform semantic analysis on video plot text. Furthermore, by processing candidate semantic feature information using a pre-set post-processing method, such as removing duplicate keywords and performing part-of-speech restoration, duplicate keywords can be avoided and subsequent similarity analysis can be facilitated.

[0083] Next, combine Figure 4 The screening process of target video keywords is described in detail.

[0084] Reference Figure 4 , shows a flowchart of the steps of a target video keyword screening method provided by an embodiment of the present application. Figure 4 As shown, the target video keyword screening method may include: step 401 and step 402.

[0085] Step 401: performing a correlation analysis on the semantic feature vector and the text encoding vector to obtain a similarity index between the video keywords and the video plot text.

[0086] In this embodiment, after obtaining the semantic feature vector corresponding to the semantic feature information of the video keyword and the text encoding vector of the video plot text, a correlation analysis can be performed on the semantic feature vector and the text encoding vector to obtain a similarity index between the video keyword and the video plot text.

[0087] In a specific implementation, the cosine distance between the semantic feature vector and the text encoding vector can be obtained and used as the similarity index between the video keywords and the video plot text. Alternatively, the Euclidean distance between the semantic feature vector and the text encoding vector can be obtained and used as the similarity index between the video keywords and the video plot text.

[0088] Of course, this is not limited to the above. In practical applications, other methods can be used to obtain the similarity index between video keywords and video plot text, such as calculating the Euclidean distance between the semantic feature vector and the text encoding vector as a similarity index. Specifically, the method for obtaining the similarity index can be determined according to business needs, and this embodiment does not limit this.

[0089] After performing correlation analysis on the semantic feature vector and the text encoding vector to obtain the similarity index between the video keywords and the video plot text, step 402 is executed.

[0090] Step 402: Filter out target video keywords whose similarity index is greater than a threshold from the video keywords.

[0091] After performing a correlation analysis on the semantic feature vector and the text encoding vector to obtain the similarity index between the video keywords and the video plot text, target video keywords with a similarity index greater than a threshold can be screened from the video keywords. For example, the video keywords include: keyword 1, keyword 2, keyword 3, and keyword 4. The similarity indexes between keyword 1, keyword 2, keyword 3, and keyword 4 and the video plot text are: 0.8, 0.9, 0.4, and 0.6, respectively. When the threshold is 0.5, keyword 1, keyword 2, and keyword 4 can be used as target video keywords.

[0092] It can be understood that the above examples are merely examples listed for a better understanding of the technical solutions of the embodiments of the present application, and are not intended to be the sole limitation on the embodiments.

[0093] The embodiment of the present application analyzes the similarity index between video keywords and video plot text to screen out target video keywords as the basis for subsequent generation of video barrage text, so that the generated video barrage text can be related to the video plot, allowing users to express their viewing experience more accurately when inputting barrage.

[0094] Next, combine Figure 5 The implementation process of generating video barrage text by combining historical barrage data is described in detail.

[0095] Reference Figure 5 , shows a flowchart of the steps of a method for obtaining video barrage text provided by an embodiment of the present application. Figure 5 As shown, the method for obtaining video barrage text may include: step 501 and step 502.

[0096] Step 501: Generate the first video barrage text of the video file in the specified video playback period according to the target video keyword.

[0097] In this embodiment, after obtaining the target video keywords, the first video barrage text of the video file in the specified video playback period can be generated based on the target video keywords. Specifically, the target video keywords can be directly used as the first video barrage text, or the target video keywords can be optimized (such as colloquial processing of the target video keywords) to obtain the first video barrage text.

[0098] After generating the first video barrage text of the video file in the specified video playback period according to the target video keyword, step 502 is executed.

[0099] Step 502: Update the first video barrage text according to the historical barrage data of the first user to obtain the video barrage text corresponding to the first user.

[0100] After generating the first video barrage text of the video file in the specified video playback period based on the target video keywords, the first video barrage text can be updated based on the historical barrage data of the first user, thereby obtaining the video barrage text corresponding to the first user. For example, if it is determined based on the historical barrage data that the user likes to use emoticons, relevant emoticons can be added to the first video barrage text (such as adding corresponding emoticons based on the emotion of the first video barrage text, etc.) to generate the video barrage text, etc.

[0101] It can be understood that the above examples are merely examples listed for a better understanding of the technical solutions of the embodiments of the present application, and are not intended to be the sole limitation on the embodiments.

[0102] The embodiment of the present application generates corresponding video barrage text by combining the user's historical barrage data, so that the generated video barrage text can be made consistent with the user's barrage input style on the basis of being relevant to the video plot, thereby improving the user's barrage interaction experience.

[0103] Next, combine Figure 6 The implementation process of generating video barrage text based on user needs is described in detail.

[0104] Reference Figure 6 , shows a flowchart of another method for obtaining video barrage text provided by an embodiment of the present application. Figure 6 As shown, the method for obtaining video barrage text may include: step 601 and step 602.

[0105] Step 601: Generate a second video bullet screen text for the video file during the specified video playback period based on the target video keyword.

[0106] In this embodiment, after obtaining the target video keywords, the second video barrage text of the video file in the specified video playback period can be generated based on the target video keywords. Specifically, the target video keywords can be directly used as the second video barrage text, or the target video keywords can be optimized (such as colloquial processing of the target video keywords) to obtain the second video barrage text.

[0107] After generating the second video bullet screen text of the video file in the specified video playback period according to the target video keyword, step 602 is executed.

[0108] Step 602: Update the second video barrage text according to the barrage requirement information of the second user to obtain the video barrage text corresponding to the second user.

[0109] After generating the second video barrage text of the video file in the specified video playback period according to the target video keywords, the second video barrage text can be updated according to the barrage demand information of the second user to obtain the video barrage text corresponding to the second user. For example, when generating the video barrage text, a barrage demand input box can be displayed on the screen of the electronic device used by the second user, and several demand options can be given for the user to input the barrage demand, and the video barrage text corresponding to the second user can be generated by combining the barrage demand and the second video barrage text. Or, for example, when the barrage demand input by the user is to add special marks, symbols, etc. at the end of the barrage text, special marks, symbols, etc. can be added at the end of the second video barrage text to obtain the video barrage text, etc.

[0110] It can be understood that the above examples are merely examples listed for a better understanding of the technical solutions of the embodiments of the present application, and are not intended to be the sole limitation on the embodiments.

[0111] The embodiment of the present application generates corresponding video barrage text by combining the user's barrage needs, so that the generated video barrage text can meet the user's personalized needs on the basis of being relevant to the video plot, thereby improving the user's barrage interactive experience.

[0112] Next, combine Figure 7 A detailed description of the video comment recommendation process is given.

[0113] Reference Figure 7 , shows a flowchart of the steps of a target video barrage text recommendation method provided by an embodiment of the present application, such as Figure 7 As shown, the target video barrage text recommendation method may include: step 701, step 702, step 703 and step 704.

[0114] Step 701: During the playback of a target video file, when a bullet screen initiation operation triggered by a user is received, a target video playback time of the target video file corresponding to the bullet screen initiation operation is obtained.

[0115] In this embodiment, when a user triggers a bullet screen initiation operation (such as a user launching a bullet screen publisher) during the playback of a target video file, the target video playback time of the target video file corresponding to the bullet screen initiation operation can be obtained. For example, when a user is watching a variety show video and reaches 00:55 in the video, the video playback time of the 55th minute of the variety show video can be used as the target video playback time.

[0116] In this example, the video file mentioned in the above implementation process includes the target video file.

[0117] Step 702: Obtain a target video file identifier of the target video file.

[0118] When the user plays the target video file, the target video file identifier of the target video file can be obtained.

[0119] Step 703: According to the pre-stored correspondence between the video file identifier, the designated video playback time period and the video barrage text, obtain the target video barrage text corresponding to the time period of the target video playback moment and the target video file identifier.

[0120] After obtaining the target video playback time of the target video file corresponding to the barrage initiation operation and the target video file identifier, the target video barrage text corresponding to the time period of the target video playback time and the target video file identifier can be obtained based on the correspondence between the pre-stored video file identifier, the specified video playback time period and the video barrage text.

[0121] In this example, the number of target video barrage texts can be one or more, and specifically, it can be determined according to actual conditions, and this embodiment does not limit this.

[0122] After obtaining the target video barrage text, execute step 704.

[0123] Step 704: Recommend the target video barrage text to the user.

[0124] After obtaining the target video barrage text, the target video barrage text can be recommended to the user. In this embodiment, the target video barrage text can be displayed in a list form for barrage text recommendation, such as displaying a list in the middle area or other areas of the terminal screen, and the target video barrage text can be displayed in sequence in the list. The target video barrage text can also be displayed in a pop-up window form for barrage text recommendation, for example, a pop-up window can be displayed in the lower right corner area or the upper left corner area of ​​the terminal screen, and the target video barrage text can be displayed in the pop-up window. The target video barrage text can also be displayed in the form of a sidebar or bottom bar, for example, a fixed or expandable sidebar / bottom bar is set on one side or the bottom of the terminal screen to display the recommended target video barrage text, etc.

[0125] It can be understood that the above examples are only listed for a better understanding of the technical solutions of the embodiments of the present application. In actual applications, other methods can also be used to display the target video barrage text, such as card-type display, etc. This embodiment does not limit this.

[0126] In an embodiment of the present application, when a user inputs a barrage, a specific vocabulary is found based on the video identifier and the current playback time, and corresponding words are recommended from the vocabulary, so that the user can quickly select and complete the barrage related to the video plot and send it.

[0127] The generation process of video barrage and the recommendation process of video barrage can be combined Figure 8 This is described in detail below.

[0128] Reference Figure 8 , shows a schematic diagram of a video barrage generation and recommendation process provided by an embodiment of the present application. Figure 8 As shown, the video comment generation and recommendation process can be as follows: First, the video plot text can be collected. Specifically, for video files, the script, dialogue, narration, and other information can be collected as the video plot text. Second, the plot keyword extraction function of the BERT large language model can be used to obtain the semantic information of each keyword in the text, including the keyword's meaning, part of speech, and sentiment. Then, a specific vocabulary can be generated. Specifically, based on the semantic feature vector of the keyword and the text encoding vector of the video plot text, the similarity index between the keyword and the video plot text is calculated. Based on the similarity index, keywords with high similarity are filtered from the keywords to generate a specific vocabulary (i.e., a database containing specific comment text). The output data structure can be: <video ID, play time (minutes), vocabulary>. Finally, a comment input prompt is provided: When the user enters a comment, a specific vocabulary is found based on the video ID and the current play time, and corresponding words are recommended from the vocabulary, allowing the user to quickly select and complete a comment related to the video plot and send it.

[0129] This embodiment generates a specific vocabulary closely related to the video plot by understanding and analyzing the plot text. This allows users to more easily find expressions related to the video content when entering barrage comments, improving the interactive experience and participation in barrage comments. Furthermore, this method can dynamically generate specific vocabulary based on different video content, providing a personalized barrage experience and increasing user stickiness.

[0130] The video barrage generation method provided by the embodiment of the present application obtains the video plot text of the video file within the specified video playback period. The semantic feature information of the video keywords in the video plot text is obtained. The video plot text and the semantic feature information are respectively encoded to obtain the text encoding vector of the video plot text and the semantic feature vector corresponding to the semantic feature information. According to the semantic feature vector and the text encoding vector, the target video keyword in the video keyword whose similarity index with the video plot text is greater than the threshold is determined. According to the target video keyword, the video barrage text of the video file in the specified video playback period is generated. The video file identifier, the specified video playback period and the video barrage text of the video file are stored in association. The embodiment of the present application analyzes and understands the video plot text of the video file within the specified video playback period to generate a specific video barrage vocabulary related to the video plot, so that users can express their viewing experience more accurately when inputting barrage, thereby improving the experience of barrage interaction.

[0131] Reference Figure 9 , shows a structural diagram of a video barrage generation device provided by an embodiment of the present application, such as Figure 9 As shown, the video barrage generation device 900 may include the following modules:

[0132] The video plot text acquisition module 910 is used to acquire the video plot text of the video file within the specified video playback period;

[0133] Semantic feature information acquisition module 920, used to obtain semantic feature information of video keywords in the video plot text;

[0134] A text semantic vector acquisition module 930 is configured to encode the video plot text and the semantic feature information respectively to obtain a text encoding vector of the video plot text and a semantic feature vector corresponding to the semantic feature information;

[0135] a target keyword determination module 940 for determining, based on the semantic feature vector and the text encoding vector, a target video keyword whose similarity index with the video plot text is greater than a threshold value among the video keywords;

[0136] The video barrage text generation module 950 is used to generate the video barrage text of the video file during the specified video playback period according to the target video keywords;

[0137] The video barrage text storage module 960 is used to store the video file identifier of the video file, the specified video playback time period and the video barrage text in an associated manner.

[0138] Optionally, the target keyword determination module includes:

[0139] A similarity index obtaining unit, configured to perform a correlation analysis on the semantic feature vector and the text encoding vector to obtain a similarity index between the video keyword and the video plot text;

[0140] The target keyword screening unit is used to screen out target video keywords whose similarity index is greater than a threshold from the video keywords.

[0141] Optionally, the similarity index obtaining unit includes:

[0142] A first index acquisition subunit is configured to acquire a cosine distance between the semantic feature vector and the text encoding vector, and use the cosine distance as the similarity index;

[0143] The second index acquisition subunit is used to obtain the Euclidean distance between the semantic feature vector and the text encoding vector, and use the Euclidean distance as the similarity index.

[0144] Optionally, the video barrage text generation module includes:

[0145] A first barrage text generating unit, configured to generate a first video barrage text for the video file during the specified video playback period according to the target video keyword;

[0146] The first video barrage text acquisition unit is used to update the first video barrage text according to the historical barrage data of the first user, and obtain the video barrage text corresponding to the first user.

[0147] Optionally, the video barrage text generation module includes:

[0148] A second barrage text generating unit is configured to generate a second video barrage text for the video file during the specified video playback period according to the target video keyword;

[0149] The second video barrage text acquisition unit is used to update the second video barrage text according to the barrage requirement information of the second user, and obtain the video barrage text corresponding to the second user.

[0150] Optionally, the device further comprises:

[0151] The target playback time acquisition module is used to obtain the target video playback time of the target video file corresponding to the barrage initiation operation when a barrage initiation operation is received by the user during the playback of the target video file;

[0152] A target file identifier acquisition module, configured to acquire a target video file identifier of the target video file;

[0153] A target barrage text acquisition module is used to obtain the target video barrage text corresponding to the time period of the target video playback moment and the target video file identifier according to the pre-stored correspondence between the video file identifier, the designated video playback time period and the video barrage text;

[0154] The target barrage text recommendation module is used to recommend the target video barrage text to the user.

[0155] Optionally, the video plot text acquisition module includes:

[0156] An initial plot text obtaining unit, configured to obtain the initial plot text of the video file during the specified video playback period;

[0157] A video plot text acquisition unit, configured to process the initial plot text based on a preprocessing method to obtain the video plot text;

[0158] The preprocessing method includes: at least one of word segmentation processing, part-of-speech tagging processing and syntactic analysis processing.

[0159] Optionally, the semantic feature information acquisition module includes:

[0160] Processing the video plot text based on a pre-trained semantic analysis model to obtain candidate semantic feature information of video keywords in the video plot text;

[0161] Processing the candidate semantic feature information based on a preset post-processing method to obtain the semantic feature information of the video keyword;

[0162] The preset post-processing method includes at least one of: repeated feature removal and feature part-of-speech restoration processing.

[0163] The video barrage generation device provided by the embodiment of the present application obtains the video plot text of the video file within the specified video playback period. The semantic feature information of the video keywords in the video plot text is obtained. The video plot text and the semantic feature information are respectively encoded to obtain the text encoding vector of the video plot text and the semantic feature vector corresponding to the semantic feature information. According to the semantic feature vector and the text encoding vector, the target video keyword in the video keyword whose similarity index with the video plot text is greater than the threshold is determined. According to the target video keyword, the video barrage text of the video file in the specified video playback period is generated. The video file identifier, the specified video playback period and the video barrage text of the video file are stored in association. The embodiment of the present application analyzes and understands the video plot text of the video file within the specified video playback period to generate a specific video barrage vocabulary related to the video plot, so that users can express their viewing experience more accurately when inputting barrage, thereby improving the experience of barrage interaction.

[0164] The present application also provides an electronic device, such as Figure 10 As shown, it includes a processor 1001, a communication interface 1002, a memory 1003 and a communication bus 1004, wherein the processor 1001, the communication interface 1002, and the memory 1003 communicate with each other through the communication bus 1004.

[0165] Memory 1003, used for storing computer programs;

[0166] The processor 1001 is configured to execute the program stored in the memory 1003 by performing the following steps:

[0167] Get the video plot text of the video file within the specified video playback period;

[0168] Obtaining semantic feature information of video keywords in the video plot text;

[0169] Encoding the video plot text and the semantic feature information respectively to obtain a text encoding vector of the video plot text and a semantic feature vector corresponding to the semantic feature information;

[0170] Determining, based on the semantic feature vector and the text encoding vector, target video keywords whose similarity index with the video plot text is greater than a threshold among the video keywords;

[0171] Generate video barrage text for the video file during the specified video playback period based on the target video keywords;

[0172] The video file identifier of the video file, the designated video playback time period and the video barrage text are stored in association.

[0173] Optionally, determining, based on the semantic feature vector and the text encoding vector, target video keywords having a similarity index with the video plot text greater than a threshold among the video keywords includes:

[0174] Performing a correlation analysis on the semantic feature vector and the text encoding vector to obtain a similarity index between the video keywords and the video plot text;

[0175] Target video keywords with the similarity index greater than a threshold are screened out from the video keywords.

[0176] Optionally, performing a correlation analysis on the semantic feature vector and the text encoding vector to obtain a similarity index between the video keyword and the video plot text includes:

[0177] Obtaining a cosine distance between the semantic feature vector and the text encoding vector, and using the cosine distance as the similarity index; or

[0178] The Euclidean distance between the semantic feature vector and the text encoding vector is obtained, and the Euclidean distance is used as the similarity index.

[0179] Optionally, generating the video barrage text of the video file during the specified video playback period according to the target video keyword includes:

[0180] Generate the first video barrage text of the video file during the specified video playback period according to the target video keyword;

[0181] According to the historical barrage data of the first user, the first video barrage text is updated to obtain the video barrage text corresponding to the first user.

[0182] Optionally, generating the video barrage text of the video file during the specified video playback period according to the target video keyword includes:

[0183] Generate a second video barrage text for the video file during the specified video playback period according to the target video keyword;

[0184] According to the barrage requirement information of the second user, the barrage text of the second video is updated to obtain the video barrage text corresponding to the second user.

[0185] Optionally, after the associated storage of the video file identifier of the video file, the designated video playback time period, and the video commentary text, the method further includes:

[0186] During the playback of the target video file, when a user-triggered bullet screen initiation operation is received, obtaining the target video playback time of the target video file corresponding to the bullet screen initiation operation;

[0187] Obtain a target video file identifier of the target video file;

[0188] According to the pre-stored correspondence between the video file identifier, the designated video playback time period, and the video barrage text, obtaining the target video barrage text corresponding to the time period of the target video playback moment and the target video file identifier;

[0189] Recommend the target video barrage text to the user.

[0190] Optionally, obtaining the video plot text of the video file within a specified video playback period includes:

[0191] Obtaining the initial plot text of the video file during the specified video playback period;

[0192] Processing the initial plot text based on a preprocessing method to obtain the video plot text;

[0193] The preprocessing method includes: at least one of word segmentation processing, part-of-speech tagging processing and syntactic analysis processing.

[0194] Optionally, obtaining semantic feature information of video keywords in the video plot text includes:

[0195] Processing the video plot text based on a pre-trained semantic analysis model to obtain candidate semantic feature information of video keywords in the video plot text;

[0196] Processing the candidate semantic feature information based on a preset post-processing method to obtain the semantic feature information of the video keyword;

[0197] The preset post-processing method includes at least one of: repeated feature removal and feature part-of-speech restoration processing.

[0198] The communication bus mentioned in the above terminal can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, only one thick line is used in the figure, but this does not mean that there is only one bus or only one type of bus.

[0199] The communication interface is used for communication between the above terminal and other devices.

[0200] The memory may include random access memory (RAM) or non-volatile memory, such as at least one disk storage. Alternatively, the memory may be at least one storage device located away from the processor.

[0201] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0202] In another embodiment provided by the present application, a computer-readable storage medium is also provided, which stores instructions. When the computer-readable storage medium is run on a computer, it enables the computer to execute the video barrage generation method described in any of the above embodiments.

[0203] In another embodiment provided by the present application, a computer program product comprising instructions is also provided, on which a computer program is stored. When the computer program is run on a computer, the computer executes any of the above-mentioned methods for generating video barrages.

[0204] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When software is used for implementation, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server or data center by wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) mode to another website, computer, server or data center. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrations. The available medium may be a magnetic medium (eg, a floppy disk, a hard disk, a magnetic tape), an optical medium (eg, a DVD), or a semiconductor medium (eg, a solid state drive (SSD)).

[0205] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.

[0206] Each embodiment in this specification is described in a related manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiment is generally similar to the method embodiment, so the description is relatively simple. For related parts, refer to the description of the method embodiment.

[0207] The above description is only a preferred embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application are included in the scope of protection of the present application.

Claims

1. A video barrage generation method, characterized in that: include: Get the video plot text of the video file within the specified video playback period; Obtaining semantic feature information of video keywords in the video plot text; Encoding the video plot text and the semantic feature information respectively to obtain a text encoding vector of the video plot text and a semantic feature vector corresponding to the semantic feature information; Determining, based on the semantic feature vector and the text encoding vector, target video keywords whose similarity index with the video plot text is greater than a threshold among the video keywords; Generate video barrage text for the video file during the specified video playback period based on the target video keywords; The video file identifier of the video file, the designated video playback time period and the video barrage text are stored in association.

2. The method according to claim 1, characterized in that The step of determining, based on the semantic feature vector and the text encoding vector, a target video keyword having a similarity index greater than a threshold value with the video plot text among the video keywords comprises: Performing a correlation analysis on the semantic feature vector and the text encoding vector to obtain a similarity index between the video keywords and the video plot text; Target video keywords with the similarity index greater than a threshold are screened out from the video keywords.

3. The method according to claim 2, characterized in that The performing correlation analysis on the semantic feature vector and the text encoding vector to obtain a similarity index between the video keyword and the video plot text includes: Obtaining a cosine distance between the semantic feature vector and the text encoding vector, and using the cosine distance as the similarity index; or The Euclidean distance between the semantic feature vector and the text encoding vector is obtained, and the Euclidean distance is used as the similarity index.

4. The method according to claim 1, wherein Generating the video barrage text of the video file during the specified video playback period according to the target video keyword includes: Generate the first video barrage text of the video file during the specified video playback period according to the target video keyword; According to the historical barrage data of the first user, the first video barrage text is updated to obtain the video barrage text corresponding to the first user.

5. The method according to claim 1, wherein Generating the video barrage text of the video file during the specified video playback period according to the target video keyword includes: Generate a second video barrage text for the video file during the specified video playback period according to the target video keyword; According to the barrage requirement information of the second user, the barrage text of the second video is updated to obtain the video barrage text corresponding to the second user.

6. The method according to claim 1, characterized in that After the video file identifier of the video file, the designated video play time period, and the video commentary text are stored in association, the method further includes: During the playback of the target video file, when a user-triggered bullet screen sending operation is received, obtaining the target video playback time of the target video file corresponding to the bullet screen sending operation; Obtain a target video file identifier of the target video file; According to the pre-stored correspondence between the video file identifier, the designated video playback time period, and the video barrage text, obtaining the target video barrage text corresponding to the time period of the target video playback moment and the target video file identifier; Recommend the target video barrage text to the user.

7. The method according to claim 1, characterized in that The step of obtaining the video plot text of the video file within the specified video playback period includes: Obtaining the initial plot text of the video file during the specified video playback period; Processing the initial plot text based on a preprocessing method to obtain the video plot text; The preprocessing method includes: at least one of word segmentation processing, part-of-speech tagging processing and syntactic analysis processing.

8. The method according to claim 1, characterized in that The obtaining of semantic feature information of video keywords in the video plot text includes: Processing the video plot text based on a pre-trained semantic analysis model to obtain candidate semantic feature information of video keywords in the video plot text; Processing the candidate semantic feature information based on a preset post-processing method to obtain the semantic feature information of the video keyword; The preset post-processing method includes at least one of: repeated feature removal and feature part-of-speech restoration processing.

9. A video barrage generation device, characterized in that: include: A video plot text acquisition module is used to obtain the video plot text of a video file within a specified video playback period; A semantic feature information acquisition module, used to acquire semantic feature information of video keywords in the video plot text; A text semantic vector acquisition module is used to encode the video plot text and the semantic feature information respectively to obtain a text encoding vector of the video plot text and a semantic feature vector corresponding to the semantic feature information; a target keyword determination module, configured to determine, based on the semantic feature vector and the text encoding vector, a target video keyword whose similarity index with the video plot text is greater than a threshold value among the video keywords; A video barrage text generation module, configured to generate a video barrage text for the video file during the specified video playback period according to the target video keywords; The video barrage text storage module is used to associate and store the video file identifier of the video file, the specified video playback time period and the video barrage text.

10. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus; Memory for storing computer programs; A processor, configured to implement the method steps described in any one of claims 1 to 8 when executing a program stored in a memory.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.

12. A computer program product comprising instructions, on which a computer program is stored, characterized in that When the computer program is run on a computer, the computer is caused to execute the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Bullet screen association input method and apparatus, and computer readable storage medium

    CN108924658A

  • Bullet screen emotion analysis method and device based on bidirectional long and short term neural network

    CN114791943A