Title generation method, apparatus, electronic device and readable storage medium

By extracting dialogue text and visual features from video data to generate video titles, the problem of low matching degree between video titles and video content is solved, and more efficient video title generation is achieved.

CN115795092BActive Publication Date: 2026-01-30BEIJING QIYI CENTURY SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211657686.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-22
Publication Date
2026-01-30
Estimated Expiration
2042-12-22

AI Technical Summary

Technical Problem

In existing technologies, the matching degree between video titles and video content is low, resulting in a large workload and low efficiency in generating video titles.

Method used

By acquiring the dialogue text and video footage from the video data, feature extraction is performed to generate text features and video features. Based on these two features, a video title is generated, taking into account information from multiple modalities.

Benefits of technology

It improves the matching degree between video titles and video data, enriches the information content of video titles, and reduces the workload of generating video titles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115795092B_ABST
    Figure CN115795092B_ABST
Patent Text Reader

Abstract

This invention provides a title generation method, apparatus, electronic device, and readable storage medium. The method includes: acquiring video data of a target video, the video data including dialogue text and video frames; extracting features from the video data to obtain text features and frame features, wherein the text features characterize the semantics of the dialogue text, and the frame features characterize the content of the video frames; and generating a video title for the target video based on the text features and the frame features. In this way, the information in the text features and the information in the frame features complement each other during the video title generation process, enabling the generated video title to comprehensively consider information from multiple modalities, enriching the information content of the video title, and improving the matching degree between the video title and the video data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video technology, and in particular to a title generation method, apparatus, electronic device, and readable storage medium. Background Technology

[0002] With the rapid development of the short video industry, the number of video works being created is also increasing. Each video usually includes a title so viewers can quickly understand its content. However, with such a large number of videos, even the creators themselves may not remember the content of some. This necessitates reviewing the videos again during editing to determine suitable titles, resulting in a significant workload for title generation.

[0003] Currently, the title of a video is usually obtained by extracting lines from the video. While this improves the efficiency of title generation, the resulting title often has a low degree of matching with the video content, and sometimes the title deviates from the theme that the video is trying to express.

[0004] It is evident that existing technologies suffer from a low degree of matching between video titles and video content. Summary of the Invention

[0005] The purpose of this invention is to provide a title generation method, apparatus, electronic device, and readable storage medium to solve the problem of low matching degree between video titles and video content. The specific technical solution is as follows:

[0006] In a first aspect of this invention, a title generation method is provided, comprising:

[0007] Acquire video data of the target video, the video data including dialogue text and video footage;

[0008] Feature extraction is performed on the video data to obtain text features and image features, wherein the text features are used to characterize the semantics of the dialogue text, and the image features are used to characterize the content of the video images;

[0009] Based on the text features and the image features, a video title for the target video is generated.

[0010] In a second aspect of the invention, a title generation apparatus is provided, comprising:

[0011] The acquisition module is used to acquire video data of the target video, the video data including dialogue text and video footage;

[0012] The feature extraction module is used to extract features from the video data to obtain text features and image features, wherein the text features are used to characterize the semantics of the dialogue text, and the image features are used to characterize the content of the video images;

[0013] The generation module is used to generate a video title for the target video based on the text features and the image features.

[0014] In a third aspect of the present invention, an electronic device is provided, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus.

[0015] Memory, used to store programs;

[0016] When a processor executes a program stored in memory, it implements the method described in the first aspect.

[0017] In a fourth aspect of the invention, a readable storage medium is provided having a program stored thereon that, when executed by a processor, implements the method described in the first aspect.

[0018] In this embodiment of the application, video data of the target video is acquired, and feature extraction is performed on the video data to obtain text features and image features. Based on the text features and image features, a video title of the target video is generated. In this way, the information in the text features and the information in the image features complement each other in the process of generating the video title, so that the generated video title comprehensively considers information from multiple modalities, enriches the information content of the video title, and improves the matching degree between the video title and the video data. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below.

[0020] Figure 1 This is a flowchart illustrating a title generation method according to an embodiment of the present invention;

[0021] Figure 2 This is a schematic diagram of the structure of a title generation device according to an embodiment of the present invention;

[0022] Figure 3 This is a structural diagram of an electronic device according to an embodiment of the present invention. Detailed Implementation

[0023] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0024] Before discussing the exemplary embodiments in more detail, it should be noted that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the steps as sequential processes, many of these steps can be performed in parallel, concurrently, or simultaneously. Furthermore, the order of the steps can be rearranged. A process can be terminated when its operation is complete, but it may also have additional steps not included in the figures. A process can correspond to a method, function, procedure, subroutine, subroutine, etc.

[0025] This application provides a title generation method, such as... Figure 1 As shown, the steps of this method include:

[0026] Step 101: Obtain video data of the target video, the video data including dialogue text and video footage;

[0027] The target video can be a long video, a short video, or a video clip—any video to be processed that requires a video title. The dialogue text can be lines from the target video, such as conversations or monologues between characters, presented in text modality. The video footage can include multiple frames of images, which may contain information such as objects and the environment; for example, a boy riding a bicycle, presented in video modality. A portion of the target video's plot can be understood through either the dialogue text or the video footage.

[0028] Step 102: Extract features from the video data to obtain text features and image features, wherein the text features are used to characterize the semantics of the dialogue text, and the image features are used to characterize the content of the video images;

[0029] Among them, text features can be used to characterize the semantics of a line or a paragraph of dialogue. A word or phrase in the dialogue text can be used as a feature term, the matching degree between the feature term and the entire dialogue text can be calculated, and the feature term with a high matching degree can be identified as a text feature.

[0030] Among them, image features can be used to characterize the content of one or more images. The colors of one or more images in a video frame can be standardized (or normalized), the image can be divided into multiple cell units, the gray-level gradient of each cell can be calculated, and one or more cells can be grouped into a block based on the gray-level gradient. The features (descriptors) of all cells in a block are concatenated to obtain the Histogram of Oriented Gradient (HOG) features of that block. Thus, the corresponding image features can be obtained by concatenating the HOG features of all blocks in one or more images.

[0031] The dialogue text contains numerous words or phrases, and may include non-critical textual content (e.g., when the generated video title does not need to reflect time, background, or other information, non-critical textual content can be words or phrases used to annotate time, background, etc.). This results in a high computational load when generating video titles from dialogue text. Therefore, text features obtained by feature extraction from video data can be key information from the dialogue text (e.g., mapping or transformation methods can be used to transform the dialogue text into its most representative features) to simplify subsequent calculations and improve the matching degree between the generated video title and the dialogue text. Similarly, video frames may include non-critical image content (e.g., when the generated video title does not need to reflect environmental information, non-critical image content can be environmental information in the image). This results in a high computational load when generating video titles from video frames. Therefore, frame features obtained by feature extraction from video data can be key information from the video frames (e.g., HOG feature extraction methods can be used to extract features from the character information contained in the video frames) to simplify subsequent calculations and improve the matching degree between the generated video title and the video frames.

[0032] Step 103: Generate the video title of the target video based on the text features and the image features.

[0033] The video titles generated based on text features and image features take into account both text and video modalities, in order to integrate information from text features and image features, thereby improving the matching degree between video titles and video data.

[0034] In this embodiment, video data of the target video is acquired, and feature extraction is performed on the video data to obtain text features and image features. Based on the text features and image features, a video title of the target video is generated. In this way, the information in the text features and the information in the image features complement each other in the process of generating the video title, so that the generated video title comprehensively considers information from multiple modalities, enriches the information content of the video title, and improves the matching degree between the video title and the video data.

[0035] Optionally, in step 103, generating the video title of the target video based on the text features and the image features includes:

[0036] A first subtitle for the target video is generated based on either the text feature or the image feature.

[0037] The first subtitle is adjusted based on the second subtitle to obtain the video title of the target video. The second subtitle is a subtitle of the target video generated based on the other of the text features and the image features.

[0038] In one embodiment, a first subtitle of the target video can be generated based on text features, that is, the first subtitle can be a subtitle obtained by summarizing the dialogue text; then, a second subtitle of the target video can be generated based on screen features, that is, the second subtitle can be a subtitle obtained by summarizing the video screen, and the information included in the second subtitle can make up for the missing information in the first subtitle, so that the generated video title of the target video comprehensively considers information from multiple modalities, enriches the information content of the video title, and improves the matching degree between the video title and the video data.

[0039] For example, after feature extraction from the dialogue text, the feature terms that match the entire dialogue text well may include "Xiaoming", "school", and "first day". Therefore, the first subtitle of the target video generated based on the text features could be "Xiaoming's () first day of school". After HOG feature extraction from the video frame, the characters in one or more images corresponding to the video frame can be obtained. The facial expressions of the characters can be further obtained through Local Binary Pattern (LBP) feature extraction, thereby determining that the characters in the video frame have the feature of "happy". Thus, the second subtitle of the target video generated based on the frame features could be "() happy ()", and the video title of the target video could be "Xiaoming's happy day" or "Xiaoming's happy first day of school".

[0040] In another embodiment, a first subtitle for the target video can be generated based on screen features, that is, the first subtitle can be a subtitle obtained by summarizing the video screen; then, a second subtitle for the target video can be generated based on text features, that is, the second subtitle can be a subtitle obtained by summarizing the dialogue text, and the information included in the second subtitle can make up for the missing information in the first subtitle, so that the generated video title of the target video comprehensively considers information from multiple modalities, enriches the information content of the video title, and improves the matching degree between the video title and the video data.

[0041] For example, the first subtitle of the target video generated based on image features could be "Food ()", and the second subtitle of the target video generated based on text features could be "() Strategy Collection". Then the video title of the target video could be "Food Strategy Collection".

[0042] Optionally, in the above-described adjustment of the first subheading based on the second subheading, the adjustment method of the first subheading includes at least one of the following:

[0043] Extract the first keyword from the first subtitle and the second keyword from the second subtitle. The first keyword and the second keyword have the same first semantic meaning. The video title of the target video includes the first semantic meaning.

[0044] Extract the third keyword from the first subtitle and the fourth keyword from the second subtitle. The third keyword and the fourth keyword have different second and third semantics. The video title of the target video includes the second semantic and the third semantic.

[0045] There are generally eight sentence components in modern Chinese: subject, predicate, object, verb, attributive, adverbial, complement, and headword.

[0046] In one example, adjusting the first subtitle based on the second subtitle requires deleting or merging words or phrases that repeat in the subject, predicate, object, verb, attributive, adverbial, complement, and headword. For instance, after feature extraction from the dialogue text, features with high matching degrees to the entire text might include "cake," "cheese," "strawberry chocolate," and "making." Therefore, based on these features and logically supplemented according to sentence components in modern Chinese, the first subtitle of the target video could be "How to Make Desserts," or in other words, the first keyword in the first subtitle could be "desserts." Further HOG feature extraction from the video footage yields target objects in one or more images corresponding to the video footage. These target objects might include "eggs," "milk," and "strawberries." Therefore, based on these target objects, the second keyword in the generated second subtitle could be "ingredients." Thus, when the first keyword in the first subtitle and the second keyword in the second subtitle express the same first semantic (i.e., food), words or phrases corresponding to the first semantic can be added to the subject of the video title, making the target video title include the first semantic. For example, the target video title could be "How to Make Food." It should be noted that in other examples, the first semantic element can be added to the position of the corresponding element in the video title based on the corresponding element in Chinese, which can achieve the same technical effect. This will not be elaborated on here.

[0047] In another example, after feature extraction from the dialogue text, features with a high degree of matching with the entire dialogue text can include "refrigeration," "3 to 8 degrees Celsius," and "refrigerator." Therefore, based on these features and logically supplemented according to sentence components in modern Chinese, the first subtitle of the target video can be "saving method." In other words, the third keyword in the first subtitle can be "saving." Further HOG feature extraction from the video frame yields the target objects in one or more images corresponding to the video frame. These target objects can include "cheese," "chocolate," and "cake." Therefore, based on these target objects, the fourth keyword in the generated second subtitle can be "dessert." Thus, when the third keyword in the first subtitle and the fourth keyword in the second subtitle express different second and third semantics, the word or phrase corresponding to the second semantic (i.e., "saving") can be added to the predicate position of the video title, and the word or phrase corresponding to the third semantic (i.e., "dessert") can be added to the subject position of the video title, so that the video title of the target video includes both the second and third semantics. For example, the video title of the target video could be "dessert saving method."

[0048] In this way, by combining the first and second subheadings during the video title generation process, the generated video title takes into account information from multiple modalities, enriches the information content of the video title, and improves the matching degree between the video title and the video data. At the same time, it can delete or merge repeated words or phrases in the same position in the subject, predicate, object, verb, attributive, adverbial, complement, and headword, thereby improving the fluency of the video title.

[0049] Optionally, in step 102, the feature extraction of the video data to obtain text features and image features includes:

[0050] The video data is input into a pre-trained first network model to obtain the text features and the image features;

[0051] The first network model is used to extract dialogue text at at least one first target location and video footage at at least one second target location from the video data, and to extract the text features from the extracted dialogue text at each first target location, and to extract the video footage features from the extracted video footage at each second target location.

[0052] In this embodiment, a pre-trained first network model can be used to extract features from video data to improve the accuracy of the obtained text features in representing the semantics of dialogue and the accuracy of the image features in representing the content of video images.

[0053] Specifically, the first target location can be determined based on the type of dialogue text in the video data, such as dialogue text or monologue text. This first target location can be the text feature extraction location determined after training the first network model. Correspondingly, the second target location can be determined based on the type of video frame in the video data, such as people, plants, or environment. This second target location can be the frame feature extraction location determined after training the first network model. In this way, by inputting the video data into the pre-trained first network model to obtain text and frame features, the accuracy of text features in representing the semantics of the dialogue text and the accuracy of frame features in representing the content of the video frames are improved. This increases the matching degree between the video title and the video data in the process of generating video titles based on text and frame features. Furthermore, the information in the text features and the information in the frame features complement each other in the process of generating video titles, ensuring that the generated video titles comprehensively consider information from multiple modalities, thus enriching the information content of the video titles.

[0054] Optionally, before the above step of inputting the video data into a pre-trained first network model, the method further includes:

[0055] The video sample data is input into the network model to be trained for N iterations of training. The video sample data includes sample dialogue text and sample video frames, where N is a positive integer.

[0056] If the network model to be trained after the Nth iteration meets the preset conditions, the network model to be trained after the Nth iteration is determined as the first network model.

[0057] The preset conditions include a matching degree between the target text features and preset text features that is greater than or equal to a first matching value, and a matching degree between the target screen features and preset screen features that is greater than or equal to a second matching value. The target text features are features generated by the network model to be trained based on the sample dialogue text. The preset text features are features preset in the sample dialogue text. The target screen features are features generated by the network model to be trained based on the sample video screen. The preset screen features are features preset in the sample video screen.

[0058] In this embodiment, the network model to be trained may include a Vector Space Model (VSM) and a HOG feature extraction model to be trained. Sample dialogue text is labeled with preset text features, and sample video footage is labeled with preset scene features. The sample dialogue text is input into the VSM model to be trained for text feature extraction, and the sample video footage is input into the HOG feature extraction model to be trained, and N iterations of training are performed for each.

[0059] Based on the VSM model trained for the Nth iteration, feature extraction is performed on the sample dialogue text to obtain target text features. If the matching degree (or similarity) between the target text features and preset text features is greater than or equal to a first matching value, the VSM model trained for the Nth iteration is considered complete. The first matching value can be adjusted according to the actual situation to determine a suitable feature extraction accuracy. Similarly, based on the HOG feature extraction model trained for the Nth iteration, feature extraction is performed on the sample video frames to obtain target frame features. If the matching degree (or similarity) between the target frame features and preset frame features is greater than or equal to a second matching value, the HOG feature extraction model trained for the Nth iteration is considered complete. The second matching value can be adjusted according to the actual situation to determine a suitable feature extraction accuracy.

[0060] The first network model can include a trained VSM model and a trained HOG feature extraction model. By inputting video data into the pre-trained first network model, textual and visual features are obtained, improving the matching accuracy during feature extraction and thus enhancing the matching accuracy between the video title and the video data.

[0061] It should be noted that other network models can be used for text feature extraction, and other network models can be used for image feature extraction, all of which can achieve the same technical effect. To avoid repetition, these will not be elaborated here.

[0062] Optionally, during one iteration of the N iterations of training, the network model to be trained is used to extract sample dialogue text at the first target position and sample video footage at the second target position from the video sample data, and to extract the target text features from the extracted sample dialogue text at the first target position and to extract the target video footage features from the extracted sample video footage at the second target position.

[0063] Wherein, the preset text feature is the preset feature of the sample dialogue text at the first preset position in the video sample data, and the preset screen feature is the preset feature of the sample video screen at the second preset position in the video sample data;

[0064] The preset conditions also include that the deviation between the first target position and the first preset position is less than or equal to a first deviation value; and the deviation between the second target position and the second preset position is less than or equal to a second deviation value.

[0065] In this embodiment, the target video can be of various types, such as historical videos, action videos, and biographical videos. The type of dialogue text also varies depending on the video type. Therefore, at least one dialogue text for a first target location can be obtained based on a specific video type, thereby improving the matching degree between the dialogue text and the target video. Furthermore, the type of video footage also varies depending on the video type. Therefore, at least one video footage for a second target location can be obtained based on a specific video type, thereby improving the matching degree between the video footage and the target video. For example, if the target video is an action video, the dialogue text for the first location may include text with many predicate verbs, and the video footage for the second target location may include images of fighting figures.

[0066] Thus, during N iterations of training, the accuracy of the training model in identifying the first and second target positions can be further improved. Based on the VSM model trained after the Nth iteration, when feature extraction of the sample dialogue text yields the target text features for the first target position, if the matching degree between the target text features and the preset text features is greater than or equal to a first matching value, and the deviation between the first target position and the first preset position is less than or equal to a first deviation value, then the VSM model trained after the Nth iteration can be considered complete.

[0067] The training of the HOG feature extraction model, based on the Nth iteration of training, is considered complete when the target image features at the second target location are obtained by extracting features from the sample video images. This is because the matching degree between the target image features and the preset image features is greater than or equal to a second matching value, and the deviation between the second target location and the second preset location is less than or equal to a second deviation value. This further improves the matching degree during feature extraction, thereby enhancing the matching degree between the video title and the video data.

[0068] Optionally, in the above step: adjusting the first subtitle based on the second subtitle to obtain the video title of the target video includes:

[0069] The first subtitle and the second subtitle are input into a pre-trained second network model to obtain the video title of the target video;

[0070] The second network model is used to align the timeline sequence corresponding to the second subtitle in the video data with the timeline sequence corresponding to the first subtitle in the video data, and then adjust the first subtitle based on the second subtitle to generate the video title.

[0071] In this embodiment, a second network model can be used to align the first modality corresponding to the dialogue text with the second modality corresponding to the video frame (mapping the text modality and video modality to the same space). This improves the correlation between the second and first subtitles during the process of adjusting the first subtitle based on the second subtitle to obtain the target video title. It avoids situations where the second and first subtitles correspond to different timeline sequences, facilitating the fusion of information from text features and image features. This allows the generated video title to comprehensively consider information from multiple modalities, enriching the information content of the video title and improving the matching degree between the video title and the video data.

[0072] The training process of the pre-trained second network model can be described as follows:

[0073] The sample dialogue text and sample video footage are input into the modality alignment model to be trained for M iterations. The sample dialogue text is labeled with a first preset timeline sequence, and the sample video footage is labeled with a second preset timeline sequence. M is a positive integer.

[0074] If the training modality alignment model after the Mth iteration meets the preset conditions, the training modality alignment model after the Mth iteration is determined as the second network model.

[0075] The preset conditions include that the first difference between the first target timeline sequence and the second target timeline sequence is less than or equal to the second difference between the first preset timeline sequence and the second preset timeline sequence. The first target timeline sequence is a timeline sequence generated by the modality alignment model to be trained based on sample dialogue text, and the second target timeline sequence is a timeline sequence generated by the modality alignment model to be trained based on sample video footage.

[0076] The modality alignment model to be trained can be a vision-language pre-training (VLP) model.

[0077] like Figure 2 As shown, this embodiment of the invention also provides a title generation device 200, comprising:

[0078] The acquisition module 201 is used to acquire video data of the target video, wherein the video data includes dialogue text and video footage;

[0079] The feature extraction module 202 is used to extract features from the video data to obtain text features and image features, wherein the text features are used to characterize the semantics of the dialogue text, and the image features are used to characterize the content of the video image;

[0080] The generation module 203 is used to generate a video title for the target video based on the text features and the image features.

[0081] Optionally, the generation module 203 includes:

[0082] A generation submodule is used to generate a first subtitle for the target video based on one of the text features and the image features;

[0083] The adjustment submodule is used to adjust the first subtitle based on the second subtitle to obtain the video title of the target video. The second subtitle is the subtitle of the target video generated based on the other of the text features and the image features.

[0084] Optionally, the adjustment method for the first subheading includes at least one of the following:

[0085] Extract the first keyword from the first subtitle and the second keyword from the second subtitle. The first keyword and the second keyword have the same first semantic meaning. The video title of the target video includes the first semantic meaning.

[0086] Extract the third keyword from the first subtitle and the fourth keyword from the second subtitle. The third keyword and the fourth keyword have different second and third semantics. The video title of the target video includes the second semantic and the third semantic.

[0087] Optionally, the feature extraction module 202 includes:

[0088] The input submodule is used to input the video data into a pre-trained first network model to obtain the text features and the image features;

[0089] The first network model is used to extract dialogue text at at least one first target location and video footage at at least one second target location from the video data, and to extract the text features from the extracted dialogue text at each first target location, and to extract the video footage features from the extracted video footage at each second target location.

[0090] Optionally, the device further includes:

[0091] The training module is used to input video sample data into the network model to be trained for N iterations of training. The video sample data includes sample dialogue text and sample video frames, where N is a positive integer.

[0092] The determination module is used to determine the network model to be trained after the Nth iteration of training as the first network model if the network model to be trained after the Nth iteration of training meets the preset conditions.

[0093] The preset conditions include a matching degree between the target text features and preset text features that is greater than or equal to a first matching value, and a matching degree between the target screen features and preset screen features that is greater than or equal to a second matching value. The target text features are features generated by the network model to be trained based on the sample dialogue text. The preset text features are features preset in the sample dialogue text. The target screen features are features generated by the network model to be trained based on the sample video screen. The preset screen features are features preset in the sample video screen.

[0094] Optionally, during one iteration of the N iterations of training, the network model to be trained is used to extract sample dialogue text at the first target position and sample video footage at the second target position from the video sample data, and to extract the target text features from the extracted sample dialogue text at the first target position and to extract the target video footage features from the extracted sample video footage at the second target position.

[0095] Wherein, the preset text feature is the preset feature of the sample dialogue text at the first preset position in the video sample data, and the preset screen feature is the preset feature of the sample video screen at the second preset position in the video sample data;

[0096] The preset conditions also include that the deviation between the first target position and the first preset position is less than or equal to a first deviation value; and the deviation between the second target position and the second preset position is less than or equal to a second deviation value.

[0097] Optionally, the adjustment submodule includes:

[0098] An input unit is used to input the first subtitle and the second subtitle into a pre-trained second network model to obtain the video title of the target video;

[0099] The second network model is used to align the timeline sequence corresponding to the second subtitle in the video data with the timeline sequence corresponding to the first subtitle in the video data, and then adjust the first subtitle based on the second subtitle to generate the video title.

[0100] The title generation device 200 provided in this embodiment of the invention can achieve Figure 1 The various processes implemented in the method embodiments shown are capable of achieving the same beneficial effects, and will not be described again here to avoid repetition.

[0101] This invention also provides an electronic device, such as... Figure 3 As shown, it includes a processor 301, a communication interface 302, a memory 303, and a communication bus 304, wherein the processor 301, the communication interface 302, and the memory 303 communicate with each other through the communication bus 304.

[0102] The memory 303 is used to store computer programs; the processor 301, when executing the program stored in the memory 303, performs the following steps:

[0103] Acquire video data of the target video, the video data including dialogue text and video footage;

[0104] Feature extraction is performed on the video data to obtain text features and image features, wherein the text features are used to characterize the semantics of the dialogue text, and the image features are used to characterize the content of the video images;

[0105] Based on the text features and the image features, a video title for the target video is generated.

[0106] Optionally, generating the video title of the target video based on the text features and the image features includes:

[0107] A first subtitle for the target video is generated based on either the text feature or the image feature.

[0108] The first subtitle is adjusted based on the second subtitle to obtain the video title of the target video. The second subtitle is a subtitle of the target video generated based on the other of the text features and the image features.

[0109] Optionally, the adjustment method for the first subheading includes at least one of the following:

[0110] Extract the first keyword from the first subtitle and the second keyword from the second subtitle. The first keyword and the second keyword have the same first semantic meaning. The video title of the target video includes the first semantic meaning.

[0111] Extract the third keyword from the first subtitle and the fourth keyword from the second subtitle. The third keyword and the fourth keyword have different second and third semantics. The video title of the target video includes the second semantic and the third semantic.

[0112] Optionally, the step of extracting features from the video data to obtain text features and image features includes:

[0113] The video data is input into a pre-trained first network model to obtain the text features and the image features;

[0114] The first network model is used to extract dialogue text at at least one first target location and video footage at at least one second target location from the video data, and to extract the text features from the extracted dialogue text at each first target location, and to extract the video footage features from the extracted video footage at each second target location.

[0115] Optionally, before inputting the video data into a pre-trained first network model, the method further includes:

[0116] The video sample data is input into the network model to be trained for N iterations of training. The video sample data includes sample dialogue text and sample video frames, where N is a positive integer.

[0117] If the network model to be trained after the Nth iteration meets the preset conditions, the network model to be trained after the Nth iteration is determined as the first network model.

[0118] The preset conditions include a matching degree between the target text features and preset text features that is greater than or equal to a first matching value, and a matching degree between the target screen features and preset screen features that is greater than or equal to a second matching value. The target text features are features generated by the network model to be trained based on the sample dialogue text. The preset text features are features preset in the sample dialogue text. The target screen features are features generated by the network model to be trained based on the sample video screen. The preset screen features are features preset in the sample video screen.

[0119] Optionally, during one iteration of the N iterations of training, the network model to be trained is used to extract sample dialogue text at the first target position and sample video footage at the second target position from the video sample data, and to extract the target text features from the extracted sample dialogue text at the first target position and to extract the target video footage features from the extracted sample video footage at the second target position.

[0120] Wherein, the preset text feature is the preset feature of the sample dialogue text at the first preset position in the video sample data, and the preset screen feature is the preset feature of the sample video screen at the second preset position in the video sample data;

[0121] The preset conditions also include that the deviation between the first target position and the first preset position is less than or equal to a first deviation value; and the deviation between the second target position and the second preset position is less than or equal to a second deviation value.

[0122] Optionally, adjusting the first subtitle based on the second subtitle to obtain the video title of the target video includes:

[0123] The first subtitle and the second subtitle are input into a pre-trained second network model to obtain the video title of the target video;

[0124] The second network model is used to align the timeline sequence corresponding to the second subtitle in the video data with the timeline sequence corresponding to the first subtitle in the video data, and then adjust the first subtitle based on the second subtitle to generate the video title.

[0125] By acquiring video data of the target video and extracting features from the video data to obtain text features and image features, the video title of the target video is generated based on the text features and image features. In this way, the information in the text features and the information in the image features complement each other in the process of generating the video title, so that the generated video title comprehensively considers information from multiple modalities, enriches the information content of the video title, and improves the matching degree between the video title and the video data.

[0126] The communication bus mentioned in the above electronic devices can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not indicate that there is only one bus or one type of bus.

[0127] The communication interface is used for communication between the aforementioned terminal and other devices.

[0128] The memory may include random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0129] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0130] In another embodiment of the present invention, a computer-readable storage medium is also provided, which stores instructions that, when executed on a computer, cause the computer to perform any of the title generation methods described in the above embodiments.

[0131] In another embodiment of the present invention, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute any of the title generation methods described in the above embodiments.

[0132] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)).

[0133] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0134] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0135] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.

Claims

1. A title generation method characterized by comprising: The method comprises: acquiring video data of a target video, the video data comprising dialogue text and video pictures; performing feature extraction on the video data to obtain text features and picture features, wherein the text features are used to represent semantics of the dialogue text, and the picture features are used to represent content of the video pictures; generating a video title of the target video based on the text features and the picture features; the generating of the video title of the target video based on the text features and the picture features comprises: generating a first sub-title of the target video based on one of the text features and the picture features; adjusting the first sub-title based on a second sub-title to obtain the video title of the target video, the second sub-title being a sub-title of the target video generated based on the other of the text features and the picture features; the adjustment manner of the first sub-title comprises at least one of the following: extracting a first keyword of the first sub-title and a second keyword in the second sub-title, the first keyword and the second keyword having a same first semantics, and the video title of the target video comprising the first semantics; extracting a third keyword of the first sub-title and a fourth keyword in the second sub-title, the third keyword and the fourth keyword having different second semantics and third semantics, and the video title of the target video comprising the second semantics and the third semantics.

2. The method of claim 1, wherein, The performing of the feature extraction on the video data to obtain the text features and the picture features comprises: inputting the video data into a pre-trained first network model to obtain the text features and the picture features; wherein the first network model is used to extract dialogue text of at least one first target position and video pictures of at least one second target position in the video data respectively, and perform feature extraction on the extracted dialogue text of each first target position to obtain the text features and perform feature extraction on the extracted video pictures of each second target position to obtain the picture features.

3. The method of claim 2, wherein, Before the inputting of the video data into the pre-trained first network model, the method further comprises: inputting video sample data into a to-be-trained network model for N times of iterative training, the video sample data comprising sample dialogue text and sample video pictures, N being a positive integer; in a case where the to-be-trained network model after the N times of iterative training meets a preset condition, determining the to-be-trained network model after the N times of iterative training as the first network model; wherein the preset condition comprises that a matching degree between target text features and preset text features is greater than or equal to a first matching value, and a matching degree between target picture features and preset picture features is greater than or equal to a second matching value, the target text features being features generated by the to-be-trained network model based on the sample dialogue text, the preset text features being preset features in the sample dialogue text, the target picture features being features generated by the to-be-trained network model based on the sample video pictures, and the preset picture features being preset features in the sample video pictures.

4. The method of claim 3, wherein, In a process of one iteration training of the N times of iteration training, the network model to be trained is used to extract sample script text of a first target position and sample video picture of a second target position in the video sample data respectively, and perform feature extraction on the extracted sample script text of the first target position to obtain the target text feature, and perform feature extraction on the extracted sample video picture of the second target position to obtain the target picture feature; The preset text feature is a preset feature of sample script text of a first preset position in the video sample data, and the preset picture feature is a preset feature of sample video picture of a second preset position in the video sample data. The preset condition further includes that a deviation value between the first target position and the first preset position is less than or equal to a first deviation value, and a deviation value between the second target position and the second preset position is less than or equal to a second deviation value.

5. The method of claim 1, wherein, The adjusting the first sub-title based on the second sub-title to obtain the video title of the target video includes: inputting the first sub-title and the second sub-title into a second network model pre-trained to obtain the video title of the target video; The second network model is used to align a time axis sequence corresponding to the second sub-title in the video data with a time axis sequence corresponding to the first sub-title in the video data, and then adjust the first sub-title based on the second sub-title to generate the video title.

6. A title generation apparatus characterized by comprising: It includes: The acquisition module is used to acquire video data of a target video, and the video data includes script text and video pictures. The feature extraction module is used to perform feature extraction on the video data to obtain text features and picture features, wherein the text features are used to represent semantics of the script text, and the picture features are used to represent content of the video pictures. The generation module is used to generate a video title of the target video based on the text features and the picture features. The generation module includes: The generation sub-module is used to generate a first sub-title of the target video based on one of the text features and the picture features. The adjustment sub-module is used to adjust the first sub-title based on a second sub-title to obtain the video title of the target video, wherein the second sub-title is a sub-title of the target video generated based on the other of the text features and the picture features. The adjustment manner of the first sub-title includes at least one of the following: extracting a first keyword of the first sub-title and a second keyword in the second sub-title, the first keyword and the second keyword having a same first semantics, and the video title of the target video including the first semantics; extracting a third keyword of the first sub-title and a fourth keyword in the second sub-title, the third keyword and the fourth keyword having different second semantics and third semantics, and the video title of the target video including the second semantics and the third semantics.

7. An electronic device, comprising: The application relates to a computer device, comprising a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus. The memory is used for storing a program. The processor is used for executing the program stored in the memory to realize the method in any one of claims 1-5.

8. A readable storage medium, having stored thereon a program, characterized in that, The program is executed by the processor to realize the method in any one of claims 1-5.

Citation Information

Patent Citations

  • Video topic processing method and device, electronic equipment and storage medium

    CN110489593A

  • Title generation method, model training method and device

    CN114611498A