Method and device for determining video title, storage medium, and electronic device

By combining automated popularity filtering and semantic classification with video relevance calculation, the problem of low efficiency in video title determination is solved, and efficient and accurate video title generation is achieved.

CN116975263BActive Publication Date: 2025-10-28TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211289074.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-20
Publication Date
2025-10-28
Estimated Expiration
2042-10-20

AI Technical Summary

Technical Problem

The current technology for determining video titles is inefficient, mainly relying on manual input, which leads to low efficiency.

Method used

By acquiring a set of target videos and candidate texts, and utilizing the popularity and semantic information of the candidate texts, combined with relevance parameters, the video title is automatically determined. This includes filtering candidate texts by popularity, classifying them semantically, and calculating their relevance to the video. Texts that meet the criteria are then selected as the video title.

Benefits of technology

It enables the efficient production of video titles that match the target video, improving the efficiency and accuracy of video title determination and optimizing video production efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116975263B_ABST
    Figure CN116975263B_ABST
Patent Text Reader

Abstract

This application discloses a method, apparatus, storage medium, and electronic device for determining video titles. The method includes: acquiring a target video and a set of candidate texts, wherein the candidate text set includes texts associated with the target video and allowing interaction; determining a first text set based on the popularity information of each candidate text in the candidate text set; performing a classification operation on the first text set to obtain a second text set; calculating a relevance parameter between each text in the second text set and the target video; and determining the text whose relevance parameter satisfies a preset relevance condition as the target text, wherein the target text is text that can be configured as the video title of the target video. This application can be applied to scenarios including but not limited to generating video titles based on artificial intelligence, and it solves the technical problem of low efficiency in determining video titles in related technologies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computers, and more specifically, to a method and apparatus for determining video titles, a storage medium, and an electronic device. Background Technology

[0002] In the process of video footage production, videos are often divided into segments or kept as a whole for content creators to use, based on scenes and plots (such as the intensity of a fight). A good video clip often needs a suitable and accurate title to reflect the content. Currently, the main method for generating titles is still manual input. When there are many video clips, manual title generation leads to low efficiency.

[0003] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention

[0004] This application provides a method and apparatus for determining video titles, a storage medium, and an electronic device, to at least solve the technical problem of low efficiency in determining video titles in related technologies.

[0005] According to one aspect of the embodiments of this application, a method for determining a video title is provided, comprising: acquiring a target video and a candidate text set, wherein the candidate text set includes text associated with the target video and allowing interaction; determining a first text set based on the popularity information of each candidate text in the candidate text set, wherein the first text set includes text in the candidate text set whose popularity information satisfies a first preset popularity condition, the popularity information including interaction parameters of the corresponding candidate text; performing a classification operation on the first text set to obtain a second text set, wherein the second text set includes text whose semantic information satisfies a preset content condition; calculating a relevance parameter between each text in the second text set and the target video, and determining the text whose relevance parameter satisfies a preset relevance condition as the target text, wherein the target text is text that can be configured as the video title of the target video.

[0006] According to another aspect of the embodiments of this application, a video title determination apparatus is also provided, comprising: an acquisition module, configured to acquire a target video and a candidate text set, wherein the candidate text set includes text associated with the target video and allowing interaction; a determination module, configured to determine a first text set based on the popularity information of each candidate text in the candidate text set, wherein the first text set includes text in the candidate text set whose popularity information satisfies a first preset popularity condition, the popularity information including interaction parameters of the corresponding candidate text; a classification module, configured to perform a classification operation on the first text set to obtain a second text set, wherein the second text set includes text whose semantic information satisfies a preset content condition; and a processing module, configured to calculate a relevance parameter between each text in the second text set and the target video, and determine the text whose relevance parameter satisfies a preset relevance condition as target text, wherein the target text is text that can be configured as a video title of the target video.

[0007] Optionally, the device is configured to perform a classification operation on the first text set to obtain a second text set by: performing a first encoding operation on each text in the first text set to obtain a first text representation vector set, wherein each text representation vector in the first text representation vector set includes semantic features representing the semantic information of the corresponding text; performing a classification operation on the first text representation vector set to obtain a first text representation vector subset, wherein the representation vectors in the first text representation vector subset include semantic features representing semantic information that satisfies the preset content conditions; and determining the text corresponding to the first text representation vector subset as the second text set.

[0008] Optionally, the apparatus is further configured to: perform a second encoding operation on the popularity information of each text in the first text set to obtain a popularity representation vector set, wherein each popularity representation vector in the popularity representation vector set includes a popularity feature representing the popularity information; concatenate the first text representation vector set with the text representation vectors and popularity representation vectors corresponding to the same text in the popularity representation vector set to obtain a second text representation vector set; perform the classification operation on the second text representation vector set to obtain a second text representation vector subset, wherein the representation vectors in the second text representation vector subset include semantic features representing semantic information satisfying the preset content conditions and popularity features representing popularity information satisfying the second preset popularity conditions; and determine the text corresponding to the second text representation vector subset as the second text set.

[0009] Optionally, the device is configured to perform encoding operations on the popularity information of each text in the first text set respectively to obtain a popularity representation vector set by: mapping the popularity information corresponding to each text in the first text set to a popularity level set according to a preset popularity level, wherein different values ​​of the popularity level correspond to different value ranges of the popularity information; performing the second encoding operation on each popularity level in the popularity level set respectively to obtain the popularity representation vector set, wherein the popularity feature is used to represent the popularity level of the corresponding text.

[0010] Optionally, the apparatus is configured to calculate a relevance parameter between each text in the second text set and the target video, and determine the text whose relevance parameter satisfies a preset relevance condition as the target text, comprising: performing a third encoding operation on the target video to obtain a video representation vector; calculating a similarity between each text representation vector in the first text representation vector subset and the video representation vector to obtain a similarity set, wherein the relevance parameter includes the similarity; and determining the text in the similarity set that corresponds to the similarity satisfying the preset similarity condition as the target text, wherein the preset relevance condition includes the preset similarity condition.

[0011] Optionally, the apparatus is configured to perform a third encoding operation on the target video to obtain a video representation vector by performing a frame extraction operation on the target video to obtain a set of target image sequences; and performing the third encoding operation on the set of target image sequences to obtain the video representation vector.

[0012] Optionally, the device is used to determine the text corresponding to the similarity that meets the preset similarity conditions in the similarity set as the target text in the following manner: sorting the similarity set in descending order, and determining the text corresponding to the similarity that ranks in the top N positions as the target text, where N is a positive integer; or, determining the text corresponding to the similarity that exceeds the preset threshold in the similarity set as the target text.

[0013] Optionally, the device is configured to perform a classification operation on the first text representation vector set to obtain a first text representation vector subset by performing a classification operation on each text representation vector in the first text representation vector set in the following manner: each text representation vector that performs the classification operation is regarded as a target text representation vector, and the target text representation vector is used to represent the representation text in the first text set: determining whether the format of the representation text meets a preset format condition based on the target text representation vector; determining whether the semantic information of the representation text is used to describe the plot based on the target text representation vector; determining whether the representation text is fluent based on the target text representation vector; and adding the target text representation vector to the first text representation vector subset when the format of the representation text meets the preset format condition, the semantic information of the representation text is used to describe the plot, and the representation text is fluent.

[0014] Optionally, the device is configured to determine a first text set based on the popularity information of each candidate text in the candidate text set through at least one of the following methods: obtaining the number of likes for each candidate text, wherein the interaction parameter includes the number of likes for the candidate text; determining candidate texts whose number of likes exceeds a preset threshold as texts in the first text set; obtaining the number of citations for each candidate text, wherein the interaction parameter includes the number of citations for the candidate text; determining candidate texts whose number of citations exceeds a preset threshold as texts in the first text set; obtaining the number of shares for each candidate text, wherein the interaction parameter includes the number of shares for the candidate text; determining candidate texts whose number of shares exceeds a preset threshold as texts in the first text set; obtaining the number of rewards for each candidate text, wherein the interaction parameter includes the number of rewards for the candidate text; determining candidate texts whose number of rewards exceeds a preset threshold as texts in the first text set.

[0015] Optionally, the device is further configured to: filter the candidate text set to obtain a third text set, wherein each text in the third text set satisfies at least one of the following conditions: the text content of each text in the third text set is different, the text length of each text in the third text set is within a preset length range, and each text in the third text set does not contain sensitive words; and determine a first text set based on the popularity information of each text in the third text set.

[0016] Optionally, the apparatus is configured to acquire a target video and a set of candidate texts by at least one of the following methods: acquiring the target video and a set of bullet screen texts generated during the playback of the target video, wherein the set of candidate texts includes the set of bullet screen texts; acquiring the target video and a set of comment texts associated with the target video, wherein the set of candidate texts includes the set of comment texts; acquiring the target video and a set of subtitle texts associated with the target video, wherein the set of candidate texts includes the set of subtitle texts.

[0017] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, wherein a computer program is stored in the computer program, and the computer program is configured to execute the above-described method for determining the video title when it is run.

[0018] According to another aspect of the embodiments of this application, a computer program product or computer program is provided, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the method for determining the video title as described above.

[0019] According to another aspect of the embodiments of this application, an electronic device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the above-described method for determining video titles through the computer program.

[0020] In this embodiment, a target video and a candidate text set are obtained. The candidate text set includes texts associated with the target video and that are interactive. A first text set is determined based on the popularity information of each candidate text in the candidate text set. The first text set includes texts whose popularity information meets a first preset popularity condition. The popularity information includes the interaction parameters of the corresponding candidate texts. A classification operation is performed on the first text set to obtain a second text set. The second text set includes texts whose semantic information meets a preset content condition. A relevance parameter is calculated between each text in the second text set and the target video, and the relevance is... Text that meets preset relevance conditions is identified as target text. Target text is the text that can be configured as the video title of the target video. By combining text with its relevance popularity, text selection is performed. Simultaneously, the text is compared with the video to calculate the relevance between the text and the video. Then, the text is sorted according to the relevance and the text with the required matching degree is used as the video title. This achieves the goal of efficiently producing video titles that match the target video, thereby improving the efficiency and accuracy of video title determination, optimizing video production efficiency, and solving the technical problem of low efficiency in video title determination in related technologies. Attached Figure Description

[0021] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0022] Figure 1 This is a schematic diagram of an application environment for an optional method for determining video titles according to an embodiment of this application;

[0023] Figure 2 This is a flowchart illustrating an optional method for determining a video title according to an embodiment of this application;

[0024] Figure 3 This is a schematic diagram of an optional method for determining a video title according to an embodiment of this application;

[0025] Figure 4 This is a schematic diagram of another optional method for determining a video title according to an embodiment of this application;

[0026] Figure 5 This is a schematic diagram of another optional method for determining a video title according to an embodiment of this application;

[0027] Figure 6 This is a schematic diagram of another optional method for determining a video title according to an embodiment of this application;

[0028] Figure 7 This is a schematic diagram of another optional method for determining a video title according to an embodiment of this application;

[0029] Figure 8 This is a schematic diagram of another optional method for determining a video title according to an embodiment of this application;

[0030] Figure 9 This is a schematic diagram of another optional method for determining a video title according to an embodiment of this application;

[0031] Figure 10 This is a schematic diagram of another optional method for determining a video title according to an embodiment of this application;

[0032] Figure 11 This is a schematic diagram of the structure of an optional video title determination device according to an embodiment of this application;

[0033] Figure 12 This is a structural schematic diagram of a product for determining an optional video title according to an embodiment of this application;

[0034] Figure 13 This is a schematic diagram of the structure of an optional electronic device according to an embodiment of this application. Detailed Implementation

[0035] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0036] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0037] First, some nouns or terms that appear in the description of the embodiments of this application shall be interpreted as follows:

[0038] Deep learning model frameworks: Commonly used deep learning frameworks in academia and industry include Café, Theano, Keras, Tensorflow, PyTorch, PaddlePaddle, MXNet, etc. These deep learning frameworks have been applied to fields such as computer vision, speech recognition, natural language processing, and bioinformatics, and have achieved excellent results.

[0039] Attention: The attention mechanism is widely used in deep learning, allowing the model to assign different attention weights to different parts of the data.

[0040] Transformer: A sequence model based on the attention mechanism, consisting of two parts: an encoder and a decoder.

[0041] Embedding: In the field of natural language processing, word embedding refers to mapping words in text into a high-dimensional representation vector, which is convenient for input into deep learning models for processing.

[0042] The present application will be described below with reference to embodiments:

[0043] According to one aspect of the embodiments of this application, a method for determining a video title is provided. Optionally, in this embodiment, the above-described method for determining a video title can be applied to, for example, Figure 1 The hardware environment shown consists of server 101 and terminal device 103. For example... Figure 1As shown, server 101 connects to terminal 103 via a network and can be used to provide services for terminal devices or applications installed on terminal devices. Applications can be video applications, instant messaging applications, browser applications, educational applications, game applications, etc. Database 105 can be set up on the server or independently of the server to provide data storage services for server 101, such as a video data storage server. The network can include, but is not limited to, wired networks and wireless networks. The wired network includes local area networks (LANs), metropolitan area networks (MANs), and wide area networks (WANs). The wireless network includes Bluetooth, Wi-Fi, and other networks that enable wireless communication. Terminal device 103 can be a terminal configured with applications and can include, but is not limited to, at least one of the following: mobile phones (such as Android phones, iOS phones, etc.), laptops, tablets, PDAs, MIDs (Mobile Internet Devices), PADs, desktop computers, smart TVs, smart voice interaction devices, smart home appliances, vehicle terminals, aircraft, and other computer devices. The server can be a single server, a server cluster consisting of multiple servers, or a cloud server.

[0044] Combination Figure 1 As shown, the method for determining the video title described above can be implemented on terminal device 103 through the following steps:

[0045] S1, the target video and a candidate text set are obtained on the terminal device 103, wherein the candidate text set includes texts that are associated with the target video and are allowed to be interacted with;

[0046] S2, on the terminal device 103, a first text set is determined based on the popularity information of each candidate text in the candidate text set. The first text set includes texts in the candidate text set whose popularity information meets the first preset popularity condition. The popularity information includes the interaction parameters of the corresponding candidate text.

[0047] S3, perform a classification operation on the first text set on the terminal device 103 to obtain a second text set, wherein the second text set includes texts whose semantic information satisfies preset content conditions;

[0048] S4, on the terminal device 103, the correlation parameters of each text in the second text set are calculated with the target video, and the text whose correlation parameters meet the preset correlation conditions are determined as the target text, wherein the target text is the text that can be configured as the video title of the target video.

[0049] Optionally, in this embodiment, the method for determining the video title described above can also be implemented via a server, for example, Figure 1It is implemented in server 101 shown; or it is implemented jointly by the terminal device and the server.

[0050] The above is merely an example, and this embodiment does not impose any specific limitations.

[0051] Alternatively, as an alternative implementation method, such as Figure 2 As shown, the methods for determining the video title include:

[0052] S202, Obtain the target video and a set of candidate texts, wherein the set of candidate texts includes texts that are associated with the target video and are allowed to be interacted with;

[0053] In an exemplary embodiment, the above-described method for determining video titles can be applied to application scenarios such as the production and recommendation of video materials, providing functions such as video title generation and video title recommendation, greatly increasing processing efficiency and reducing labor costs.

[0054] Optionally, in this embodiment, the target video may include, but is not limited to, complete video footage, or one or more video segments obtained by dividing a video. The candidate text set may include, but is not limited to, bullet screen text, comment text, and subtitle text associated with the video footage. The interactive text may include, but is not limited to, text that allows quoting, text that allows liking, and text that allows copying.

[0055] For example, the aforementioned candidate text set may include, but is not limited to, bullet screen text that allows liking, subtitle text that allows copying, and comment text that allows quoting.

[0056] For example, Figure 3 This is a schematic diagram of an optional method for determining a video title according to an embodiment of this application, such as... Figure 3 As shown, during the playback of the target video, users can input relevant bullet screen text based on the content of the video. Other users watching the video can also see the bullet screen text and perform relevant interactive operations on the bullet screen text, such as liking, tipping, and following. Each interactive operation performed on the bullet screen text is counted as popularity information, and finally all popularity information related to the bullet screen text is counted for subsequent processing.

[0057] S204, determine a first text set based on the popularity information of each candidate text in the candidate text set, wherein the first text set includes texts in the candidate text set whose popularity information satisfies a first preset popularity condition, and the popularity information includes the interaction parameters of the corresponding candidate text.

[0058] In an exemplary embodiment, determining the first text set based on the popularity information of each candidate text in the candidate text set can be implemented using, but is not limited to, a rule-based classification model. The popularity information reflects the popularity of the candidate text and the user's approval of it. The interaction parameters can include, but are not limited to, the number of likes and rewards, and can include, but are not limited to, a type of UGC (User Generated Content) interaction attribute value. Generally speaking, taking a bullet screen text as an example where the candidate text is a bullet screen comment and the interaction parameter is the number of likes, the higher the number of likes, the better it reflects the matching degree between the bullet screen text and the target video, as well as the sophistication of the bullet screen text. Setting a popularity threshold as the first preset popularity condition can be included, but is not limited to, filtering out some candidate texts with lower popularity, allowing the first text set that meets the first preset popularity condition to enter the subsequent processing flow.

[0059] It should be noted that there may be multiple candidate texts with the same content. This is determined by both user actions and the attributes of the video playback publisher. Before completing the popularity filtering, the popularity of candidate texts with the same content can be aggregated to ensure that the content of each candidate text is different.

[0060] Additionally, further filtering of candidate texts can be performed, including but not limited to: removing candidate texts that do not meet the word count requirement, removing emoji characters, removing candidate texts containing sensitive words, and removing offensive candidate texts.

[0061] For example, the aforementioned popularity information may include, but is not limited to, popularity information generated from the number of times users liked the candidate text, popularity information generated from the number of times users forwarded the candidate text, popularity information generated from the number of times users copied the candidate text, etc.

[0062] For example, Figure 4 This is a schematic diagram of another optional method for determining a video title according to an embodiment of this application, such as... Figure 4As shown, taking the candidate text as a bullet screen text as an example, during the playback of the target video, each time the user performs a related interactive operation on the bullet screen text, such as liking, tipping, or following, it is counted as an interaction parameter in the popularity information. Different interaction parameters can be counted in the popularity information in different proportions. When bullet screen text A is liked 5 times and forwarded 0 times, bullet screen text B is liked 20 times and forwarded 2 times, and bullet screen text C is liked 30 times and forwarded 1 time, and each like is equivalent to a popularity value of 1 and each forward is equivalent to a popularity value of 5, then the popularity information of the above bullet screen text A is 5, the popularity information of bullet screen text B is 30, and the popularity information of bullet screen text C is 35. The above first preset popularity condition is set to a popularity value greater than or equal to 15. At this time, bullet screen text A cannot be used as text in the first text set, while bullet screen text B and bullet screen text C can be used as text in the first text set.

[0063] S206, Perform a classification operation on the first text set to obtain a second text set, wherein the second text set includes texts whose semantic information satisfies preset content conditions;

[0064] In an exemplary embodiment, the classification operation performed on the first text set may include, but is not limited to, inputting each text in the first text set into a pre-trained neural network model for classification, and determining a second text set composed of texts in the first text set whose semantic information satisfies preset content conditions. The aforementioned semantic information may include, but is not limited to, the semantic content, text format, and text fluency of the texts in the first text set. Satisfying the preset content conditions may include, but is not limited to, the semantic content of the texts in the first text set being content used to describe the video's plot, the text format of the texts in the first text set conforming to preset format requirements, and the text fluency of the texts in the first text set being sufficiently fluent.

[0065] Optionally, in this embodiment, the pre-trained neural network model may include, but is not limited to, neural network models composed of network architectures such as Transformer, RNN (Recurrent Neural Network), and CNN (Convolutional Neural Network).

[0066] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.

[0067] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0068] Computer vision (CV) is the science that studies how to enable machines to "see." More specifically, it refers to machine vision, which uses cameras and computers to replace human eyes in recognizing and measuring targets, and then performs image processing to create images more suitable for human observation or transmission to instruments. As a scientific discipline, computer vision studies related theories and technologies, attempting to build artificial intelligence systems capable of extracting information from images or multidimensional data. Computer vision technologies typically include image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping (SLAM), and common biometric recognition technologies such as facial recognition and fingerprint recognition.

[0069] Natural Language Processing (NLP) is an important field within computer science and artificial intelligence. It studies the theories and methods for enabling effective communication between humans and computers using natural language. NLP is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field involves natural language—the language people use in daily life—and thus it has a close relationship with linguistic research. NLP techniques typically include text processing, semantic understanding, machine translation, question answering, and knowledge graphs. For example, using a Transformer neural network, each text in a first text set is input into the Transformer to obtain a first text representation vector set. This first text representation vector set is then input into a linear classifier to perform a classification operation, resulting in a subset of first text representation vectors. The texts corresponding to this subset are then identified as the second text set.

[0070] For example, Figure 5 This is a schematic diagram of another optional method for determining a video title according to an embodiment of this application. The original Transformer structure is as follows: Figure 8 As shown, Figure 8The left side is the Transformer Encoder, which is used to input the word embedding representation vectors of each text in the first text set. After passing through the Encoder, the output is a text representation vector containing the semantic features of the text. Then, a normalization operation is performed to obtain the probability set corresponding to each text representation vector in the subset of the first text representation vectors. Each probability in the above probability set is input into a linear classifier to classify the first text set and determine the second text set.

[0071] S208, calculate the correlation parameters between each text in the second text set and the target video, and determine the text whose correlation parameters meet the preset correlation conditions as the target text, wherein the target text is the text that can be configured as the video title of the target video.

[0072] In an exemplary embodiment, the above-mentioned calculation of the correlation parameters between each text in the second text set and the target video may include, but is not limited to, extracting features from each text to obtain a text representation vector, extracting features from the target video to obtain a video representation vector, and then calculating the similarity between each text representation vector and the video representation vector to determine the above-mentioned correlation parameters. The above-mentioned determination of the text whose correlation parameters satisfy the preset correlation conditions as the target text can be understood as sorting each text in the second text set according to the value of the correlation parameters, and determining the text corresponding to the top N positions of the sort as the above-mentioned target text, wherein the higher the sorting, the more similar the text representation vector and the video representation vector of the text are.

[0073] Optionally, in this embodiment, the aforementioned correlation parameters may include, but are not limited to, similarity, which may be determined based on, but is not limited to, cosine distance, Euclidean distance, etc.

[0074] For example, taking the relevance score as the relevance parameter, the text in the second text set is input into the video text relevance module, and the relevance score between the text and the video is output. The text is sorted according to the relevance score, and the top-N texts are taken as the target text (if top-N is 1, it means that the text with the highest relevance score in the second text set is taken).

[0075] It should be noted that the above-mentioned video text relevance module may include, but is not limited to, two parts: a Text Encoder and a Video Encoder. Both the Text Encoder and the Video Encoder may use neural network structures including, but not limited to, the Transformer Encoder.

[0076] For example, Figure 6 This is a schematic diagram of another optional method for determining a video title according to an embodiment of this application, such as... Figure 6As shown, taking the candidate text as a bullet screen text as an example, when the relevance score of bullet screen text A is 5 points, the relevance score of bullet screen text B is 20 points, and the relevance score of bullet screen text C is 30 points, and the selection of Top-N is set to 1, then bullet screen text C is the target text mentioned above, which can be used as the video title of the target video mentioned above.

[0077] In this embodiment, a target video and a candidate text set are obtained. The candidate text set includes texts associated with the target video and that are interactive. A first text set is determined based on the popularity information of each candidate text in the candidate text set. The first text set includes texts whose popularity information meets a first preset popularity condition. The popularity information includes the interaction parameters of the corresponding candidate texts. A classification operation is performed on the first text set to obtain a second text set. The second text set includes texts whose semantic information meets a preset content condition. A relevance parameter is calculated between each text in the second text set and the target video. Texts that meet preset relevance conditions are identified as target texts. Target texts are texts that can be configured as video titles for target videos. By combining text with its relevance popularity, text selection is performed. Simultaneously, the text is compared with the video to calculate the relevance between the text and the video. Then, the texts are sorted according to their relevance to obtain the texts that meet the required matching degree as video titles. This achieves the goal of efficiently producing video titles that match the target video, thereby improving the efficiency and accuracy of video title determination, optimizing video production efficiency, and solving the technical problem of low efficiency in video title determination in related technologies.

[0078] As an alternative approach, a classification operation is performed on the first text set to obtain a second text set, including:

[0079] Perform a first encoding operation on each text in the first text set to obtain a first text representation vector set, wherein each text representation vector in the first text representation vector set includes semantic features that represent the semantic information of the corresponding text;

[0080] A classification operation is performed on the first set of text representation vectors to obtain a subset of the first set of text representation vectors, wherein the representation vectors in the subset of the first set of text representation vectors include semantic features that represent semantic information that satisfy preset content conditions.

[0081] The text corresponding to the first text representation vector subset is determined as the second text set.

[0082] In an exemplary embodiment, the first encoding operation described above may include, but is not limited to, using various types of feature extraction models to extract features from each text in the first text set to obtain the first text representation vector. Each text representation vector includes semantic features representing the semantic information of the corresponding text; this can be understood as the values ​​of a portion of the dimensions of each text representation vector being used to represent the semantic features of the semantic information. The classification operation performed on the first text representation vector set may include, but is not limited to, using a linear classifier. The text representation vectors in the first text representation vector set are input, and based on the semantics represented by the text representation vectors in the first text representation vector set, the text representation vectors are classified into text representation vectors whose semantic information satisfies preset content conditions and text representation vectors whose semantic information does not satisfy the preset content conditions.

[0083] Optionally, in this embodiment, the aforementioned semantic information satisfying the preset content conditions may include, but is not limited to, the aforementioned semantic information indicating that the corresponding text is a plot-based text. A text having a plot needs to satisfy at least one or more of the following conditions:

[0084] (1) The semantic information of the text is in the format of “person (animal) + event”, for example: Zhang San tries to obtain the sword in order to save Li Si.

[0085] (2) The semantic information of the text is fluent.

[0086] (3) Textual semantic information describes the plot of the video.

[0087] Typical negative examples include:

[0088] Colloquial language, for example: Is this chest compression a joke?

[0089] Phrases with emotional tone, such as: These are people who are going to starve to death, haha!

[0090] First-person perspective, for example: I couldn't bear to let her do that for a man.

[0091] Describe yourself, for example: Your friend Wang Wu is online.

[0092] A commanding tone, such as: Break up with him.

[0093] The values ​​are incorrect, and it involves various types of sensitive words;

[0094] Contains discriminatory or derogatory terms, such as: deaf person, blind person.

[0095] As an optional approach, the above method also includes:

[0096] Perform a second encoding operation on the popularity information of each text in the first text set to obtain a popularity representation vector set, wherein each popularity representation vector in the popularity representation vector set includes popularity features that represent popularity information;

[0097] The first set of text representation vectors is concatenated with the text representation vectors and popularity representation vectors corresponding to the same text in the first set of text representation vectors to obtain the second set of text representation vectors.

[0098] A classification operation is performed on the second text representation vector set to obtain a subset of the second text representation vectors. The representation vectors in the subset of the second text representation vectors include semantic features that represent semantic information that satisfy the preset content conditions and popularity features that represent popularity information that satisfy the second preset popularity conditions.

[0099] The text corresponding to the subset of the second text representation vectors is determined as the second text set.

[0100] In an exemplary embodiment, the above-mentioned second encoding operation performed on the popularity information of each text in the first text set to obtain a popularity representation vector set can be understood as encoding the popularity information of each text in the first text set into popularity representation vectors to form the above-mentioned popularity representation vector set. The above-mentioned concatenation of the first text representation vector set with the text representation vectors and popularity representation vectors corresponding to the same text in the popularity representation vector set to obtain a second text representation vector set can be understood as concatenating the text representation vectors and popularity representation vectors corresponding to the same text into a second text representation vector to form the above-mentioned second text representation vector set. The classification operation performed on the second set of text representation vectors can include, but is not limited to, using a linear classifier. The input is a text representation vector from the second set of text representation vectors. Based on the semantics represented by the text representation vectors in the second set of text representation vectors, the text representation vectors are classified as follows: text representation vectors whose semantic information satisfies a preset content condition and whose popularity information satisfies a second preset popularity condition; text representation vectors whose semantic information satisfies the preset content condition but whose popularity information does not satisfy the second preset popularity condition; text representation vectors whose semantic information does not satisfy the preset content condition but whose popularity information satisfies the second preset popularity condition; and text representation vectors whose semantic information does not satisfy either the preset content condition or the second preset popularity condition.

[0101] Optionally, in this embodiment, the text representation vectors whose semantic information satisfies the preset content conditions and whose popularity information satisfies the second preset popularity conditions, and the text representation vectors whose semantic information satisfies the preset content conditions but whose popularity information does not satisfy the second preset popularity conditions, can be determined as the representation vectors in the above-mentioned second text representation vector subset.

[0102] For example, if text A satisfies the preset content condition and the second preset popularity condition, text B does not satisfy the preset content condition but satisfies the second preset popularity condition, text C satisfies the preset content condition but does not satisfy the second preset popularity condition, and text D does not satisfy either the preset content condition or the second preset popularity condition, then text A can be determined as the representation vector in the aforementioned second text representation vector subset, and texts B, C, and D can be filtered out.

[0103] For example, Figure 7 This is a schematic diagram of another optional method for determining a video title according to an embodiment of this application, such as... Figure 7 As shown, it includes a Transformer Encoder part, a popularity level mapping part, and a classifier normalization part. It takes a text representation vector in the first text representation vector set as input, outputs the output probability of the text representation vector, and inputs the output probability into the classifier to determine the classification result of the text corresponding to the text representation vector. The classification result includes the text corresponding to the text representation vector having semantic features that represent semantic information that meet the preset content conditions and popularity features that meet the second preset popularity conditions, or the text corresponding to the text representation vector not having semantic features that represent semantic information that meet the preset content conditions or popularity features that represent popularity information that meet the second preset popularity conditions.

[0104] Figure 8 This is a schematic diagram of another optional method for determining a video title according to an embodiment of this application, such as... Figure 8 As shown, this includes the specific structure of the Transformer Encoder, which contains a multi-head attention module, a residual and normalization module, a feedforward network module, and another residual and normalization module. The overall formula is expressed as:

[0105] This indicates the Attention calculation method, where i represents the index of the multi-head attention module;

[0106] MH(QKV) = Concat(Att1,Att2,…), which means concatenating multiple Attention mechanisms, i.e., multi-head attention mechanism;

[0107] This refers to the Add&Norm method, the residual and normalization module;

[0108] Indicates the feedforward network module;

[0109] This indicates the second Add&Norm module, which handles residuals and normalization.

[0110] Where X is the word embedding representation of the input text (corresponding to the text representation vector in the aforementioned first set of text representation vectors). In the multi-head attention mechanism, X is simultaneously calculated as Q, K, and V, and is the output of the text after passing through the Encoder. It contains the semantic features of the corresponding text.

[0111] As an optional approach, the popularity information of each text in the first text set is encoded to obtain a set of popularity representation vectors, including:

[0112] Based on the pre-set popularity levels, the popularity information corresponding to each text in the first text set is mapped to a set of popularity levels, where different values ​​of the popularity level correspond to different value ranges of the popularity information.

[0113] Perform a second encoding operation on each popularity level in the popularity level set to obtain a popularity representation vector set, where the popularity features are used to represent the popularity level of the corresponding text.

[0114] Optionally, in this embodiment, the aforementioned pre-set popularity level may include, but is not limited to, configuring corresponding popularity levels for popularity intervals represented by different popularity information. Multiple popularity levels together form the aforementioned popularity level set. The aforementioned execution of the second encoding operation on each popularity level in the popularity level set may include, but is not limited to, performing a mapping operation on the popularity information. For example, mapping the popularity information to a predefined popularity level. The popularity level may be non-linear.

[0115] For example, a popularity score of 50-150 corresponds to popularity level 1, 150-400 to level 2, 400-800 to level 3, 800-1500 to level 4, 1500-3000 to level 5, and above 3000 to level 6. The discretized popularity levels can be mapped to embedding features, which are then concatenated with the text representation vector from the Transformer Encoder for subsequent processing.

[0116] As an optional approach, each text in the second text set is compared with the target video to calculate a correlation parameter, and the text whose correlation parameters satisfy a preset correlation condition is identified as the target text, including:

[0117] Perform a third encoding operation on the target video to obtain the video representation vector;

[0118] The similarity between each text representation vector in the first text representation vector subset and the video representation vector is calculated to obtain a similarity set, wherein the relevance parameter includes similarity.

[0119] The texts corresponding to the similarity scores in the similarity set that meet the preset similarity conditions are identified as the target texts. The preset relevance conditions include the preset similarity conditions.

[0120] In an exemplary embodiment, the third encoding operation described above may include, but is not limited to, feature extraction operations performed on the video, by inputting the target video into the encoder to obtain the video representation vector. The similarity set described above includes the similarity calculated between each text representation vector in the first text representation vector subset and the video representation vector.

[0121] Optionally, in this embodiment, the third encoding operation described above can be implemented by the Video Encoder module, or it can be replaced by various network structures, including but not limited to Clip, various variants based on ViT, etc.

[0122] For example, text A corresponds to text representation vector A', text B corresponds to text representation vector B', text C corresponds to text representation vector C', and the video representation vector obtained after encoding the target video is V. Then, the similarity between A' and V, A' and V, and A' and V are calculated respectively to obtain similarity S1, S2, and S3. Then, the target text is determined according to the preset relevance conditions.

[0123] As an optional approach, a third encoding operation is performed on the target video to obtain a video representation vector, including:

[0124] Perform frame extraction on the target video to obtain a set of target image sequences;

[0125] A third encoding operation is performed on the target image sequence set to obtain the video representation vector.

[0126] For example, Figure 9 This is a schematic diagram of another optional method for determining a video title according to an embodiment of this application, such as... Figure 9 As shown, firstly, the video is frame-by-frame (e.g., one frame per second) to convert it into a set of target image sequences. Then, these sequences are fed into a Video Encoder to obtain the video representation vector. The Video Encoder can be of various types, such as ViT, which employs a structure similar to a Text Encoder, including multi-head attention, a Norm layer, a perceptron layer, and a final Norm layer.

[0127] As an optional approach, the text corresponding to the similarity scores in the similarity set that meet preset similarity conditions is identified as the target text, including:

[0128] Sort the similarity set in descending order, and determine the texts corresponding to the top N similarity scores as the target texts, where N is a positive integer; or,

[0129] The texts with similarity scores exceeding a preset threshold in the similarity set are identified as target texts.

[0130] In one exemplary embodiment, the method may include, but is not limited to, using cosine similarity.<f(text),g(video)> This represents the similarity between the video representation vector and the text representation vector. The final training loss function has the following symmetric form:

[0131]

[0132] Where B represents the total number of samples in the batch, i and j represent the sample numbers, and τ represents the hyperparameters that can be adjusted in the neural network model for determining similarity. During prediction, the similarity sim(t) between the text representation vector and the video representation vector is used. i ,v i The similarity is sorted, and the text corresponding to the top N similarity values ​​is selected as the title of the target video.

[0133] It should be noted that the aforementioned preset threshold can be understood as a similarity threshold pre-set based on experience. Texts with a similarity exceeding the similarity threshold are all considered target texts for subsequent processing.

[0134] As an optional approach, a classification operation is performed on the first set of text representation vectors to obtain a subset of the first set of text representation vectors, including:

[0135] A subset of the first text representation vectors is obtained by performing a classification operation on each text representation vector in the first text representation vector set in the following manner: Each text representation vector that has undergone the classification operation is regarded as a target text representation vector, and the target text representation vector is used to represent the represented text in the first text set.

[0136] Determine whether the format of the representation text meets the preset format conditions based on the target text representation vector;

[0137] Determine whether the semantic information of the target text is used to describe the plot based on the target text representation vector;

[0138] Determine whether the target text is fluent based on its representation vector;

[0139] If the format of the text representation satisfies the preset format conditions, the semantic information of the text representation is used to describe the plot, and the text representation is fluent, then the target text representation vector is added to the first text representation vector subset.

[0140] In an exemplary embodiment, whether the format of the characterizing text meets the preset format conditions can be understood as setting a preset format based on experience, calculating the format similarity between the format of the characterizing text and the preset format conditions, and determining the text whose format similarity meets the preset value as the text that meets the preset format conditions.

[0141] In an exemplary embodiment, whether the semantic information of the characterizing text is used to describe the plot can be understood as identifying the semantic information of the characterizing text, predicting the probability that the semantic information is used to describe the plot, and determining the text whose probability meets a preset value as the text used to describe the plot.

[0142] In an exemplary embodiment, the above-mentioned characterization of the fluency of text as fluent text can be understood as pre-setting the text arrangement order of fluent text based on experience, calculating the text order similarity between the character order of the characterization text and the pre-set text order of fluent text, and determining the text whose text order similarity meets the preset value as fluent text.

[0143] As an optional approach, a first text set is determined based on the popularity information of each candidate text in the candidate text set, including at least one of the following:

[0144] Get the number of likes for each candidate text, where the interaction parameter includes the number of likes for the candidate text; determine the candidate texts whose number of likes exceeds a preset threshold as texts in the first text set;

[0145] The number of times each candidate text is cited is obtained, where the interaction parameter includes the number of times the candidate text is cited; candidate texts whose number of citations exceeds a preset threshold are identified as texts in the first text set;

[0146] Obtain the number of times each candidate text has been forwarded, where the interaction parameter includes the number of times the candidate text has been forwarded; and determine the candidate texts whose number of forwards exceeds a preset threshold as texts in the first text set;

[0147] Get the number of times each candidate text has been tipped, where the interaction parameter includes the number of times the candidate text has been tipped; and determine the candidate text whose number of tips exceeds a preset threshold as the text in the first text set.

[0148] Optionally, in this embodiment, each of the above candidate texts may include, but is not limited to, texts that allow accounts watching the video to perform like, quote, forward, or tipping operations. Different operations can be counted together or separately, and different types of interaction parameters can be determined according to different actual needs to generate the video title most recognized by accounts watching the target video.

[0149] As an optional approach, the above method also includes:

[0150] The candidate text set is filtered to obtain a third text set, wherein each text in the third text set satisfies at least one of the following conditions: the text content of each text in the third text set is different; the text length of each text in the third text set is within a preset length range; and each text in the third text set does not contain sensitive words.

[0151] The first text set is determined based on the popularity information of each text in the third text set.

[0152] Optionally, in this embodiment, the third text set is a text set obtained by filtering the candidate text set. The filtering operation can be implemented by one or more dimensions such as text length, whether the text contains sensitive words, and whether the text is a duplicate text. After filtering, the first text set is determined based on the popularity information of each text in the third text set. This can reduce resource consumption and avoid performing related calculations based on popularity information on texts that obviously do not meet the conditions.

[0153] As an optional approach, the target video and candidate text set are obtained, including at least one of the following:

[0154] Obtain the target video and the set of bullet screen texts generated during the playback of the target video, wherein the candidate text set includes the bullet screen text set;

[0155] Obtain the target video and the set of comment texts associated with the target video, wherein the candidate text set includes the set of comment texts;

[0156] Obtain the target video and the set of subtitle texts associated with the target video, wherein the candidate text set includes the subtitle text set.

[0157] Optionally, in this embodiment, the candidate text set may include, but is not limited to, a combination of one or more of the above-mentioned bullet screen text set, comment text set, and subtitle text set. Interactive operations such as liking can be performed on the bullet screen text set, interactive operations such as forwarding and quoting can be performed on the comment text set, and interactive operations such as quoting and searching can be performed on the subtitle text set.

[0158] The following specific examples will further explain this application:

[0159] In the process of video footage production, videos can be segmented into clips based on scenes and plots (such as the intensity of a fight) for further use by content creators. Each segmented video clip often needs a suitable title. This application can utilize video comments to generate appropriate video clip titles.

[0160] This application proposes a method for generating titles based on the selection of bullet comments according to popularity and plot. Figure 10 This is a schematic diagram of another optional method for determining a video title according to an embodiment of this application, such as... Figure 10 As shown, the bullet screen data is filtered by popularity, and then sorted by popularity, plot relevance, and other modules to finally select suitable bullet screen data to generate video material titles. This method can be applied to material production, recommendation, and other businesses, providing functions such as video title candidate generation and video title recommendation, greatly increasing processing efficiency and reducing labor costs. Simple text classification cannot solve the bullet screen selection problem. The solution in this application integrates bullet screen popularity and plot features to select bullet screen candidates, and further scores and sorts them through the video text relevance module to finally generate titles suitable for video materials.

[0161] First, the heat filtering module will be explained as follows:

[0162] The number of likes on video bullet comments reflects their popularity and user approval, serving as a user-generated content (UGC) interactive attribute. Generally, a higher number of likes indicates a better match between the comment and the storyline, as well as its sophistication. By setting a threshold, some less popular bullet comments can be filtered out, allowing those meeting the popularity criteria to proceed to the subsequent filtering and sorting stages.

[0163] It is important to note that there may be multiple bullet comments with the same content. This is determined by both user actions and the attributes of the video player publisher. Before filtering by popularity, it is necessary to aggregate the popularity of bullet comments with the same content to ensure that the content of each candidate bullet comment is different.

[0164] Meanwhile, the popularity filtering module also undertakes the initial filtering function of bullet comments, specifically including: removing bullet comments that do not meet the character count requirements, removing emoji symbols, removing bullet comments containing sensitive words, removing offensive bullet comments, etc., which can be implemented using a rule + classification model.

[0165] Next, the following explanation is given regarding the popularity and plot modules:

[0166] A bullet comment must have a narrative element and meet the following conditions:

[0167] (1) The format of “person (animal) + event”, for example: Zhang San tries to obtain the sword in order to save Li Si.

[0168] (2) The sentences need to be fluent.

[0169] (3) The statement is telling the story.

[0170] Typical negative examples include:

[0171] Colloquial language, for example: Is this chest compression a joke?

[0172] Phrases with emotional tone, such as: These are people who are going to starve to death, haha!

[0173] First-person perspective, for example: I couldn't bear to let her do that for a man.

[0174] Describe yourself, for example: Your friend Wang Wu is online.

[0175] A commanding tone, such as: Break up with him.

[0176] Commenting on a character, for example: After acting in this play, actress A has been playing the role of a mother ever since;

[0177] Actor B portrayed the animal's true nature well; Actor C should be able to convey a sense of excitement; Actor D is far inferior to Actor E.

[0178] Low quality, for example: incorrect value orientation, involving various types of sensitive words; discriminatory and derogatory words, such as: deaf person, blind person.

[0179] Various profanities, such as: "Crazy and rude, why did I like this kind of woman in the first place?"; "Ru Ping is a jinx, no one who is with her will have a good time."

[0180] The problem of narrative in bullet comments can often be abstracted into a text classification problem. By combining the popularity of bullet comments, the following structure is used for comprehensive training. The model is roughly divided into three modules, including the Transformer Encoder module, the popularity level mapping module, and the linear classifier module.

[0181] (1) Transformer Encoder module:

[0182] This module is taken from the Encoder part of Transformer, which contains multi-head attention, residual and normalization, feedforward network, residual and normalization.

[0183] (2) Popularity Level Mapping Module:

[0184] This module maps popularity scores to predefined popularity levels, which can be non-linear. For example, a popularity score of 50-150 is level 1, 150-400 is level 2, 400-800 is level 3, 800-1500 is level 4, 1500-3000 is level 5, and above 3000 is level 6.

[0185] The discretized popularity level can be mapped to an embedding feature, which is then concatenated with the text features of the Transformer Encoder to form the bullet screen feature, which is then fed into the linear classifier module.

[0186] (3) Linear classifier module:

[0187] The module takes the hidden layer dimension as input and outputs the number of categories (2 if positive and negative are distinguished, and 3 if positive, negative and neutral are distinguished). After passing through Softmax, the final result is the analysis of all categories.

[0188] Secondly, the video and text relevance module is explained as follows:

[0189] After filtering by the two modules mentioned above, the number of remaining candidate bullet comments is further reduced. Theoretically, all remaining bullet comments are candidates that can be used as titles for the video footage. Next, the video text relevance module outputs the relevance score between the bullet comment text and the video footage, and sorts them according to the relevance score. The top-N bullet comments are taken as candidate video titles (if top-N is 1, it means that the bullet comment with the highest score among the candidate bullet comments is taken).

[0190] The video text relevance module consists of two parts: a Text Encoder and a Video Encoder. The Text Encoder is the same as the Transformer Encoder mentioned earlier. For the Video Encoder, the video is first processed by frame extraction (we use one frame per second), converting the video into an image sequence, which is then fed into the Video Encoder to obtain the video's feature representation. Several Video Encoders can be used; a well-known one is ViT, which uses a similar structure to the Text Encoder, including multi-head attention, a Norm layer, a perceptron layer, and a final Norm layer.

[0191] Finally, the relevance of the video and text matching can be measured using a similarity function; here we use cosine similarity.<f(text),g(video)> It represents the similarity between two feature vectors.

[0192] This application combines textual features with the popularity features of bullet comments (danmaku) as joint features for bullet comment selection. Simultaneously, it compares textual features with video features, calculating the relevance between text and video as a ranking feature for candidate bullet comments. This method can efficiently generate high-quality bullet comments as video clip titles, significantly improving the production efficiency of video clips and alleviating the problem of manually creating titles for large amounts of footage. While the overall architecture of the model used in this application is fixed, the individual modules have a high degree of flexibility. The Text Encoder module can be replaced by various network structures, including but not limited to BERT, XLNET, LSTM, GRU, etc., and the Video Encoder module can also be replaced by various network structures, including but not limited to Clip, various ViT-based variants, etc.

[0193] It is understood that in the specific embodiments of this application, data such as user information are involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0194] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0195] According to another aspect of the embodiments of this application, a device for determining a video title for implementing the above-described method for determining a video title is also provided. For example... Figure 11 As shown, the device includes:

[0196] The acquisition module 1102 is used to acquire a target video and a candidate text set, wherein the candidate text set includes text associated with the target video and that is interactive.

[0197] The determining module 1104 is used to determine a first text set based on the popularity information of each candidate text in the candidate text set, wherein the first text set includes texts in the candidate text set whose popularity information satisfies a first preset popularity condition, and the popularity information includes the interaction parameters of the corresponding candidate text.

[0198] The classification module 1106 is used to perform a classification operation on the first text set to obtain a second text set, wherein the second text set includes texts whose semantic information satisfies preset content conditions;

[0199] The processing module 1108 is used to calculate the correlation parameters between each text in the second text set and the target video, and to determine the text whose correlation parameters satisfy the preset correlation conditions as the target text, wherein the target text is text that can be configured as the video title of the target video.

[0200] As an optional approach, the device is used to perform a classification operation on the first text set to obtain a second text set in the following manner: performing a first encoding operation on each text in the first text set to obtain a first text representation vector set, wherein each text representation vector in the first text representation vector set includes semantic features representing the semantic information of the corresponding text; performing a classification operation on the first text representation vector set to obtain a first text representation vector subset, wherein the representation vectors in the first text representation vector subset include semantic features representing semantic information that satisfies the preset content conditions; and determining the text corresponding to the first text representation vector subset as the second text set.

[0201] As an optional solution, the device is further configured to: perform a second encoding operation on the popularity information of each text in the first text set to obtain a popularity representation vector set, wherein each popularity representation vector in the popularity representation vector set includes a popularity feature representing the popularity information; concatenate the first text representation vector set with the text representation vectors and popularity representation vectors corresponding to the same text in the popularity representation vector set to obtain a second text representation vector set; perform the classification operation on the second text representation vector set to obtain a second text representation vector subset, wherein the representation vectors in the second text representation vector subset include semantic features representing that the semantic information satisfies the preset content conditions and popularity features representing that the popularity information satisfies the second preset popularity conditions; and determine the text corresponding to the second text representation vector subset as the second text set.

[0202] As an optional solution, the device is used to perform encoding operations on the popularity information of each text in the first text set in the following manner to obtain a popularity representation vector set: mapping the popularity information corresponding to each text in the first text set to a popularity level set according to a preset popularity level, wherein different values ​​of the popularity level correspond to different value ranges of the popularity information; performing the second encoding operation on each popularity level in the popularity level set to obtain the popularity representation vector set, wherein the popularity feature is used to represent the popularity level of the corresponding text.

[0203] As an optional solution, the device is used to calculate a relevance parameter between each text in the second text set and the target video, and to determine the text whose relevance parameter satisfies a preset relevance condition as the target text, including: performing a third encoding operation on the target video to obtain a video representation vector; calculating a similarity between each text representation vector in the first text representation vector subset and the video representation vector to obtain a similarity set, wherein the relevance parameter includes the similarity; and determining the text in the similarity set that satisfies the preset similarity condition as the target text, wherein the preset relevance condition includes the preset similarity condition.

[0204] As an optional approach, the apparatus is used to perform a third encoding operation on the target video to obtain a video representation vector by performing a frame extraction operation on the target video to obtain a set of target image sequences; and performing the third encoding operation on the set of target image sequences to obtain the video representation vector.

[0205] As an optional solution, the device is used to determine the text corresponding to the similarity that meets the preset similarity conditions in the similarity set as the target text in the following manner: sorting the similarity set in descending order, and determining the text corresponding to the similarity that ranks in the top N positions as the target text, where N is a positive integer; or, determining the text corresponding to the similarity that exceeds the preset threshold in the similarity set as the target text.

[0206] As an optional approach, the device is configured to perform a classification operation on the first set of text representation vectors to obtain a first subset of text representation vectors by performing a classification operation on each text representation vector in the first set of text representation vectors in the following manner: each text representation vector that has undergone the classification operation is considered a target text representation vector, and the target text representation vector is used to represent the representation text in the first set of texts: determining whether the format of the representation text meets a preset format condition based on the target text representation vector; determining whether the semantic information of the representation text is used to describe the plot based on the target text representation vector; determining whether the representation text is fluent based on the target text representation vector; and adding the target text representation vector to the first subset of text representation vectors when the format of the representation text meets the preset format condition, the semantic information of the representation text is used to describe the plot, and the representation text is fluent.

[0207] As an optional solution, the device is configured to determine a first text set based on the popularity information of each candidate text in the candidate text set through at least one of the following methods: obtaining the number of likes for each candidate text, wherein the interaction parameter includes the number of likes for the candidate text; determining candidate texts whose number of likes exceeds a preset threshold as texts in the first text set; obtaining the number of citations for each candidate text, wherein the interaction parameter includes the number of citations for the candidate text; determining candidate texts whose number of citations exceeds a preset threshold as texts in the first text set; obtaining the number of shares for each candidate text, wherein the interaction parameter includes the number of shares for the candidate text; determining candidate texts whose number of shares exceeds a preset threshold as texts in the first text set; obtaining the number of rewards for each candidate text, wherein the interaction parameter includes the number of rewards for the candidate text; determining candidate texts whose number of rewards exceeds a preset threshold as texts in the first text set.

[0208] As an optional solution, the device is further configured to: filter the candidate text set to obtain a third text set, wherein each text in the third text set satisfies at least one of the following conditions: the text content of each text in the third text set is different, the text length of each text in the third text set is within a preset length range, and each text in the third text set does not contain sensitive words; and determine a first text set based on the popularity information of each text in the third text set.

[0209] As an optional solution, the device is configured to acquire a target video and a set of candidate texts by at least one of the following methods: acquiring the target video and a set of bullet screen texts generated during the playback of the target video, wherein the candidate text set includes the set of bullet screen texts; acquiring the target video and a set of comment texts associated with the target video, wherein the candidate text set includes the set of comment texts; acquiring the target video and a set of subtitle texts associated with the target video, wherein the candidate text set includes the set of subtitle texts.

[0210] According to one aspect of this application, a computer program product is provided, comprising a computer program / instructions containing program code for performing the methods shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via communication section 1209, and / or installed from removable medium 1211. When the computer program is executed by central processing unit 1201, it performs various functions provided in embodiments of this application.

[0211] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0212] Figure 12 A schematic block diagram of a computer system architecture for implementing an electronic device according to embodiments of the present application is shown.

[0213] It should be noted that, Figure 12 The computer system 1200 of the electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0214] like Figure 12 As shown, the computer system 1200 includes a central processing unit (CPU) 1201, which can perform various appropriate actions and processes based on programs stored in read-only memory (ROM) 1202 or programs loaded from storage section 1208 into random access memory (RAM). The RAM 1203 also stores various programs and data required for system operation. The CPU 1201, ROM 1202, and RAM 1203 are interconnected via a bus 1204. An input / output interface 1205 (I / O interface) is also connected to the bus 1204.

[0215] The following components are connected to the input / output interface 1205: an input section 1206 including a keyboard, mouse, etc.; an output section 1207 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 1208 including a hard disk, etc.; and a communication section 1209 including a network interface card such as a local area network card, modem, etc. The communication section 1209 performs communication processing via a network such as the Internet. A drive 1210 is also connected to the input / output interface 1205 as needed. A removable medium 1211, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 1210 as needed so that computer programs read from it can be installed into the storage section 1208 as needed.

[0216] Specifically, according to embodiments of this application, the processes described in the various method flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 1209, and / or installed from removable medium 1211. When the computer program is executed by central processing unit 1201, it performs various functions defined in the system of this application.

[0217] According to another aspect of the embodiments of this application, an electronic device for implementing the above-described method for determining video titles is also provided. This electronic device may be... Figure 1 The terminal device or server shown. This embodiment uses this electronic device as an example for illustration. Figure 13 As shown, the electronic device includes a memory 1302 and a processor 1304. The memory 1302 stores a computer program, and the processor 1304 is configured to execute the steps of any of the above method embodiments through the computer program.

[0218] Optionally, in this embodiment, the aforementioned electronic device may be located in at least one of a plurality of network devices in a computer network.

[0219] Optionally, in this embodiment, the processor may be configured to execute the following steps through a computer program:

[0220] S1, obtain the target video and a set of candidate texts, wherein the set of candidate texts includes texts that are associated with the target video and are allowed to be interacted with;

[0221] S2, determine the first text set based on the popularity information of each candidate text in the candidate text set, wherein the first text set includes texts in the candidate text set whose popularity information meets the first preset popularity condition, and the popularity information includes the interaction parameters of the corresponding candidate text;

[0222] S3, perform a classification operation on the first text set to obtain a second text set, wherein the second text set includes texts whose semantic information satisfies the preset content conditions;

[0223] S4, calculate the relevance parameters of each text in the second text set with the target video, and determine the text whose relevance parameters meet the preset relevance conditions as the target text, wherein the target text is the text that can be configured as the video title of the target video.

[0224] Alternatively, as those skilled in the art will understand, Figure 13The structure shown is for illustrative purposes only. Electronic devices can also be smartphones (such as Android phones, iOS phones, etc.), tablets, PDAs, mobile internet devices (MIDs), PADs, and other terminal devices. Figure 13 This does not limit the structure of the aforementioned electronic devices or electronic equipment. For example, electronic devices or electronic equipment may also include components that are more... Figure 13 The more or fewer components shown (such as network interfaces, etc.), or having the same Figure 13 The different configurations shown.

[0225] The memory 1302 can be used to store software programs and modules, such as the program instructions / modules corresponding to the video title determination method and apparatus in this embodiment. The processor 1304 executes various functional applications and data processing by running the software programs and modules stored in the memory 1302, thereby realizing the aforementioned video title determination method. The memory 1302 may include high-speed random access memory and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 1302 may further include memory remotely located relative to the processor 1304, and these remote memories can be connected to the terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof. Specifically, the memory 1302 may be used, but is not limited to, to store the aforementioned candidate text and other information. As an example, such as... Figure 13 As shown, the memory 1302 may include, but is not limited to, the acquisition module 1102, determination module 1104, classification module 1106, and processing module 1108 of the video title determination device. Furthermore, it may include, but is not limited to, other module units of the video title determination device, which will not be described in detail in this example.

[0226] Optionally, the transmission device 1306 described above is used to receive or send data via a network. Specific examples of the network described above may include wired networks and wireless networks. In one example, the transmission device 1306 includes a Network Interface Controller (NIC), which can be connected to other network devices and a router via a network cable to communicate with the Internet or a local area network. In another example, the transmission device 1306 is a Radio Frequency (RF) module, used for wireless communication with the Internet.

[0227] In addition, the aforementioned electronic device also includes: a display 1308 for displaying the target video; and a connection bus 1310 for connecting the various module components in the aforementioned electronic device.

[0228] In other embodiments, the aforementioned terminal device or server can be a node in a distributed system, wherein the distributed system can be a blockchain system, which is a distributed system formed by connecting multiple nodes through network communication. The nodes can form a peer-to-peer (P2P) network, and any form of computing device, such as a server, terminal, or other electronic device, can become a node in the blockchain system by joining this peer-to-peer network.

[0229] According to one aspect of this application, a computer-readable storage medium is provided, from which a processor of a computer device reads computer instructions, and the processor executes the computer instructions, causing the computer device to perform a video title determination method provided in various alternative implementations of the video title determination aspect described above.

[0230] Optionally, in this embodiment, the computer-readable storage medium may be configured to store a computer program for performing the following steps:

[0231] S1, obtain the target video and a set of candidate texts, wherein the set of candidate texts includes texts that are associated with the target video and are allowed to be interacted with;

[0232] S2, determine the first text set based on the popularity information of each candidate text in the candidate text set, wherein the first text set includes texts in the candidate text set whose popularity information meets the first preset popularity condition, and the popularity information includes the interaction parameters of the corresponding candidate text;

[0233] S3, perform a classification operation on the first text set to obtain a second text set, wherein the second text set includes texts whose semantic information satisfies the preset content conditions;

[0234] S4, calculate the relevance parameters of each text in the second text set with the target video, and determine the text whose relevance parameters meet the preset relevance conditions as the target text, wherein the target text is the text that can be configured as the video title of the target video.

[0235] Optionally, in this embodiment, those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0236] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0237] If the integrated units in the above embodiments are implemented as software functional units and sold or used as independent products, they can be stored in the aforementioned computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause one or more computer devices (which may be personal computers, servers, or network devices, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application.

[0238] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0239] In the several embodiments provided in this application, it should be understood that the disclosed client can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection between units or modules, and may be electrical or other forms.

[0240] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0241] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0242] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A method for determining a video title, characterized in that, include: Obtain a target video and a set of candidate texts, wherein the set of candidate texts includes texts associated with the target video and that are interactive. A first text set is determined based on the popularity information of each candidate text in the candidate text set, wherein the first text set includes texts in the candidate text set whose popularity information satisfies a first preset popularity condition, and the popularity information includes the interaction parameters of the corresponding candidate text. A classification operation is performed on the first text set to obtain a second text set, wherein the second text set includes texts whose semantic information satisfies preset content conditions; Each text in the second text set is compared with the target video to calculate a relevance parameter, and the text whose relevance parameter satisfies a preset relevance condition is determined as the target text, wherein the target text is the text that can be configured as the video title of the target video.

2. The method according to claim 1, characterized in that, The step of performing a classification operation on the first text set to obtain a second text set includes: Perform a first encoding operation on each text in the first text set to obtain a first text representation vector set, wherein each text representation vector in the first text representation vector set includes semantic features that represent the semantic information of the corresponding text; A classification operation is performed on the first set of text representation vectors to obtain a subset of first text representation vectors, wherein the representation vectors in the subset of first text representation vectors include semantic features that represent semantic information that satisfies the preset content conditions. The text corresponding to the first text representation vector subset is determined as the second text set.

3. The method according to claim 2, characterized in that, The method further includes: Perform a second encoding operation on the popularity information of each text in the first text set to obtain a popularity representation vector set, wherein each popularity representation vector in the popularity representation vector set includes a popularity feature representing the popularity information; The first set of text representation vectors is concatenated with the text representation vectors and heat representation vectors corresponding to the same text in the set of heat representation vectors to obtain the second set of text representation vectors. Perform the classification operation on the second text representation vector set to obtain a second text representation vector subset, wherein the representation vectors in the second text representation vector subset include semantic features that represent semantic information that satisfies the preset content conditions and popularity features that represent popularity information that satisfies the second preset popularity conditions. The text corresponding to the second text representation vector subset is determined as the second text set.

4. The method according to claim 3, characterized in that, The encoding operation is performed on the popularity information of each text in the first text set to obtain a set of popularity representation vectors, including: The popularity information corresponding to each text in the first text set is mapped to a set of popularity levels according to the preset popularity levels, wherein different values ​​of the popularity levels correspond to different value ranges of the popularity information. The second encoding operation is performed on each of the popularity levels in the popularity level set to obtain the popularity representation vector set, wherein the popularity feature is used to represent the popularity level of the corresponding text.

5. The method according to claim 2, characterized in that, The step of calculating the correlation parameters between each text in the second text set and the target video, and determining the text whose correlation parameters satisfy a preset correlation condition as the target text, includes: Perform a third encoding operation on the target video to obtain a video representation vector; The similarity between each text representation vector in the first text representation vector subset and the video representation vector is calculated to obtain a similarity set, wherein the relevance parameter includes the similarity. The text corresponding to the similarity in the similarity set that meets the preset similarity conditions is determined as the target text, wherein the preset relevance conditions include the preset similarity conditions.

6. The method according to claim 5, characterized in that, The third encoding operation performed on the target video to obtain a video representation vector includes: Perform frame extraction on the target video to obtain a set of target image sequences; The third encoding operation is performed on the target image sequence set to obtain the video representation vector.

7. The method according to claim 5, characterized in that, The step of determining the text corresponding to the similarity scores in the similarity set that meet the preset similarity conditions as the target text includes: Sort the similarity set in descending order, and determine the text corresponding to the top N similarity scores as the target text, where N is a positive integer; or, The texts with similarity scores exceeding a preset threshold in the similarity set are identified as the target texts.

8. The method according to claim 2, characterized in that, The step of performing a classification operation on the first set of text representation vectors to obtain a subset of the first set of text representation vectors includes: A first subset of text representation vectors is obtained by performing a classification operation on each text representation vector in the first text representation vector set in the following manner: each text representation vector that has undergone the classification operation is regarded as a target text representation vector, and the target text representation vector is used to represent the represented text in the first text set. Determine whether the format of the represented text meets the preset format conditions based on the target text representation vector; Determine whether the semantic information of the represented text is used to describe the plot based on the target text representation vector; Determine whether the represented text is fluent based on the target text representation vector; If the format of the representation text satisfies the preset format conditions, the semantic information of the representation text is used to describe the plot, and the representation text is fluent, then the target text representation vector is added to the first text representation vector subset.

9. The method according to any one of claims 1 to 8, characterized in that, The step of determining the first text set based on the popularity information of each candidate text in the candidate text set includes at least one of the following: The number of likes for each candidate text is obtained, wherein the interaction parameter includes the number of likes for the candidate text; candidate texts whose number of likes exceeds a preset threshold are identified as texts in the first text set; The number of times each candidate text is cited is obtained, wherein the interaction parameter includes the number of times the candidate text is cited; candidate texts whose number of citations exceeds a preset threshold are determined as texts in the first text set; The number of times each candidate text is forwarded is obtained, wherein the interaction parameter includes the number of times the candidate text is forwarded; candidate texts whose number of forwards exceeds a preset threshold are determined as texts in the first text set; The number of times each candidate text has been tipped is obtained, wherein the interaction parameter includes the number of times the candidate text has been tipped; candidate texts whose number of tipped times exceeds a preset threshold are identified as texts in the first text set.

10. The method according to any one of claims 1 to 8, characterized in that, The method further includes: The candidate text set is filtered to obtain a third text set, wherein each text in the third text set satisfies at least one of the following conditions: the text content of each text in the third text set is different; the text length of each text in the third text set is within a preset length range; and each text in the third text set does not contain sensitive words. The first text set is determined based on the popularity information of each text in the third text set.

11. The method according to any one of claims 1 to 8, characterized in that, The acquisition of the target video and candidate text set includes at least one of the following: Obtain the target video and the set of bullet screen texts generated during the playback of the target video, wherein the candidate text set includes the set of bullet screen texts; Obtain the target video and a set of comment texts associated with the target video, wherein the candidate text set includes the set of comment texts; Obtain the target video and a set of subtitle texts associated with the target video, wherein the candidate text set includes the set of subtitle texts.

12. A device for determining a video title, characterized in that, include: An acquisition module is used to acquire a target video and a set of candidate texts, wherein the set of candidate texts includes texts associated with the target video and that are interactive. The determining module is used to determine a first text set based on the popularity information of each candidate text in the candidate text set, wherein the first text set includes texts in the candidate text set whose popularity information satisfies a first preset popularity condition, and the popularity information includes the interaction parameters of the corresponding candidate text; The classification module is used to perform a classification operation on the first text set to obtain a second text set, wherein the second text set includes texts whose semantic information satisfies preset content conditions; The processing module is used to calculate the relevance parameters of each text in the second text set with the target video, and to determine the text whose relevance parameters satisfy a preset relevance condition as the target text, wherein the target text is text that can be configured as the video title of the target video.

13. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein the program can be executed by a terminal device or computer at runtime as described in any one of claims 1 to 11.

14. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method described in any one of claims 1 to 11.

15. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to execute the method described in any one of claims 1 to 11 through the computer program.

Citation Information

Patent Citations

  • Title generation method and device, computer readable storage medium and computer device

    CN110866391A

  • Video title generation method and device and storage medium

    CN114298018A