Label Prediction Method, Training Method, Device and Equipment of Text Prediction Model

By extracting video content on the video platform and matching it with the tag library, the video tag is directly determined, and the problem of adding new tags in the existing technology requires retraining the model, achieving efficient and compatible new tags is achieved, and the adaptability and efficiency of the model is improved.

CN115631496BActive Publication Date: 2025-07-04TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211385978.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-07
Publication Date
2025-07-04
Estimated Expiration
2042-11-07

AI Technical Summary

Technical Problem

In the prior art, the neural network model needs to be retrained after adding new tags, resulting in a long time to train the model and inefficient efficiency.

Method used

By extracting the video frames, audio and text content of the target video, the target text prediction model is used to determine the video attention vector, and match it with the description text attention vector in the label library, the tag of the target video is directly determined, avoiding retraining the model.

Benefits of technology

It improves the efficiency of text detection models to be compatible with new tags, reduces model training time, reduces manual and training costs, and ensures the timeliness and user experience of tags.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115631496B_ABST
    Figure CN115631496B_ABST
Patent Text Reader

Abstract

The present application provides a label prediction method, a training method for a text prediction model, a device and a device, belonging to the field of artificial intelligence technology. The method includes: extracting the target video content of the target video; determining the target video attention vector of the target video content through the target text prediction model, and the target text prediction model is trained based on multiple groups of sample pairs, and each group of sample pairs includes the sample video content of the sample video and the sample description text; obtaining the text attention vector of each description text in the label library, and the text attention vector is determined through the target text prediction model; through the target text prediction model, respectively matching the text attention vector of each description text with the target video attention vector, and using the description text whose matching result meets the target condition as the target description text of the target video; using the label corresponding to the target description text as the target label of the target video, which improves the efficiency of the text detection model in accommodating newly added labels.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and particularly relates to a method for predicting labels, a method for training a text prediction model, an apparatus, and a device. Background Art

[0002] With the rapid development of multimedia technology, people can select their favorite videos to watch on video platforms, and video platforms can also recommend videos to people. In order to facilitate video recommendation, it is necessary to label videos in advance. The label is used to describe the content of the video, and a label is usually a phrase, including the name of the video, the name of the person, the theme, the subject matter, the plot, etc.

[0003] In related technologies, generally, model training is performed based on videos and labels, and the neural network model obtained through training is used to label videos. And whenever a new label is added, it is necessary to re-train the model based on the new label in order to make the trained neural network model compatible with the new label. The above solution results in a long training time for the model. Summary of the Invention

[0004] Embodiments of this application provide a method for predicting labels, a method for training a text prediction model, an apparatus, and a device, which improve the efficiency of the text detection model in being compatible with new labels. The technical solution is as follows:

[0005] On the one hand, a method for predicting labels is provided. The method includes:

[0006] Extract the target video content of the target video, where the target video content includes video frames, audio, and text in the target video;

[0007] Determine the target video attention vector of the target video content through a target text prediction model. The target text prediction model is trained based on multiple groups of sample pairs, and each group of sample pairs includes the sample video content of the sample video and the sample description text, and the sample description text is used to summarize the sample video;

[0008] Obtain the text attention vector of each description text in the label library. The text attention vector is determined through the target text prediction model. The label library stores multiple description texts and the labels respectively corresponding to the multiple description texts;

[0009] Match the text attention vector of each description text with the target video attention vector respectively, and use the description text whose matching result meets the target condition as the target description text of the target video;

[0010] Use the label corresponding to the target description text as the target label of the target video.

[0011] On the other hand, a training method for a text prediction model is provided, and the method includes:

[0012] Obtain multiple groups of sample pairs, each group of sample pairs including sample video content of a sample video and a sample description text, where the sample video content includes sample video frames, sample audio, and sample text in the sample video, and the sample description text is used to summarize the sample video;

[0013] Based on the multiple groups of sample pairs, iteratively execute the following steps to train a text prediction model to obtain a target text prediction model, where the target text prediction model is used to determine, based on video content, a description text that matches the video content, and a label corresponding to the description text is used to determine a label of a video corresponding to the video content:

[0014] For a sample pair used in the iterative process, determine, through the text prediction model, a video attention vector of the sample video content in the sample pair and a text attention vector of the sample description text;

[0015] Based on the video attention vector of the sample video content and the text attention vector of the sample description text, determine a first loss value, where the first loss value is used to represent the gap between the sample description text and the true description text of the sample video content;

[0016] Based on the first loss value of the sample pair, adjust the model parameters of the text prediction model.

[0017] On the other hand, a label prediction device is provided, and the device includes:

[0018] An extraction module, configured to extract target video content of a target video, where the target video content includes video frames, audio, and text in the target video;

[0019] An attention vector determination module, configured to determine, through a target text prediction model, a target video attention vector of the target video content, where the target text prediction model is trained based on multiple groups of sample pairs, each group of sample pairs including sample video content of a sample video and a sample description text, and the sample description text is used to summarize the sample video;

[0020] An attention vector acquisition module, configured to acquire a text attention vector of each description text in a label library, where the text attention vector is determined through the target text prediction model, and the label library stores multiple description texts and labels respectively corresponding to the multiple description texts;

[0021] A matching module, configured to, through the target text prediction model, respectively match the text attention vectors of each description text with the target video attention vector, and use the description text whose matching result meets the target condition as the target description text of the target video;

[0022] A label determination module, configured to use the label corresponding to the target description text as the target label of the target video.

[0023] In some embodiments, the matching module is configured to:

[0024] Input the text attention vector of each description text and the target video attention vector into the target text prediction model;

[0025] Through the target text prediction model, for each description text, determine the similarity between the text attention vector of the description text and the target video attention vector, determine the prediction probability corresponding to the similarity, and use the description text with a prediction probability greater than the probability threshold as the target description text, where the prediction probability is used to represent the matching degree between the description text and the target video content.

[0026] In some embodiments, the attention vector determination module is configured to:

[0027] Respectively determine the target video frame feature vector of the video frame, the target audio feature vector of the audio, and the target text feature vector of the text;

[0028] Input the target video frame feature vector, the target audio feature vector, and the target text feature vector into the target text prediction model respectively. Through the target text prediction model, respectively perform attention feature extraction on the target video frame feature vector, the target audio feature vector, and the target text feature vector to obtain a target video frame attention vector, a target audio attention vector, and a target text attention vector, and fuse the target video frame attention vector, the target audio attention vector, and the target text attention vector to obtain the target video attention vector.

[0029] In some embodiments, the apparatus further includes:

[0030] A feature vector determination module, configured to, for each description text, determine the text feature vector of the description text;

[0031] A feature extraction module, configured to input the text feature vector of the description text into the target text prediction model, and through the target text prediction model, perform attention feature extraction on the text feature vector of the description text to obtain the text attention vector of the description text.

[0032] In some embodiments, the device further comprises:

[0033] A first acquisition module, configured to acquire a plurality of tags;

[0034] A first text interpretation module, configured to respectively perform text interpretation on the plurality of tags to obtain description texts respectively corresponding to the plurality of tags;

[0035] A tag library determination module, configured to obtain the tag library based on the plurality of tags and the description texts respectively corresponding to the plurality of tags.

[0036] In some embodiments, the device further comprises:

[0037] A second acquisition module, configured to periodically acquire newly added tags;

[0038] A second text interpretation module, configured to perform text interpretation on the newly added tags to obtain a description text corresponding to the newly added tags;

[0039] A storage module, configured to store the newly added tags and the description text corresponding to the newly added tags into the tag library.

[0040] On the other hand, there is provided a training device for a text prediction model, the device comprising:

[0041] An acquisition module, configured to acquire multiple groups of sample pairs, each group of sample pairs including sample video content of a sample video and a sample description text, the sample video content including sample video frames, sample audio, and sample text in the sample video, and the sample description text being used to summarize the sample video;

[0042] A training module, configured to iteratively perform the following steps based on the multiple groups of sample pairs to train a text prediction model to obtain a target text prediction model, the target text prediction model being used to determine, based on video content, a description text that matches the video content, and a tag corresponding to the description text being used to determine a tag of a video corresponding to the video content: for a sample pair used in the iterative process, determine, through the text prediction model, a video attention vector of the sample video content in the sample pair and a text attention vector of the sample description text; determine a first loss value based on the video attention vector of the sample video content and the text attention vector of the sample description text, the first loss value being used to represent a gap between the sample description text and a true description text of the sample video content; and adjust model parameters of the text prediction model based on the first loss value of the sample pair.

[0043] In some embodiments, the training module is configured to:

[0044] Determine the video frame feature vector of the sample video frame, the audio feature vector of the sample audio, and the text feature vector of the sample text respectively;

[0045] Input the video frame feature vector, the audio feature vector, and the text feature vector into the text prediction model respectively. Through the text prediction model, perform attention feature extraction on the video frame feature vector, the audio feature vector, and the text feature vector respectively to obtain a video frame attention vector, an audio attention vector, and a text attention vector, and fuse the video frame attention vector, the audio attention vector, and the text attention vector to obtain the video attention vector;

[0046] Determine the text feature vector of the sample description text;

[0047] Input the text feature vector of the sample description text into the text prediction model. Through the text prediction model, perform attention feature extraction on the text feature vector of the sample description text to obtain the text attention vector of the sample description text.

[0048] In some embodiments, in each iteration process, train the text prediction model based on a target number of sample pairs. The training module is further configured to:

[0049] For each group of sample pairs in the target number of sample pairs, combine the sample video content in the sample pair with the sample description texts in the remaining sample pairs in the target number of sample pairs respectively to form negative sample pairs, and obtain multiple groups of negative sample pairs;

[0050] Train the text prediction model based on the target number of sample pairs and the multiple groups of negative sample pairs. The target number of sample pairs is positive sample pairs.

[0051] In some embodiments, the training module is configured to:

[0052] Fuse the first loss values of the target number of sample pairs and the multiple groups of negative sample pairs to obtain a second loss value, and adjust the model parameters of the text prediction model based on the second loss value.

[0053] In some embodiments, the training module is configured to:

[0054] Determine the sample similarity between the video attention vector of the sample video content and the text attention vector of the sample description text;

[0055] Determine the sample prediction probability corresponding to the sample similarity. The sample prediction probability is used to represent the matching degree between the sample description text and the sample video content;

[0056] When the sample pair is a positive sample pair, determine the first loss value based on the sample prediction probability and the first reference probability, where the first reference probability is the reference probability corresponding to the positive sample pair;

[0057] When the sample pair is a negative sample pair, determine the first loss value based on the sample prediction probability and the second reference probability, where the second reference probability is the reference probability corresponding to the negative sample pair.

[0058] In some embodiments, the description text is obtained by text interpretation of the label corresponding to the description text.

[0059] On the other hand, a computer device is provided, which includes a processor and a memory. The memory is used to store at least one segment of computer program, and the at least one segment of computer program is loaded and executed by the processor to implement the label prediction method or the training method of the text prediction model in the embodiments of the present application.

[0060] On the other hand, a computer-readable storage medium is provided, in which at least one segment of computer program is stored, and the at least one segment of computer program is loaded and executed by a processor to implement the label prediction method or the training method of the text prediction model as in the embodiments of the present application.

[0061] On the other hand, a computer program product is provided, which includes a computer program stored in a computer-readable storage medium. The processor of the computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the label prediction method or the training method of the text prediction model described in any of the above implementation manners.

[0062] The embodiment of the present application provides a label prediction method. This method determines the target video attention vector of the target video content and the text attention vector of the description text through a target text prediction model. Since this target text prediction model is trained based on the sample video content and sample description text of the sample video, the model learns the matching rule between the video content and the description text. Then, the target text prediction model is used to match the text attention vector and the target video attention vector, and based on the matching result, it can be determined whether the description text matches the target video content. Since each description text corresponds to a label, when any description text matches the target video content, the label corresponding to this description text can be used as the label of the target video corresponding to the target video content. This method enables, after adding a new label, only adding the description text corresponding to the new label to the label library, so that the text detection model can be compatible with the new label, and there is no need to re-train the model based on the new label. This method improves the efficiency of the text detection model in being compatible with new labels. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0064] Figure 1 is a schematic diagram of an implementation environment provided by the embodiment of the present application;

[0065] Figure 2 is a flowchart of a training method for a text prediction model provided by the embodiment of the present application;

[0066] Figure 3 is a flowchart of another training method for a text prediction model provided by the embodiment of the present application;

[0067] Figure 4 is a schematic diagram of a sample pair provided by the embodiment of the present application;

[0068] Figure 5 is a schematic diagram of a text prediction model provided by the embodiment of the present application;

[0069] Figure 6 is a flowchart of another training method for a text prediction model provided by the embodiment of the present application;

[0070] Figure 7 is a flowchart of a label prediction method provided by the embodiment of the present application;

[0071] Figure 8It is a flowchart of another label prediction method provided by an embodiment of the present application;

[0072] Figure 9 It is a flowchart of another label prediction method provided by an embodiment of the present application;

[0073] Figure 10 It is a block diagram of a label prediction device provided by an embodiment of the present application;

[0074] Figure 11 It is a block diagram of a training device for a text prediction model provided by an embodiment of the present application;

[0075] Figure 12 It is a block diagram of a terminal provided by an embodiment of the present application;

[0076] Figure 13 It is a block diagram of a server provided by an embodiment of the present application. Detailed implementation manners

[0077] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.

[0078] In the present application, terms such as "first" and "second" are used to distinguish identical or similar items with basically the same functions and effects. It should be understood that there is no logical or chronological dependency between "first", "second", and "nth", nor are the quantity and execution order limited.

[0079] In the present application, the term "at least one" means one or more, and the meaning of "a plurality" means two or more.

[0080] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.), and signals involved in the present application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions. For example, the sample pairs involved in the present application are obtained under full authorization.

[0081] The following explains the terms involved in the present application.

[0082] Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making.

[0083] Artificial intelligence technology is an interdisciplinary subject that involves a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0084] Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers in natural language. Natural language processing is a science that combines linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language people use in daily life, so it has a close connection with the research of linguistics. Natural language processing technologies usually include technologies such as text processing, semantic understanding, machine translation, robot question answering, and knowledge graphs.

[0085] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning from demonstration.

[0086] Next, the implementation environment involved in this application will be introduced:

[0087] The label prediction method provided by the embodiments of the present application can be executed by a computer device. In some embodiments, the computer device is at least one of a terminal and a server. The following introduces a schematic diagram of the implementation environment of the label prediction method provided by the embodiments of the present application. Refer to Figure 1 , the implementation environment includes a terminal 101 and a server 102. The terminal 101 and the server 102 can be directly or indirectly connected through wired or wireless communication methods, which are not limited in this application. In some embodiments, the terminal 101 is used to obtain a video and send it to the server 102, and the server 102 is used to determine the label of the video. Alternatively, the terminal 101 is used to obtain a video, and the terminal 101 is used to determine the label of the video.

[0088] In some embodiments, the terminal 101 can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, an aircraft, etc., but is not limited thereto. In some embodiments, the server 102 is an independent server or can also be a server cluster or a distributed system composed of multiple servers, and can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. In some embodiments, the server 102 undertakes the main computing work, and the terminal 101 undertakes the secondary computing work; or, the server 102 undertakes the secondary computing work, and the terminal 101 undertakes the main computing work; or, the server 102 and the terminal 101 adopt a distributed computing architecture for collaborative computing.

[0089] Figure 2 is a flowchart of a training method for a text prediction model provided by the embodiments of the present application. Refer to Figure 2 , in the embodiments of the present application, taking being executed by a computer device as an example, the method includes the following steps:

[0090] 201. The computer device obtains multiple groups of sample pairs. Each group of sample pairs includes the sample video content of the sample video and the sample description text. The sample video content includes the sample video frames, sample audio, and sample text in the sample video, and the sample description text is used to summarize the sample video.

[0091] In the embodiments of the present application, the sample video can be any type of video, such as a news current affairs video, a TV drama video, an entertainment short video, etc.

[0092] In an embodiment of the present application, the sample audio is the audio signal in the sample video. The sample text is the text in the sample video, including at least one of the title, subtitles, and explanatory text of the sample video, etc.

[0093] 202. The computer device iteratively trains the text prediction model based on the multiple groups of sample pairs to obtain a target text prediction model, which is used to determine, based on the video content, a description text that matches the video content, and the label corresponding to the description text is used to determine the label of the video corresponding to the video content.

[0094] In an embodiment of the present application, for the sample pairs used in the iterative process, the computer device determines, through the text prediction model, the video attention vector of the sample video content in the sample pair and the text attention vector of the sample description text; based on the video attention vector of the sample video content and the text attention vector of the sample description text, a first loss value is determined, and the first loss value is used to represent the gap between the sample description text and the true description text of the sample video content; the computer device adjusts the model parameters of the text prediction model based on the first loss value of the sample pair. In an embodiment of the present application, it is necessary to determine whether the iteration stop condition is reached in each iteration process. If the iteration stop condition is not reached during the iteration process, the model parameters are adjusted for the next iteration; if the iteration stop condition is reached during the iteration process, the text prediction model used in this iteration is output as the target text prediction model, and the iteration stop condition may be that the number of iterations reaches a preset number, the loss value reaches a preset value, or the loss value reaches a convergence state.

[0095] The embodiment of the present application provides a training method for a text prediction model. This method iteratively trains based on the sample video content and the sample description text of the sample video. Since during the iteration process, the loss value is determined based on the video attention vector of the sample video content and the text attention vector of the sample description text, and the model parameters are adjusted based on the loss value, the target text prediction model learns the matching rule between the video content and the description text. Thus, for any video content, the description text that matches it can be determined through the target text prediction model. Since each description text corresponds to a label, when any description text matches the video content, the label corresponding to the description text can be used as the label of the video corresponding to the video content. This method enables, after adding a new label, only by matching the text attention vector of the description text corresponding to the new label with the video attention vector of the video content through the target text prediction model, the text detection model can be made compatible with the new label, and there is no need to re-train the model based on the new label. This method improves the efficiency of the text detection model in being compatible with new labels.

[0096] The above Figure 2For the basic process of training a text prediction model, the following further elaborates on the training process of the text prediction model based on Figure 3 See Figure 3 , and in the embodiments of the present application, it is described by taking the execution by a computer device as an example. The method includes the following steps:

[0097] 301. The computer device obtains multiple groups of sample pairs. Each group of sample pairs includes the sample video content of the sample video and the sample description text. The sample video content includes the sample video frames, sample audio, and sample text in the sample video, and the sample description text is used to summarize the sample video.

[0098] In the embodiments of the present application, there are multiple sample video frames. The computer device can evenly extract frames from the sample video based on a preset time interval to obtain multiple sample video frames. The number of the multiple sample video frames can be set and changed as needed, such as 8, 14, etc.

[0099] In the embodiments of the present application, the process by which the computer device obtains the multiple groups of sample pairs includes the following steps: The computer device obtains multiple sample videos, then obtains the sample description text and the sample video content of the multiple sample videos, combines the sample description text and the sample video content of each sample video to obtain a group of sample pairs, and further obtains the multiple groups of sample pairs based on the multiple sample videos.

[0100] In the embodiments of the present application, the multiple sample videos can be from videos historically published on various video applications. Among them, if the sample video frames, sample audio, and sample text of the sample video are all stored in the database associated with the computer device, the sample video content can be directly obtained; if the sample video frames, sample audio, and sample text of the sample video are not stored in the database associated with the computer device, the video content of the sample video is extracted to obtain the sample video content of the sample video. In the embodiments of the present application, the sample description text can be obtained based on manual annotation.

[0101] See Figure 4 , Figure 4 is a schematic diagram of a group of sample pairs provided by the embodiments of the present application, which includes the sample video content and the sample description text of the sample video. Among them, the sample video content includes sample video frames, sample audio, and sample text, and the sample description text summarizes the sample video.

[0102] In the embodiments of the present application, the computer device iteratively executes the following steps 302-306 to train the text prediction model to obtain the target text prediction model. In the embodiments of the present application, in each iteration process, the computer device trains based on the target number of sample pairs in the above multiple groups of sample pairs, and the target number can be set and changed as needed.

[0103] 302. In each iteration process, the computer device selects a target number of sample pairs from multiple groups of sample pairs, and the sample pairs with the target number are positive sample pairs.

[0104] In the embodiment of the present application, the computer device randomly selects a target number of sample pairs from multiple groups of sample pairs.

[0105] 303. For each group of sample pairs among the sample pairs with the target number, the computer device combines the sample video content in the sample pair with the sample description texts in the remaining sample pairs among the sample pairs with the target number respectively to form negative sample pairs, and obtains multiple groups of negative sample pairs.

[0106] In the embodiment of the present application, since the sample video content in each group of positive sample pairs is combined with the sample description texts in the remaining sample pairs respectively to obtain negative sample pairs, if the target number of positive sample pairs is N, then the number of negative sample pairs corresponding to the sample video content in each group of positive sample pairs is N - 1. Correspondingly, the number of multiple groups of negative sample pairs corresponding to N groups of positive sample pairs is N*(N - 1). Furthermore, the total number of positive sample pairs and negative sample pairs is N*N, where N is an integer greater than 1.

[0107] In the embodiment of the present application, the sample video content in each group of sample pairs is combined with the sample description texts of the remaining sample pairs to obtain negative sample pairs, and the sample description texts in the remaining sample pairs are obviously not matched with the sample video content in this sample pair. The negative sample pairs obtained in this way are accurate and reasonable; and obtaining negative sample pairs based on positive sample pairs also improves the efficiency of obtaining negative sample pairs. Since multiple groups of positive sample pairs and negative sample pairs are determined in this way, the comprehensiveness of training samples is improved. Furthermore, training the text prediction model based on positive sample pairs and negative sample pairs can improve the model training effect.

[0108] 304. For each group of sample pairs among the sample pairs with the target number and the multiple groups of negative sample pairs, the computer device determines the video attention vector of the sample video content in the sample pair and the text attention vector of the sample description text through the text prediction model.

[0109] In the embodiment of the present application, the computer device determines the video attention vector of the sample video content in the sample pair and the text attention vector of the sample description text through the text prediction model, including the following steps A1 - A2.

[0110] A1: The computer device respectively determines the video frame feature vector of the sample video frame, the audio feature vector of the sample audio, and the text feature vector of the sample text; the computer device respectively inputs the video frame feature vector, the audio feature vector, and the text feature vector into the text prediction model, and through the text prediction model, respectively performs attention feature extraction on the video frame feature vector, the audio feature vector, and the text feature vector to obtain the video frame attention vector, the audio attention vector, and the text attention vector, and fuses the video frame attention vector, the audio attention vector, and the text attention vector to obtain the video attention vector.

[0111] In the embodiments of the present application, the video frame feature vector, the audio feature vector, and the text feature vector respectively represent the features of the sample video frame, the features of the sample audio, and the features of the sample text.

[0112] In the embodiments of the present application, if there are multiple sample video frames, then there are multiple video frame feature vectors, and one sample video frame corresponds to one video frame feature vector. The dimension of the video frame feature vector can be set and changed as needed, such as 1024.

[0113] In the embodiments of the present application, the computer device digitizes the sample audio to obtain multiple audio feature vectors. The number of the audio feature vectors represents the coding length of the sample audio. The number of the multiple audio feature vectors and the dimension of each audio feature vector can be set and changed as needed, such as there are 200 of the multiple audio feature vectors and the dimension of the audio feature vector is 64.

[0114] In the embodiments of the present application, if there are multiple sample texts, then there are multiple text feature vectors, and one sample text corresponds to one text feature vector. The number of the text feature vectors represents the coding length of the sample text. The dimension of the text feature vector can be set and changed as needed, such as 1024.

[0115] See Figure 5 , Figure 5 is a schematic diagram of a text prediction model provided by the embodiments of the present application. The text prediction model includes a video content module 501. The video content module 501 includes an image encoder 5011, an audio encoder 5012, a text encoder 5013, and a multimodal decoder 5014. The image encoder 5011, the audio encoder 5012, the text encoder 5013, and the multimodal decoder 5014 are all transformers (an encoder for extracting attention features) encoders. In the embodiments of the present application, any encoder can input multiple vectors at the same time, and the maximum number of vectors that each encoder can input can be set and changed as needed, such as the number is 100.

[0116] Correspondingly, based on Figure 5The process of obtaining the video attention vector by the provided text prediction model includes the following steps: The computer device inputs the video frame feature vector into the image encoder 5011. Through the image encoder 5011, attention features are extracted from the video frame feature vector to obtain the video frame attention vector; The computer device inputs the audio feature vector into the audio encoder 5012. Through the audio encoder 5012, attention features are extracted from the audio feature vector to obtain the audio attention vector; The computer device inputs the text feature vector into the text encoder 5013. Through the text encoder 5013, attention features are extracted from the text feature vector to obtain the text attention vector. The computer device inputs the video frame attention vector, the audio attention vector, and the text attention vector into the multi-modal decoder 5014. Through the multi-modal decoder 5014, the video frame attention vector, the audio attention vector, and the text attention vector are fused to obtain the video attention vector.

[0117] In the embodiment of the present application, the video frame attention vector, the audio attention vector, and the text attention vector respectively represent the context information features of the sample video frame, the sample audio, and the sample text. The dimensions of the video frame attention vector, the audio attention vector, and the text attention vector can be set and changed as needed, such as 1024. The dimension of the video attention vector can be set and changed as needed, such as 1024.

[0118] In the embodiment of the present application, each feature vector corresponds to an attention vector. Correspondingly, if the number of video frame feature vectors is A, the number of audio feature vectors is B, and the number of text feature vectors is C, then the input of the multi-modal decoder is M attention vectors, where M = A + B + C. The output of the multi-modal decoder is a single attention vector, that is, the video attention vector.

[0119] A2: The computer device determines the text feature vector of the sample description text; The computer device inputs the text feature vector of the sample description text into the text prediction model. Through the text prediction model, attention features are extracted from the text feature vector of the sample description text to obtain the text attention vector of the sample description text.

[0120] Continue to refer to Figure 5 The text prediction model further includes a description text module 502. The description text module 502 includes a description text encoder 5021. The description text encoder 5021 is a transformer encoder. Correspondingly, based on Figure 5The process of the provided text prediction model obtaining the text attention vector includes the following steps: The computer device inputs the text feature vector of the sample description text into the description text encoder 5021. Through the description text encoder 5021, attention features are extracted from the text feature vector of the sample description text to obtain the text attention vector of the sample description text.

[0121] In the embodiments of the present application, the sample description text includes multiple characters. Then, there are multiple text feature vectors of the sample description text, with one text feature vector corresponding to one character, representing the feature of that character. If the number of the multiple characters is D and the dimension of the text feature vector is 1024, the input of the description text encoder is D 1024-dimensional feature vectors, and the output of the description text encoder is a single attention vector, that is, the text attention vector. In the embodiments of the present application, the dimension of the text attention vector is the same as that of the video attention vector, which is convenient for subsequent matching of attention vectors.

[0122] In the embodiments of the present application, due to the fusion of the video frame attention vector, the audio attention vector, and the text attention vector, the video attention vector can comprehensively represent the sample video content and is convenient for subsequent matching of the same dimension based on the fused video attention vector and the text attention vector.

[0123] It should be noted that the numbers of steps A1 - A2 are only for convenience of description. The computer device can execute steps A1 - A2 sequentially or synchronously. In the embodiments of the present application, taking the synchronous execution of steps A1 - A2 as an example for description to improve the efficiency of determining the video attention vector and the text attention vector.

[0124] 305. For each sample pair among the target number of sample pairs and the multiple groups of negative sample pairs, the computer device determines a first loss value based on the video attention vector of the sample video content and the text attention vector of the sample description text in the sample pair.

[0125] In the embodiments of the present application, the process of the above computer device determining the first loss value based on the video attention vector of the sample video content and the text attention vector of the sample description text includes the following steps B1 - B3.

[0126] B1: The computer device determines the sample similarity between the video attention vector of the sample video content and the text attention vector of the sample description text.

[0127] In some embodiments, the sample similarity is the cosine similarity. The cosine similarity refers to the cosine value of the angle between two vectors, and the value range of the cosine similarity is [-1, 1]. Correspondingly, the computer device obtains the sample similarity through the following formula (1).

[0128]

[0129] Among them, A represents the video attention vector, B represents the text attention vector, and cosθ represents the sample similarity.

[0130] B2: The computer device determines the sample prediction probability corresponding to the sample similarity, and this sample prediction probability is used to represent the matching degree between the sample description text and the sample video content.

[0131] In the embodiment of the present application, the computer device performs normalization processing on the sample similarity based on the activation function to obtain the sample prediction probability. The value range of this sample prediction probability is [0, 1], and this sample prediction probability can be expressed as p.

[0132] B3: When the sample pair is a positive sample pair, the computer device determines the first loss value based on the sample prediction probability and the first reference probability, where the first reference probability is the reference probability corresponding to the positive sample pair; when the sample pair is a negative sample pair, the computer device determines the first loss value based on the sample prediction probability and the second reference probability, where the second reference probability is the reference probability corresponding to the negative sample pair.

[0133] In the embodiment of the present application, the first reference probability and the second reference probability can be set and changed as needed. For example, the first reference probability is 1 and the second reference probability is 0.

[0134] In some embodiments, the cross-entropy loss function is used to determine the first loss value, and this cross-entropy loss function is shown in the following formula (2).

[0135] L = -[y i *log(p i ) + (1 - y i )log(1 - p i )](2);

[0136] Among them, L represents the first loss value, y i represents the reference probability corresponding to the i-th sample pair, and p i represents the sample prediction probability corresponding to the i-th sample pair. Correspondingly, when the sample pair is a positive sample pair and the first reference probability is 1, the first loss value is -log(p); when the sample pair is a negative sample pair and the second reference probability is 0, the first loss value is -log(1 - p).

[0137] In the embodiment of the present application, the prediction probability is determined based on the similarity between the attention vectors. Since the similarity can effectively reflect the matching situation between two vectors, and then the loss value is determined based on the prediction probability, which improves the rationality and accuracy of determining the loss value.

[0138] 306. The computer device combines the first loss value of the target number of sample pairs and the multiple groups of negative sample pairs to obtain a second loss value, and adjusts the model parameters of the text prediction model based on the second loss value.

[0139] In an embodiment of the present application, the computer device takes the sum of the first loss values of the target number of sample pairs and the multiple groups of negative sample pairs as the second loss value; alternatively, the computer device takes the mean of the first loss values of the target number of sample pairs and the multiple groups of negative sample pairs as the second loss value. This method combines the loss values of multiple sample pairs to adjust the model parameters, improving the efficiency of model training and reducing the error of model training.

[0140] In an embodiment of the present application, the computer device adjusts the model parameters of the text prediction model based on the second loss value using the gradient descent method. Among them, the computer device determines the derivative of the model parameters corresponding to the second loss value, and subtracts the product of the derivative and a preset ratio from the model parameters to obtain the updated model parameters.

[0141] In an embodiment of the present application, the target text prediction model is used to determine a description text that matches the video content based on the video content. The label corresponding to the description text is used to determine the label of the video corresponding to the video content, and the description text is obtained by text interpretation of the label corresponding to the description text. Since the label itself contains less information, and by text interpretation of the label, the label is transformed into a piece of text, expanding the content of the label, and thus obtaining a more comprehensive description text. In this way, after the text prediction model learns the matching rule between the video content and the description text, it can determine the label based on the description text that actually matches the video content.

[0142] The method provided in the embodiments of the present application does not regard the label as a single label, but regards the label as a piece of text, and trains a model to judge the matching between the video content and a piece of text. In this way, after training this model, it can judge whether any video content matches a piece of text. The model trained by the method provided in the embodiments of the present application can, when a new label is added, transform the problem of judging whether to add a label into the problem of judging whether the video content matches the description text corresponding to the label, without having to retrain the model based on the new label. This method obviously reduces the complexity of the model's compatibility with the new label, simplifies the process of the model's compatibility with the new label, and thus improves the efficiency of the model's compatibility with the new label.

[0143] See Figure 6 , Figure 6A training method for a text prediction model provided by an embodiment of the present application. The execution subject of this method is a computer device, and the computer device iteratively executes the following steps based on multiple groups of samples to train the text prediction model and obtain a target text prediction model. 601. Randomly sample N groups of sample pairs from multiple groups of sample pairs. 602. For each group of sample pairs, perform attention feature extraction on the sample video content in the sample pair. Among them, for the sample video frames, image attention features are extracted through an image encoder to obtain video frame attention vectors; for the sample audio, audio attention features are extracted through an audio encoder to obtain audio attention vectors; for the sample text, text attention features are extracted through a text encoder to obtain text attention vectors; these three attention vectors are used by a multi-modal decoder to extract video content attention features to obtain video attention vectors. 603. Perform attention feature extraction on the sample description text. Among them, description text attention features are extracted through a description text encoder to obtain text attention vectors. 604. Combine the loss values determined based on the video attention vectors and text attention vectors of N groups of sample pairs. 605. Based on the loss value, adjust the model parameters of the text prediction model through the gradient descent method. 606. After reaching the iteration stop condition, save the model parameters in this iteration process to obtain a target text prediction model.

[0144] The embodiment of the present application provides a training method for a text prediction model. This method iteratively trains based on the sample video content of the sample video and the sample description text. Since during the iteration process, the loss value is determined based on the video attention vector of the sample video content and the text attention vector of the sample description text, and the model parameters are adjusted based on the loss value, the target text prediction model thus learns the matching rule between the video content and the description text. Furthermore, for any video content, the corresponding description text can be determined through this target text prediction model. Since each description text corresponds to a label, when any description text matches the video content, the label corresponding to this description text can be used as the label of the video corresponding to this video content. This method enables, after adding a new label, only by matching the text attention vector of the description text corresponding to the new label with the video attention vector of the video content through this target text prediction model, the text detection model can be made compatible with the new label, and there is no need to re-train the model based on the new label. This method improves the efficiency of the text detection model in being compatible with new labels.

[0145] The above Figure 2-3 is the process for training the text prediction model, Figure 7 which is a flowchart of a label prediction method provided by an embodiment of the present application and is implemented based on the target text prediction model trained according to any of the above embodiments. Refer to Figure 7, in the embodiments of the present application, taking being executed by a computer device as an example for illustration. The method includes the following steps:

[0146] 701. The computer device extracts the target video content of the target video, and the target video content includes video frames, audio, and text in the target video.

[0147] In the embodiments of the present application, the target video is the video to be tagged. The video frames, audio, and text in the target video are the same as the sample video frames, sample audio, and sample text in step 301, and will not be elaborated here.

[0148] 702. The computer device determines the target video attention vector of the target video content through the target text prediction model, and the target text prediction model is obtained by training based on multiple groups of samples. Each group of sample pairs includes the sample video content of the sample video and the sample description text, and the sample description text is used to summarize the sample video.

[0149] 703. The computer device obtains the text attention vector of each description text in the tag library. The text attention vector is determined through the target text prediction model, and the tag library stores multiple description texts and the tags respectively corresponding to the multiple description texts.

[0150] In the embodiments of the present application, the computer device extracts the attention features of the description text through the target text prediction model to obtain the text attention vector of the description text.

[0151] 704. The computer device matches the text attention vector of each description text with the target video attention vector respectively through the target text prediction model, and uses the description text whose matching result meets the target condition as the target description text of the target video.

[0152] 705. The computer device uses the tag corresponding to the target description text as the target tag of the target video.

[0153] In the embodiments of the present application, there may be multiple description texts whose matching results meet the target condition, and correspondingly, there may be multiple target tags.

[0154] An embodiment of the present application provides a label prediction method. This method determines the target video attention vector of the target video content and the text attention vector of the description text through a target text prediction model. Since this target text prediction model is trained based on the sample video content and sample description text of the sample video, the model learns the matching rule between the video content and the description text. Then, the target text prediction model is used to match the text attention vector and the target video attention vector, and based on the matching result, it can be determined whether the description text matches the target video content. Since each description text corresponds to a label, when any description text matches the target video content, the label corresponding to this description text can be used as the label of the target video corresponding to the target video content. This method enables, after adding a new label, only adding the description text corresponding to the new label to the label library, so that the text detection model can be compatible with the new label, and there is no need to retrain the model based on the new label. This method improves the efficiency of the text detection model in being compatible with new labels.

[0155] The above Figure 7 is the basic process for predicting labels. Next, based on Figure 8 the process of predicting labels will be further elaborated. Refer to Figure 8 , and in the embodiment of the present application, it is described by taking the execution by a computer device as an example. This method includes the following steps:

[0156] 801. The computer device extracts the target video content of the target video, and the target video content includes video frames, audio, and text in the target video.

[0157] This step is the same as step 701 and will not be elaborated here.

[0158] 802. The computer device determines the target video attention vector of the target video content through a target text prediction model. This target text prediction model is trained based on multiple groups of sample pairs, and each group of sample pairs includes the sample video content of the sample video and the sample description text, and the sample description text is used to summarize the sample video.

[0159] In the embodiment of the present application, since the target video content includes video frames, audio, and text of the target video, correspondingly, the process by which the computer device determines the target video attention vector of the target video content through the target text prediction model includes the following steps:

[0160] The computer device respectively determines the target video frame feature vector of the video frame, the target audio feature vector of the audio, and the target text feature vector of the text; the computer device inputs the target video frame feature vector, the target audio feature vector, and the target text feature vector into the target text prediction model respectively, and through the target text prediction model, performs attention feature extraction on the target video frame feature vector, the target audio feature vector, and the target text feature vector respectively to obtain the target video frame attention vector, the target audio attention vector, and the target text attention vector, and fuses the target video frame attention vector, the target audio attention vector, and the target text attention vector to obtain the target video attention vector. In the embodiment of the present application, due to the fusion of the target video frame attention vector, the target audio attention vector, and the target text attention vector, the target video attention vector can comprehensively represent the target video content, and it is convenient for subsequent matching in the same dimension based on the fused target video attention vector and the text attention vector.

[0161] In the embodiment of the present application, the target video frame feature vector, the target audio feature vector, and the target text feature vector are the same as the video frame feature vector, the audio feature vector, and the text feature vector in step 304, and will not be elaborated here. The computer device determines the target video attention vector through the image encoder 5011, the audio encoder 5012, the text encoder 5013, and the multimodal decoder 5014 of the target text prediction model, and this process is the same as the process of determining the video attention vector in step 304, and will not be elaborated here.

[0162] 803. The computer device obtains the text attention vector of each description text in the tag library, and this text attention vector is determined by the target text prediction model. The tag library stores multiple description texts and the tags respectively corresponding to the multiple description texts.

[0163] In the embodiment of the present application, the process by which the computer device determines the text attention vector of each description text in the tag library through the target text prediction model includes the following steps: for each description text, the computer device determines the text feature vector of the description text; the computer device inputs the text feature vector of the description text into the target text prediction model, and through the target text prediction model, performs attention feature extraction on the text feature vector of the description text to obtain the text attention vector of the description text. In the embodiment of the present application, since the text attention vector is determined, it is convenient for subsequent matching based on the text attention vector.

[0164] In the embodiment of the present application, the computer device determines the text attention vector of the description text through the description text encoder 5021 of the target text prediction model, and this process is the same as the process of determining the text attention vector of the sample description text in step 304, and will not be elaborated here.

[0165] In an embodiment of the present application, the process of determining the tag library includes the following steps: The computer device obtains a plurality of tags; the computer device respectively performs text interpretation on the plurality of tags to obtain description texts corresponding to the plurality of tags respectively; the computer device obtains a tag library based on the plurality of tags and the description texts corresponding to the plurality of tags respectively.

[0166] For example, for the tag Zhao**, its description text can be "actress, entered the entertainment industry in ** year, won ** awards". It should be noted that the tag itself contains less information, and by performing text interpretation on the tag, the tag is transformed into a piece of text, expanding the content of the tag, and thus obtaining a more comprehensive description text.

[0167] In an embodiment of the present application, the text attention vector of each description text in the tag library can be generated offline. For example, after the computer device obtains the description text corresponding to any tag, it obtains the text attention vector of the description text through the target text prediction model. Optionally, the text attention vector of each description text can be stored in the tag library corresponding to the description text. In an embodiment of the present application, since the text attention vectors of each description text in the tag library are determined in advance, the efficiency of obtaining the text attention vectors is improved.

[0168] It should be noted that when a new Internet hot event occurs or a new episode is added to a TV drama, new tags need to be added quickly and specifically, that is, the addition of tags occurs frequently. Then, in some embodiments, the computer device periodically obtains the newly added tags; the computer device performs text interpretation on the newly added tags to obtain the description texts corresponding to the newly added tags; the computer device stores the newly added tags and the description texts corresponding to the newly added tags in the tag library. This method facilitates subsequent prediction of tags for videos based on the newly added tags, and improves the application timeliness of the newly added tags.

[0169] 804. The computer device inputs the text attention vector of each description text and the target video attention vector into the target text prediction model. Through the target text prediction model, for each description text, it determines the similarity between the text attention vector of the description text and the target video attention vector, and determines the prediction probability corresponding to the similarity. This prediction probability is used to represent the matching degree between the description text and the target video content.

[0170] In an embodiment of the present application, the process for the computer device to determine the prediction probability is the same as the process for determining the sample prediction probability in step 305, and will not be elaborated here.

[0171] 805. The computer device uses the target text prediction model to use the description text with a prediction probability greater than the probability threshold as the target description text.

[0172] In the embodiments of the present application, the probability threshold can be set and changed as needed. For example, the probability threshold is 0.5.

[0173] In the embodiments of the present application, when the predicted probability of the description text is greater than the probability threshold, it indicates that there is a high degree of matching between the description text and the target video content, thereby improving the accuracy of determining the target description text.

[0174] 806. The computer device uses the label corresponding to the target description text as the target label of the target video.

[0175] In the embodiments of the present application, there is no need to obtain a sample video based on the newly added labels manually annotated for the video, and then retrain the model based on the sample video. Instead, only the description text needs to be added to the newly added label, and the newly added label and its description text are added to the label library. Subsequently, it can be automatically determined whether the newly added label is added to the video based on the video content. This method enables the newly added label to be compatible with the model in a short time, such as 5 minutes. Compared with the 2 days or 3 days required for retraining the model, it improves the efficiency of the newly added label being compatible with the model. Moreover, since the method provided in the embodiments of the present application avoids manual annotation of data and retraining of the model, it reduces the labor and training costs, achieving cost reduction and efficiency improvement. And this method enables the newly added label to be applied to the video recommendation system in a timely manner, ensuring the timeliness of the newly added label, and thus effectively improving the user experience of the product.

[0176] See Figure 9 , Figure 9The flowchart of a label prediction method provided by an embodiment of this application. Among them, a computer device deploys a service for label prediction on the GPU (Graphics Processing Unit) of the computer device. This service deploys a target text prediction model, which includes a video content module, a description text module, and a matching module. The video content module is used to obtain the target video attention vector of the target video content of the target video; the description text module is used to obtain the text attention vector of each description text in the label library; the matching module is used to match the text attention vector of each description text in the label library with the target video attention vector respectively to obtain the target description text, and then obtain the target label. Optionally, the target video can be stored in a video library, which is used to store multiple videos and the video content corresponding to each of the multiple videos. The computer device, in response to receiving a label prediction request carrying the identifier of the target video, obtains the target video content of the target video from the video library, and then determines the target video attention vector. Or, the label prediction request carries the target video, and the computer device extracts the video content of the target video through the video content module, and then determines the target video attention vector. Among them, the video library supports the addition, deletion, modification, and query of videos. The label library supports the addition, deletion, modification, and query of labels. When adding a new label to the label library, it is necessary to add the description text corresponding to the label to the label library at the same time. This method can output the target label of the target video through the label prediction service provided above.

[0177] An embodiment of this application provides a label prediction method. This method determines the target video attention vector of the target video content and the text attention vector of the description text through a target text prediction model. Since this target text prediction model is trained based on the sample video content and sample description text of the sample video, the model learns the matching rule between the video content and the description text. Then, through this target text prediction model, the text attention vector and the target video attention vector are matched, and based on the matching result, it can be determined whether the description text matches the target video content. Since each description text corresponds to a label, then in the case where any description text matches the target video content, the label corresponding to this description text can be used as the label of the target video corresponding to the target video content. This method enables, after adding a new label, only adding the description text corresponding to the new label to the label library, so that the text detection model can be compatible with the new label, and there is no need to retrain the model based on the new label. This method improves the efficiency of the text detection model in being compatible with new labels.

[0178] Figure 10 It is a block diagram of a label prediction device provided by an embodiment of this application. This device is used to execute the steps in the above label prediction method. Refer to Figure 10 and the device includes:

[0179] An extraction module 1001 is configured to extract target video content of a target video, where the target video content includes video frames, audio, and text in the target video;

[0180] An attention vector determination module 1002 is configured to determine a target video attention vector of the target video content through a target text prediction model, and the target text prediction model is obtained by training based on multiple groups of sample pairs. Each group of sample pairs includes sample video content of a sample video and a sample description text, and the sample description text is used to summarize the sample video;

[0181] An attention vector acquisition module 1003 is configured to acquire a text attention vector of each description text in a tag library. The text attention vector is determined through the target text prediction model, and the tag library stores multiple description texts and tags respectively corresponding to the multiple description texts;

[0182] A matching module 1004 is configured to, through the target text prediction model, respectively match the text attention vector of each description text with the target video attention vector, and use the description text whose matching result meets the target condition as the target description text of the target video;

[0183] A tag determination module 1005 is configured to use the tag corresponding to the target description text as the target tag of the target video.

[0184] In some embodiments, the matching module 1004 is configured to:

[0185] Input the text attention vector of each description text and the target video attention vector into the target text prediction model;

[0186] Through the target text prediction model, for each description text, determine the similarity between the text attention vector of the description text and the target video attention vector, determine the prediction probability corresponding to the similarity, and the prediction probability is used to represent the matching degree between the description text and the target video content. Use the description text with a prediction probability greater than the probability threshold as the target description text.

[0187] In some embodiments, the attention vector determination module 1002 is configured to:

[0188] Respectively determine a target video frame feature vector of the video frame, a target audio feature vector of the audio, and a target text feature vector of the text;

[0189] Input the target video frame feature vector, the target audio feature vector, and the target text feature vector into the target text prediction model respectively. Through the target text prediction model, perform attention feature extraction on the target video frame feature vector, the target audio feature vector, and the target text feature vector respectively to obtain the target video frame attention vector, the target audio attention vector, and the target text attention vector. Then fuse the target video frame attention vector, the target audio attention vector, and the target text attention vector to obtain the target video attention vector.

[0190] In some embodiments, the apparatus further includes:

[0191] A feature vector determination module, configured to determine the text feature vector of each description text.

[0192] A feature extraction module 1001, configured to input the text feature vector of the description text into the target text prediction model, and perform attention feature extraction on the text feature vector of the description text through the target text prediction model to obtain the text attention vector of the description text.

[0193] In some embodiments, the apparatus further includes:

[0194] A first acquisition module, configured to acquire a plurality of tags.

[0195] A first text interpretation module, configured to perform text interpretation on the plurality of tags respectively to obtain the description texts corresponding to the plurality of tags respectively.

[0196] A tag library determination module, configured to obtain a tag library based on the plurality of tags and the description texts corresponding to the plurality of tags respectively.

[0197] In some embodiments, the apparatus further includes:

[0198] A second acquisition module, configured to periodically acquire newly added tags.

[0199] A second text interpretation module, configured to perform text interpretation on the newly added tags to obtain the description text corresponding to the newly added tags.

[0200] A storage module, configured to store the newly added tags and the description texts corresponding to the newly added tags into the tag library.

[0201] An embodiment of the present application provides a label prediction device, which determines the target video attention vector of the target video content and the text attention vector of the description text through a target text prediction model. Since the target text prediction model is trained based on the sample video content and the sample description text of the sample video, the model learns the matching rule between the video content and the description text. Then, the text attention vector and the target video attention vector are matched through the target text prediction model, and based on the matching result, it can be determined whether the description text matches the target video content. Since each description text corresponds to a label, when any description text matches the target video content, the label corresponding to the description text can be used as the label of the target video corresponding to the target video content. This method enables, after adding a new label, only adding the description text corresponding to the new label to the label library, so that the text detection model can be compatible with the new label, and there is no need to retrain the model based on the new label, improving the efficiency of the text detection model in being compatible with the new label.

[0202] Figure 11 It is a block diagram of a training device for a text prediction model according to an embodiment of the present application. The device is used to perform the steps when executing the above-mentioned training method of the text prediction model. Refer to Figure 11 The device includes:

[0203] An acquisition module 1101, configured to acquire multiple groups of sample pairs. Each group of sample pairs includes the sample video content of the sample video and the sample description text. The sample video content includes the sample video frames, sample audio, and sample text in the sample video, and the sample description text is used to summarize the sample video;

[0204] A training module 1102, configured to iteratively execute the following steps based on multiple groups of sample pairs to train the text prediction model to obtain a target text prediction model. The target text prediction model is used to determine the description text that matches the video content based on the video content, and the label corresponding to the description text is used to determine the label of the video corresponding to the video content: For the sample pair used in the iteration process, determine the video attention vector of the sample video content and the text attention vector of the sample description text in the sample pair through the text prediction model; determine a first loss value based on the video attention vector of the sample video content and the text attention vector of the sample description text. The first loss value is used to represent the gap between the sample description text and the true description text of the sample video content; adjust the model parameters of the text prediction model based on the first loss value of the sample pair.

[0205] In some embodiments, the training module 1102 is configured to:

[0206] Respectively determine the video frame feature vector of the sample video frame, the audio feature vector of the sample audio, and the text feature vector of the sample text;

[0207] Input the video frame feature vector, audio feature vector, and text feature vector into the text prediction model respectively. Through the text prediction model, perform attention feature extraction on the above-mentioned video frame feature vector, audio feature vector, and text feature vector respectively to obtain a video frame attention vector, an audio attention vector, and a text attention vector. Then fuse the video frame attention vector, audio attention vector, and text attention vector to obtain a video attention vector;

[0208] Determine the text feature vector of the sample description text;

[0209] Input the text feature vector of the sample description text into the text prediction model. Through the text prediction model, perform attention feature extraction on the text feature vector of the sample description text to obtain the text attention vector of the sample description text.

[0210] In some embodiments, in each iteration process, based on a target number of sample pairs, train the text prediction model. The training module 1102 is further configured to:

[0211] For each group of sample pairs in the target number of sample pairs, combine the sample video content in the sample pair with the sample description texts in the remaining sample pairs in the target number of sample pairs respectively to form negative sample pairs, and obtain multiple groups of negative sample pairs;

[0212] Based on the target number of sample pairs and multiple groups of negative sample pairs, train the text prediction model. The target number of sample pairs is positive sample pairs.

[0213] In some embodiments, the training module 1102 is configured to:

[0214] Fuse the first loss values of the target number of sample pairs and multiple groups of negative sample pairs to obtain a second loss value. Based on the second loss value, adjust the model parameters of the text prediction model.

[0215] In some embodiments, the training module 1102 is configured to:

[0216] Determine the sample similarity between the video attention vector of the sample video content and the text attention vector of the sample description text;

[0217] Determine the sample prediction probability corresponding to the sample similarity. The sample prediction probability is used to represent the matching degree between the sample description text and the sample video content;

[0218] In the case where the sample pair is a positive sample pair, based on the sample prediction probability and the first reference probability, determine the first loss value. The first reference probability is the reference probability corresponding to the positive sample pair;

[0219] In the case where the sample pair is a negative sample pair, a first loss value is determined based on the sample prediction probability and the second reference probability, and the second reference probability is the reference probability corresponding to the negative sample pair.

[0220] In some embodiments, the description text is obtained by text interpretation based on the label corresponding to the description text.

[0221] The embodiment of the present application provides a training device for a text prediction model, which performs iterative training based on the sample video content and the sample description text of the sample video. Since in the iterative process, the loss value is determined based on the video attention vector of the sample video content and the text attention vector of the sample description text, and the model parameters are adjusted based on the loss value, the target text prediction model learns the matching rule between the video content and the description text. Thus, for any video content, the corresponding description text can be determined through the target text prediction model. Since each description text corresponds to a label, when any description text matches the video content, the label corresponding to the description text can be used as the label of the video corresponding to the video content. This method enables, after adding a new label, only by matching the text attention vector of the description text corresponding to the new label and the video attention vector of the video content through the target text prediction model, to achieve that the text detection model is compatible with the new label, and there is no need to re-train the model based on the new label, improving the efficiency of the text detection model being compatible with the new label.

[0222] In the embodiment of the present application, the computer device can be a terminal or a server. When the computer device is a terminal, the terminal is used as the execution subject to implement the technical solution provided by the embodiment of the present application; when the computer device is a server, the server is used as the execution subject to implement the technical solution provided by the embodiment of the present application; or, the technical solution provided by the present application is implemented through the interaction between the terminal and the server, and the embodiment of the present application does not limit this.

[0223] Figure 12 The block diagram of the terminal 1200 provided by an exemplary embodiment of the present application is shown.

[0224] Generally, the terminal 1200 includes: a processor 1201 and a memory 1202.

[0225] The processor 1201 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor 1201 may be implemented in at least one of the following hardware forms: DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 1201 may also include a main processor and a co-processor. The main processor is used to process data in the wake state and is also referred to as the CPU (Central Processing Unit); the co-processor is a low-power processor used to process data in the standby state. In some embodiments, the processor 1201 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1201 may further include an AI (Artificial Intelligence) processor, which is used to process computational operations related to machine learning.

[0226] The memory 1202 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 1202 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1202 is used to store at least one program code, and the at least one program code is used to be executed by the processor 1201 to implement the label prediction method or the training method of the text prediction model provided in the method embodiments of the present application.

[0227] In some embodiments, the terminal 1200 may further optionally include: a peripheral device interface 1203 and at least one peripheral device. The processor 1201, the memory 1202, and the peripheral device interface 1203 may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 1203 through a bus, signal lines, or a circuit board. Specifically, the peripheral device includes at least one of the following: a radio frequency circuit 1204, a display screen 1205, a camera module 1206, an audio circuit 1207, and a power supply 1208.

[0228] The peripheral device interface 1203 can be used to connect at least one I / O (Input / Output) related peripheral device to the processor 1201 and the memory 1202. In some embodiments, the processor 1201, the memory 1202, and the peripheral device interface 1203 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 1201, the memory 1202, and the peripheral device interface 1203 can be implemented on a separate chip or circuit board, and this embodiment does not limit this.

[0229] The radio frequency circuit 1204 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 1204 communicates with a communication network and other communication devices through electromagnetic signals. The radio frequency circuit 1204 converts an electrical signal into an electromagnetic signal for transmission, or converts a received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 1204 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, and so on. The radio frequency circuit 1204 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: the World Wide Web, a metropolitan area network, an intranet, various generations of mobile communication networks (2G, 3G, 4G, and 5G), a wireless local area network, and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 1204 may further include a circuit related to NFC (Near Field Communication), and this application does not limit this.

[0230] The display screen 1205 is used to display the UI (User Interface). The UI may include graphics, label descriptions, icons, videos, and any combination thereof. When the display screen 1205 is a touch display screen, the display screen 1205 also has the ability to collect touch signals on or above the surface of the display screen 1205. The touch signals can be input to the processor 1201 as control signals for processing. At this time, the display screen 1205 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 1205, which is provided on the front panel of the terminal 1200; in other embodiments, there may be at least two display screens 1205, which are respectively provided on different surfaces of the terminal 1200 or are in a folding design; in other embodiments, the display screen 1205 may be a flexible display screen, which is provided on the curved surface or the folding surface of the terminal 1200. Even, the display screen 1205 can also be set to an irregular non-rectangular shape, that is, a special-shaped screen. The display screen 1205 can be prepared from materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).

[0231] The camera module 1206 is used to collect images or videos. Optionally, the camera module 1206 includes a front camera and a rear camera. Generally, the front camera is provided on the front panel of the terminal, and the rear camera is provided on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth-of-field camera, a wide-angle camera, and a telephoto camera, so as to realize the function of background blurring by fusing the main camera and the depth-of-field camera, the function of panoramic shooting by fusing the main camera and the wide-angle camera, and the VR (Virtual Reality) shooting function or other fusion shooting functions. In some embodiments, the camera module 1206 may also include a flash. The flash can be a single-color-temperature flash or a two-color-temperature flash. The two-color-temperature flash refers to the combination of a warm-light flash and a cold-light flash, which can be used for light compensation under different color temperatures.

[0232] The audio circuit 1207 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 1201 for processing, or input to the radio frequency circuit 1204 to enable voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the terminal 1200. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signal from the processor 1201 or the radio frequency circuit 1204 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signal into sound waves audible to humans, but also convert the electrical signal into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 1207 may further include a headphone jack.

[0233] The power supply 1208 is used to supply power to each component in the terminal 1200. The power supply 1208 may be alternating current, direct current, a primary battery or a rechargeable battery. When the power supply 1208 includes a rechargeable battery, the rechargeable battery may be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery charged through a wired line, and a wireless rechargeable battery is a battery charged through a wireless coil. The rechargeable battery may also be used to support fast charging technology.

[0234] In some embodiments, the terminal 1200 further includes one or more sensors 1209. The one or more sensors 1209 include but are not limited to: an acceleration sensor 1210, a gyroscope sensor 1211, a pressure sensor 1212, an optical sensor 1213, and a proximity sensor 1214.

[0235] The acceleration sensor 1210 can detect the magnitude of acceleration on the three coordinate axes of the coordinate system established with the terminal 1200. For example, the acceleration sensor 1210 can be used to detect the components of the gravitational acceleration on the three coordinate axes. The processor 1201 can control the display screen 1205 to display the user interface in a landscape view or a portrait view according to the gravitational acceleration signal collected by the acceleration sensor 1210. The acceleration sensor 1210 can also be used for collecting game or user's motion data.

[0236] The gyroscope sensor 1211 can detect the body direction and rotation angle of the terminal 1200. The gyroscope sensor 1211 can cooperate with the acceleration sensor 1210 to collect the 3D actions of the user on the terminal 1200. Based on the data collected by the gyroscope sensor 1211, the processor 1201 can implement the following functions: motion sensing (such as changing the UI according to the user's tilt operation), image stabilization during shooting, game control, and inertial navigation.

[0237] The pressure sensor 1212 can be disposed on the side frame of the terminal 1200 and / or the lower layer of the display screen 1205. When the pressure sensor 1212 is disposed on the side frame of the terminal 1200, it can detect the holding signal of the user on the terminal 1200, and the processor 1201 performs left / right hand recognition or quick operation according to the holding signal collected by the pressure sensor 1212. When the pressure sensor 1212 is disposed on the lower layer of the display screen 1205, the processor 1201 controls the operable controls on the UI interface according to the pressure operation of the user on the display screen 1205. The operable controls include at least one of a button control, a scroll bar control, an icon control, and a menu control.

[0238] The optical sensor 1213 is used to collect the ambient light intensity. In one embodiment, the processor 1201 can control the display brightness of the display screen 1205 according to the ambient light intensity collected by the optical sensor 1213. Specifically, when the ambient light intensity is high, the display brightness of the display screen 1205 is increased; when the ambient light intensity is low, the display brightness of the display screen 1205 is decreased. In another embodiment, the processor 1201 can also dynamically adjust the shooting parameters of the camera assembly 1206 according to the ambient light intensity collected by the optical sensor 1213.

[0239] The proximity sensor 1214, also known as a distance sensor, is usually disposed on the front panel of the terminal 1200. The proximity sensor 1214 is used to collect the distance between the user and the front of the terminal 1200. In one embodiment, when the proximity sensor 1214 detects that the distance between the user and the front of the terminal 1200 is gradually decreasing, the processor 1201 controls the display screen 1205 to switch from the lit state to the off state; when the proximity sensor 1214 detects that the distance between the user and the front of the terminal 1200 is gradually increasing, the processor 1201 controls the display screen 1205 to switch from the off state to the lit state.

[0240] Those skilled in the art can understand that Figure 12 the structure shown in

[0241] Figure 13It is a schematic structural diagram of a server provided by an embodiment of the present application. The server 1300 may vary greatly due to different configurations or performances, and may include one or more Central Processing Units (CPUs) 1301 and one or more memories 1302. Among them, the memory 1302 is used to store executable program codes, and the processor 1301 is configured to execute the above-mentioned executable program codes to implement the label prediction method or the training method of the text prediction model provided by each of the above method embodiments. Of course, the server may also have components such as wired or wireless network interfaces, keyboards, and input / output interfaces for input / output. The server may also include other components for implementing device functions, which will not be elaborated here.

[0242] An embodiment of the present application also provides a computer-readable storage medium, which is used to store at least one segment of computer program, and the at least one segment of computer program is used to implement the label prediction method or the training method of the text prediction model in any of the above implementation manners.

[0243] An embodiment of the present application also provides a computer program product, which includes a computer program. The computer program is stored in a computer-readable storage medium, and the processor of the computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the label prediction method or the training method of the text prediction model in any of the above implementation manners.

[0244] In some embodiments, the computer program product involved in the embodiments of the present application may be deployed to be executed on a computer device, or on multiple computer devices located at one place, or on multiple computer devices distributed at multiple places and interconnected through a communication network. The multiple computer devices distributed at multiple places and interconnected through a communication network may form a blockchain system.

[0245] The above are only optional embodiments of the present application, and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A label prediction method, characterized in that, The method includes: extracting the target video content of the target video, where the target video content includes video frames, audio, and text in the target video; determining a target video attention vector of the target video content through a target text prediction model, where the target text prediction model is trained based on multiple groups of sample pairs, and each group of sample pairs includes the sample video content of a sample video and a sample description text for summarizing the sample video; obtaining a text attention vector of each description text in a tag library, where the text attention vector is determined through the target text prediction model, and the tag library stores multiple description texts and tags respectively corresponding to the multiple description texts; matching the text attention vector of each description text with the target video attention vector respectively through the target text prediction model, and using the description text whose matching result meets the target condition as the target description text of the target video; using the tag corresponding to the target description text as the target tag of the target video; wherein, the method further includes: obtaining multiple tags; respectively performing text interpretation on the multiple tags to convert the multiple tags into a piece of text respectively, so as to obtain the description texts respectively corresponding to the multiple tags; obtaining the tag library based on the multiple tags and the description texts respectively corresponding to the multiple tags; periodically obtaining newly added tags; performing text interpretation on the newly added tags to convert the newly added tags into a piece of text, so as to obtain the description text corresponding to the newly added tag; storing the newly added tag and the description text corresponding to the newly added tag into the tag library.

2. The method according to claim 1, characterized in that, The step of matching the text attention vector of each description text with the target video attention vector respectively and using the description text whose matching result meets the target condition as the target description text of the target video includes: inputting the text attention vector of each description text and the target video attention vector into the target text prediction model; through the target text prediction model, for each description text, determining the similarity between the text attention vector of the description text and the target video attention vector, determining the prediction probability corresponding to the similarity, and using the description text with a prediction probability greater than a probability threshold as the target description text, where the prediction probability is used to represent the matching degree between the description text and the target video content.

3. The method according to claim 1, characterized in that, The step of determining a target video attention vector of the target video content through a target text prediction model includes: respectively determining a target video frame feature vector of the video frame, a target audio feature vector of the audio, and a target text feature vector of the text; Input the target video frame feature vector, the target audio feature vector, and the target text feature vector into the target text prediction model respectively. Through the target text prediction model, perform attention feature extraction on the target video frame feature vector, the target audio feature vector, and the target text feature vector respectively to obtain a target video frame attention vector, a target audio attention vector, and a target text attention vector. Then fuse the target video frame attention vector, the target audio attention vector, and the target text attention vector to obtain the target video attention vector.

4. The method according to claim 1, wherein The method further includes: For each descriptive text, determine the text feature vector of the descriptive text; Input the text feature vector of the descriptive text into the target text prediction model. Through the target text prediction model, perform attention feature extraction on the text feature vector of the descriptive text to obtain the text attention vector of the descriptive text.

5. A training method for a text prediction model, characterized in that, The method includes: Obtain multiple groups of sample pairs. Each group of sample pairs includes the sample video content of a sample video and a sample descriptive text. The sample video content includes sample video frames, sample audio, and sample text in the sample video. The sample descriptive text is used to summarize the sample video; Based on the multiple groups of sample pairs, iteratively execute the following steps to train a text prediction model to obtain a target text prediction model. The target text prediction model is used to determine, based on the video content, a descriptive text that matches the video content from a tag library. The tag corresponding to the descriptive text is used to determine the tag of the video corresponding to the video content. The tag library stores multiple descriptive texts and the tags corresponding to the multiple descriptive texts respectively: For the sample pair used in the iteration process, through the text prediction model, determine the video attention vector of the sample video content in the sample pair and the text attention vector of the sample descriptive text; Based on the video attention vector of the sample video content and the text attention vector of the sample descriptive text, determine a first loss value. The first loss value is used to represent the gap between the sample descriptive text and the true descriptive text of the sample video content; Based on the first loss value of the sample pair, adjust the model parameters of the text prediction model; Wherein, the method further includes: obtaining multiple tags; respectively performing text interpretation on the multiple tags to convert the multiple tags into a piece of text respectively to obtain the descriptive texts corresponding to the multiple tags respectively; based on the multiple tags and the descriptive texts corresponding to the multiple tags respectively, obtain the tag library; periodically obtain newly added tags; perform text interpretation on the newly added tags to convert the newly added tags into a piece of text to obtain the descriptive text corresponding to the newly added tags; store the newly added tags and the descriptive texts corresponding to the newly added tags into the tag library.

6. The method according to claim 5, wherein The step of, through the text prediction model, determining the video attention vector of the sample video content in the sample pair and the text attention vector of the sample descriptive text includes: Determine the video frame feature vector of the sample video frame, the audio feature vector of the sample audio, and the text feature vector of the sample text respectively; Input the video frame feature vector, the audio feature vector, and the text feature vector into the text prediction model respectively. Through the text prediction model, perform attention feature extraction on the video frame feature vector, the audio feature vector, and the text feature vector respectively to obtain a video frame attention vector, an audio attention vector, and a text attention vector, and fuse the video frame attention vector, the audio attention vector, and the text attention vector to obtain the video attention vector; Determine the text feature vector of the sample description text; Input the text feature vector of the sample description text into the text prediction model. Through the text prediction model, perform attention feature extraction on the text feature vector of the sample description text to obtain the text attention vector of the sample description text.

7. The method according to claim 5, characterized in that In each iteration process, train the text prediction model based on a target number of sample pairs. The process of training the text prediction model based on a target number of sample pairs includes: For each group of sample pairs in the target number of sample pairs, combine the sample video content in the sample pair with the sample description texts in the remaining sample pairs in the target number of sample pairs respectively to form negative sample pairs, and obtain multiple groups of negative sample pairs; Train the text prediction model based on the target number of sample pairs and the multiple groups of negative sample pairs. The target number of sample pairs is positive sample pairs.

8. The method according to claim 7, characterized in that, Adjust the model parameters of the text prediction model based on the first loss value of the sample pairs, including: Fuse the first loss values of the target number of sample pairs and the multiple groups of negative sample pairs to obtain a second loss value, and adjust the model parameters of the text prediction model based on the second loss value.

9. The method according to claim 7, wherein Determine the first loss value based on the video attention vector of the sample video content and the text attention vector of the sample description text, including: Determine the sample similarity between the video attention vector of the sample video content and the text attention vector of the sample description text; Determine the sample prediction probability corresponding to the sample similarity. The sample prediction probability is used to represent the matching degree between the sample description text and the sample video content; In the case where the sample pair is a positive sample pair, determine the first loss value based on the sample prediction probability and the first reference probability. The first reference probability is the reference probability corresponding to the positive sample pair; In the case where the sample pair is a negative sample pair, determine the first loss value based on the sample prediction probability and the second reference probability. The second reference probability is the reference probability corresponding to the negative sample pair.

10. A label prediction device, characterized in that, The device includes: An extraction module, configured to extract target video content of a target video. The target video content includes video frames, audio, and text in the target video; An attention vector determination module, configured to determine a target video attention vector of the target video content through a target text prediction model, where the target text prediction model is trained based on multiple groups of samples, and each group of samples includes sample video content of a sample video and a sample description text for summarizing the sample video; An attention vector acquisition module, configured to acquire a text attention vector of each description text in a tag library, where the text attention vector is determined through the target text prediction model, and the tag library stores multiple description texts and tags respectively corresponding to the multiple description texts; A matching module, configured to match the text attention vector of each description text with the target video attention vector respectively, and use the description text whose matching result meets the target condition as the target description text of the target video; A tag determination module, configured to use the tag corresponding to the target description text as the target tag of the target video; A first acquisition module, configured to acquire multiple tags; A first text interpretation module, configured to respectively perform text interpretation on the multiple tags to convert the multiple tags into a text segment each, and obtain description texts respectively corresponding to the multiple tags; A tag library determination module, configured to obtain the tag library based on the multiple tags and the description texts respectively corresponding to the multiple tags; A second acquisition module, configured to periodically acquire newly added tags; A second text interpretation module, configured to perform text interpretation on the newly added tags to convert the newly added tags into a text segment, and obtain a description text corresponding to the newly added tags; A storage module, configured to store the newly added tags and the description texts corresponding to the newly added tags into the tag library.

11. The device according to claim 10, characterized in that, The matching module is configured to: Input the text attention vector of each description text and the target video attention vector into the target text prediction model; Through the target text prediction model, for each description text, determine the similarity between the text attention vector of the description text and the target video attention vector, determine the prediction probability corresponding to the similarity, and use the description text with a prediction probability greater than a probability threshold as the target description text, where the prediction probability is used to represent the matching degree between the description text and the target video content.

12. The device according to claim 10, characterized in that, The attention vector determination module is configured to: Respectively determine a target video frame feature vector of the video frame, a target audio feature vector of the audio, and a target text feature vector of the text; Input the target video frame feature vector, the target audio feature vector, and the target text feature vector into the target text prediction model respectively. Through the target text prediction model, perform attention feature extraction on the target video frame feature vector, the target audio feature vector, and the target text feature vector respectively to obtain a target video frame attention vector, a target audio attention vector, and a target text attention vector, and fuse the target video frame attention vector, the target audio attention vector, and the target text attention vector to obtain the target video attention vector.

13. The device according to claim 10, characterized in that, The device further includes: A feature vector determination module, configured to determine a text feature vector of each description text. A feature extraction module, configured to input the text feature vector of the description text into the target text prediction model, and perform attention feature extraction on the text feature vector of the description text through the target text prediction model to obtain a text attention vector of the description text.

14. A training device for a text prediction model, characterized in that, The device includes: An acquisition module, configured to acquire multiple groups of sample pairs, each group of sample pairs including sample video content of a sample video and a sample description text, where the sample video content includes sample video frames, sample audio, and sample text in the sample video, and the sample description text is used to summarize the sample video. A training module, configured to iteratively execute the following steps based on the multiple groups of sample pairs to train a text prediction model to obtain a target text prediction model, where the target text prediction model is used to determine, based on video content, a description text matching the video content from a tag library, and a tag corresponding to the description text is used to determine a tag of a video corresponding to the video content, and the tag library stores multiple description texts and tags respectively corresponding to the multiple description texts: for a sample pair used in the iteration process, determine a video attention vector of the sample video content and a text attention vector of the sample description text in the sample pair through the text prediction model; determine a first loss value based on the video attention vector of the sample video content and the text attention vector of the sample description text, where the first loss value is used to represent a gap between the sample description text and a true description text of the sample video content; and adjust model parameters of the text prediction model based on the first loss value of the sample pair. A first acquisition module, configured to acquire multiple tags. A first text interpretation module, configured to respectively perform text interpretation on the multiple tags to convert the multiple tags into a paragraph of text, thereby obtaining description texts respectively corresponding to the multiple tags. A tag library determination module, configured to obtain the tag library based on the multiple tags and the description texts respectively corresponding to the multiple tags. A second acquisition module, configured to periodically acquire newly added tags. A second text interpretation module, configured to perform text interpretation on the newly added tags to convert the newly added tags into a paragraph of text, thereby obtaining a description text corresponding to the newly added tags. A storage module, configured to store the newly added tags and the description texts corresponding to the newly added tags into the tag library.

15. The device according to claim 14, characterized in that, The training module is configured to: Respectively determine a video frame feature vector of the sample video frame, an audio feature vector of the sample audio, and a text feature vector of the sample text. Input the video frame feature vector, the audio feature vector, and the text feature vector into the text prediction model respectively. Through the text prediction model, perform attention feature extraction on the video frame feature vector, the audio feature vector, and the text feature vector respectively to obtain a video frame attention vector, an audio attention vector, and a text attention vector. Then, fuse the video frame attention vector, the audio attention vector, and the text attention vector to obtain the video attention vector; Determine the text feature vector of the sample description text; Input the text feature vector of the sample description text into the text prediction model. Through the text prediction model, perform attention feature extraction on the text feature vector of the sample description text to obtain the text attention vector of the sample description text.

16. The device according to claim 14, characterized in that, In each iteration process, train the text prediction model based on a target number of sample pairs. The training module is further configured to: For each group of sample pairs in the target number of sample pairs, combine the sample video content in the sample pair with the sample description texts in the remaining sample pairs in the target number of sample pairs respectively to form negative sample pairs, and obtain multiple groups of negative sample pairs; Based on the target number of sample pairs and the multiple groups of negative sample pairs, train the text prediction model. The target number of sample pairs is positive sample pairs.

17. The device according to claim 16, characterized in that, The training module is configured to: Fuse the first loss values of the target number of sample pairs and the multiple groups of negative sample pairs to obtain a second loss value, and based on the second loss value, adjust the model parameters of the text prediction model.

18. The device according to claim 16, characterized in that, The training module is configured to: Determine the sample similarity between the video attention vector of the sample video content and the text attention vector of the sample description text; Determine the sample prediction probability corresponding to the sample similarity. The sample prediction probability is used to represent the matching degree between the sample description text and the sample video content; In the case where the sample pair is a positive sample pair, based on the sample prediction probability and a first reference probability, determine the first loss value. The first reference probability is the reference probability corresponding to the positive sample pair; In the case where the sample pair is a negative sample pair, based on the sample prediction probability and a second reference probability, determine the first loss value. The second reference probability is the reference probability corresponding to the negative sample pair.

19. A computer device, characterized in that, The computer device includes a processor and a memory. The memory is used to store at least one segment of computer program, and the at least one segment of computer program is loaded and executed by the processor to perform the label prediction method according to any one of claims 1 to 4 or the training method of the text prediction model according to any one of claims 5 to 9.

20. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store at least one segment of computer program, and the at least one segment of computer program is used to perform the label prediction method according to any one of claims 1 to 4 or the training method of the text prediction model according to any one of claims 5 to 9.

21. A computer program product, characterized in that, The computer program product includes a computer program which is stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the label prediction method described in any one of claims 1 to 4 or the training method of the text prediction model described in any one of claims 5 to 9.

Citation Information

Patent Citations

  • Data processing method and apparatus, electronic device and storage medium

    CN109522424A

  • Video label obtaining method and device, storage medium and server

    CN111695422A

  • Text classification method and device, text processing method and device, computer equipment and storage medium

    CN114443847A