Label generation model training method, label generation method, device and equipment

By training a label generation model and utilizing the encoding and prediction techniques of sample quadruples and the label generation model, the problem of video recommendation not matching the interests of the viewers is solved, resulting in more efficient video recommendation and an improved viewing experience.

CN115905621BActive Publication Date: 2026-05-01TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2022-11-15
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing technologies, when adding person tags to videos, may recommend person tags that viewers are not interested in, resulting in recommended videos that do not match their interests, thus reducing recommendation efficiency and viewing experience.

Method used

By obtaining sample quadruplets, including sample videos, sample person names, and facial features, a label generation model is used for encoding and prediction. The model is trained to determine suitable video labels, and then combined with a general label model to generate suitable input video labels.

Benefits of technology

It improves the accuracy and efficiency of video recommendations, making the recommended videos more in line with the interests of the viewers and enhancing the viewing experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115905621B_ABST
    Figure CN115905621B_ABST
Patent Text Reader

Abstract

The application provides a label generation model training method, a label generation method, a device and equipment, and belongs to the technical field of Internet. The method comprises the following steps: acquiring a sample quadruple; encoding a sample video based on a label generation model to obtain a sample video feature, wherein the label generation model is used for generating a character label for a video based on the video, a character name of a character in the video and a face feature; predicting the sample video feature, the sample character name and the sample face feature based on the label generation model to obtain a prediction result; and training the label generation model based on the difference between the prediction result and sample supervision information. The above technical solution enables the trained label generation model to determine a label suitable for a video, and then enables the label to be used for video recommendation for a viewing object, so that the recommended video is more in line with the interest of the viewing object, the recommendation efficiency is improved, and the experience of the viewing object is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of Internet technology, and in particular to a training method, a label generation method, an apparatus, and a device for a label generation model. Background Technology

[0002] With the development of multimedia technology, watching videos has become a popular form of entertainment. If a viewer interacts positively with a video, such as liking, saving, or repeatedly watching it, then similar videos can be recommended based on the video's character tags, such as videos featuring the same characters. Therefore, accurately adding character tags to videos is a problem that needs to be solved.

[0003] Currently, the typical approach is to first detect faces in the video, then extract facial features from the detected face regions, match these facial features with a face retrieval database, and finally use the matched person names as tags for the video. The face retrieval database includes multiple sets of facial sample features linked to person names.

[0004] The problem with the above technical solution is that a video may include multiple people. If a person tag is added to the video for a person that the viewer is not interested in, and then recommendations are made to the viewer based on that person tag, the recommended videos will be videos that the viewer is not interested in, resulting in a poor viewing experience and low recommendation efficiency. Summary of the Invention

[0005] This application provides a training method, a tag generation method, an apparatus, and a device for a tag generation model. The trained tag generation model can determine suitable tags for videos, and then recommend videos to viewers based on these tags. This makes the recommended videos more aligned with the viewers' interests, improves recommendation efficiency, and ultimately enhances the viewers' experience. The technical solution is as follows:

[0006] On the one hand, a training method for a label generation model is provided, the method comprising:

[0007] Obtain a sample quadruple, which includes a sample video, the sample person's name in the sample video, the sample face features of the sample person, and sample supervision information. The sample supervision information is used to indicate whether the sample person's name is selected as the person's tag in the sample video.

[0008] The sample video is encoded based on the tag generation model to obtain sample video features. The tag generation model is used to generate character tags for the video based on the video, the names of the people in the video, and facial features.

[0009] Based on the tag generation model, the features of the sample video, the name of the sample person, and the facial features of the sample are predicted to obtain a prediction result. The prediction result is used to represent the probability that the name of the sample person is selected as the person tag of the sample video.

[0010] The label generation model is trained based on the difference between the prediction results and the sample supervision information.

[0011] On the other hand, a tag generation method is provided, the method comprising:

[0012] The input video is processed based on a general labeling model to obtain at least one face region image and at least one general person label. The general labeling model is used to add corresponding person labels to the people in the input video. The general person label is the name of the person corresponding to the face region image.

[0013] Based on the label generation model, the input video, the at least one general person label, and the at least one face region image are predicted to obtain at least one target label probability. The label generation model is trained using the above-mentioned label generation model training method.

[0014] Based on the probability of the at least one target tag, at least one target person tag is determined from the at least one general person tag of the input video.

[0015] On the other hand, a training apparatus for a label generation model is provided, the apparatus comprising:

[0016] The acquisition module is used to acquire sample quadruples, which include sample videos, sample person names of sample people in the sample videos, sample facial features of the sample people, and sample supervision information. The sample supervision information is used to indicate whether the sample person name is selected as a person tag for the sample video.

[0017] The first encoding module is used to encode the sample video based on the tag generation model to obtain sample video features. The tag generation model is used to generate character tags for the video based on the video, the names of the people in the video, and facial features.

[0018] The prediction module is used to predict the features of the sample video, the name of the sample person, and the facial features of the sample based on the tag generation model, and obtain a prediction result. The prediction result is used to represent the probability that the name of the sample person is selected as the person tag of the sample video.

[0019] The training module is used to train the label generation model based on the difference between the prediction results and the sample supervision information.

[0020] In some embodiments, the prediction module includes:

[0021] The first mapping unit is used to map the sample person names to a first person vector;

[0022] The second mapping unit is used to map the sample facial features into a second person vector.

[0023] The first encoding unit is used to perform auto-encoding and cross-encoding on the first person vector, the second person vector, and the sample video features based on the tag generation model to obtain sample person features.

[0024] The prediction unit is used to predict the features of the sample person based on the label generation model, and obtain the prediction result.

[0025] In some embodiments, the first encoding unit is configured to perform self-encoding on the first person vector and the second person vector based on the tag generation model to obtain self-encoded person features; and to perform mutual encoding on the self-encoded person features and the sample video features based on the tag generation model to obtain the sample person features.

[0026] In some embodiments, the first encoding module includes:

[0027] The first acquisition unit is used to acquire multiple sample video frames from the sample video for any sample quadruple;

[0028] The second encoding unit is used to encode the plurality of sample video frames and the video titles of the sample videos based on the tag generation model, so as to obtain video frame features and video title features.

[0029] The splicing unit is used to splice the video frame features and the video title features to obtain the sample video features.

[0030] In some embodiments, the apparatus further includes:

[0031] The second encoding module is used to encode the audio signal of the sample video based on the tag generation model to obtain audio features;

[0032] The splicing unit is used to splice the video frame features, the video title features, and the audio features to obtain the sample video features.

[0033] In some embodiments, the training module includes:

[0034] The second acquisition unit is used to acquire multiple first training losses for multiple sample quadruplets, wherein the first training loss is used to represent the difference between the prediction result of the corresponding sample quadruplet and the sample supervision information.

[0035] A determining unit is used to determine the average value of the plurality of first training losses as a second training loss;

[0036] The update unit is used to update the multiple model parameters based on the derivative of the second training loss with respect to multiple model parameters of the label generation model and the model learning rate.

[0037] In some embodiments, the updating unit is configured to determine an intermediate value as the product of the derivative of the second training loss with respect to multiple model parameters of the label generation model and the model learning rate; and for any model parameter, to determine the difference between the model parameter and the intermediate value as the updated model parameter.

[0038] On the other hand, a label generation apparatus is provided, the apparatus comprising:

[0039] The processing module is used to process the input video based on a general labeling model to obtain at least one face region image and at least one general person label. The general labeling model is used to add corresponding person labels to the people in the input video. The general person label is the name of the person corresponding to the face region image.

[0040] The prediction module is used to predict the input video, the at least one general person label, and the at least one face region image based on the label generation model to obtain at least one target label probability. The label generation model is trained using the training method described above.

[0041] A determination module is configured to determine at least one target person label from the at least one general person label of the input video based on the at least one target label probability.

[0042] In some embodiments, the prediction module is configured to encode the video information of the input video based on the label generation model to obtain target video features; encode the at least one general person label, the at least one face region image, and the target video features based on the label generation model to obtain at least one target person feature; and predict the at least one target label probability based on the label generation model.

[0043] In some embodiments, the determining module is configured to determine, from the at least one general person tag, at least one target person tag whose target tag probability is greater than a probability threshold; or, from the at least one general person tag, determine the target person tag whose target tag probability ranks first.

[0044] On the other hand, a computer device is provided, the computer device including a processor and a memory, the memory being used to store at least one computer program, the at least one computer program being loaded and executed by the processor to implement the training method of the label generation model in the embodiments of this application.

[0045] On the other hand, a computer device is provided, the computer device including a processor and a memory, the memory being used to store at least one computer program, the at least one computer program being loaded and executed by the processor to implement the tag generation method in the embodiments of this application.

[0046] On the other hand, a computer-readable storage medium is provided for storing at least one computer program for executing to implement a training method for a label generation model as described in the embodiments of this application.

[0047] On the other hand, a computer-readable storage medium is provided for storing at least one computer program for executing to implement the tag generation method as described in the embodiments of this application.

[0048] On the other hand, a computer program product is provided, including a computer program that is executed by a processor to implement the training method of the label generation model provided in the embodiments of this application.

[0049] On the other hand, a computer program product is provided, including a computer program that is executed by a processor to implement the label generation method provided in the embodiments of this application.

[0050] This application provides a method for training a tag generation model. The model encodes sample videos to obtain sample video features, enabling these features to represent the overall characteristics of the video. Then, the tag generation model predicts the sample video features, sample person names, and sample facial features, achieving prediction from both the overall video and the individual person's features. This yields the probability that a person's name will be selected as a tag for the video. The difference between the predicted probability and the sample supervision information reflects the accuracy of the tag generation model's predictions. This allows the trained tag generation model to determine suitable tags for videos, enabling video recommendations based on these tags. This makes the recommended videos more aligned with the viewer's interests, improving recommendation efficiency and ultimately enhancing the viewer's experience. Attached Figure Description

[0051] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0052] Figure 1 This is a schematic diagram of the implementation environment for a training method of a label generation model provided in an embodiment of this application;

[0053] Figure 2 This is a flowchart of a training method for a label generation model according to an embodiment of this application;

[0054] Figure 3 This is a flowchart of a label generation method provided according to an embodiment of this application;

[0055] Figure 4 This is a flowchart of another label generation method provided according to an embodiment of this application;

[0056] Figure 5 This is a schematic diagram of a sample quadruple provided according to an embodiment of this application;

[0057] Figure 6 This is a block diagram of a tag generation model provided according to an embodiment of this application;

[0058] Figure 7 This is a flowchart illustrating the training process of a label generation model according to an embodiment of this application;

[0059] Figure 8 This is an application diagram illustrating a general label model and a label generation model provided in the embodiments of this application;

[0060] Figure 9 This is a schematic diagram illustrating a method for recommending videos to viewers based on target person tags, according to an embodiment of this application.

[0061] Figure 10 This is a block diagram of a training device for a label generation model according to an embodiment of this application;

[0062] Figure 11 This is a block diagram of a training apparatus for another label generation model provided according to an embodiment of this application;

[0063] Figure 12 This is a block diagram of a label generation apparatus provided according to an embodiment of this application;

[0064] Figure 13 This is a structural block diagram of a terminal provided according to an embodiment of this application;

[0065] Figure 14 This is a schematic diagram of the structure of a server according to an embodiment of this application. Detailed Implementation

[0066] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0067] In this application, the terms "first," "second," etc., are used to distinguish identical or similar items with essentially the same function. It should be understood that there is no logical or temporal dependency between "first," "second," and "nth," nor are there any restrictions on quantity or execution order.

[0068] In this application, the term "at least one" means one or more, and "multiple" means two or more.

[0069] It should be noted that all information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in this application have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the videos involved in this application were all obtained with full authorization.

[0070] The following is an explanation of the terms used in this application.

[0071] Facial features: By encoding a specific face using a neural network, a facial feature can be obtained. This feature can be stored in a database, and then similarity can be calculated to determine whether a new face belongs to the same person.

[0072] Character Tags: Tags that describe the people appearing in the video; the tag content is the person's name.

[0073] Face detection: A method for detecting whether there is a human face in an image and where the face appears.

[0074] The training method and label generation method for the label generation model provided in this application can be executed by a computer device. In some embodiments, the computer device is a terminal or a server. Figure 1 This is a schematic diagram illustrating the implementation environment of a training method for a label generation model according to an embodiment of this application. See also... Figure 1 The implementation environment includes terminal 101 and server 102. Optionally, the implementation environment of the label generation method is similar to that of the training method of the label generation model, and will not be described again.

[0075] Terminal 101 and server 102 can be connected directly or indirectly via wired or wireless communication, and this application does not impose any restrictions on this.

[0076] In some embodiments, terminal 101 may be a smartphone, tablet, laptop, desktop computer, smart speaker, smart TV, smartwatch, in-vehicle device, etc., but is not limited thereto. Terminal 101 has an application installed and running that supports video playback.

[0077] In some embodiments, server 102 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. Server 102 is used to provide background services for applications that support video playback. In some embodiments, server 102 undertakes the main computing work, and terminal 101 undertakes the secondary computing work; or, server 102 undertakes the secondary computing work, and terminal 101 undertakes the main computing work; or, server 102 and terminal 101 collaborate on computing using a distributed computing architecture.

[0078] Figure 2 This is a flowchart of a training method for a label generation model according to an embodiment of this application, such as... Figure 2 As shown in the illustration, this application embodiment uses server execution as an example for explanation. The method includes the following steps:

[0079] 201. The server obtains the sample quadruple, which includes the sample video, the sample person's name in the sample video, the sample person's facial features, and the sample supervision information. The sample supervision information is used to indicate whether the sample person's name is selected as the person's label in the sample video.

[0080] In this embodiment, the server can obtain sample quadruples as training data to train the label generation model. Each sample quadruple includes a sample video, a sample person's name, sample facial features, and sample supervision information. The sample video may include multiple people; therefore, multiple sample quadruples can be constructed from a single sample video. The sample facial features represent the facial features of the sample person in the sample video. The sample supervision information indicates whether the manually labeled sample person's name was selected as the person's label for the sample video.

[0081] Optionally, the data in this sample quadruple can be obtained using a general labeling model. This general labeling model can perform face detection on the input video, extract facial features from the detected face regions, and output the names of the people corresponding to the facial features as person labels.

[0082] 202. The server encodes the sample video based on the tag generation model to obtain the sample video features. The tag generation model is used to generate person tags for the video based on the video, the names of the people in the video, and facial features.

[0083] In this embodiment, the server trains the label generation model based on the obtained sample quadruples. Correspondingly, the server can encode the sample videos in the sample quadruples based on the label generation model to obtain sample video features. These sample video features can reflect the overall characteristics of the sample videos.

[0084] 203. The server predicts the features of the sample video, the names of the sample people, and the facial features of the sample based on the label generation model, and obtains the prediction results. The prediction results are used to represent the probability that the names of the sample people are selected as the person labels of the sample video.

[0085] In this embodiment, the sample video features are the overall features of the sample video, and the name and facial features of the sample person are the local features of the sample video. The server uses a label generation model to predict the overall and local features of the sample video, obtaining the probability that the sample person's name is the person's label in the sample video, i.e., the prediction result. Since the predicted probability is based on the features of the person in the sample video, it reflects the degree of association between the person and the sample video. Accordingly, for any sample person in the sample video, if the sample person is a main character in the sample video, appears for a long time or has many shots, the probability of predicting the sample person's name is high; if the sample person is a non-main character in the sample video, appears for a short time or has few shots, the probability of predicting the sample person's name is low. By predicting the sample video features, the name of the sample person, and facial features based on a label generation model, the prediction result of the sample person's name as the person's label can be obtained from both the overall features of the sample video and the features of the sample person in the sample video, making the prediction result more accurate.

[0086] 204. The server trains the label generation model based on the difference between the prediction results and the sample supervision information.

[0087] In this embodiment, the prediction result is obtained by the label generation model, while the sample supervision information is obtained through manual annotation. The server can determine the training loss value based on the difference between the prediction result and the sample supervision information. Since the magnitude of this difference reflects the accuracy of the prediction result, the server can train the label generation model based on this difference.

[0088] This application provides a method for training a tag generation model. The model encodes sample videos to obtain sample video features, enabling these features to represent the overall characteristics of the video. Then, the tag generation model predicts the sample video features, sample person names, and sample facial features, achieving prediction from both the overall video and the individual person's features. This yields the probability that a person's name will be selected as a tag for the video. The difference between the predicted probability and the sample supervision information reflects the accuracy of the tag generation model's predictions. This allows the trained tag generation model to determine suitable tags for videos, enabling video recommendations based on these tags. This makes the recommended videos more aligned with the viewer's interests, improving recommendation efficiency and ultimately enhancing the viewer's experience.

[0089] Figure 3This is a flowchart of a label generation method provided according to an embodiment of this application, such as... Figure 3 As shown in the illustration, this application embodiment uses server execution as an example for explanation. The method includes the following steps:

[0090] 301. The server processes the input video based on a general labeling model to obtain at least one face region image and at least one general person label. The general labeling model is used to add corresponding person labels to the people in the input video. The general person label is the name of the person corresponding to the face region image.

[0091] In this embodiment, when adding person tags to the input video, the general tagging model adds corresponding person tags to all people appearing in the video. If there are people in the video that the viewer is not interested in, their tags will also be added to the video. This could lead the server to recommend videos to the viewer based on the person tags of people the viewer is not interested in, resulting in recommended videos that do not match the viewer's interests, leading to low recommendation efficiency and a poor viewer experience. Therefore, the server can combine the general tagging model with the tag generation model trained in steps 201 to 204 and deploy it in a GPU (Graphics Processing Unit) environment to generate suitable person tags for the input video. Accordingly, the server can first process the input video based on the general tagging model to obtain the face region image and corresponding general person tag for each person in the input video. By processing the input video based on the general tagging model to obtain the face region image and general face tag, preparation is made for the tag generation model to generate tags suitable for the input video.

[0092] 302. The server predicts the input video, at least one general person label, and at least one face region image based on the label generation model to obtain at least one target label probability. The label generation model is trained using the training method of the label generation model in steps 201 to 204 above.

[0093] In this embodiment, the server uses a label generation model to predict the face region image and general person labels obtained from the aforementioned general label model, as well as the input video, to obtain the probability that each general person label will be selected as the person label for the input video, i.e., the target label probability. The label generation model is a trained model capable of generating suitable person labels for the video. By predicting the general person labels and face region image based on the label generation model to obtain the target label probability, the probability of each general person label being used as the person label for the input video can be determined, thereby identifying suitable person labels for the input video.

[0094] 303. The server determines at least one target person label for the input video from at least one general person label based on at least one target label probability.

[0095] In this embodiment, the server determines the target person tag for the input video from general person tags based on the target tag probability. The target person tag is the name of a person in the general person tags who has a high degree of relevance to the input video. Accordingly, if the input video is an interview video, the target person tag can be the name of the interviewee; if the input video is a cover song video, the target person tag can be the name of the cover singer; if the input video is a ranking video, the target person tag can be the name of the highest-ranking person. By determining the target person tag based on the target tag probability, the determined target person tag is suitable as the person tag for the input video, thereby enabling the recommendation of videos that match the viewer's interests based on the target person tag.

[0096] This application provides a tag generation method. Since a general tag model adds corresponding tags to all people in the input video, it may recommend videos of non-interesting characters to the viewer, making the viewer uninterested in the recommended videos. Therefore, a trained tag generation model can be used to encode and predict the general character tags and face region images obtained from the general tag model, determine the target character tags suitable as tags for the input video from the general character tags, and recommend the target videos with the target character tags to the viewer, so that the viewer has a better viewing experience and the recommendation effect is good.

[0097] Figure 4 This is a flowchart of another label generation method provided according to an embodiment of this application, such as... Figure 4 As shown in the illustration, this application embodiment uses server execution as an example for explanation. The method includes the following steps:

[0098] 401. The server obtains the sample quadruple, which includes the sample video, the sample person's name in the sample video, the sample person's facial features, and the sample supervision information. The sample supervision information is used to indicate whether the sample person's name is selected as the person's label in the sample video.

[0099] In this embodiment, the server obtains sample quadruples to train the sample generation model. The sample supervision information in the sample quadruples represents the result of manual annotation based on label selection rules, indicating whether the names of the sample characters are selected as character tags for the sample video (selected or not). These label selection rules are preset criteria for selecting character tags. If any character in the sample video meets these rules, the manual annotation process selects that character's name as a character tag for the sample video. It should be noted that these label selection rules can be customized according to actual application scenarios, providing a basis for manual annotation of character tags in sample videos. By obtaining the sample quadruples, the label generation model can be trained based on them.

[0100] For example, the process of obtaining sample supervision information is illustrated using the following six label screening rules.

[0101] Firstly, if a sample person appears throughout the entire sample video, their name will be used as the sample video's character tag, i.e., the sample supervision information will be selected. If the sample video is a video narration video, and its main content describes any sample person, then that sample person's name will be used as the sample video's character tag, i.e., the sample supervision information will be selected. If the sample video is an interview video, then the interviewee's name will be used as the sample video's character tag, i.e., the sample supervision information will be selected. If the sample video compares two types of content, both types of content are the main content of the sample video, and the names of the sample people in both types of content will be used as the sample video's character tags, i.e., the sample supervision information will be selected. If the sample video is a cover song video, then the name of the sample person performing the cover song will be used as the sample video's character tag, i.e., the sample supervision information will be selected, but the sample person being covered is not the focus of the sample video, i.e., not used as a character tag, and the sample supervision information will not be selected.

[0102] Article 2: If the sample video is a ranking video, the name of the highest-ranking sample person will be used as the person tag of the sample video, that is, the sample supervision information will be selected, and the remaining sample people will not be used as the person tags of the sample video, and the sample supervision information will be not selected.

[0103] Thirdly, if a person has been manually tagged with at least one person according to the first two rules, no further tagging will be performed; otherwise, the next rule will be executed.

[0104] Article 4. In the sample video, if the video frames containing the sample person account for 50% or more of all sample video frames, or if the text information containing the sample person accounts for 50% or more of all text information, then the name of the sample person shall be used as the person tag of the sample video, that is, the sample supervision information shall be selected.

[0105] Article 5. If the title of the sample video includes the name of the sample person, and the sample person appears in the sample video for more than 5 seconds, then the name of the sample person will be used as the person tag of the sample video, that is, the sample supervision information will be selected.

[0106] Article 6: After manual labeling according to the above five labeling rules, no further manual labeling will be performed.

[0107] In some embodiments, since the data in the sample quadruple can be obtained through a general labeling model, the general labeling model can extract facial features from the detected face regions and use the average value of the extracted facial features as the sample facial features of the sample person in the sample quadruple.

[0108] It should be noted that the general labeling model can also obtain person labels based on text. That is, the title of the sample video is segmented into words, and the names of the people in the segmented titles are matched. The matched names are then used as the person labels for the sample video. If the sample quadruple is obtained by the general labeling model based on text, meaning that the sample person's name is only detected in the title of the sample video, and no face region is detected in the sample video using the above method, then the sample face feature in the sample quadruple is a zero vector.

[0109] For example, Figure 5 This is a schematic diagram of a sample quadruple provided according to an embodiment of this application. For example... Figure 5 As shown, the sample video includes three sample individuals, thus yielding three sample quadruples. The first sample quadruple is the sample video, individual A, and the corresponding sample facial features; individual A was manually selected as the person label for the sample video. The second sample quadruple is the sample video, individual B, and the corresponding sample facial features; individual B was not manually selected as the person label for the sample video. The third sample quadruple is the sample video, individual C, and the corresponding sample facial features; individual C was manually selected as the person label for the sample video.

[0110] 402. The server encodes the sample video based on the tag generation model to obtain the sample video features. The tag generation model is used to generate person tags for the video based on the video, the names of the people in the video, and facial features.

[0111] In this embodiment, the server encodes the sample video based on a tag generation model, that is, it encodes multiple aspects of information in the sample video, such as image information, sound information, and text information, to obtain sample video features that can represent the overall characteristics of the sample. Optionally, the process of encoding the sample video can be implemented by a video content encoder in the tag generation model. By encoding the sample video based on the tag generation model, the overall characteristics of the sample video can be obtained.

[0112] In some embodiments, the server can encode sample video frames and video titles using a tag generation model to obtain video frame features and video title features, which are then concatenated to obtain sample video features. Correspondingly, for any sample quadruple, the server obtains multiple sample video frames from the sample video. Then, based on the tag generation model, the server encodes the multiple sample video frames and the video title of the sample video respectively to obtain video frame features and video title features. The server then concatenates the video frame features and video title features to obtain the sample video features. The sample video frames are obtained by the server uniformly sampling frames from the sample video; they can be continuous or discontinuous video frames, and this embodiment does not limit the number of sample video frames.

[0113] Optionally, the server can also encode the audio signal of the sample video based on the video content encoder in the label generation model to obtain audio features, and then concatenate the video frame features, video title features, and the audio features to obtain the sample video features. Since the audio signal of the sample video may contain relevant information about the sample person, the degree of association between the sample person and the sample video can be better determined, thereby enabling more accurate training of the model.

[0114] For example, suppose the server obtains eight non-contiguous sample video frames from a sample video, and inputs them, along with the sample video's title and audio signal, into a tag generation model for encoding by a video content encoder. Specifically, encoding a sample video frame yields a 1024-dimensional vector, encoding a single character in the title yields a 1024-dimensional vector, and encoding a single sample point in the audio signal yields a 1024-dimensional vector. The server concatenates the encoded results of the sample video frames, title, and audio signal to obtain the sample video feature, which is an M*1024 matrix. Here, M is a positive integer representing the sum of the number of sample video frames, characters in the title, and sample points in the audio signal. The threshold for this sample video feature is 500 1024-dimensional vectors, meaning the maximum value of M is 500.

[0115] 403. The server maps the sample person's name to a first person vector and the sample person's facial features to a second person vector.

[0116] In some embodiments, the server can map sample person names to a first person vector using an embedding layer in the label generation model, and map sample facial features to a second person vector using a linear layer in the label generation model. Optionally, the first person vector and the second person vector are both 1024-dimensional vectors. By mapping sample person names and sample facial features, the label generation model can be trained based on the obtained person vectors.

[0117] 404. The server performs auto-encoding and cross-encoding on the first person vector, the second person vector, and the sample video features based on the tag generation model to obtain the sample person features.

[0118] In this embodiment, since the first person vector represents the name of the sample person and the second person vector represents the facial features of the sample person, the server encodes the first and second person vectors to obtain the features of the sample person. Then, the server encodes the features of the sample person and the features of the sample video to obtain the features of the sample person in the sample video, i.e., the sample person features. Optionally, the autoencoding process of the first and second person vectors can be implemented by a person autoencoder in the label generation model, and the mutual encoding process of the autoencoder's encoding result and the sample video features can be implemented by a person mutual encoder in the label generation model. By obtaining the sample person features based on the sample generation model, the local features of the sample video can be obtained, i.e., the features exhibited by the sample person in the sample video can be obtained.

[0119] In some embodiments, taking a single autoencoding and inter-encoding process as an example, the server first autoencodes the first and second person vectors based on the person autoencoder in the label generation model, and then inter-encodes the autoencoded result and sample video features based on the person inter-encoder to obtain sample person features. Correspondingly, the server autoencodes the first and second person vectors based on the label generation model to obtain autoencoded person features. Then, the server inter-encodes the autoencoded person features and sample video features based on the label generation model to obtain sample person features. By autoencoding and inter-encoding the first person vector, the second person vector, and the sample video features, the features of the sample person in the sample video can be obtained, thereby enabling the determination of suitable person labels for the sample video.

[0120] It should be noted that the process of autoencoding and cross-encoding the first person vector, the second person vector, and the sample video features based on the label generation model can be repeated multiple times, for example, 5, 10, or 15 times. This embodiment does not limit the number of repetitions of the autoencoding and cross-encoding process. Accordingly, the server inputs the cross-encoding result of the previous group into the person autoencoder to obtain the autoencoded person features. Then, the server inputs the autoencoded person features into the person cross-encoder to obtain the cross-encoded person features. Then, the server inputs the cross-encoded person features into the next group of person autoencoders, and repeats the autoencoding and cross-encoding process again. Through multiple autoencoding and cross-encoding processes, the label generation model can learn simultaneously through self-learning and combining it with sample videos, improving the model training effect.

[0121] For example, the server inputs the first and second person vectors into the person autoencoder. If the sum of the first and second person vectors is less than 100, the server pads the missing parts with zero vectors, resulting in a 100*1024 matrix, which is then input into the person autoencoder. The output of the person autoencoder is the autoencoded person feature, which is a 100*1024 dimensional feature. Then, the server inputs the autoencoded person feature and the sample video feature into the person cross-encoder, outputting a 100*1024 dimensional feature. The autoencoded person feature is a 100*1024 dimensional feature, and the sample video feature is a vector of up to 500*1024 dimensions.

[0122] 405. The server predicts the features of the sample person based on the tag generation model and obtains the prediction result. The prediction result is used to represent the probability that the sample person's name is selected as the person tag of the sample video.

[0123] In this embodiment, the server obtains the probability, or prediction result, of a sample person's name as a person's label in a sample video by analyzing the sample person's features using a label generation model. Optionally, the process of predicting the sample person's features can be implemented by the predictor in the label generation model. These sample person features reflect the characteristics exhibited by the sample person in the sample video, and the predicted probability is based on these features, reflecting the degree of association between the person and the sample video. By predicting the sample person's features using a label generation model to obtain the prediction result of the sample person's name as a person's label, predictions can be made based on the characteristics exhibited by the person in the video, making the prediction results more accurate.

[0124] For example, the predictor in a tag generation model takes sample character features as input and outputs a prediction result. The sample character feature is a 1024-dimensional feature. The prediction result is the probability within a preset range [0, 1]. Assuming the server inputs the sample character features of the main character in the sample video into the predictor, the prediction result is a selection probability of 0.8 and a non-selection probability of 0.2. Similarly, assuming the server inputs the sample character features of non-main characters in the sample video into the predictor, the prediction result is also a selection probability of 0.2 and a non-selection probability of 0.8.

[0125] 406. The server trains the label generation model based on the difference between the prediction results and the sample supervision information.

[0126] In this embodiment, the prediction result is obtained by the label generation model based on the features of the person in the sample video. The sample supervision information is manually labeled according to label selection rules. The server can determine the training loss based on the difference between the prediction result and the sample supervision information. Since this difference reflects the accuracy of the prediction result, the server can train the label generation model based on this difference.

[0127] In some embodiments, the server obtains a second training loss for the label generation model and updates the model parameters based on the derivative of the second training loss with respect to the model parameters and the model learning rate. Correspondingly, the server obtains multiple first training losses for multiple sample quadruples. Then, the server determines the average of the multiple first training losses as the second training loss. The server then updates the multiple model parameters based on the derivative of the second training loss with respect to the multiple model parameters of the label generation model and the model learning rate. Here, the first training loss is used to represent the difference between the prediction result of the corresponding sample quadruple and the sample supervision information. By updating the model parameters based on the derivative of the second training loss with respect to the model parameters and the model learning rate, the label generation model is trained.

[0128] In some embodiments, for any sample quadruple, if the sample supervision information is to select the name of the person in the sample as the person label of the sample video, the server can determine the first training loss of the sample quadruple by the following formula (1).

[0129] x1=log(P) (1)

[0130] Where x1 represents the first training loss; P represents the prediction result.

[0131] In some embodiments, for any sample quadruple, if the sample supervision information is that the name of the person in the sample is not selected as the person label of the sample video, the server can determine the first training loss of the sample quadruple by the following formula (2).

[0132] x2=-log(P) (2)

[0133] Where x2 represents the first training loss; P represents the prediction result.

[0134] In some embodiments, the server determines an intermediate value and then obtains updated model parameters based on the model parameters and the intermediate value. Accordingly, the server determines the intermediate value as the product of the derivative of the second training loss with respect to multiple model parameters of the label generation model and the model learning rate. Then, for any model parameter, the difference between the model parameter and the intermediate value is determined as the updated model parameter.

[0135] In some embodiments, the server can determine the updated model parameters using the following formula (3).

[0136] y = ab * c (3)

[0137] Where y represents the updated model parameters; a represents the current model parameters; b represents the derivative of the second training loss with respect to the model parameters; and c represents the model learning rate, which can be 0.1, 0.2, or 0.3, and this embodiment does not limit this.

[0138] For example, Figure 6 This is a block diagram of a tag generation model provided according to an embodiment of this application. For example... Figure 6 As shown, the sample video is input into the video content encoder, and person A and their corresponding facial features are input into the person autoencoder. Then, the autoencoded person features A1 obtained from the person autoencoder and the sample video features A2 obtained from the video content encoder are input into the person cross-encoder. After N sets of encoding processes by the person autoencoder and the person cross-encoder, the sample person features A3 are obtained. Then, the sample person features A3 are input into the predictor to obtain the prediction result A4. Finally, the first training loss A6 is obtained by comparing the prediction result A4 with the sample supervision information A5. The processing of other sample quadruples is the same as described above and will not be repeated.

[0139] It should be noted that when the number of training iterations of the label generation model reaches the preset number, or when the second training loss is less than the set threshold, the training of the label generation model ends, and the server will save the updated model parameters. Otherwise, the server will obtain new sample quadruplets and continue to execute the above steps 401 to 406 to train the label generation model.

[0140] For example, Figure 7 This is a flowchart of the training process for a tag generation model according to an embodiment of this application. 701. Obtain sample quadruples. 702. Obtain the probability that a sample person's name is a tag for a sample video using the video content encoder, person autoencoder, person cross-encoder, and predictor in the tag generation model. 703. Obtain the second training loss of the tag generation model based on the first training loss of multiple sets of sample quadruples. 704. Update the model parameters of the tag generation model based on the derivative of the second training loss with respect to the model parameters and the model learning rate. 705. Determine whether the model training iterations have reached the preset number. If the tag generation model has reached the preset number of training iterations, proceed to 706; otherwise, proceed to 701. 706. Save the updated model parameters of the tag generation model.

[0141] 407. The server processes the input video based on a general labeling model to obtain at least one face region image and at least one general person label. The general labeling model is used to add corresponding person labels to the people in the input video. The general person label is the name of the person corresponding to the face region image.

[0142] In this embodiment, the server performs face detection on the input video using a general labeling model to obtain at least one face region image. Then, the server extracts facial features from the at least one face region image using the general labeling model. Next, the server performs feature matching on the extracted facial features and uses the names of the people corresponding to the matched facial features as general person tags. Since the general labeling model adds corresponding general person tags to all people appearing in the input video, the server can determine suitable tags for the input video from the general person tags using a trained tag generation model.

[0143] 408. The server predicts the input video, at least one general person label, and at least one face region image based on the label generation model to obtain at least one target label probability. The label generation model is trained using the training method of the label generation model in steps 401 to 406 above.

[0144] In this embodiment, the server processes the generic person labels and face region images obtained from the generic label generation model using a trained label generation model to determine the probability of each generic person label being used as a label for the input video. Correspondingly, the server encodes the video information of the input video based on the label generation model to obtain target video features. Then, the server encodes at least one generic person label, at least one face region image, and target video features based on the label generation model to obtain at least one target person feature. Then, the server predicts the at least one target person feature based on the label generation model to obtain at least one target label probability. By predicting the target label probability based on the generic person labels and face region images using the label generation model, the probability of each generic person label being used as a person label for the input video can be determined, thereby enabling the identification of suitable person labels for the input video.

[0145] 409. The server determines at least one target person label from at least one general person label of the input video based on at least one target label probability.

[0146] In this embodiment, the server determines the target person tag for the input video from general person tags based on the target tag probability. The target person tag is the name of a person in the general person tags who has a high degree of relevance to the input video. By determining the target person tag based on the target tag probability, the determined target person tag is suitable as the person tag for the input video. Compared to the general tag model that adds a corresponding tag to every person, this method enables tag addition based on video information and the characteristics of the person in the video, thereby achieving personalized tag generation.

[0147] In some embodiments, the server may determine the target person tag in the following two ways.

[0148] Method 1: If the server does not limit the number of target person tags, then the server can determine at least one target person tag whose probability is greater than the probability threshold from at least one general person tag.

[0149] Method 2: If the server limits the number of target person tags, then the server can determine the target person tag with the highest probability ranking from at least one general person tag.

[0150] For example, Figure 8 This is an application diagram illustrating a general label model and label generation model provided according to embodiments of this application. For example... Figure 8As shown, a combination of a general labeling model and a label generation model is used to generate target person labels. First, the general labeling model processes the input video to obtain at least one general person label and at least one face region image. Then, the general person label and face region image are input into a person autoencoder, and the input video is input into a video content encoder. Next, the encoding results from the person autoencoder and the video content encoder are input into a person cross-encoder. Then, through autoencoding and cross-encoding of multiple sets of person autoencoders and person cross-encoders, sample person features are obtained. Finally, the sample person features are input into a predictor to obtain prediction results. Finally, the target person label is determined based on the prediction results.

[0151] It should be noted that steps 401 to 409 above are illustrated by example using the addition of person tags to a video. Optionally, the server can also add corresponding scene tags, time tags, animal tags, or item tags to the video based on the scenes, times, animals, or objects that appear in the video.

[0152] For example, for any given video, suppose it contains multiple animals, which can be categorized as birds, cats, and dogs. The server first obtains the animal tags and images of the animals in the video using a general tagging model. The animal tags are birds, cats, and dogs, respectively. Then, the server inputs the animal tags, animal images, and the video into a tag generation model. If birds are not the main animals in the video, the tag generation model outputs that the animal tag "bird" has a low probability of being selected; if animals categorized as cats and dogs are the main animals in the video, the tag generation model outputs that the animal tags "cats" and "dogs" have a high probability of being selected, and thus, cats and dogs are chosen as the animal tags for the video.

[0153] For example, Figure 9 This is a schematic diagram illustrating a method for recommending videos to viewers based on target person tags, according to an embodiment of this application. For example... Figure 9 As shown, if a viewer interacts positively with video 1, such as liking, saving, or repeatedly watching it, the server can input video 1 into the general labeling model and the label generation model to obtain the predicted labels for the two characters in video 1, ultimately determining the target character label for video 1 as character D. Then, the server can recommend video 2, which has the same target character label as video 1, to the viewer.

[0154] This application provides a tag generation method. Since a general tag model adds corresponding character tags to all people in the input video, it may recommend videos of non-interesting characters to the viewer, making the viewer uninterested in the recommended videos. Therefore, a trained tag generation model can be used to encode and predict the general character tags and face region images obtained from the general tag model, determine the target character tags suitable as tags for the input video from the general character tags, and recommend the target videos with the target character tags to the viewer, so that the viewer has a better viewing experience and the recommendation effect is good.

[0155] Figure 10 This is a block diagram of a training apparatus for a label generation model according to an embodiment of this application. The apparatus is used to perform the steps of the above-described method, see [link to relevant documentation]. Figure 10 The device includes: an acquisition module 1001, a first encoding module 1002, a prediction module 1003, and a training module 1004.

[0156] The acquisition module 1001 is used to acquire sample quadruplets. The sample quadruplets include sample video, sample person name of sample person in sample video, sample facial features of sample person and sample supervision information. The sample supervision information is used to indicate whether the sample person name is selected as person label of sample video.

[0157] The first encoding module 1002 is used to encode the sample video based on the label generation model to obtain the sample video features. The label generation model is used to generate person labels for the video based on the video, the names of the people in the video, and facial features.

[0158] The prediction module 1003 is used to predict the features of the sample video, the names of the sample people and the features of the sample face based on the label generation model, and to obtain the prediction result. The prediction result is used to represent the probability that the names of the sample people are selected as the person labels of the sample video.

[0159] Training module 1004 is used to train the label generation model based on the difference between the prediction results and sample supervision information.

[0160] In some embodiments, Figure 11 This is a block diagram of a training apparatus for another label generation model provided according to an embodiment of this application. See also Figure 11 As shown.

[0161] In some embodiments, the prediction module 1003 includes:

[0162] The first mapping unit 1101 is used to map the sample person name to the first person vector;

[0163] The second mapping unit 1102 is used to map the sample face features into a second person vector;

[0164] The first coding unit 1103 is used to perform auto-encoding and cross-encoding on the first person vector, the second person vector and the sample video features based on the label generation model to obtain the sample person features.

[0165] Prediction unit 1104 is used to predict the features of the sample person based on the label generation model and obtain the prediction result.

[0166] In some embodiments, the first encoding unit 1103 is used to autoencode the first person vector and the second person vector based on the tag generation model to obtain autoencoded person features; and to inter-encode the autoencoded person features and the sample video features based on the tag generation model to obtain sample person features.

[0167] In some embodiments, the first encoding module 1002 includes:

[0168] The first acquisition unit 1105 is used to acquire multiple sample video frames from the sample video for any sample quadruple;

[0169] The second encoding unit 1106 is used to encode multiple sample video frames and the video titles of sample videos based on the label generation model to obtain video frame features and video title features.

[0170] The splicing unit 1107 is used to splice video frame features and video title features to obtain sample video features.

[0171] In some embodiments, see Figure 11 As shown, the device also includes:

[0172] The second encoding module 1005 is used to encode the audio signal of the sample video based on the label generation model to obtain audio features;

[0173] The splicing unit 1107 is used to splice video frame features, video title features, and audio features to obtain sample video features.

[0174] In some embodiments, the training module 1004 includes:

[0175] The second acquisition unit 1108 is used to acquire multiple first training losses of multiple sample quadruplets, where the first training loss is used to represent the difference between the prediction result of the corresponding sample quadruplet and the sample supervision information.

[0176] The determining unit 1109 is used to determine the average of multiple first training losses as the second training loss;

[0177] The update unit 1110 is used to update multiple model parameters based on the derivative of the second training loss with respect to multiple model parameters of the label generation model and the model learning rate.

[0178] In some embodiments, the updating unit 1110 is used to determine the intermediate value as the product of the derivative of the second training loss with respect to multiple model parameters of the label generation model and the model learning rate; and for any model parameter, the difference between the model parameter and the intermediate value is determined as the updated model parameter.

[0179] This application provides a training apparatus for a tag generation model. The model encodes sample videos to obtain sample video features, enabling these features to represent the overall characteristics of the video. Then, the tag generation model predicts the sample video features, sample person names, and sample facial features, achieving prediction from both the overall video and the individual person's features. This yields the probability that a person's name will be selected as a tag for the video. The difference between the predicted probability and the sample supervision information reflects the accuracy of the tag generation model's predictions. This allows the trained tag generation model to determine suitable tags for videos, enabling video recommendations based on these tags. This makes the recommended videos more aligned with the viewer's interests, improving recommendation efficiency and ultimately enhancing the viewer's experience.

[0180] It should be noted that the label generation model training device provided in the above embodiments is only illustrated by the division of the above functional modules when running the application. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the label generation model training device and the label generation model training method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0181] Figure 12 This is a block diagram of a label generation apparatus according to an embodiment of this application. The apparatus is used to perform the steps of the method described above, see [link to relevant documentation]. Figure 12 The device includes: a processing module 1201, a prediction module 1202, and a determination module 1203.

[0182] Processing module 1201 is used to process the input video based on a general labeling model to obtain at least one face region image and at least one general person label. The general labeling model is used to add corresponding person labels to the people in the input video. The general person label is the name of the person corresponding to the face region image.

[0183] The prediction module 1202 is used to predict the input video, at least one general person label and at least one face region image based on the label generation model to obtain at least one target label probability. The label generation model is trained by the label generation model training method.

[0184] The determination module 1203 is used to determine at least one target person label from at least one general person label of the input video based on at least one target label probability.

[0185] In some embodiments, the prediction module 1202 is used to encode the video information of the input video based on the label generation model to obtain target video features; to encode at least one general person label, at least one face region image and target video features based on the label generation model to obtain at least one target person feature; and to predict at least one target label probability based on the label generation model.

[0186] In some embodiments, the determining module 1203 is configured to determine at least one target person tag from at least one general person tag whose target tag probability is greater than a probability threshold; or, from at least one general person tag, determine the target person tag whose target tag probability ranks first.

[0187] This application provides a tag generation device. Since a general tag model adds corresponding tags to all people in the input video, it may recommend videos of non-interesting characters to the viewer, making the viewer uninterested in the recommended videos. Therefore, a trained tag generation model can be used to encode and predict the general character tags and face region images obtained from the general tag model, determine the target character tags suitable as tags for the input video from the general character tags, and recommend the target videos with the target character tags to the viewer, so that the viewer has a better viewing experience and the recommendation effect is good.

[0188] It should be noted that the label generation device provided in the above embodiments is only illustrated by the division of the above functional modules when the application is running. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the label generation device and the label generation method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0189] In the embodiments of this application, the computer device can be configured as a terminal or a server. When the computer device is configured as a terminal, the terminal can act as the execution subject to implement the technical solutions provided in the embodiments of this application. When the computer device is configured as a server, the server can act as the execution subject to implement the technical solutions provided in the embodiments of this application. Alternatively, the technical solutions provided in this application can be implemented through the interaction between the terminal and the server. The embodiments of this application do not limit this.

[0190] Figure 13 This is a structural block diagram of a terminal 1300 provided according to an embodiment of this application.

[0191] Typically, terminal 1300 includes a processor 1301 and a memory 1302.

[0192] Processor 1301 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 1301 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 1301 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 1301 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content required to be displayed on the screen. In some embodiments, processor 1301 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.

[0193] The memory 1302 may include one or more computer-readable storage media, which may be non-transitory. The memory 1302 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1302 are used to store at least one computer program, which is executed by the processor 1301 to implement the training method or label generation method of the label generation model provided in the method embodiments of this application.

[0194] In some embodiments, the terminal 1300 may also optionally include a peripheral device interface 1303 and at least one peripheral device. The processor 1301, memory 1302, and peripheral device interface 1303 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 1303 via a bus, signal line, or circuit board. Specifically, the peripheral device includes at least one of the following: a radio frequency circuit 1304, a display screen 1305, a camera assembly 1306, an audio circuit 1307, and a power supply 1308.

[0195] Peripheral device interface 1303 can be used to connect at least one I / O (Input / Output) related peripheral device to processor 1301 and memory 1302. In some embodiments, processor 1301, memory 1302 and peripheral device interface 1303 are integrated on the same chip or circuit board; in some other embodiments, any one or two of processor 1301, memory 1302 and peripheral device interface 1303 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.

[0196] The radio frequency (RF) circuit 1304 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 1304 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 1304 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. In some embodiments, the RF circuit 1304 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a programming chipset, a user identity module card, etc. The RF circuit 1304 can communicate with other terminals through at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to: the World Wide Web, metropolitan area networks, intranets, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 1304 may also include circuitry related to NFC (Near Field Communication), which is not limited in this application.

[0197] Display screen 1305 is used to display a UI (User Interface). This UI may include graphics, text, icons, videos, and any combination thereof. When display screen 1305 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 1301 for processing. In this case, display screen 1305 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 1305, disposed on the front panel of terminal 1300; in other embodiments, there may be at least two display screens, disposed on different surfaces of terminal 1300 or in a folded design; in still other embodiments, display screen 1305 may be a flexible display screen, disposed on a curved or folded surface of terminal 1300. Furthermore, display screen 1305 may be configured as a non-rectangular, irregular shape, i.e., a non-rectangular screen. The display screen 1305 can be made of materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).

[0198] The camera assembly 1306 is used to acquire images or videos. In some embodiments, the camera assembly 1306 includes a front-facing camera and a rear-facing camera. Typically, the front-facing camera is located on the front panel of the terminal, and the rear-facing camera is located on the back of the terminal. In some embodiments, there are at least two rear-facing cameras, which are any one of a main camera, a depth-sensing camera, a wide-angle camera, and a telephoto camera, to achieve background blurring by fusion of the main camera and the depth-sensing camera, panoramic shooting by fusion of the main camera and the wide-angle camera, VR (Virtual Reality) shooting, or other fusion shooting functions. In some embodiments, the camera assembly 1306 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm-light flash and a cool-light flash, which can be used for light compensation at different color temperatures.

[0199] The audio circuit 1307 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, converting the sound waves into electrical signals that are input to the processor 1301 for processing, or input to the radio frequency circuit 1304 for voice communication. For stereo sound acquisition or noise reduction purposes, multiple microphones may be used, each located at a different part of the terminal 1300. The microphone may also be an array microphone or an omnidirectional microphone. The speaker is used to convert electrical signals from the processor 1301 or the radio frequency circuit 1304 into sound waves. The speaker may be a conventional diaphragm speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can convert electrical signals not only into audible sound waves but also into inaudible sound waves for purposes such as distance measurement. In some embodiments, the audio circuit 1307 may also include a headphone jack.

[0200] Power supply 1308 is used to power the various components in terminal 1300. Power supply 1308 can be AC ​​power, DC power, a disposable battery, or a rechargeable battery. When power supply 1308 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery that is charged via a wired line, and a wireless rechargeable battery is a battery that is charged via a wireless coil. The rechargeable battery can also be used to support fast charging technology.

[0201] In some embodiments, the terminal 1300 further includes one or more sensors 1309. The one or more sensors 1309 include, but are not limited to: an acceleration sensor 1310, a gyroscope sensor 1311, a pressure sensor 1312, an optical sensor 1313, and a proximity sensor 1314.

[0202] Accelerometer 1310 can detect the magnitude of acceleration along the three coordinate axes of a coordinate system established by terminal 1300. For example, accelerometer 1310 can be used to detect the components of gravitational acceleration along the three coordinate axes. Processor 1301 can control display screen 1305 to display the user interface in either a landscape or portrait view based on the gravitational acceleration signal acquired by accelerometer 1310. Accelerometer 1310 can also be used for games or for acquiring user motion data.

[0203] The gyroscope sensor 1311 can detect the orientation and rotation angle of the terminal 1300. The gyroscope sensor 1311 can work in conjunction with the accelerometer sensor 1310 to acquire the user's 3D movements on the terminal 1300. Based on the data acquired by the gyroscope sensor 1311, the processor 1301 can perform the following functions: motion sensing (e.g., changing the UI based on the user's tilt), image stabilization during shooting, game control, and inertial navigation.

[0204] The pressure sensor 1312 can be disposed on the side bezel of the terminal 1300 and / or on the lower layer of the display screen 1305. When the pressure sensor 1312 is disposed on the side bezel of the terminal 1300, it can detect the user's grip signal on the terminal 1300, and the processor 1301 can generate or perform quick operations based on the grip signal collected by the pressure sensor 1312. When the pressure sensor 1312 is disposed on the lower layer of the display screen 1305, the processor 1301 can control the operable controls on the UI interface based on the user's pressure operation on the display screen 1305. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.

[0205] Optical sensor 1313 is used to collect ambient light intensity. In one embodiment, processor 1301 can control the display brightness of display screen 1305 based on the ambient light intensity collected by optical sensor 1313. Specifically, when the ambient light intensity is high, the display brightness of display screen 1305 is increased; when the ambient light intensity is low, the display brightness of display screen 1305 is decreased. In another embodiment, processor 1301 can also dynamically adjust the shooting parameters of camera assembly 1306 based on the ambient light intensity collected by optical sensor 1313.

[0206] The proximity sensor 1314, also known as a distance sensor, is typically located on the front panel of the terminal 1300. The proximity sensor 1314 is used to detect the distance between the user and the front of the terminal 1300. In one embodiment, when the proximity sensor 1314 detects that the distance between the user and the front of the terminal 1300 is gradually decreasing, the processor 1301 controls the display screen 1305 to switch from a screen-on state to a screen-off state; when the proximity sensor 1314 detects that the distance between the user and the front of the terminal 1300 is gradually increasing, the processor 1301 controls the display screen 1305 to switch from a screen-off state to a screen-on state.

[0207] Those skilled in the art will understand that Figure 13 The structure shown does not constitute a limitation on terminal 1300 and may include more or fewer components than shown, or combine certain components, or use different component arrangements.

[0208] Figure 14This is a schematic diagram of a server structure according to an embodiment of this application. The server 1400 can vary significantly due to different configurations or performance. It may include one or more Central Processing Units (CPUs) 1401 and one or more memories 1402. The memory 1402 stores at least one computer program, which is loaded and executed by the processor 1401 to implement the training method or label generation method of the label generation model provided in the above-described method embodiments. Of course, the server may also have wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. The server may also include other components for implementing device functions, which will not be elaborated here.

[0209] This application also provides a computer-readable storage medium storing at least one computer program. This computer program is loaded and executed by a processor of a computer device to implement the operations performed by the computer device in the training method or label generation method of the label generation model described above. For example, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, or optical data storage device, etc.

[0210] In some embodiments, the computer program involved in the present application embodiments may be deployed and executed on a computer device, or executed on multiple computer devices located in one location, or executed on multiple computer devices distributed in multiple locations and interconnected through a communication network. Multiple computer devices distributed in multiple locations and interconnected through a communication network may constitute a blockchain system.

[0211] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the training method or label generation method of the label generation model provided in the various optional implementations described above.

[0212] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0213] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A training method for a label generation model, characterized in that, The method includes: Obtain a sample quadruple, which includes a sample video, a sample person's name in the sample video, a sample person's facial features, and sample supervision information. The sample supervision information indicates whether the sample person's name is selected as the person's label for the sample video. The sample person's name and the sample facial features reflect the local features of the sample video. The sample person's name is obtained by performing face detection on the sample video and extracting facial features from the detected face regions. It is the person's label corresponding to the extracted facial features. The sample person's name is the general person's label for the sample video. The sample video is encoded based on a tag generation model to obtain sample video features. The tag generation model is used to generate character tags for the video based on the video, the names of the people in the video, and facial features. The sample video features reflect the overall features of the sample video. Based on the tag generation model, the features of the sample video, the name of the sample person, and the facial features of the sample are predicted to obtain a prediction result. The prediction result is used to represent the probability that the name of the sample person is selected as the person tag of the sample video. The probability reflects the degree of association between the corresponding sample person and the sample person. The greater the degree of association, the greater the probability. The label generation model is trained based on the difference between the prediction results and the sample supervision information; Based on the trained tag generation model, target person tags suitable as video tags are determined from the general person tags of the video.

2. The method according to claim 1, characterized in that, The prediction based on the tag generation model of the sample video features, the sample person names, and the sample facial features, to obtain the prediction result, includes: Map the sample person names to a first person vector; The sample facial features are mapped to a second person vector; Based on the tag generation model, the first person vector, the second person vector, and the sample video features are auto-encoded and cross-encoded to obtain the sample person features. Based on the label generation model, the characteristics of the sample individuals are predicted to obtain the prediction results.

3. The method according to claim 2, characterized in that, The process of autoencoding and cross-encoding the first person vector, the second person vector, and the sample video features based on the tag generation model to obtain sample person features includes: Based on the tag generation model, the first person vector and the second person vector are self-encoded to obtain self-encoded person features. Based on the tag generation model, the self-encoded character features and the sample video features are mutually encoded to obtain the sample character features.

4. The method according to claim 1, characterized in that, The label-based generation model encodes the sample video to obtain sample video features, including: For any sample quad, obtain multiple sample video frames from the sample video; Based on the tag generation model, the multiple sample video frames and the video titles of the sample videos are encoded respectively to obtain video frame features and video title features; The video frame features and the video title features are concatenated to obtain the sample video features.

5. The method according to claim 4, characterized in that, The method further includes: Based on the tag generation model, the audio signal of the sample video is encoded to obtain audio features; The step of concatenating the video frame features and the video title features to obtain the sample video features includes: The video frame features, the video title features, and the audio features are concatenated to obtain the sample video features.

6. The method according to claim 1, characterized in that, The step of training the label generation model based on the difference between the prediction result and the sample supervision information includes: Multiple first training losses are obtained for multiple sample quadruples, where the first training loss is used to represent the difference between the prediction result of the corresponding sample quadruple and the sample supervision information; The average value of the plurality of first training losses is determined as the second training loss; The multiple model parameters are updated based on the derivative of the second training loss with respect to multiple model parameters of the label generation model and the model learning rate.

7. The method according to claim 6, characterized in that, The step of updating the multiple model parameters based on the derivative of the second training loss with respect to multiple model parameters of the label generation model and the model learning rate includes: The intermediate value is determined by multiplying the derivative of the second training loss with respect to multiple model parameters of the label generation model and the model learning rate; For any model parameter, the difference between the model parameter and the intermediate value is determined as the updated model parameter.

8. A label generation method, characterized in that, The method includes: The input video is processed based on a general labeling model to obtain at least one face region image and at least one general person label. The general labeling model is used to add corresponding person labels to the people in the input video. The general person label is the name of the person corresponding to the face region image. Based on the label generation model, the input video, the at least one general person label, and the at least one face region image are predicted to obtain at least one target label probability. The label generation model is trained by the training method of the label generation model according to any one of claims 1 to 7. Based on the probability of the at least one target tag, at least one target person tag is determined from the at least one general person tag of the input video.

9. The method according to claim 8, characterized in that, The label-based generation model predicts at least one target label probability from the input video, the at least one general person label, and the at least one face region image, including: The video information of the input video is encoded based on the tag generation model to obtain the target video features; Based on the tag generation model, the at least one general person tag, the at least one face region image, and the target video features are encoded to obtain at least one target person feature; Based on the label generation model, the features of the at least one target person are predicted to obtain the probability of the at least one target label.

10. The method according to claim 8, characterized in that, Determining at least one target person tag from the at least one general person tag of the input video based on the at least one target tag probability includes: From the at least one general person tag, determine at least one target person tag whose probability of the target tag is greater than a probability threshold; or... From the at least one general person tag, determine the target person tag that ranks first in probability for the target tag.

11. A training device for a label generation model, characterized in that, The device includes: The acquisition module is used to acquire sample quadruples, which include sample videos, sample person names of sample people in the sample videos, sample facial features of the sample people, and sample supervision information. The sample supervision information is used to indicate whether the sample person name is selected as the person label of the sample video. The sample person name and the sample facial features reflect the local features of the sample video. The sample person name is obtained by performing face detection on the sample video and extracting facial features from the detected face regions. It is the person label corresponding to the extracted facial features. The sample person name is the general person label of the sample video. The first encoding module is used to encode the sample video based on the tag generation model to obtain sample video features. The tag generation model is used to generate character tags for the video based on the video, the names of the people in the video, and facial features. The sample video features reflect the overall features of the sample video. The prediction module is used to predict the features of the sample video, the name of the sample person, and the facial features of the sample based on the tag generation model, and obtain a prediction result. The prediction result is used to represent the probability that the name of the sample person is selected as the person tag of the sample video. The probability reflects the degree of association between the corresponding sample person and the sample person. The greater the degree of association, the greater the probability. The training module is used to train the label generation model based on the difference between the prediction results and the sample supervision information; The module is used to perform the following steps: Based on the trained tag generation model, determine the target person tags that are suitable as tags for the video from the general person tags of the video.

12. The apparatus according to claim 11, characterized in that, The prediction module includes: The first mapping unit is used to map the sample person names to a first person vector; The second mapping unit is used to map the sample facial features into a second person vector. The first encoding unit is used to perform auto-encoding and cross-encoding on the first person vector, the second person vector, and the sample video features based on the tag generation model to obtain sample person features. The prediction unit is used to predict the features of the sample person based on the label generation model, and obtain the prediction result.

13. The apparatus according to claim 12, characterized in that, The first encoding unit is used for: Based on the tag generation model, the first person vector and the second person vector are self-encoded to obtain self-encoded person features. Based on the tag generation model, the self-encoded character features and the sample video features are mutually encoded to obtain the sample character features.

14. The apparatus according to claim 11, characterized in that, The first encoding module includes: The first acquisition unit is used to acquire multiple sample video frames from the sample video for any sample quadruple; The second encoding unit is used to encode the plurality of sample video frames and the video titles of the sample videos based on the tag generation model, so as to obtain video frame features and video title features. The splicing unit is used to splice the video frame features and the video title features to obtain the sample video features.

15. The apparatus according to claim 14, characterized in that, The device further includes: The second encoding module is used to encode the audio signal of the sample video based on the tag generation model to obtain audio features; The splicing unit is used to splice the video frame features, the video title features, and the audio features to obtain the sample video features.

16. The apparatus according to claim 11, characterized in that, The training module includes: The second acquisition unit is used to acquire multiple first training losses for multiple sample quadruplets, wherein the first training loss is used to represent the difference between the prediction result of the corresponding sample quadruplet and the sample supervision information. A determining unit is used to determine the average value of the plurality of first training losses as a second training loss; The update unit is used to update the multiple model parameters based on the derivative of the second training loss with respect to multiple model parameters of the label generation model and the model learning rate.

17. The apparatus according to claim 16, characterized in that, The update unit is used for: The intermediate value is determined by multiplying the derivative of the second training loss with respect to multiple model parameters of the label generation model and the model learning rate; For any model parameter, the difference between the model parameter and the intermediate value is determined as the updated model parameter.

18. A label generating apparatus, characterized in that, The device includes: The processing module is used to process the input video based on a general labeling model to obtain at least one face region image and at least one general person label. The general labeling model is used to add corresponding person labels to the people in the input video. The general person label is the name of the person corresponding to the face region image. The prediction module is used to predict the input video, the at least one general person label, and the at least one face region image based on the label generation model to obtain at least one target label probability. The label generation model is trained by the training method of the label generation model according to any one of claims 1 to 7. A determination module is configured to determine at least one target person label from the at least one general person label of the input video based on the at least one target label probability.

19. The apparatus according to claim 18, characterized in that, The prediction module is used for: The video information of the input video is encoded based on the tag generation model to obtain the target video features; Based on the tag generation model, the at least one general person tag, the at least one face region image, and the target video features are encoded to obtain at least one target person feature; Based on the label generation model, the features of the at least one target person are predicted to obtain the probability of the at least one target label.

20. The apparatus according to claim 18, characterized in that, The determining module is used for: From the at least one general person tag, determine at least one target person tag whose probability of the target tag is greater than a probability threshold; or... From the at least one general person tag, determine the target person tag that ranks first in probability for the target tag.

21. A computer device, characterized in that, The computer device includes a processor and a memory, the memory being used to store at least one computer program, the at least one computer program being loaded by the processor and executed as a training method for the label generation model according to any one of claims 1 to 7, or the at least one computer program being loaded by the processor and executed as a label generation method according to any one of claims 8 to 10.

22. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store at least one computer program, which is used to perform the training method of the label generation model according to any one of claims 1 to 7, or the at least one computer program is used to perform the label generation method according to any one of claims 8 to 10.

23. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the training method of the label generation model as described in any one of claims 1 to 7, or, when the computer program is executed by the processor, it implements the label generation method as described in any one of claims 8 to 10.

Citation Information

Patent Citations

  • Video label obtaining method and device, storage medium and server

    CN111695422A