A video material screening method and device, computer equipment and storage medium
By recalling unpushed content and utilizing the semantic and visual content represented by tags, combined with the continuous bag-of-words model and gradient boosting decision tree model, the problem of low efficiency in video content screening was solved, achieving efficient screening of high-quality content and reducing the screening cost for optimization specialists.
Patent Information
- Application Number
- CN202210010684.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-06
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2042-01-06
AI Technical Summary
In existing technologies, the efficiency of video material selection is low, the cost for optimization specialists to select video materials is high, and the process of first trying out the push and then adjusting based on user feedback is cumbersome.
By recalling materials that have not been pushed to the client, using optimization evaluation indicators as the target, and combining the semantic representation of tags and the visual content of video data, the continuous bag-of-words model and gradient boosting decision tree model are used for screening, and push tasks are generated for optimization specialists to select and push to the client.
It improved the efficiency of video material selection, reduced the number of selections required by optimization specialists, saved costs, and simplified the operation process by learning about the quality of materials and quickly selecting high-quality materials.
Smart Images

Figure CN114417058B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer processing, and particularly relates to a video material screening method and device, a computer device and a storage medium. BACKGROUND
[0002] With the continuous progress of Internet technology, Internet media has almost covered all aspects of people's life. Since the amount of information on the Internet is very large, the search efficiency of users is low. Therefore, in order to provide users with higher quality materials, specific information will be pushed to users in the operation process.
[0003] For new materials, the quality of the materials is currently evaluated by an optimizer according to experience. First, materials with high quality evaluation are pushed to the client in a trial manner. Dynamic tracking and adjustment are performed according to user feedback on the materials.
[0004] However, the forms of materials for different businesses are various, especially video materials, the content of which is particularly rich, and the quantity of which reaches hundreds of thousands or millions. The efficiency of the optimizer in selecting materials is greatly challenged. In addition, the operation is relatively cumbersome and the cost is increased by trying to push first and then adjusting according to user feedback. SUMMARY
[0005] The present application provides a video material screening method, device, computer device and storage medium to solve the problem of how to reduce the selected materials, improve the efficiency of the optimizer and reduce the cost.
[0006] In a first aspect, an embodiment of the present application provides a material screening method, comprising:
[0007] Recalling materials not pushed to the client as first candidate materials, the materials containing video data and being marked with labels;
[0008] Screening part of the first candidate materials according to the semantics represented by the labels as second candidate materials, with an optimization evaluation index as a target;
[0009] Screening part of the second candidate materials according to the semantics represented by the labels and the visual content of the video data as third candidate materials, with the optimization evaluation index as a target;
[0010] Generating a push task for the third candidate materials, the push task being used for screening part of the third candidate materials by a user in a role of an optimizer and pushing to the client;
[0011] The evaluation index is data formed by statistics on operations triggered by the client on the materials after the materials are pushed to the client.
[0012] In a second aspect, the embodiments of the present application further provide a video material screening device, comprising:
[0013] a recall module configured to recall a material that is not pushed to a client as a first candidate material, the material containing video data and being marked with a label;
[0014] a coarse screening module configured to screen part of the first candidate material as a second candidate material according to semantics represented by the label, with an optimization evaluation index as a target;
[0015] a fine screening module configured to screen part of the second candidate material as a third candidate material according to the semantics represented by the label and visual content of the video data, with the optimization evaluation index as a target;
[0016] a task generation module configured to generate a pushing task for the third candidate material, the pushing task being used to screen part of the third candidate material by a user as an optimizationist and push the part of the third candidate material to the client;
[0017] wherein the evaluation index is data formed by counting operations triggered by the client on the material after the material is pushed to the client.
[0018] In a third aspect, the embodiments of the present application further provide a computer device, comprising:
[0019] one or more processors;
[0020] a memory configured to store one or more programs,
[0021] when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the video material screening method according to the first aspect.
[0022] In a fourth aspect, the embodiments of the present application further provide a computer readable storage medium, the computer readable storage medium storing a computer program, the computer program being executed by a processor to implement the video material screening method according to the first aspect.
[0023] In the embodiment, the material not pushed to the client is recalled as the first candidate material, the material contains video data and is marked with a label; part of the first candidate material is screened out as the second candidate material according to the semantic represented by the label, with the optimization evaluation index as the target; part of the second candidate material is screened out as the third candidate material according to the semantic represented by the label and the visual content of the video data, with the optimization evaluation index as the target; the third candidate material is generated to push the task, and the push task is used to screen out part of the third candidate material by the user of the role of the optimization and push to the client; wherein the evaluation index is the data formed by the operation triggered by the client to the material pushed to the client. The embodiment filters the material through the three links of recall, rough sorting and fine sorting, provides high-quality materials for the optimization on the basis of relatively high efficiency, greatly reduces the number of materials screened by the optimization, saves the cost, and learns the advantages and disadvantages of the material in the two stages of rough sorting and fine sorting with the evaluation index as the target, accumulates knowledge, and thus quickly selects high-quality materials, avoids the mode of trying to push part of the material first and then adjusting, greatly improves the simplicity of operation and improves the efficiency of pushing the material. BRIEF DESCRIPTION OF DRAWINGS
[0024] Figure 1 A flowchart of a video material screening method provided for the embodiment one of the present application;
[0025] Figure 2 A process schematic diagram of screening material provided for the embodiment one of the present application;
[0026] Figure 3 A flowchart of a video material screening method provided for the embodiment two of the present application;
[0027] Figure 4 A structure schematic diagram of a content extraction network and a content understanding network provided for the embodiment two of the present application;
[0028] Figure 5 A structure schematic diagram of a video material screening device provided for the embodiment three of the present application;
[0029] Figure 6 A structure schematic diagram of a computer device provided for the embodiment four of the present application. DETAILED DESCRIPTION
[0030] The present application will be further described below in conjunction with the drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present application, but not to limit the present application. In addition, it should be noted that, for the convenience of description, only the parts related to the present application are shown in the drawings, but not all the structures.
[0031] Embodiment one
[0032] Figure 1 This is a flowchart of a video material filtering method provided in Embodiment 1 of the present invention. This embodiment is applicable to situations where materials are filtered based on the semantics of tags and the visual content of video data. The method can be executed by a video material filtering device, which can be implemented by software and / or hardware and can be configured in a computer device of a multimedia platform, such as a server, workstation, personal computer, etc. The specific steps include:
[0033] Step 101: Recall the materials that have not been pushed to the client and use them as the first candidate materials.
[0034] like Figure 2 As shown in this embodiment, a media library, such as a distributed database, can be pre-established in the multimedia platform. The media library stores a large number of media, which are identified by unique IDs. The relevant information of the media, such as tags, video data, etc., can be queried in the media library by the media ID.
[0035] The content can be content that has already been pushed to the client or content that has not been pushed to the client. Content has a lifespan, which is generally less than one month, but the lifespan of selected content can be as long as six months or longer.
[0036] Specifically, materials can include audio data, video data, text data, image data, Uniform Resource Locator (URL), JSON (JavaScript Object Notation), and so on. The form of materials will vary depending on the business scenario.
[0037] For example, in the news media field, the material can be news data; in multimedia entertainment, the material can be short videos; in the e-commerce (EC) field, the material can be advertising data, and so on.
[0038] Although various materials carry business characteristics under different business scenarios, their essence is still data.
[0039] In addition, the materials are tagged. Some tags are carried when the materials are generated, including material parameter characteristics (such as material effects, background music, video dimensions, video duration, video quality, material description, etc.) and material creator characteristics (such as creator's country, creator's tags, creator's level, creator's age, creator's gender, etc.). Some tags are generated after being pushed to the client, such as exposure count, click count, video playback count, share count, comment count, completion rate, etc.
[0040] The multimedia platform, upon receiving the request of the client, can recall part of the materials from the material library according to different business needs (such as recalling high-quality (non-personalized) video data, recalling video data meeting the personalized needs of the user, etc.) using different recall strategies, denoted as first candidate materials, for different business scenarios, and waiting for rough sorting and fine sorting.
[0041] The request of the client can be triggered by the user, for example, the user inputs a keyword in the client and requests the multimedia platform to search for materials related to the keyword, the user pulls down the list of existing materials and requests the multimedia platform to refresh the materials, etc. The request of the client can also not be triggered by the user, for example, the client requests the multimedia platform to push high-quality materials when displaying the home page, the client requests the multimedia platform to push related materials before the video data in the current material ends playing, etc. The present embodiment does not limit this.
[0042] In one example, the recall strategy includes but is not limited to:
[0043] Popular recall (recalling a plurality of materials with the highest click rate or play rate), online recall (recalling a live program (material) hosted by an online anchor user), subscription recall (recalling materials of a column (such as a certain game, catering, etc.) subscribed by the user), same country recall (recalling materials of the same country as the user), same language recall (recalling materials of the same language as the user), collaborative filtering recall (recalling materials using a collaborative filtering algorithm), preference recall (recalling materials of the same preference as the user), and similar recall (recalling other materials similar to the recalled materials).
[0044] Step 102: filtering part of the first candidate materials according to the semantics represented by the label as second candidate materials, with the optimization of the evaluation index as the goal.
[0045] Generally, the number of the first candidate materials recalled is large, usually reaching the order of magnitude of ten thousand or one thousand, and the algorithm used in fine sorting can be relatively complex. In order to improve the speed of sorting, a rough sorting link can be added between recall and fine sorting.
[0046] In the present embodiment, an evaluation index can be set, wherein the evaluation index is data formed by counting the operations triggered by the client on the materials when the materials are pushed to the client.
[0047] For example, if the material is the title of news data, it contains a URL pointing to the page where the news data is located, and the evaluation index can be the exposure rate of the page.
[0048] For example, if the material is used to showcase an application and contains a URL that points to the application's download address, then the evaluation metric could be the probability of installing the application.
[0049] For example, if the content is used to display a product and contains a URL that points to the address of a product, then the evaluation metric can be the conversion rate of users placing an order to purchase the product.
[0050] like Figure 2 As shown, during the coarse ranking process, a small number of semantic features representing the labels are extracted and loaded into a simple ranking model, such as LR (Logistic Regression) model, GBDT (Gradient Boost Decision Tree) model, etc., with the aim of optimizing the evaluation metrics. The first candidate materials recalled are roughly ranked, and the first candidate materials with higher ranking are selected as the second candidate materials. That is, pushing the second candidate materials to the client is more conducive to optimizing the evaluation metrics than pushing other first candidate materials to the client.
[0051] Rough layout can further reduce the amount of material in fine layout while ensuring a certain level of accuracy. Generally, the amount of material can be reduced to the level of thousands or hundreds.
[0052] Step 103: With the goal of optimizing the evaluation indicators, select some second candidate materials based on the semantics represented by the tags and the visual content of the video data, and use them as third candidate materials.
[0053] like Figure 2 As shown, during the fine ranking process, more semantic and visual features of the video data are extracted from the tags and loaded into a more complex ranking model, such as CNN (Convolutional Neural Networks), RNN (Recurrent Neural Network), etc., with the aim of optimizing the evaluation metrics. The second candidate materials in the coarse ranking are then precisely ranked, and the second candidate materials with higher rankings are selected as the third candidate materials. That is, pushing the third candidate materials to the client is more conducive to optimizing the evaluation metrics than pushing other second candidate materials to the client.
[0054] Fine sorting can improve the accuracy of sorting as much as possible, and further reduce the number of materials sent to the client. Generally, the number of materials can be reduced to the hundreds or tens.
[0055] Step 104: Generate a push task for the third candidate material.
[0056] In this embodiment, for the third candidate material, a push task can be generated, which is used to filter out part of the third candidate material by the user with the role of an optimizer and push it to the client, that is, the push task is assigned to the user with the role of an optimizer, the user with the role of an optimizer logs in the client using account, password and the like, executes the push task, and selects part of the third candidate material according to the business requirements, and the selected part of the third candidate material can be scattered and then maintained in a number (such as hundreds or tens) to push the client for display.
[0057] Of course, in addition to part of the third candidate material, part of the material that has been pushed to the client belongs to relatively high-quality material, and the user with the role of an optimizer can also select part of the material that has been pushed to the client according to the business requirements, and the selected part of the third candidate material and part of the material that has been pushed to the client can be scattered and then maintained in a number (such as hundreds or tens) to push the client for display.
[0058] Among them, scattering is also called rearrangement, that is, reordering the material globally to make various types of material more evenly distributed.
[0059] In this embodiment, the material that has not been pushed to the client is recalled as the first candidate material, the material contains video data and is marked with a label; part of the first candidate material is selected as the second candidate material according to the semantic represented by the label, with the optimization evaluation index as the target; part of the second candidate material is selected as the third candidate material according to the semantic represented by the label and the visual content of the video data, with the optimization evaluation index as the target; a push task is generated for the third candidate material, which is used to filter out part of the third candidate material by the user with the role of an optimizer and push it to the client; wherein the evaluation index is the data formed by the operation triggered by the client to the material after the material is pushed to the client. This embodiment selects the material through the three links of recall, rough sorting and fine sorting, provides high-quality material for the optimizer on the basis of high efficiency, greatly reduces the number of materials selected by the optimizer, saves cost, and learns the advantages and disadvantages of the material in the rough sorting and fine sorting stages with the evaluation index as the target, accumulates knowledge, thereby quickly selects high-quality material, avoids the mode of trying to push part of the material first and then adjusting, greatly improves the operation convenience and improves the efficiency of pushing the material.
[0060] Embodiment three
[0061] Figure 3 A flowchart of a video material selection method provided by the second embodiment of the application, based on the foregoing embodiment, further refines the rough sorting and fine sorting operations, and the method specifically includes the following steps:
[0062] Step 301, recalling the material not pushed to the client as the first candidate material.
[0063] The material contains video data and is labeled with a label.
[0064] Step 302, extracting a first material feature representing semantics from the label of the first candidate material.
[0065] For the first candidate material, one or more labels tagged on the material can be found in the material library, and the one or more labels are processed in natural language, so as to extract the feature of the one or more labels in semantics, which is recorded as the first material feature.
[0066] In an embodiment of the present application, step 302 can include the following steps:
[0067] Step 3021, determining a continuous bag of words model.
[0068] In this embodiment, the continuous bag of words model (CBOW) can be pre-trained, and the structure and parameters of the continuous bag of words model are stored in the database. When the first candidate material is roughly sorted, the continuous bag of words model and its parameters are loaded into the memory for running.
[0069] The continuous bag of words model can predict the target word by the context word around the target word.
[0070] In an embodiment of the present application, step 3021 can further include the following steps:
[0071] Step 30211, obtaining the material pushed to the client as historical material.
[0072] Whether the material is pushed to the client can be recorded as an item of information in the material library. Then, the item of information is queried in the material library, and the material pushed to the client is extracted, which is recorded as historical material.
[0073] Generally, the continuous bag of words model can be trained using real historical material pushed to the client, but the real historical material pushed to the client is relatively sparse. In order to ensure the performance of the continuous bag of words model, the historical material pushed to the client can be constructed on the basis of the real historical material pushed to the client.
[0074] In the process of constructing the historical material, the material not pushed to the client can be obtained as the first original material. The label of the first original material is the label carried by the first original material when it is generated, and lacks the label generated after the first original material is pushed to the client.
[0075] At this time, the first original material similar to the historical material is recalled as the second original material. For the convenience of calculation, the tags can be used to evaluate whether the materials are similar. For example, if the country (tag) of the historical material is the same as the country (tag) of the first original material, it can be considered that the historical material is similar to the first original material. If the creator (tag) of the historical material is similar to the behavior of the creator (tag) of the first original material, it can be considered that the historical material is similar to the first original material, and so on.
[0076] The second original material is clustered around the historical material by algorithms such as K-means (K-means clustering) to obtain a material cluster.
[0077] In the material cluster, the part of the tags generated by the historical material after being pushed to the client are shared to the second original materials closest to the historical material, so that the second original materials become new historical materials.
[0078] The present embodiment constructs the first original material not pushed to the client into a new historical material by referring to the historical material. Since the behaviors of the client to similar materials are similar, the tags shared to the first original material can have a certain accuracy, thereby greatly increasing the number of historical materials and ensuring the performance of the continuous bag-of-words model.
[0079] Step 30212, the tags of the historical material are processed by word segmentation to obtain a plurality of word groups.
[0080] The tags of the historical material are processed by word segmentation, and the long tags in the historical material are disassembled to obtain a plurality of word groups.
[0081] Among them, the language of the tags is different, and the corresponding word segmentation processing is also different. For example, for English tags, the word segmentation processing is to split the words, and for Chinese tags, the word segmentation processing is jieba word segmentation, etc.
[0082] Step 30213, the plurality of word groups are encoded into a first word vector.
[0083] For the word groups of each tag in the historical material, one-hot (one-hot encoding) and other encoding methods can be used to encode the vector, denoted as the first word vector, that is, the word groups in each tag of the historical material are represented in the form of a vector, and the analysis of each tag in the historical material is simplified to vector operation in the vector space.
[0084] Wherein, one-hot is also called one-bit effective coding, N-bit state register is used to encode N states, each state has independent register bit, and only one bit is effective at any time. The vector of one-hot coding is a classification variable, which is represented as a binary vector, which requires to map the classification value to an integer value, and each integer value is represented as a binary vector, except the index of the integer, which is marked as 1.
[0085] Assuming that there are V word groups to be coded, the first word vector of the current word group is x i , and the dimension of the first word vector is 1xV.
[0086] Step 30214, for the current word group, input the first word vector of the other word groups belonging to the context into the continuous bag-of-words model, and map it into the second word vector of the current word group.
[0087] Traverse each word group in each label of the historical material, and regard it as the current word group in turn according to the order, determine the other word groups belonging to the context of the current word group, and the other word groups are generally the word groups before the current word group in the order and the word groups after the current word group in the order, input the first word vector of the context into the continuous bag-of-words model, and the continuous bag-of-words model maps the first word vector of the context into the second word vector of the current word group.
[0088] Further, the continuous bag-of-words model has an input layer (Input Layer), one or more hidden layers (Hidden Layer), and an output layer (Output Layer).
[0089] The input of the input layer is the other first word vector belonging to the context of the current first word vector, and for the hidden layer, the output h1 of the first hidden layer is calculated, the shared matrix W input , the dimension of h1 is 1xN, and then:
[0090]
[0091] Wherein, window is the window for selecting the first word vector as x i The context window.
[0092] After n layers of hidden layers, the output of the output layer is output, wherein h n represents the last hidden layer, the dimension of which is N, the shared matrix W output , the dimension of output is 1xV, and then:
[0093] output=h n ×W output
[0094] The output vector is normalized by using an activation function such as Softmax to obtain a vector of dimension 1xV, and the position corresponding to the number with the maximum probability in the V values is determined as the second word vector of the current word group.
[0095] wherein Softmax is a logistic function that can compress a K-dimensional vector z containing any real number into another K-dimensional real vector σ(z) such that the range of each element is between (0, 1) and the sum of all elements is 1, and the function is commonly used in multi-classification problems.
[0096] Step 30215, calculating a label loss value based on the second word vector.
[0097] For the same word group, the predicted second word vector and the real label are substituted into a preset loss function (Loss Function) such as cross entropy to calculate the loss value LOSS between the predicted second word vector and the real label, which is recorded as the label loss value.
[0098] wherein the real label is a vector of dimension 1xV, and one of the V values is 1 and the others are 0.
[0099] Step 30216, updating the continuous bag-of-words model according to the label loss value.
[0100] After completing the forward propagation, the continuous bag-of-words model can be back-propagated, and the label loss value can be substituted into an optimization algorithm such as SGD (stochastic gradient descent) or Adam (Adaptive momentum) to calculate the amplitude of updating the parameters in the continuous bag-of-words model, and the parameters (shared matrix W input , shared matrix W output ) in the continuous bag-of-words model are updated according to the amplitude.
[0101] Step 30217, determining whether the bag-of-words training condition is met; if yes, performing step 30218, and if no, returning to step 30214.
[0102] Step 30218, determining that the training of the continuous bag-of-words model is completed.
[0103] In this embodiment, the bag-of-words training condition can be set in advance as a condition for stopping training the continuous bag-of-words model, for example, the number of iterations reaches a threshold value, the change amplitude of the label loss value for a plurality of times is less than a threshold value, and the like. In each round of iterative training, it is determined whether the bag-of-words training condition is met.
[0104] If the bag-of-words training condition is met, it can be considered that the continuous bag-of-words model training is completed, at this time, the parameters in the continuous bag-of-words model are output and are persisted in the database.
[0105] If the bag-of-words training condition is not met, the next round of iterative training can be entered, and steps 3014-3016 are re-executed, and the iterative training is cycled in this way until the continuous bag-of-words model training is completed.
[0106] Further, the continuous bag-of-words model can be independently trained, or can be fine-tuned on a pre-trained continuous bag-of-words model using historical materials as samples, that is, on the basis of the pre-trained continuous bag-of-words model, the historical materials are used as samples of the target task to continue training, which is not limited by the present embodiment.
[0107] Step 3022, performing word segmentation processing on the labels of the first candidate material to obtain a plurality of word groups.
[0108] The labels of the first candidate material are subjected to word segmentation processing, and the labels with longer lengths in the first candidate material are disassembled to obtain a plurality of word groups.
[0109] Step 3023, encoding the plurality of word groups into a first word vector.
[0110] For the word groups of each label in the first candidate material, one-hot (one-hot encoding) or other encoding methods can be used to encode into a vector, denoted as a first word vector, that is, the word groups in each label of the first candidate material are represented in the form of a vector, and the analysis and simplification of each label in the first candidate material are simplified into vector operations in a vector space, laying a data foundation for material screening.
[0111] Step 3024, for the current word group, inputting the first word vectors of other word groups belonging to the context into the continuous bag-of-words model to map the second word vector of the current word group as the first material feature.
[0112] Each word group in each label of the first candidate material is traversed, and is sequentially regarded as a current word group according to the order, and other word groups belonging to the context of the current word group are determined, which are generally word groups before the current word group in the order and word groups after the current word group in the order. The first word vectors of the context are input into the continuous bag-of-words model, and the continuous bag-of-words model maps the first word vectors of the context into the second word vector of the current word group.
[0113] Step 303, calculating the importance of the first candidate material for the evaluation index according to the first material feature as a first score.
[0114] For the semantic feature of the first candidate material (i.e., the first material feature), the historical material pushing to the client can be mined by machine learning or deep learning, so as to learn the importance of the first candidate material to the evaluation index in semantics, recorded as the first score.
[0115] In an embodiment of the present application, step 303 can further include the following steps:
[0116] Step 3031, determining the first gradient boosting decision tree and the first feature set.
[0117] In this embodiment, the Light Gradient Boosting Machine (LightGBM) can be trained in advance for the semantic feature of the label of the material, recorded as the first gradient boosting decision tree, the structure and parameters of the first gradient boosting decision tree are stored in the database, and the first gradient boosting decision tree and its parameters are loaded into the memory for running when the first candidate material is roughly sorted.
[0118] Gradient Boosting Decision Tree (GBDT) is a weak classifier (decision tree) iterative training to get the optimal model, that is, the boosting decision tree is composed of multiple decision trees, and the conclusions of all decision trees are accumulated to make the final result. This model has the advantages of good training effect, not easy to overfit, etc. LightGBM is a framework that implements the GBDT algorithm, supports efficient parallel training, and has the advantages of faster training speed, lower memory consumption, better accuracy, supports distributed and can quickly process massive data, etc.
[0119] In addition, the first feature set can be trained in advance for the semantic feature of the label of the material, and the first feature set records the first sample feature screened by the first decision boosting tree according to the importance to the evaluation index. The way of constructing the first sample feature is consistent with the way of constructing the first material feature, that is, the first sample feature is the semantic feature of the label (phrase) of the material.
[0120] In a specific implementation, the material that has been pushed to the client is obtained as historical material.
[0121] In the case of sparse historical materials, the materials not pushed to the client can be obtained as first original materials, the first original materials similar to the historical materials are recalled as second original materials, the second original materials are clustered around the historical materials to obtain material clusters, and the part of the labels generated after the historical materials are pushed to the client are shared to the second original materials closest to the historical materials in the material clusters, so that the second original materials become new historical materials. Since the client's behavior on similar materials is similar, the accuracy of the labels shared to the first original materials can be ensured, thereby greatly increasing the number of historical materials and ensuring the performance of the first gradient boosting decision tree.
[0122] The first material feature representing the semantic is extracted from the label of the historical material.
[0123] Specifically, a continuous bag-of-words model is determined, the label of the historical material is segmented, a plurality of word groups are obtained, the plurality of word groups are encoded into a first word vector, and for the current word group, the first word vector of the other word groups belonging to the context is input into the continuous bag-of-words model and mapped into a second word vector of the current word group as the first material feature.
[0124] For historical materials, they can be divided into two categories: positive samples and negative samples. Due to the real reason of material pushing to the client, the positive samples and the negative samples in the historical materials are not balanced. If the unbalanced phenomenon is ignored, the first gradient boosting decision tree will be biased to the category with more quantity.
[0125] Therefore, the ratio between the number of positive samples and the number of negative samples in the historical materials is used to determine the weight of the positive samples and the weight of the negative samples, that is, the weight of the positive samples is negatively correlated with the proportion of the number of positive samples, and the weight of the negative samples is negatively correlated with the proportion of the number of positive samples.
[0126] After weighting, when determining the branch point of the sub-tree in the training process of the first gradient boosting decision tree, the category with larger weight is emphasized, thereby ensuring the performance of the first gradient boosting decision tree.
[0127] For the first gradient boosting decision tree, the index for evaluating the pros and cons can be set as AUC (Area Under Curve), binary_logloss (binary classification log loss), etc.
[0128] The AUC is the area surrounded by the ROC (Receiver Operating Characteristic Curve) curve and the coordinate axis.
[0129] In addition, the parameters for training are set, including:
[0130] 1、core parameters, mainly for index type, task type, training target, model type, iteration number, learning rate, leaf node number, etc.
[0131] 2、learning control parameters, mainly for the depth of the decision tree, the minimum number of data on one leaf (to reduce overfitting), the proportion of randomly selected data without resampling, the number of Bagging (an important ensemble learning method), the proportion of randomly selected features in each iteration, L1 regularization, L2 regularization, etc.
[0132] 3、other parameters, mainly for dataset parameters, prediction parameters, etc.
[0133] Then, using the first material feature to train the first gradient boosting decision tree for the purpose of optimizing the evaluation index, in order to find the best parameters, GridSearchCV (detailed search of specified parameter values for machine learning models) in scikit-sklearn (a free software machine learning library in python programming language) can be used to perform grid search on the specified parameters to find the optimal solution.
[0134] The first gradient boosting decision tree can automatically adjust the parameters of the first material feature, thereby highlighting the weight of the first material feature that is more important for the evaluation index and weakening the weight of the first material feature that has small correlation with the evaluation target.
[0135] The first gradient boosting decision tree outputs the importance of the first material feature for the evaluation index when the training is completed, i.e., the correlation between the first material feature and the evaluation index.
[0136] The first material feature with an importance greater than a preset first threshold is selected as the first sample feature and written into the first feature set.
[0137] Using the first sample feature to train the first gradient boosting decision tree again for the purpose of optimizing the evaluation index, the parameters of the first gradient boosting decision tree are saved when the training is completed again, and through two rounds of training, the accuracy of the first gradient boosting decision tree in calculating the importance of the first material feature for the evaluation index can be greatly improved.
[0138] Step 3032, screening the first material feature identical to the first sample feature as the first target feature.
[0139] In this embodiment, the first material feature can be compared with the first sample feature in the first feature set. If the first material feature is identical to the first sample feature in the first feature set, the first material feature is retained and recorded as the first target feature. If the first material feature is different from the first sample feature in the first feature set, the first material feature is filtered out.
[0140] Step 3033, input the first target feature into the first gradient boosting decision tree to calculate the importance of the first candidate material for the evaluation index as the first score.
[0141] For the first material feature in the first candidate material, after all the first target features are combined, input into the first gradient boosting decision tree, the first gradient boosting decision tree calculates the importance of the first candidate material for the evaluation index, denoted as the first score.
[0142] Step 304, select the first candidate material with the highest first score as the second candidate material.
[0143] According to the descending order of the first score of the first candidate material, select the top k (k is a positive integer) first candidate materials as the second candidate material.
[0144] Step 305, determine the first score.
[0145] In the process of fine sorting, the first score generated in the rough sorting can be queried, and the first score represents the importance of the first candidate material for the evaluation index under the semantic of the label representation.
[0146] Step 306, extracting second material features representing visual content from video data of the second candidate material.
[0147] For the second candidate material, its contained video data can be searched in the material library, and computer vision processing is performed on the video data to extract features of the video data on the visual content, denoted as the second material feature.
[0148] In an embodiment of the present application, step 306 can include the following steps:
[0149] Step 3061, determining a content extraction network and a content understanding network.
[0150] In this embodiment, the content extraction network and the content understanding network can be pre-trained, and the content extraction network and the content understanding network both belong to a deep learning model, and the structure and parameters of the content extraction network and the structure and parameters of the content understanding network are stored in a database. When fine sorting the second candidate material, the content extraction network and parameters and the content understanding network and parameters are loaded into the memory for running.
[0151] The content extraction network is used to extract the content features of the video data in the second candidate material, and the content features are irrelevant to the evaluation index.
[0152] The content understanding network is used to map the content features so as to be related to the evaluation index.
[0153] In an embodiment of the present application, step 3061 can further include the following steps:
[0154] Step 30611, obtaining the material pushed to the client as historical material.
[0155] In the case of sparse historical material, the material not pushed to the client can be obtained as the first original material, the first original material similar to the historical material is recalled as the second original material, the second original material is clustered around the historical material to obtain a material cluster, and the part of the label of the historical material generated after being pushed to the client is shared to the second original material closest to the historical material in the material cluster, so that the second original material becomes new historical material. Since the client's behavior on similar materials is similar, the label shared to the first original material can have a certain accuracy, thereby greatly increasing the number of historical materials and ensuring the performance of the content understanding network.
[0156] Step 30612, extracting multiple frames of image data from the video data of the historical material.
[0157] In the present embodiment, multiple frames of image data can be extracted from the video data of the historical material at a preset frequency (such as 1 FPS (Frames Per Second, frames per second)) to form a sequence.
[0158] Step 30613, inputting the image data into the content extraction network pre-trained for image classification to extract the first image feature irrelevant to the evaluation index.
[0159] Since the amount of video data in the historical material is limited, the entire model (i.e., the content extraction network and the content understanding network) cannot be directly trained, and training the entire model (the content extraction network and the content understanding network) after updating the video data of the historical material each time is too costly. Therefore, in the present embodiment, a two-stage fine-tuning method is used to divide the entire model into the content extraction network and the content understanding network.
[0160] The content extraction network is a model pre-trained for image classification, and its parameters are fixed. The task of training the content extraction network is not necessarily related to the evaluation index.
[0161] Therefore, the image data in the sequence is input into the content extraction network, and the content extraction network extracts the feature irrelevant to the evaluation index, which is denoted as the first image feature.
[0162] In one example, as Figure 4As shown, the content extraction network includes a residual neural network (ResNet), a temporal shift network (TSM), i.e., the content extraction network is a network structure in which a residual neural network and a temporal shift network are fused, including a 2D CNN (a two-dimensional convolution layer, a 2D convolution is performed on the convolution layer to extract features from a local neighborhood on the feature map of the previous layer), a residual block (a residual block is a group of layers, which is set in a way that the output of the layer is added to another layer deeper in the block. After adding it to the output of the corresponding layer in the main path, a nonlinear operation is applied, and this bypass connection is called a shortcut or a skip connection), a time shift module, a BN (batch normalization layer, also called batch normalization, which is a normalization method that transforms and restructures the input of the layer to solve the problem of changes in the data distribution of the intermediate layer during the training process, making the network training faster and more stable) layer, an activation function layer, a pooling layer (a pooling layer is a dimensionality reduction method that simulates the human visual system to represent an image with summarized features), a fully connected layer (generally located at the end of the entire convolutional neural network, responsible for converting the two-dimensional feature map output by convolution into a one-dimensional vector, implementing an end-to-end learning process of the network), and the like. The loss function uses cross-entropy loss and triple loss (triple loss compares the anchor sample with the positive sample and the negative sample, minimizes the distance between the anchor sample and the positive sample, and maximizes the distance between the anchor sample and the negative sample), which can be trained using image data in historical material video data. Image data of the same origin (video data) is considered to be of the same class and has the same class label.
[0163] Among them, the residual neural network is a deep residual learning framework for image recognition, which is easier to optimize using a residual network structure and can achieve higher accuracy from significantly increased depth, which can specifically include ResNet11, ResNet18, etc.
[0164] The temporal shift network is applied to video understanding in deep learning, which can achieve the performance of a 3D CNN (a three-dimensional convolution layer, a 3D convolution is implemented by convolving a 3D kernel to a cube formed by stacking multiple consecutive frames together to calculate features from spatial and temporal dimensions) while maintaining the complexity of a 2D CNN. The TSM moves part of the channels along the time dimension, thereby facilitating information exchange between adjacent frames.
[0165] Then, the image data in the sequence is input into the residual neural network to extract the features irrelevant to the evaluation index, denoted as residual features, and the residual features are input into the temporal shift network to extract the features irrelevant to the evaluation index, denoted as first image features.
[0166] Step 30614, the first image features are input into the content understanding network to extract second image features related to the evaluation index.
[0167] The input layer of the content understanding network is connected to the last layer (generally a fully connected layer) of the content extraction network, the first image features in the sequence are input into the content understanding network, and the content understanding network extracts features related to the evaluation index, denoted as second image features.
[0168] In one example, as shown in Figure 4 the content understanding network includes a first fully connected layer FC and a second fully connected layer FC, then the first image features are input into the first fully connected layer to be mapped into features related to the evaluation index, denoted as fully connected features, and the fully connected features are input into the second fully connected layer to be mapped into features related to the evaluation index, denoted as second image features, which are used as second material features.
[0169] Step 30615, based on the second image features, the content loss value is calculated according to the evaluation index.
[0170] For the same frame of image data, the predicted second image features and the real label are substituted into the preset loss function to calculate the loss value between the predicted second image features and the real label, denoted as the content loss value.
[0171] Further, for different types of evaluation indexes, the loss function is different, that is, there is a mapping relationship between the evaluation index and the loss function, for example, if the evaluation index is whether to install an application, the loss function is cross-entropy, if the evaluation index is the installation rate and conversion rate of the application, the loss function is mean square error, etc.
[0172] Step 30616, the content understanding network is updated according to the content loss value.
[0173] After completing the forward propagation, the content loss value can be back propagated, and the content loss value can be substituted into the SGD, Adam, etc. optimization algorithm to calculate the amplitude of updating the parameters in the content loss value, and the parameters in the content loss value are updated according to the amplitude.
[0174] Step 30617, it is judged whether the content training condition is met; if yes, step 30618 is executed, and if no, step 30613 is returned.
[0175] Step 30618, it is determined that the content understanding network training is completed.
[0176] In the embodiment, the content training condition can be set in advance as a condition for stopping training the content understanding network, for example, the number of iterations reaches a threshold, the change amplitude of the content loss value for a plurality of times is less than a threshold, and the like. In each round of iterative training, it is determined whether the content training condition is met.
[0177] If the content training condition is met, it can be considered that the content understanding network is trained, at which time the parameters in the content understanding network are output and persisted in the database.
[0178] If the content training condition is not met, the next round of iterative training can be entered, and steps 30613-30616 are re-executed, and the iterative training is cycled until the content understanding network is trained.
[0179] Further, the content extraction network and the content understanding network can be independently trained, or fine-tuned on the pre-trained content extraction network and the content understanding network for image classification using historical materials as samples, that is, on the basis of the pre-trained content extraction network and the content understanding network for image classification, the historical materials are used as samples of the target task to continue training, which is not limited in the embodiment.
[0180] Step 3062, extracting a plurality of frames of image data from the video data of the second candidate material.
[0181] In the embodiment, a plurality of frames of image data can be extracted from the video data of the second candidate material at a preset frequency (such as 1 FPS) to form a sequence.
[0182] Step 3063, inputting the image data into the content extraction network to extract the first image feature irrelevant to the evaluation index.
[0183] The image data in the sequence is input into the content extraction network, and the content extraction network extracts the feature irrelevant to the evaluation index, which is denoted as the first image feature.
[0184] In one example, the content extraction network includes a residual neural network and a temporal shift network, and in this example, the image data in the sequence is input into the residual neural network to extract the feature irrelevant to the evaluation index, which is denoted as a residual feature, and the residual feature is input into the temporal shift network to extract the feature irrelevant to the evaluation index, which is denoted as the first image feature.
[0185] Step 3064, inputting the first image feature into the content understanding network to extract the second image feature related to the evaluation index as the second material feature.
[0186] The first image feature in the sequence is input into the content understanding network, and the content understanding network extracts a feature related to the evaluation index, denoted as a second image feature.
[0187] In one example, the content understanding network includes a first fully connected layer and a second fully connected layer. In this example, the first image feature is input into the first fully connected layer to be mapped into a fully connected feature related to the evaluation index, and the fully connected feature is input into the second fully connected layer to be mapped into the second image feature related to the evaluation index as the second material feature.
[0188] Step 307: Calculate the importance of the second material for the evaluation index according to the second material feature, and obtain a second score.
[0189] For the feature (i.e., the second material feature) represented by the visual content of the second candidate material, the historical cases of pushing the material to the client can be mined through machine learning or deep learning, so as to learn the importance of the second candidate material for the evaluation index on the visual content, denoted as a second score.
[0190] In one embodiment of the present application, step 307 can further include the following steps:
[0191] Step 3071: Determine a second gradient boosting decision tree and a second feature set.
[0192] In this embodiment, the gradient boosting decision tree can be pre-trained for the feature represented by the visual content of the video data of the material, denoted as a first gradient boosting decision tree, and the structure and parameters of the second gradient boosting decision tree are stored in a database. When the second candidate material is refined, the second gradient boosting decision tree and its parameters are loaded into the memory for running.
[0193] In addition, the second feature set can be pre-trained for the feature represented by the visual content of the video data of the material. The second feature set records the second sample features screened by the second decision boosting tree according to the importance for the evaluation index. The way of constructing the second sample features is consistent with the way of constructing the second material features, i.e., the second sample features are the features represented by the visual content of the video data (image data) of the material.
[0194] In a specific implementation, the material that has been pushed to the client is obtained as historical material.
[0195] In the case of sparse historical material, the material not pushed to the client can be obtained as the first original material, the first original material similar to the historical material is recalled as the second original material, the second original material is clustered around the historical material, the material cluster is obtained, and the part of the label generated after the historical material is pushed to the client is shared to the plurality of second original materials closest to the historical material in the material cluster, so that the plurality of second original materials become new historical materials. Since the client's behavior on similar materials is similar, the shared label to the first original material can have a certain accuracy, thereby greatly increasing the number of historical materials and ensuring the performance of the second gradient boosting decision tree.
[0196] Second material features representing visual content are extracted from historical video data.
[0197] Specifically, a content extraction network and a content understanding network are determined, multi-frame image data is extracted from video data of historical material, the image data is input into the content extraction network to extract first image features irrelevant to evaluation indicators, and the first image features are input into the content understanding network to extract second image features related to evaluation indicators as second material features.
[0198] In one example, the content extraction network includes a residual neural network and a temporal shift network. In this example, the image data is input into the residual neural network to extract residual features irrelevant to evaluation indicators, and the residual features are input into the temporal shift network to extract first image features irrelevant to evaluation indicators.
[0199] In another example, the content understanding network includes a first fully connected layer and a second fully connected layer. In this example, the first image features are input into the first fully connected layer to be mapped into fully connected features related to evaluation indicators, and the fully connected features are input into the second fully connected layer to be mapped into second image features related to evaluation indicators as second material features.
[0200] The historical material is divided into positive samples and negative samples, the weight of the positive samples is negatively correlated with the proportion of the number of positive samples, and the weight of the negative samples is negatively correlated with the proportion of the number of positive samples.
[0201] After weighting, when the second gradient boosting decision tree determines the branch point of the sub-tree in the training process, it will emphasize the categories with larger weights, thereby ensuring the performance of the second gradient boosting decision tree.
[0202] The second material features are used to train the second gradient boosting decision tree with the goal of optimizing the evaluation indicators, and the second gradient boosting decision tree outputs the importance of the second material features to the evaluation indicators when the training is completed.
[0203] Screening the second material feature whose importance degree is greater than the preset second threshold value as the second sample feature and writing the second sample feature into the second feature set.
[0204] Using the second sample feature to retrain the second gradient boosting decision tree with the optimization evaluation index as the target, saving the parameters of the second gradient boosting decision tree when the retraining of the second gradient boosting decision tree is completed, and through the two rounds of training, the accuracy of the second gradient boosting decision tree in calculating the importance degree of the second material feature for the evaluation index can be greatly improved.
[0205] Step 3072, screening the second material feature from the second sample feature as the second target feature.
[0206] In this embodiment, the second material feature can be compared with the second sample feature in the second feature set. If the second material feature is the same as the second sample feature in the second feature set, the second material feature is retained and recorded as the second target feature. If the second material feature is different from the second sample feature in the second feature set, the second material feature is filtered out.
[0207] Step 3073, inputting the second target feature into the second gradient boosting decision tree to calculate the importance degree of the second candidate material for the evaluation index as the second score.
[0208] After all the second target features are combined for the second material feature in the second candidate material, the second target features are input into the second gradient boosting decision tree. The second gradient boosting decision tree calculates the importance degree of the second candidate material for the evaluation index and records the importance degree as the second score.
[0209] Step 308, fusing the first score and the second score into a third score.
[0210] In this embodiment, the first score and the second score are fused into a third score by comprehensively referring to the features represented by the labels of the reference material in the semantic and the features represented by the video data of the material in the visual content, improving the feature dimension of the third score, and thus improving the accuracy of the evaluation material for the evaluation index.
[0211] The fusion can be linear fusion or nonlinear fusion, which is not limited in this embodiment.
[0212] Taking linear fusion as an example, the first score can be multiplied by the first weight matched with the label to obtain a first weight-adjusted value, the second score can be multiplied by the second weight matched with the video data to obtain a second weight-adjusted value, and thus the sum value between the first weight-adjusted value and the second weight-adjusted value is calculated as the third score.
[0213] The size relationship between the first weight and the second weight is different for different services, in some cases, the first weight can be greater than the second weight, in other cases, the first weight can be equal to the second weight, and in yet other cases, the first weight can be less than the second weight.
[0214] For example, for advertising data, the first weight is greater than the second weight.
[0215] Step 309, selecting the second candidate material with the highest third score as the third candidate material.
[0216] The second candidate materials are sorted in descending order according to the third score, and the n(n is a positive integer, n < k) second candidate materials with the highest ranking are selected as the third candidate materials.
[0217] Step 310, generating a push task for the third candidate material.
[0218] The push task is used to filter out part of the third candidate material by the user with the role of an optimizer and push it to the client.
[0219] It should be noted that, for the method embodiment, in order to simply describe, it is expressed as a series of action combinations, but those skilled in the art should know that the embodiments of the present application are not limited by the order of the described actions, because according to the embodiments of the present application, certain steps can be performed in other order or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification all belong to preferred embodiments, and the actions involved are not necessarily necessary for the embodiments of the present application.
[0220] Embodiment three
[0221] Figure 5 A structural block diagram of a video material screening device provided for the third embodiment of the present application, which can specifically include the following modules:
[0222] The recall module 501 is used to recall the material not pushed to the client as the first candidate material, and the material contains video data and is marked with a label;
[0223] The rough sorting module 502 is used to filter out part of the first candidate material as the second candidate material according to the semantic represented by the label, with the optimization evaluation index as the target;
[0224] The fine sorting module 503 is used to filter out part of the second candidate material as the third candidate material according to the semantic represented by the label and the visual content of the video data, with the optimization evaluation index as the target;
[0225] Task generation module 504 is used to generate push tasks for the third candidate materials. The push tasks are used by a user with the role of an optimizer to select some of the third candidate materials and push them to the client.
[0226] The evaluation metric is the data generated by pushing the material to the client and counting the operations triggered by the client on the material.
[0227] The video material screening device provided in this embodiment of the invention can execute the video material screening method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method.
[0228] Example 4
[0229] Figure 6 This is a schematic diagram of the structure of a computer device provided in Embodiment 4 of the present invention. Figure 6 A block diagram of an exemplary computer device 12 suitable for implementing embodiments of the present invention is shown. Figure 6 The computer device 12 shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of the present invention.
[0230] like Figure 6 As shown, the computer device 12 is represented in the form of a general-purpose computing device. The components of the computer device 12 may include, but are not limited to: one or more processors or processing units 16, system memory 28, and a bus 18 connecting different system components (including system memory 28 and processing unit 16).
[0231] Bus 18 represents one or more of several bus architectures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the various bus architectures. For example, these architectures include, but are not limited to, the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MAC) bus, the Enhanced ISA bus, the Video Electronics Standards Association (VESA) local bus, and the Peripheral Component Interconnect (PCI) bus.
[0232] Computer device 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by computer device 12, including volatile and non-volatile media, removable and non-removable media.
[0233] System memory 28 can include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. Computer device 12 can further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 34 can be provided for reading from and writing to a non-removable, non-volatile magnetic media (not shown and typically called a "hard drive"). Figure 6 Although not shown, a magnetic disk drive can also be utilized for reading from and writing to a removable, non-volatile magnetic disk (e.g., a "floppy disk"), and an optical disk drive can be utilized for reading from or writing to a removable, non-volatile optical disk (e.g., a CD-ROM, DVD-ROM or other optical media). In such instances, each drive can be connected to bus 18 by one or more data media interfaces. Storage 28 can include at least one program product having a set (e.g., at least one) of program modules that are configured to carry out the functions of embodiments of the application. Figure 6 Although not shown, a magnetic disk drive can also be utilized for reading from and writing to a removable, non-volatile magnetic disk (e.g., a "floppy disk"), and an optical disk drive can be utilized for reading from or writing to a removable, non-volatile optical disk (e.g., a CD-ROM, DVD-ROM or other optical media). In such instances, each drive can be connected to bus 18 by one or more data media interfaces. Storage 28 can include at least one program product having a set (e.g., at least one) of program modules that are configured to carry out the functions of embodiments of the application.
[0234] Program / utility 40, having a set (at least one) of program modules 42, can be stored in, for example, storage 28 by way of example, and not limitation, as well as an operating system, one or more application programs, other program modules, and program data, each or some combination thereof, can include an implementation of a network environment. Program modules 42 generally carry out the functions and / or methodologies of embodiments of the application as described herein.
[0235] Computer device 12 can also communicate with one or more external devices 14 such as a keyboard, a pointing device, a display 24, etc.; one or more devices that enable a user to interact with computer device 12; and / or one or more devices that enable computer device 12 to communicate with one or more other computing devices. Such communication can be via input / output (I / O) interfaces 22. Further, computer device 12 can communicate with one or more networks such as a local area network (LAN), a wide area network (WAN), and / or the Internet through network adapter 20. As depicted, network adapter 20 communicates with the other components of computer device 12 via bus 18. It should be appreciated that although not shown, other hardware and / or software modules could be utilized in conjunction with computer device 12 such as, but not limited to, microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archival storage systems, etc.
[0236] Processing unit 16 executes various program applications and data processing by running programs stored in system memory 28, such as to implement the method of screening video materials provided by embodiments of the application.
[0237] Embodiment five
[0238] Embodiment five of the present application also provides a computer readable storage medium, and the computer readable storage medium stores a computer program, the computer program is executed by a processor to realize each process of the video material screening method, and the same technical effects can be achieved. To avoid repetition, it will not be described here.
[0239] The computer readable storage medium may, for example, include but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any combination of the above. More specific examples (a non-exhaustive list) of the computer readable storage medium include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device.
[0240] Note that the above are only preferred embodiments of the present application and the principles of technology applied. Those skilled in the art will understand that the present application is not limited to the specific embodiments described herein, and those skilled in the art can make various obvious changes, readjustments and substitutions without departing from the scope of the present application. Therefore, although the present application has been described in more detail through the above embodiments, the present application is not limited to the above embodiments, and can include more other equivalent embodiments without departing from the concept of the present application, and the scope of the present application is determined by the scope of the appended claims.
Claims
1. A method for selecting video materials, characterized in that, include: Recall materials that have not been pushed to the client as the first candidate materials; the materials contain video data and are tagged. With the goal of optimizing the evaluation indicators, some of the first candidate materials are selected as the second candidate materials based on the semantics represented by the tags. With the goal of optimizing evaluation metrics, some of the second candidate materials are selected as the third candidate materials based on the semantics represented by the tags and the visual content of the video data. A push task is generated for the third candidate material. The push task is used by a user with the role of an optimizer to select some of the third candidate materials and push them to the client. The evaluation metric is the data generated by pushing the material to the client and counting the operations triggered by the client on the material.
2. The method according to claim 1, characterized in that, The step of selecting a portion of the first candidate materials as second candidate materials based on the semantic representation of the tags, with the goal of optimizing evaluation indicators, includes: Extract the semantic features of the first material from the tags of the first candidate material; The importance of the first candidate material to the evaluation index is calculated based on the characteristics of the first material, and this is used as the first score. The first candidate material with the highest score is selected as the second candidate material.
3. The method according to claim 2, characterized in that, The step of extracting the first material features representing semantics from the tags of the first candidate material includes: Determine the continuous bag-of-words model; The tags of the first candidate material are segmented to obtain multiple word groups; Encode the multiple phrases into a first word vector; For the current word group, the first word vector of other word groups belonging to the context is input into the continuous bag-of-words model and mapped to the second word vector of the current word group, which serves as the first material feature.
4. The method according to claim 3, characterized in that, The determination of the continuous bag-of-words model includes: Retrieve materials that have already been pushed to the client and use them as historical materials; The tags of the historical materials are segmented to obtain multiple word groups; Encode the multiple phrases into a first word vector; For the current word group, the first word vector of other word groups belonging to the context is input into the continuous bag-of-words model and mapped to the second word vector of the current word group; Calculate the label loss value based on the second word vector; Update the continuous bag-of-words model according to the label loss value; Determine whether the bag-of-words training conditions are met; if yes, determine that the continuous bag-of-words model training is complete; if no, return to execute the step of inputting the first word vector of other word groups belonging to the context into the continuous bag-of-words model and mapping it to the second word vector of the current word group.
5. The method according to claim 2, characterized in that, The step of calculating the importance of the first candidate material to the evaluation index based on the characteristics of the first material, as the first score, includes: Determine a first gradient boosting decision tree and a first feature set. The first feature set records the first sample features selected by the first gradient boosting decision tree according to their importance to the evaluation index. Select the first material features that are identical to the first sample features as the first target features; The first target feature is input into the first gradient boosting decision tree to calculate the importance of the first candidate material to the evaluation index, which is then used as the first score.
6. The method according to claim 5, characterized in that, The determination of the first gradient boosting decision tree and the first feature set includes: Retrieve materials that have already been pushed to the client and use them as historical materials; Extract the first material features representing semantics from the tags of the historical materials; With the goal of optimizing the evaluation index, a first gradient boosting decision tree is trained using the first material features. When the first gradient boosting decision tree is completed, it outputs the importance of the first material features to the evaluation index. The first material features whose importance is greater than a preset first threshold are selected and written into the first feature set as first sample features; With the goal of optimizing the evaluation index, the first gradient boosting decision tree is trained using the features of the first sample.
7. The method according to claim 6, characterized in that, The historical materials are divided into positive samples and negative samples. The weight of the positive samples is negatively correlated with the proportion of positive samples, and the weight of the negative samples is negatively correlated with the proportion of positive samples.
8. The method according to any one of claims 1-7, characterized in that, The step of selecting a portion of the second candidate materials as the third candidate materials, with the goal of optimizing evaluation indicators, based on the semantics represented by the tags and the visual content of the video data, includes: A first score is determined, which represents the importance of the first candidate material to the evaluation index under the semantics represented by the label; Extract second material features representing visual content from the video data of the second candidate material; Calculate the importance of the second material to the evaluation index based on the characteristics of the second material, and obtain the second score; The first score and the second score are combined to form a third score; The second candidate material with the highest third score is selected as the third candidate material.
9. The method according to claim 8, characterized in that, The step of extracting second material features representing visual content from the video data of the second candidate material includes: Identify content extraction networks and content understanding networks; Extract multiple frames of image data from the video data of the second candidate material; The image data is input into the content extraction network to extract a first image feature that is unrelated to the evaluation metric; The first image features are input into the content understanding network to extract second image features related to the evaluation metrics, which are then used as second material features.
10. The method according to claim 9, characterized in that, The determined content extraction network and content understanding network include: Retrieve materials that have already been pushed to the client and use them as historical materials; Extract multiple frames of image data from the video data of the historical materials; The image data is input into a pre-trained content extraction network for image classification to extract first image features that are unrelated to the evaluation metrics. The first image features are input into the content understanding network to extract the second image features related to the evaluation metrics; Based on the second image features, the content loss value is calculated according to the evaluation index; Update the content understanding network according to the content loss value; Determine whether the content training conditions are met; if yes, then determine that the content understanding network training is complete; if not, return to the step of inputting the image data into the pre-trained content extraction network for image classification to extract the first image features that are unrelated to the evaluation metrics.
11. The method according to claim 9, characterized in that, The content extraction network includes a residual neural network and a temporal shift network; The step of inputting the image data into the content extraction network to extract a first image feature that is unrelated to the evaluation metric includes: The image data is input into the residual neural network to extract residual features that are unrelated to the evaluation index; The residual features are input into the temporal shift network to extract the first image features that are independent of the evaluation metrics.
12. The method according to claim 9, characterized in that, The content understanding network includes a first fully connected layer and a second fully connected layer; The step of inputting the first image features into the content understanding network to extract second image features related to the evaluation metrics, as the second material features, includes: The first image features are input into the first fully connected layer and mapped to fully connected features related to the evaluation metric. The fully connected features are input into the second fully connected layer and mapped to second image features related to the evaluation metric, which are then used as second material features.
13. The method according to claim 8, characterized in that, The step of calculating the importance of the second material to the evaluation index based on the characteristics of the second material, and obtaining the second score, includes: Determine the second gradient boosting decision tree and the second feature set. The second feature set records the second sample features selected by the second gradient boosting decision tree according to their importance to the evaluation index. The second material features that match the second sample features are selected as the second target features; The second target feature is input into the second gradient boosting decision tree to calculate the importance of the second candidate material to the evaluation index, which is then used as the second score.
14. The method according to claim 13, characterized in that, The determination of the second gradient boosting decision tree and the second feature set includes: Retrieve materials that have already been pushed to the client and use them as historical materials; Extract second material features representing visual content from the historical video data; With the goal of optimizing the evaluation index, a second gradient boosting decision tree is trained using the second material features. When the training is completed, the second gradient boosting decision tree outputs the importance of the second material features to the evaluation index. The second material features whose importance is greater than a preset second threshold are selected and written into the second feature set as second sample features; With the goal of optimizing the evaluation index, the second sample features are used to train the second gradient boosting decision tree.
15. The method according to claim 8, characterized in that, The step of merging the first score and the second score into a third score includes: The first score is multiplied by a first weight that matches the label to obtain a first adjustment value; The second score is multiplied by a second weight that matches the video data to obtain a second adjustment value; Calculate the sum between the first adjustment weight and the second adjustment weight, and use it as the third score; Wherein, the first weight is greater than the second weight.
16. The method according to any one of claims 4, 6, 10, and 14, characterized in that, Also includes: Obtain materials that have not yet been pushed to the client and use them as the first original materials; Recall the first original material that is similar to the historical material, and use it as the second original material; Cluster the second original material with the historical material as the center to obtain material clusters; In the material cluster, some tags generated after the historical material is pushed to the client are shared with multiple second original materials that are closest to the historical material, so that the multiple second original materials become new historical materials.
17. A video material screening device, characterized in that, include: The recall module is used to recall materials that have not been pushed to the client as the first candidate materials. The materials contain video data and are tagged. The coarse ranking module is used to select a portion of the first candidate materials as second candidate materials based on the semantics represented by the tags, with the goal of optimizing the evaluation indicators; The fine-ranking module is used to select some of the second candidate materials as the third candidate materials based on the semantics represented by the tags and the visual content of the video data, with the goal of optimizing the evaluation indicators. The task generation module is used to generate push tasks for the third candidate materials. The push tasks are used by a user with the role of an optimizer to select some of the third candidate materials and push them to the client. The evaluation metric is the data generated by pushing the material to the client and counting the operations triggered by the client on the material.
18. A computer device, characterized in that, The computer device includes: One or more processors; Memory, used to store one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the video material filtering method as described in any one of claims 1-16.
19. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, which, when executed by a processor, implements the method for selecting video material as described in any one of claims 1-16.
Citation Information
Patent Citations
Recommendation method and device of video data and server
CN108509465A
Video optimization recommendation method and device and readable storage medium
CN109831684A