Click-through rate prediction model training method and system for video push

By acquiring and cross-combining video and user feature data, generating combined feature data and training a click-through rate estimate model, the problem of insufficient model accuracy in the prior art is solved, and a video push model with high accuracy and generalization is achieved.

CN114510620BActive Publication Date: 2025-05-13SHANGHAI BILIBILI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202011165348.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-10-27
Publication Date
2025-05-13
Estimated Expiration
2040-10-27

AI Technical Summary

Technical Problem

How to train machine learning models to obtain high-accuracy click-through rate estimate models for video push, there is currently the problem of insufficient accuracy in training models.

Method used

By obtaining multiple sets of training data, including video feature data and user feature data, and generating more combined feature data through cross-combination, the feature vector is constructed to train the click-through rate estimate model.

Benefits of technology

It improves the accuracy and generalization of the click-through rate estimate model, enhances the recommendation effect of video push, and ensures the applicability of the model among different user groups.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114510620B_ABST
    Figure CN114510620B_ABST
Patent Text Reader

Abstract

The present application provides a method for training a click rate prediction model for video push, including obtaining multiple sets of training data; wherein each set of training data includes multiple video feature data and multiple user feature data; obtaining multiple combined feature data of each set of training data; obtaining a feature vector corresponding to each set of training data according to each set of training data and the multiple combined feature data of each set of training data; and training a model to be trained according to the feature vector corresponding to each set of training data to obtain a click rate prediction model for video push. The method described in the present application can train a highly accurate click rate prediction model for video push.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of computer technology, and in particular to a method, system, device and computer-readable storage medium for training a click-through rate prediction model for video push. Background Art

[0002] With the development of the Internet, people have begun to use online platforms for entertainment and various transactions. How to filter and push differentiated data (such as product or service data) to each user has become a concern for all parties. With the development of machine learning, people have begun to use machine learning to filter and push data. For example, the click-through rate prediction model can be used to predict the user click-through rate of each video, and the appropriate video can be pushed to the appropriate user.

[0003] How to train the model to obtain a highly accurate click-through rate prediction model for video push has become one of the current problems to be solved. Summary of the invention

[0004] The purpose of the embodiments of the present application is to provide a click-through rate prediction model training method, system, computer device and computer-readable storage medium for video push, which are used to solve the technical problem of how to train a machine learning model to obtain a high-accuracy click-through rate prediction model for video push.

[0005] One aspect of an embodiment of the present application provides a method for training a click-through rate prediction model for video push, the method comprising: obtaining multiple groups of training data; wherein each group of training data comprises multiple video feature data and multiple user feature data; obtaining multiple combined feature data for each group of training data; wherein the multiple combined feature data comprises multiple first combined feature data, each first combined feature data being a combination of one video feature data and one user feature data in the group of training data; obtaining a feature vector corresponding to each group of training data based on the each group of training data and the multiple combined feature data of the each group of training data; and training the model to be trained based on the feature vector corresponding to each group of training data to obtain a click-through rate prediction model for video push.

[0006] Optionally, the multiple video feature data include multiple video type feature data; obtaining multiple combined feature data for each group of training data includes: cross-combining each user feature data in the i-th group of training data with at least one of the multiple video type feature data in the i-th group of training data to obtain multiple combined feature data corresponding to the i-th group of training data; wherein the i-th group of training data is any group of training data among the multiple groups of training data.

[0007] Optionally, the video is a promotional video, each group of training data includes a promotional resource position, and the multiple video feature data include multiple video type feature data; obtaining multiple combined feature data of each group of training data, including: cross-combining each user feature data in the i-th group of training data with at least one of the multiple video type feature data in the i-th group of training data, so as to obtain multiple first combined feature data corresponding to the i-th group of training data; wherein the i-th group of training data is any group of training data among the multiple groups of training data; cross-combining the promotional resource position in the i-th group of training data with at least one of the multiple video type feature data in the i-th group of training data, so as to obtain second combined feature data corresponding to the i-th group of training data; and obtaining multiple combined feature data corresponding to the i-th group of training data based on the multiple first combined feature data corresponding to the i-th group of training data and the second combined feature data corresponding to the i-th group of training data.

[0008] Optionally, the multiple video feature data include the multiple video type feature data and the multiple video numerical feature data; and according to the each group of training data and the multiple combined feature data of the each group of training data, obtaining a feature vector corresponding to the each group of training data, including: according to the multiple video numerical feature data in the i-th group of training data and the multiple combined feature data obtained according to the i-th group of training data, obtaining the i-th feature vector corresponding to the i-th group of training data.

[0009] Optionally, according to the multiple video numerical feature data in the i-th group of training data and the multiple combined feature data obtained according to the i-th group of training data, an i-th feature vector corresponding to the i-th group of training data is obtained, including: discretizing the multiple video numerical feature data into equidistant buckets to obtain a corresponding multiple discretized values; performing hash coding on each discretized value and each combined feature data to obtain multiple hash coding values; and constructing the i-th feature vector according to the multiple hash coding values.

[0010] Optionally, the multiple video type feature data include at least the following items: video tag, partition and video uploader identifier; and the multiple video numerical feature data include video play time and video play times.

[0011] Optionally, the model to be trained is trained according to the feature vector corresponding to each group of training data, including: introducing an L1 regularization term or an L2 regularization term to train the model to be trained; and during the training process, removing non-important features not selected by the L1 regularization term or the L2 regularization term.

[0012] One aspect of an embodiment of the present application provides a click-through rate prediction model training system for video push, including: a first acquisition module, used to acquire multiple groups of training data; wherein each group of training data includes multiple video feature data and multiple user feature data; a second acquisition module, used to acquire multiple combined feature data for each group of training data; wherein the multiple combined feature data include multiple first combined feature data, each first combined feature data is obtained by combining one video feature data and one user feature data in the group of training data; a third acquisition module, used to acquire a feature vector corresponding to each group of training data based on each group of training data and the multiple combined feature data of each group of training data; and a training module, used to train the model to be trained based on the feature vector corresponding to each group of training data to obtain a click-through rate prediction model for video push.

[0013] One aspect of an embodiment of the present application provides a computer device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor is used to implement the steps of the above-mentioned click-through rate prediction model training method for video push when executing the computer program.

[0014] One aspect of an embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. The computer program can be executed by at least one processor to enable the at least one processor to perform the steps of the above-mentioned click-through rate prediction model training method for video push.

[0015] The click-through rate prediction model training method, system, device and computer-readable storage medium for video push provided in the embodiments of the present application can cross-examine user feature data and video feature data to obtain more combined feature data and use them in model training, so as to train a highly accurate click-through rate prediction model for video push. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 A schematic diagram of an environmental application according to an embodiment of the present application is schematically shown;

[0017] Figure 2 An exemplary training framework is schematically shown;

[0018] Figure 3 A flowchart of a click rate prediction model training method for video push according to Embodiment 1 of the present application is schematically shown;

[0019] Figure 4 for Figure 3 Flow chart of sub-steps of step S302;

[0020] Figure 5 The characteristic cross combination scheme for the example;

[0021] Figure 6 for Figure 3 Another sub-step flow chart of step S302;

[0022] Figure 7 Another characteristic cross-combination scheme for the example;

[0023] Figure 8 for Figure 3 Flow chart of sub-steps of step S304;

[0024] Fig. 9 The numerical features of the cross-combination of the features of the example;

[0025] Fig.10 for Figure 8 Flow chart of sub-steps of step S800;

[0026] Fig.11 for Figure 3 Flow chart of sub-steps of step S306;

[0027] Fig.12 A block diagram schematically shows a method for training a click rate prediction model for video push according to the second embodiment of the present application; and

[0028] Fig.13 The hardware architecture diagram of a computer device suitable for implementing a click-through rate prediction model training method for video push according to the third embodiment of the present application is schematically shown. DETAILED DESCRIPTION

[0029] In order to make the purpose, technical solutions and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present application.

[0030] It should be noted that the descriptions of "first", "second", "third", etc. in this application are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In addition, the technical solutions between the various embodiments can be combined with each other, but they must be based on the ability of ordinary technicians in this field to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that such combination of technical solutions does not exist and is not within the scope of protection required by this application.

[0031] The click rate prediction model plays an important role in video recommendation (e.g., promotional video recommendation) and other fields. The applicant believes that:

[0032] (1) The relevant data at the video level are generally massive structured discrete and categorical data, and high-dimensional data brings challenges.

[0033] (2) The existing technology lacks some attribute features of the video and cannot provide users with accurate information for judging clicks at the video level. The applicant believes that: effective features for users to click on promotional videos can be mined in complex video data through clever feature engineering, and important features can be screened out through embedding training of machine learning in feature selection, which can better characterize video information to improve users' click preferences for videos, significantly improve the accuracy of model estimation, ensure the generalization of the model, and optimize the recommendation effect of videos (such as promotional videos).

[0034] The following are some explanations of terms in this application:

[0035] Click-through rate (CTR), for example, means the actual number of clicks on a video (promotional video) divided by the number of impressions of the video (promotional video).

[0036] Embedded method: Use machine learning algorithms and model training to obtain the weight coefficients of each feature, and select features based on the coefficients in descending order.

[0037] FTRL: Follow the regularized Leader, an online algorithm proposed by Google, which has been greatly optimized in engineering implementation.

[0038] LR: Logistic Regression, a linear classification model.

[0039] AUC: When a positive and negative sample is randomly selected from the positive and negative sample sets respectively, the probability that the predicted value of the positive sample is greater than that of the negative sample.

[0040] L1 regularization: refers to the sum of the absolute values ​​of each element in the weight vector w, which can produce a sparse model for feature selection.

[0041] L2 regularization: refers to the sum of the squares of each element in the weight vector w and then taking the square root. By limiting the size of the norm of w, the model space is restricted, thereby avoiding overfitting to a certain extent.

[0042] Incremental learning: A model training method that adjusts the model in a timely manner based on online feedback data to reflect online changes and improve the accuracy of online predictions.

[0043] Promotional video, can be an advertisement.

[0044] Partition, which can be a video section partition or other partitions.

[0045] Figure 1 A schematic diagram of an environment for training a click-through rate prediction model for video push is shown.

[0046] like Figure 1 As shown, the environment diagram may include a database server 100, a network 200, and a server 300. The database server 100 may be located in a data center such as a single location, or distributed in different geographical locations (e.g., in multiple locations). The database server 100 may provide services via the network 200. The network 200 includes various network devices, such as routers, switches, multiplexers, hubs, modems, bridges, repeaters, firewalls, proxy devices, and / or the like. The network 9 may include physical links, such as coaxial cable links, twisted pair cable links, optical fiber links, combinations thereof, and the like. The network 200 may also include wireless links, such as cellular links, satellite links, Wi-Fi links, and the like.

[0047] The database server 100 can store various types of training data, such as offline training data.

[0048] The server 300 can provide various services. For example, the server 300 can obtain offline training data from the database server 100, and perform model training based on the offline training data to obtain a click rate prediction model. The server 300 can also obtain online training data, and perform model training based on the online training data to obtain a click rate prediction model.

[0049] As an example, Figure 2As shown, in order to improve the video business indicators, videos with better relevance are recommended to users in multiple dimensions. The overall training process of the click-through rate prediction model can be as follows: (1) Through logs and other information, a large number of users' exposure and click behavior data (exposure means video display in the list, and click is the same) are obtained. After being processed by HIVE and Python, these data can become training data for model training. (2) The baseline model and the experimental model are trained by user feature data of user features, and the experimental model is trained by user feature data of user features and video feature data of video features (such as Example 1 below), and the model output is the score of the video. (3) After offline evaluation, the experimental model that meets the training standards is deployed online to perform online estimation and video score ranking. It should be noted that the baseline model is a control of the experimental model and can be used to evaluate the training effect of the experimental model; when the video in this application is a promotional video, in addition to some inherent features of the promotional video, the features mined for the unique attributes of the video are video features. Therefore, the experimental model training can add relevant features of the promotional video dimension and relevant features of the video dimension to the baseline model training.

[0050] It should be noted that the click-through rate prediction model training method for video push provided in the embodiment of the present application can be executed by the server 300, and the click-through rate prediction model training system for video push can be set in the server 300.

[0051] It should be noted that Figure 1 The number of database servers 100, networks 200 and servers 300 in the embodiment is only for illustration. Any number of database servers, networks and servers may be provided as required. It should be noted that, in the case where training data is stored in the server 300, the database server 300 may not be provided.

[0052] Embodiment 1

[0053] Figure 3 The flowchart of the click rate prediction model training method for video push according to the first embodiment of the present application is schematically shown. It can be understood that the flowchart in the embodiment of the present application is not used to limit the order of executing the steps.

[0054] like Figure 3 As shown, the click rate prediction model training method for video push may include steps S300 to S306, wherein:

[0055] Step S300, obtaining multiple sets of training data; wherein each set of training data includes multiple video feature data and multiple user feature data.

[0056] It should be noted that the multiple sets of training data may be offline data or online data.

[0057] As an example, the plurality of video feature data may include data related to the following features: (1) statistical features of the video, such as: play time, play times, like times, favorite times, etc.; (2) category features such as video tags, such as: video tags, video types, video attributes, partitions, video uploader identifiers, etc. Among them, video tags include, for example, "game", "PUBG", "beauty", "summer makeup", "chicken eating", etc.

[0058] As an example, the multiple user feature data may include relevant data of the following features: portrait interests, age, gender, membership level, screen type, province, city, and model of device used, such as Huawei Mate 40, etc.

[0059] Step S302, obtaining multiple combined feature data of each set of training data.

[0060] The plurality of combined feature data include a plurality of first combined feature data, each of which is obtained by combining one of the video feature data and one of the user feature data in the group of training data.

[0061] Optimize the video features of users on the video, use single features and combined features for subsequent model training, and improve the accuracy of the model.

[0062] As an example, the multiple video feature data include multiple video type feature data, which may include video tag data, video type data, video attribute data, partition data, video uploader identification, and the like.

[0063] like Figure 4 As shown, to improve the accuracy of the model, the step S302 may include step S400: cross-combining each user feature data in the i-th group of training data with at least one of the multiple video type feature data in the i-th group of training data to obtain multiple combined feature data corresponding to the i-th group of training data; wherein the i-th group of training data is any one of the multiple groups of training data. For ease of understanding, as Figure 5As shown, the following provides a preferred example: through the following combination of features, corresponding multiple combined feature data can be obtained: portrait interest-video tag, age-video tag, age-district, age-video uploader ID, gender-video tag, gender-district, gender-video uploader ID, membership level-video tag, screen type-video tag, province-video tag, province-district, province-video uploader ID, city-video tag, city-district, city-video uploader ID, model-video tag.

[0064] As an example, the video is a promotional video, each set of training data includes a promotional resource location, and the plurality of video feature data includes a plurality of video type feature data.

[0065] like Figure 6 As shown, in order to improve the accuracy of the model, the step S302 may include steps S600 to S604, wherein: step S600, cross-combining each user feature data in the i-th group of training data with at least one of the multiple video type feature data in the i-th group of training data, to obtain multiple first combined feature data corresponding to the i-th group of training data; wherein the i-th group of training data is any one of the multiple groups of training data; step S602, cross-combining the promotion resource bits in the i-th group of training data with at least one of the multiple video type feature data in the i-th group of training data, to obtain second combined feature data corresponding to the i-th group of training data; and step S604, obtaining multiple combined feature data corresponding to the i-th group of training data based on the multiple first combined feature data corresponding to the i-th group of training data and the second combined feature data corresponding to the i-th group of training data. For ease of understanding, as Figure 7 As shown, the following provides a preferred example: through the following combination of features, corresponding multiple combined feature data can be obtained: portrait interest-video tag, age-video tag, age-district, age-video uploader ID, gender-video tag, gender-district, gender-video uploader ID, membership level-video tag, screen type-video tag, province-video tag, province-district, province-video uploader ID, city-video tag, city-district, city-video uploader ID, model-video tag, video tag-promotion resource position.

[0066] like Figure 7As shown in the figure, single features with continuous values ​​(such as the number of video plays and the length of time played) with a large value range and a floating point type need to be appropriately transformed, and the number of video plays and the length of time played are not used in the feature crossover operation. User feature data with category attributes, promotion resource positions, and video feature data with category attributes are used for feature crossover processing. The real-time promotion resource positions in the context features are determined when the promotion video is requested. Crossing with the video feature data with category attributes can distinguish the promotion video characteristics of different promotion resource positions and improve the training effect. Crossing the video feature data and the user feature data to form combined feature data can help characterize the user's feature representation of different videos and their changes over time, and improve the generalization of the model to videos.

[0067] Step S304: acquiring a feature vector corresponding to each set of training data according to each set of training data and a plurality of combined feature data of each set of training data.

[0068] In one embodiment, taking the i-th group of training data as an example:

[0069] At least part of the data can be selected from the following data, and the i-th feature vector corresponding to the i-th group of training data can be obtained based on the at least part of the data: multiple video feature data in the i-th group of training data, multiple user feature data in the i-th group of training data, and multiple combined feature data corresponding to the i-th group of training data.

[0070] In another embodiment, continuing to take the i-th group of training data as an example:

[0071] The i-th group of training data includes multiple video type feature data and multiple video numerical feature data. The multiple video type feature data include at least the following items: video label, partition and video uploader identifier; and the multiple video numerical feature data include video play time and video play times. Each video type feature data is combined with one or more of the multiple user feature data to form combined feature data. Figure 8 As shown, in order to reduce the data processing burden and thus not use separate video type feature data and separate user feature data as input data for the training model, the step S304 can be obtained through step S800: according to the multiple video numerical feature data in the i-th group of training data and the multiple combined feature data obtained according to the i-th group of training data, the i-th feature vector corresponding to the i-th group of training data is obtained.

[0072] In another embodiment, before inputting into the model for training, appropriate transformation and processing of the numerical features can bring better indicators, such as normalization and logarithm. In this embodiment, the continuous values ​​of the video (multiple video numerical feature data, such as video playback time and video playback times) are discretized into equidistant buckets, and cross-combinations and hash codes are performed between category features such as various video type feature data (video tags, video attributes, video types), such as Fig. 9 As shown, user feature data and video feature data can be IDed (digitized) for storage and calculation, such as video tag A corresponds to 154551, partition B corresponds to 13659884, and video uploader ID corresponds to 31864; user A's interest portrait corresponds to 11296, user A's province corresponds to 11657551, user A's city corresponds to 1207642, and user A's model corresponds to 6942. The IDed video tags and each user feature data are crossed to obtain the corresponding combined feature data (the combined feature data of video tag A and user A's interest portrait is 154551_11296), and the obtained combined feature data is hashed (such as the hash code value of 154551_11296 is 4619129638092800), which reduces the size of the model and the processing speed without losing accuracy. As an example, Fig.10 As shown, the step S800 may include steps S1000 to S1004, wherein: step S1000, discretizing the multiple video numerical feature data into equidistant buckets to obtain a corresponding multiple discretized numerical values; step S1002, hash coding each discretized numerical value and each combined feature data to obtain a plurality of hash code values; each hash code value corresponds to a discretized numerical value or a combined feature data; and step S1004, constructing the i-th feature vector according to the multiple hash code values.

[0073] Step S306: training the model to be trained according to the feature vector corresponding to each set of training data to obtain a click rate prediction model for video push.

[0074] The model to be trained may be an LR model or other models.

[0075] like Fig.11As shown, step S306 may include S1100 to S1102, wherein: step S1100, introducing L1 regularization term or L2 regularization term to train the model to be trained; step S1100, during the training process, removing non-important features not selected by L1 regularization term or L2 regularization term. The L1 regularization term and L2 regularization term generate a sparse weight matrix for selecting features and preventing overfitting, thereby improving the generalization of the model, and accurately measuring the user's click preference for the video. In specific implementation, the model to be trained can be trained by the L1 embedding method based on FTRL, automatically screening features, introducing a feature exit mechanism in the training framework, adjusting the influence of low-frequency sparse parameters on the model, and removing features that have not been updated for a long time (non-important features) to control the model size within a reasonable range.

[0076] Of course, in order to further improve the accuracy of the click-through rate prediction model, online iterative training methods can be used to train the model.

[0077] The technical solution described in the first embodiment has at least the following technical effects:

[0078] (1) It can improve the prediction accuracy of the trained click rate prediction model for videos.

[0079] (2) It can improve the generalization and AUC index of the trained click rate prediction model for videos.

[0080] (3) When the video is a promotional video, the combined features formed by the promotional resource position and the video features are further added to further improve the accuracy of the click-through rate prediction model for the video promotion video.

[0081] (4) Automatically select features to control the model size within a reasonable range and prevent model overfitting.

[0082] (5) The model size and processing speed can be reduced without losing accuracy.

[0083] Embodiment 2

[0084] Fig.12 A block diagram of a click-through rate prediction model training method for video push according to the second embodiment of the present application is schematically shown. The click-through rate prediction model training method for video push can be divided into one or more program modules, and one or more program modules are stored in a storage medium and executed by one or more processors to complete the embodiment of the present application. The program module referred to in the embodiment of the present application refers to a series of computer program instruction segments that can perform specific functions. The following description will specifically introduce the functions of each program module of this embodiment.

[0085] like Fig.12As shown, the click rate prediction model training system 1200 for video push may include a first acquisition module 1210, a first acquisition module 1220, a first acquisition module 1230 and a training module 1240, wherein:

[0086] The first acquisition module 1210 is used to acquire multiple sets of training data; wherein each set of training data includes multiple video feature data and multiple user feature data;

[0087] The second acquisition module 1220 acquires a plurality of combined feature data of each set of training data; wherein the plurality of combined feature data includes a plurality of first combined feature data, each of which is obtained by combining one video feature data and one user feature data in the set of training data;

[0088] A third acquisition module 1230 is used to acquire a feature vector corresponding to each set of training data according to each set of training data and a plurality of combined feature data of each set of training data; and

[0089] The training module 1240 is used to train the to-be-trained model according to the feature vector corresponding to each set of training data, so as to obtain a click rate prediction model for video push.

[0090] In an exemplary embodiment, the multiple video feature data include multiple video type feature data; the second acquisition module 1220 is also used to cross-combine each user feature data in the i-th group of training data with at least one of the multiple video type feature data in the i-th group of training data to obtain multiple combined feature data corresponding to the i-th group of training data; wherein the i-th group of training data is any group of training data among the multiple groups of training data.

[0091] In an exemplary embodiment, the video is a promotional video, each group of training data includes a promotional resource position, and the multiple video feature data include multiple video type feature data; the second acquisition module 1220 is also used to cross-combine each user feature data in the i-th group of training data with at least one of the multiple video type feature data in the i-th group of training data to obtain multiple first combined feature data corresponding to the i-th group of training data; wherein the i-th group of training data is any group of training data among the multiple groups of training data; the promotional resource position in the i-th group of training data is cross-combined with at least one of the multiple video type feature data in the i-th group of training data to obtain second combined feature data corresponding to the i-th group of training data; and multiple combined feature data corresponding to the i-th group of training data are obtained based on the multiple first combined feature data corresponding to the i-th group of training data and the second combined feature data corresponding to the i-th group of training data.

[0092] In an exemplary embodiment, the multiple video feature data include the multiple video type feature data and the multiple video numerical feature data; the third acquisition module 1230 is also used to obtain the i-th feature vector corresponding to the i-th group of training data based on the multiple video numerical feature data in the i-th group of training data and the multiple combined feature data obtained based on the i-th group of training data.

[0093] In an exemplary embodiment, the third acquisition module 1230 is also used to discretize the multiple video numerical feature data into equidistant buckets to obtain a corresponding multiple discretized numerical values; hash encode each discretized numerical value and each combined feature data to obtain multiple hash code values; and construct the i-th feature vector based on the multiple hash code values.

[0094] In an exemplary embodiment, the plurality of video type feature data include at least the following items: video tag, partition, and video uploader identifier; and the plurality of video numerical feature data include video play duration and video play count.

[0095] In an exemplary embodiment, the training module 1240 is further used to introduce an L1 regularization term or an L2 regularization term to train the model to be trained; and during the training process, remove non-important features that are not selected by the L1 regularization term or the L2 regularization term.

[0096] Embodiment 3

[0097] Fig.13 The schematic diagram of the hardware architecture of a computer device suitable for implementing a click rate prediction model training method for video push according to the third embodiment of the present application is shown schematically. The computer device 1300 may be a server 300 or a part of a server 300. In the present embodiment, the computer device 1300 is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions. For example, it may be a rack server, a blade server, a tower server, or a cabinet server (including an independent server, or a server cluster consisting of multiple servers), etc. Fig.13 As shown, the computer device 1300 includes at least but is not limited to: a memory 1310, a processor 1320, and a network interface 1330 which can communicate with each other through a system bus. Among them:

[0098] The memory 1310 includes at least one type of computer-readable storage medium, and the readable storage medium includes a flash memory, a hard disk, a multimedia card, a card-type memory (e.g., an SD or DX memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory 1310 may be an internal storage module of the computer device 1300, such as a hard disk or a memory of the computer device 1300. In other embodiments, the memory 1310 may also be an external storage device of the computer device 1300, such as a plug-in hard disk equipped on the computer device 1300, a smart memory card (Smart Media Card, referred to as SMC), a secure digital (Secure Digital, referred to as SD) card, a flash card, etc. Of course, the memory 1310 may also include both the internal storage module of the computer device 1300 and its external storage device. In the embodiment of the present application, the memory 1310 is generally used to store the operating system and various application software installed on the computer device 1300, such as the program code of the click rate prediction model training method for video push, etc. In addition, the memory 1310 can also be used to temporarily store various data that have been output or will be output.

[0099] In some embodiments, the processor 1320 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor 1320 is generally used to control the overall operation of the computer device 1300, such as performing control and processing related to data interaction or communication with the computer device 1300. In this embodiment, the processor 1320 is used to run the program code stored in the memory 1310 or process data.

[0100] The network interface 1330 may include a wireless network interface or a wired network interface, and the network interface 1330 is generally used to establish a communication link between the computer device 1300 and other computer devices. For example, the network interface 1330 is used to connect the computer device 1300 to an external terminal through a network, and to establish a data transmission channel and a communication link between the computer device 1300 and the external terminal. The network may be a wireless or wired network such as an intranet, the Internet, the Global System of Mobile communication (GSM), Wideband Code Division Multiple Access (WCDMA), 4G network, 5G network, Bluetooth, Wi-Fi, etc.

[0101] It should be pointed out that Fig.13 Only a computer device having components 1310 - 1330 is shown, but it should be understood that implementing all of the components shown is not a requirement, and more or fewer components may alternatively be implemented.

[0102] In this embodiment, the click-through rate prediction model training method for video push stored in the memory 1310 can also be divided into one or more program modules and executed by one or more processors (processor 1320 in this embodiment) to complete this application.

[0103] Embodiment 4

[0104] This embodiment also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the click-through rate prediction model training method for video push in the embodiment are implemented.

[0105] In this embodiment, the computer-readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (for example, SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, disk, optical disk, etc. In some embodiments, the computer-readable storage medium may be an internal storage unit of a computer device, such as a hard disk or memory of the computer device. In other embodiments, the computer-readable storage medium may also be an external storage device of a computer device, such as a plug-in hard disk, a smart memory card (Smart Media Card, referred to as SMC), a secure digital (Secure Digital, referred to as SD) card, a flash card (Flash Card), etc. equipped on the computer device. Of course, the computer-readable storage medium may also include both an internal storage unit of a computer device and an external storage device thereof. In this embodiment, the computer-readable storage medium is generally used to store an operating system and various application software installed on a computer device, such as the program code of a click-through rate prediction model training method for video push in the embodiment, etc. In addition, the computer-readable storage medium may also be used to temporarily store various types of data that have been output or are to be output.

[0106] Obviously, those skilled in the art should understand that the modules or steps of the above-mentioned embodiments of the present application can be implemented by a general computing device, they can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices, and optionally, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, and in some cases, the steps shown or described can be executed in a different order from that herein, or they can be made into individual assembled circuit modules, or multiple modules or steps therein can be made into a single assembled circuit module for implementation. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.

[0107] It should be noted that video is one of the main ways of information dissemination. Realizing automated and intelligent video content processing and distribution is the only way to handle massive amounts of video. Similar to this are product recommendations and content recommendations, which recommend items of interest to users based on item information and user clicks, favorites, and other behaviors. In Internet recommendation and promotion videos, feature engineering maps the original data (user and item) space to a new feature vector space, enabling the model to better learn the patterns in the data. The complex network of interactions between users and videos, promotional video hosts, and the environment is the original data. Mining effective features and learning the probability distribution of clicks between users and items through models is an effective strategy to improve business indicators. In this application, a large number of video features that have been overlooked by the industry have been mined, greatly improving the accuracy of estimates.

[0108] The above are only preferred embodiments of the present application, and are not intended to limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A click rate prediction model training method for video push, characterized in that: The method comprises: Acquire multiple sets of training data; wherein each set of training data includes multiple video feature data and multiple user feature data; Acquire multiple combined feature data of each set of training data; wherein the multiple combined feature data include multiple first combined feature data, each first combined feature data is obtained by combining one video feature data and one user feature data in the set of training data; According to each set of training data and the multiple combined feature data of each set of training data, obtaining a feature vector corresponding to each set of training data; and The model to be trained is trained according to the feature vector corresponding to each set of training data to obtain a click rate prediction model for video push; Wherein, the video is a promotional video, each set of training data includes a promotional resource position, and the plurality of video feature data includes a plurality of video type feature data; and obtaining a plurality of combined feature data of each set of training data includes: Cross-combining each user feature data in the i-th group of training data with at least one of the multiple video type feature data in the i-th group of training data to obtain multiple first combined feature data corresponding to the i-th group of training data; wherein the i-th group of training data is any one group of training data in the multiple groups of training data; Cross-combining the promotion resource position in the i-th group of training data with at least one of the plurality of video type feature data in the i-th group of training data to obtain second combined feature data corresponding to the i-th group of training data; and A plurality of combined feature data corresponding to the i-th group of training data is obtained according to the plurality of first combined feature data corresponding to the i-th group of training data and the second combined feature data corresponding to the i-th group of training data.

2. The click rate prediction model training method for video push according to claim 1, characterized in that: The plurality of video feature data include a plurality of video type feature data; Get multiple combined feature data for each set of training data, including: Each user feature data in the i-th group of training data is cross-combined with at least one of the multiple video type feature data in the i-th group of training data to obtain multiple combined feature data corresponding to the i-th group of training data; wherein the i-th group of training data is any one group of training data among the multiple groups of training data.

3. The click rate prediction model training method for video push according to claim 1 or 2, characterized in that: The plurality of video feature data include the plurality of video type feature data and a plurality of video value feature data; According to each set of training data and the multiple combined feature data of each set of training data, obtaining a feature vector corresponding to each set of training data includes: According to the plurality of video numerical feature data in the i-th group of training data and the plurality of combined feature data obtained according to the i-th group of training data, an i-th feature vector corresponding to the i-th group of training data is obtained.

4. The click rate prediction model training method for video push according to claim 3 is characterized in that: Acquiring an i-th feature vector corresponding to the i-th group of training data according to the plurality of video numerical feature data in the i-th group of training data and the plurality of combined feature data obtained according to the i-th group of training data, comprising: The plurality of video numerical feature data are respectively discretized into equidistant buckets to obtain a plurality of corresponding discretized values; Performing hash coding on each discretized value and each combined feature data to obtain multiple hash coding values; and The i-th feature vector is constructed according to the multiple hash code values.

5. The click rate prediction model training method for video push according to claim 3 is characterized in that: The multiple video type feature data include at least the following items: video tag, partition and video uploader identifier; and the multiple video numerical feature data include video play time and video play times.

6. The click rate prediction model training method for video push according to claim 1 or 2, characterized in that: Training the model to be trained according to the feature vector corresponding to each set of training data includes: Introducing an L1 regularization term or an L2 regularization term to train the model to be trained; and During training, unimportant features that are not selected by the L1 regularization term or the L2 regularization term are removed.

7. A click rate prediction model training system for video push, characterized in that: include: A first acquisition module is used to acquire multiple sets of training data; wherein each set of training data includes multiple video feature data and multiple user feature data; A second acquisition module is used to acquire a plurality of combined feature data of each set of training data; wherein the plurality of combined feature data includes a plurality of first combined feature data, each of which is obtained by combining one of the video feature data and one of the user feature data in the set of training data; A third acquisition module is used to acquire a feature vector corresponding to each set of training data according to each set of training data and a plurality of combined feature data of each set of training data; and A training module, used for training the to-be-trained model according to the feature vector corresponding to each set of training data, so as to obtain a click rate prediction model for video push; Wherein, the video is a promotional video, each set of training data includes a promotional resource position, and the plurality of video feature data includes a plurality of video type feature data; and obtaining a plurality of combined feature data of each set of training data includes: Cross-combining each user feature data in the i-th group of training data with at least one of the multiple video type feature data in the i-th group of training data to obtain multiple first combined feature data corresponding to the i-th group of training data; wherein the i-th group of training data is any one group of training data in the multiple groups of training data; Cross-combining the promotion resource position in the i-th group of training data with at least one of the plurality of video type feature data in the i-th group of training data to obtain second combined feature data corresponding to the i-th group of training data; and A plurality of combined feature data corresponding to the i-th group of training data is obtained according to the plurality of first combined feature data corresponding to the i-th group of training data and the second combined feature data corresponding to the i-th group of training data.

8. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, it is used to implement the steps of the click-through rate prediction model training method for video push as described in any one of claims 1 to 6.

9. A computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium, and the computer program can be executed by at least one processor so that the at least one processor performs the steps of the click-through rate prediction model training method for video push as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Click rate prediction method and device, equipment and medium

    CN110490389A