Model training method, information release method, device, equipment and medium
By training the label prediction model to predict pseudo-labels with label-free samples and combining positive and negative samples, the problem of sample selection bias in the delivery prediction model is solved, and the full-space learning is achieved, and the accuracy and accuracy of the delivery prediction model is improved.
Patent Information
- Application Number
- CN202410164026.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-02
- Publication Date
- 2025-08-05
AI Technical Summary
In the prior art, the delivery prediction model only considers the object of the multimedia information being delivered during training, and does not consider the object of the delivery, resulting in sample selection deviation and sample imbalance, resulting in poor generalization capabilities of the model, which in turn affects the delivery accuracy.
By obtaining positive samples, negative samples and label-free samples, the label prediction model is trained to predict pseudo-labels of label-free samples, and combine the pseudo-labels of positive samples, negative samples and label-free samples to train the delivery prediction model to achieve full-space learning and avoid sample selection bias.
The accuracy and delivery accuracy of the delivery prediction model are improved, ensuring the consistent data space in the model training stage and the use stage, and improving the delivery effect.
Smart Images

Figure CN120430833A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of computer technology, and particularly to a model training method, an information delivery method, a device, a device and a medium. Background Art
[0002] With the rapid development of computer technology, multimedia information is more and more widely used on the Internet. When an object browses multimedia information through the Internet, it can interact with the multimedia information, such as clicking or commenting on the multimedia information.
[0003] In the related art, the delivery objects that interact with the delivered multimedia information are used as positive samples, and the delivery objects that do not interact with the delivered multimedia information are used as negative samples to train a delivery prediction model. Subsequently, the delivery prediction model is used to determine which objects are used as delivery objects. However, since only the objects of the delivered multimedia information are considered when training the delivery prediction model, and the objects that are not delivered with multimedia information are not considered, there is a problem of sample bias, resulting in the delivery prediction model being inaccurate, and thus the delivery accuracy is poor. Summary of the Invention
[0004] Embodiments of the present application provide a model training method, an information delivery method, a device, a device and a medium, which can improve the accuracy of the delivery prediction model and thus improve the delivery accuracy. The technical solutions are as follows:
[0005] On the one hand, a model training method is provided, and the method includes:
[0006] Obtain positive samples, negative samples and unlabeled samples. The positive samples include the object features and positive sample labels of the delivery objects. The negative samples include the object features and negative sample labels of the delivery objects. The delivery object refers to the object to which the multimedia information is delivered. The positive sample label indicates that the delivery object interacts with the multimedia information. The negative sample label indicates that the delivery object does not interact with the multimedia information. The unlabeled samples include the object features of non-delivery objects. The non-delivery object refers to the object that is not delivered with the multimedia information;
[0007] Based on the positive samples and the negative samples, train a label prediction model, and the label prediction model is used to predict the probability that any object interacts with the multimedia information under the condition that the multimedia information is delivered;
[0008] Through the trained label prediction model, predict the pseudo labels of the unlabeled samples, and the pseudo labels indicate the probability that the non-delivery objects interact with the multimedia information under the condition that the multimedia information is delivered;
[0009] Train a placement prediction model based on the positive samples, the negative samples, the unlabeled samples, and the pseudo labels of the unlabeled samples. The placement prediction model is used to predict the placement score of any object, and the placement score is used to determine whether to place the multimedia information for the object.
[0010] On the other hand, an information placement method is provided. The method includes:
[0011] Obtain the object features of multiple candidate objects;
[0012] For any candidate object, determine the placement score of the candidate object based on the object features of the candidate object through the placement prediction model corresponding to the multimedia information;
[0013] Among the multiple candidate objects, determine the candidate objects whose placement scores meet the placement conditions as placement objects;
[0014] Place the multimedia information for the placement objects;
[0015] Wherein, the placement prediction model is trained by the model training method described in the above aspect.
[0016] On the other hand, a model training device is provided. The device includes:
[0017] A sample acquisition module, configured to acquire positive samples, negative samples, and unlabeled samples. The positive samples include the object features and positive sample labels of the placement objects. The negative samples include the object features and negative sample labels of the placement objects. The placement objects refer to the objects for which the multimedia information is placed. The positive sample labels indicate that the placement objects interact with the multimedia information. The negative sample labels indicate that the placement objects do not interact with the multimedia information. The unlabeled samples include the object features of non-placement objects. The non-placement objects refer to the objects for which the multimedia information is not placed;
[0018] A first training module, configured to train a label prediction model based on the positive samples and the negative samples. The label prediction model is used to predict the probability that any object interacts with the multimedia information under the condition that the multimedia information is placed for the object;
[0019] A label determination module, configured to predict the pseudo labels of the unlabeled samples through the trained label prediction model. The pseudo labels indicate the probability that the non-placement objects interact with the multimedia information under the condition that the multimedia information is placed for the non-placement objects;
[0020] A second training module, configured to train a placement prediction model based on the positive samples, the negative samples, the unlabeled samples, and the pseudo labels of the unlabeled samples, where the placement prediction model is used to predict the placement score of any object, and the placement score is used to determine whether to place the multimedia information for the object.
[0021] Optionally, the first training module is configured to:
[0022] Determine a preset label as the sample label of the unlabeled samples, where the preset label indicates that the non-placement objects do not interact with the multimedia information;
[0023] Train the label prediction model based on the positive samples, the negative samples, the unlabeled samples, and the preset labels of the unlabeled samples.
[0024] Optionally, the label prediction model includes an intervention variable prediction model, a result variable prediction model, and a causal inference model; the first training module is configured to:
[0025] Determine the positive samples, the negative samples, the unlabeled samples, and the preset labels of the unlabeled samples as a first sample set;
[0026] Train the intervention variable prediction model based on the first sample set, where the intervention variable prediction model is used to predict an intervention variable based on the object features of any object, and the intervention variable indicates whether to place the object;
[0027] Train the result variable prediction model based on the first sample set, where the result variable prediction model is used to predict a result variable based on the object features of any object, and the result variable indicates whether the object interacts with the multimedia information;
[0028] Determine the intervention variable residuals and the result variable residuals of multiple samples in the first sample set through the trained intervention variable prediction model and the result variable prediction model respectively;
[0029] Train the causal inference model based on the multiple samples in the first sample set and the intervention variable residuals and the result variable residuals of the multiple samples in the first sample set, where the causal inference model is used to predict the probability that the object interacts with the multimedia information under the condition that the object is placed with the multimedia information based on the object features, the result variable residuals, and the intervention variable residuals of any object.
[0030] Optionally, the first training module is configured to:
[0031] For any sample in the first sample set, determine the sample intervention variable of the sample, where the sample intervention variable is used to indicate whether the object in the sample is delivered with the multimedia information;
[0032] Through the initial intervention variable prediction model, based on the object characteristics in the sample, determine the first predicted intervention variable of the sample;
[0033] Based on the difference between the first predicted intervention variable and the sample intervention variable, train the initial intervention variable prediction model to obtain the intervention variable prediction model.
[0034] Optionally, the first training module is used for:
[0035] For any sample in the first sample set, determine the label of the sample as the sample result variable of the sample, where the sample result variable is used to indicate whether the object in the sample interacts with the multimedia information;
[0036] Through the initial result variable prediction model, based on the object characteristics in the sample, determine the first predicted result variable of the sample;
[0037] Based on the difference between the first predicted result variable and the sample result variable, train the initial result variable prediction model to obtain the result variable prediction model.
[0038] Optionally, the first training module is used for:
[0039] For any sample in the first sample set, determine the sample intervention variable and the sample control variable of the sample;
[0040] Through the intervention variable prediction model, based on the object characteristics in the sample, determine the second predicted intervention variable, and determine the difference between the sample intervention variable and the second predicted intervention variable as the intervention variable residual of the sample;
[0041] Through the result variable prediction model, based on the object characteristics in the sample, determine the second predicted result variable, and determine the difference between the sample result variable and the second predicted result variable as the result variable residual of the sample.
[0042] Optionally, the first training module is used for:
[0043] For any sample in the first sample set, through the initial causal inference model, based on the object features in the sample, the intervention variable residual and the outcome variable residual of the sample, determine the first causal effect of the sample, where the first causal effect represents the probability that the object in the sample interacts with the multimedia information under the condition of being exposed to the multimedia information;
[0044] Train the initial causal inference model based on the first causal effect of the sample to obtain the causal inference model.
[0045] Optionally, the label determination module is used to:
[0046] Determine the intervention variable residual and the outcome variable residual of the unlabeled sample through the intervention variable prediction model and the outcome variable prediction model;
[0047] Determine the second causal effect of the unlabeled sample through the causal inference model based on the object features in the unlabeled sample, the intervention variable residual and the outcome variable residual of the unlabeled sample;
[0048] Determine the second causal effect of the unlabeled sample as the pseudo-label of the unlabeled sample.
[0049] Optionally, the second training module is used to:
[0050] Determine the positive sample, the negative sample, the unlabeled sample, and the pseudo-label of the unlabeled sample as the second sample set;
[0051] For any sample in the second sample set, determine the sample placement score of the sample based on the label of the sample;
[0052] Determine the predicted placement score of the sample through the initial placement prediction model based on the object features in the sample;
[0053] Train the initial placement prediction model based on the difference between the predicted placement score and the sample placement score to obtain the placement prediction model.
[0054] Optionally, the second training module is used to:
[0055] When the label of the sample indicates that the object in the sample interacts with the multimedia information, take the first value as the sample placement score of the sample;
[0056] When the label of the sample indicates that the object in the sample does not interact with the multimedia information, take the second value as the sample placement score of the sample, where the first value is greater than the second value.
[0057] Optionally, the label prediction model and the placement prediction model are trained every preset period; the sample acquisition module is configured to:
[0058] Determine a plurality of exposed objects within a preset duration before the current period, where the exposed object refers to an object that has browsed the target section, and the multimedia information is to be placed in the target section;
[0059] In the case where the exposed object belongs to the placement object and the exposed object interacts with the multimedia information, determine the object feature and the positive sample label of the exposed object as the positive sample;
[0060] In the case where the exposed object belongs to the placement object and the exposed object does not interact with the multimedia information, determine the object feature and the negative sample label of the exposed object as the negative sample;
[0061] In the case where the exposed object belongs to a non-placement object, determine the object feature of the exposed object as the unlabeled sample.
[0062] Optionally, the device further includes a feature acquisition module, configured to:
[0063] For any object, obtain the attribute information and interaction information of the object, where the interaction information includes at least one of section interaction information, placement interaction information, or associated interaction information. The section interaction information represents the interaction situation between the object and the target section, the multimedia information is to be placed in the target section, the placement interaction information represents the interaction situation between the object and the multimedia information, and the associated content interaction information represents the interaction situation between the object and the associated information, and the associated information refers to other information associated with the multimedia information;
[0064] Generate the object feature of the object based on the attribute information and the interaction information of the object.
[0065] On the other hand, an information placement device is provided, and the device includes:
[0066] A feature acquisition module, configured to acquire the object features of a plurality of candidate objects;
[0067] A score prediction module, configured to, for any candidate object, determine the placement score of the candidate object based on the object feature of the candidate object through the placement prediction model corresponding to the multimedia information;
[0068] An object determination module, configured to determine, among the plurality of candidate objects, the candidate objects whose placement scores meet the placement conditions as the placement objects;
[0069] An information delivery module, configured to deliver the multimedia information to the delivery target;
[0070] Wherein, the delivery prediction model is trained based on the model training method described in the above aspect.
[0071] Optionally, the delivery prediction model is trained once every preset period; the score prediction module is configured to:
[0072] Determine the delivery prediction model obtained by training in the current period, and based on the object features of the candidate object through the delivery prediction model obtained by training in the current period, determine the delivery score of the candidate object;
[0073] The information delivery module is configured to:
[0074] Deliver the multimedia information to the delivery target in the current period.
[0075] Optionally, the feature acquisition module is configured to:
[0076] Obtain the attribute information and interaction information of the candidate object, where the interaction information includes at least one of section interaction information, delivery interaction information or associated interaction information. The section interaction information represents the interaction situation between the object and the target section, the multimedia information is used to be delivered to the target section, the delivery interaction information represents the interaction situation between the candidate object and the multimedia information, and the associated content interaction information represents the interaction situation between the candidate object and the associated information. The associated information refers to other information associated with the multimedia information;
[0077] Generate the object features of the candidate object based on the attribute information and the interaction information of the candidate object.
[0078] On the other hand, a computer device is provided. The computer device includes a processor and a memory. At least one computer program is stored in the memory, and the at least one computer program is loaded and executed by the processor to implement the operations performed by the model training method described in the above aspect, or to implement the operations performed by the information delivery method described in the above aspect.
[0079] On the other hand, a computer-readable storage medium is provided. At least one computer program is stored in the computer-readable storage medium, and the at least one computer program is loaded and executed by a processor to implement the operations performed by the model training method described in the above aspect, or to implement the operations performed by the information delivery method described in the above aspect.
[0080] On the other hand, a computer program product is provided, including a computer program which is loaded and executed by a processor to implement the operations performed by the model training method as described in the above aspects, or to implement the operations performed by the information delivery method as described in the above aspects.
[0081] In the solution provided by the embodiments of the present application, the delivery objects that interact with the multimedia information are used as positive samples, the delivery objects that do not interact with the multimedia information are used as negative samples, and the non-delivery objects that are not delivered with the multimedia information are used as unlabeled samples. The sample label indicates whether it interacts with the multimedia information. Therefore, both positive samples and negative samples have their respective sample labels, while unlabeled samples do not have sample labels. Based on this, the present application uses positive samples and negative samples to train a label prediction model, and uses the trained label prediction model to infer the probability that a non-delivery object interacts with the multimedia information under the condition of being delivered, so as to obtain the pseudo-labels of unlabeled samples. Furthermore, using positive samples, negative samples, unlabeled samples, and the sample labels corresponding to each sample to train a delivery prediction model enables the training process of the delivery prediction model to globally learn the characteristics of delivery objects and non-delivery objects, which is beneficial to improving the accuracy of the delivery prediction model and further improving the delivery accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0082] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0083] Figure 1 is a schematic diagram of an implementation environment provided by the embodiments of the present application;
[0084] Figure 2 is a flowchart of a model training method provided by the embodiments of the present application;
[0085] Figure 3 is a flowchart of another model training method provided by the embodiments of the present application;
[0086] Figure 4 is a schematic diagram of an exposure object provided by the embodiments of the present application;
[0087] Figure 5 is a flowchart of a training method for a delivery prediction model provided by the embodiments of the present application;
[0088] Figure 6 is a flowchart of an object feature determination method provided by the embodiments of the present application;
[0089] Figure 7 It is a flowchart of an information delivery method provided by an embodiment of the present application;
[0090] Figure 8 It is a flowchart of another information delivery method provided by an embodiment of the present application;
[0091] Figure 9 It is a schematic diagram of a model deployment method provided by an embodiment of the present application;
[0092] Figure 10 It is a flowchart of another information delivery method provided by an embodiment of the present application;
[0093] Figure 11 It is a schematic diagram of an AUUC index provided by an embodiment of the present application;
[0094] Figure 12 It is a schematic diagram of a Decile Chart index provided by an embodiment of the present application;
[0095] Figure 13 It is a schematic diagram of the structure of a model training device provided by an embodiment of the present application;
[0096] Figure 14 It is a schematic diagram of the structure of another model training device provided by an embodiment of the present application;
[0097] Figure 15 It is a schematic diagram of the structure of an information delivery device provided by an embodiment of the present application;
[0098] Figure 16 It is a schematic diagram of the structure of a terminal provided by an embodiment of the present application;
[0099] Figure 17 It is a schematic diagram of the structure of a server provided by an embodiment of the present application. Detailed implementation manners
[0100] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.
[0101] It can be understood that the terms "first", "second", etc. used in the present application may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the present application, the first sample set may be referred to as the second sample set, and similarly, the second sample set may be referred to as the first sample set.
[0102] Among them, "at least one" means one or more than one. For example, at least one sample can be one sample, two samples, three samples, etc., which is any integer greater than or equal to one. "Multiple" means two or more than two. For example, multiple samples can be two samples, three samples, etc., which is any integer greater than or equal to two. "Each" means each one in at least one. For example, each sample means each one in multiple samples. If there are three samples in multiple samples, then each sample means each one of the three samples.
[0103] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.), and signals (including but not limited to signals transmitted between user terminals and other devices, etc.) involved in this application are all fully authorized by the user or relevant parties. And the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards in relevant countries and regions. For example, the positive samples, negative samples, unlabeled samples, and multimedia information involved in this application are all obtained under the full knowledge and authorization of the user.
[0104] Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.
[0105] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the foundation model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0106] Machine Learning (ML) is an interdisciplinary subject involving multiple fields, such as probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory, etc. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning from demonstration.
[0107] The solution provided in the embodiments of this application, based on the machine learning technology of artificial intelligence, can implement a model training method to train a model for predicting placement objects. Using the trained placement prediction model, it is possible to determine which objects are the targeted placement objects for multimedia information.
[0108] In the scenario of targeted placement, due to the powerful learning ability of neural networks, the placement prediction method based on neural network models has been widely applied in actual business. However, in related technologies, the placement objects that have been placed with multimedia information and interacted with the multimedia information are used as positive samples, and the placement objects that have been placed with multimedia information but have not interacted with the multimedia information are used as negative samples to train the placement prediction model. But this training method has the following problems: Since only the placement objects that have been placed with multimedia information are considered during the training process, and the non-placement objects that have not been placed with multimedia information are not considered, while all objects need to be considered in actual applications, there is a problem of sample selection bias, resulting in inconsistent data spaces between the model training stage and the model usage stage, and further leading to poor generalization ability of the model. In addition, if the non-placement objects that have not been placed with multimedia information are directly used as negative samples, the probability of these non-placement objects interacting with the multimedia information will be underestimated, because these non-placement objects cannot interact with the multimedia information because they have not been placed with multimedia information, and there is still a probability of interacting with the multimedia information if they are placed with multimedia information. Therefore, there are problems of sample selection bias and sample imbalance in the model training process in related technologies.
[0109] Among them, sample selection bias refers to the model inferring general conclusions based on unrepresentative samples. The common reasons for the lack of representativeness of samples are that the sample size is too small or the samples are not randomly selected. Sample imbalance means that the number of samples in each category in the training set is uneven. Taking the binary classification problem as an example, usually when the sample ratio exceeds 4:1, it is called sample imbalance. For example, if the ratio between the number of positive samples and the number of negative samples is 1900:100, it is sample imbalance. Sample imbalance will cause the model to overfit the samples of the category with a large number during the learning process and underfit the category with insufficient samples, resulting in a low generalization ability of the model. The generalization ability refers to the adaptability of machine learning algorithms to new samples. The purpose of model learning is to learn the rules hidden behind the data. The ability of the trained model to give appropriate outputs for data outside the learning set with the same rules is the generalization ability of the model.
[0110] In view of the above problems, the embodiments of the present application propose a machine learning method that can avoid sample selection bias. By using the causal inference method, based on counterfactual learning, label prediction is performed on unlabeled samples (non-target objects that have not been delivered with multimedia information) to obtain pseudo labels, thereby achieving unbiased label imputation. Furthermore, on the basis of positive and negative samples, unlabeled samples and their pseudo labels are fused for full-space learning, so as to alleviate the problem of poor model performance caused by sample selection bias. The implementation manners of the method provided by the embodiments of the present application are described in detail in the following embodiments.
[0111] The model training method and information delivery method provided by the embodiments of the present application can both be executed by a computer device. Optionally, the computer device is a terminal or a server. Optionally, the server is an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. Optionally, the terminal is a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, etc., but is not limited thereto.
[0112] In some embodiments, the computer program involved in the embodiments of the present application can be deployed to be executed on one computer device, or on multiple computer devices located at one place, or on multiple computer devices distributed at multiple places and interconnected through a communication network. The multiple computer devices distributed at multiple places and interconnected through a communication network can form a blockchain system.
[0113] In some embodiments, the computer device is provided as a server. Figure 1 It is a schematic diagram of an implementation environment provided by an embodiment of the present application. Refer to Figure 1 , this implementation environment includes a terminal 101 and a server 102. The terminal 101 and the server 102 are connected through a wireless or wired network.
[0114] The terminal 101 is used to provide samples for the server 102, and the server 102 is used to train a label prediction model and a placement prediction model based on the samples provided by the terminal 101. Subsequently, the server 102 can predict the placement object of the multimedia information through the placement prediction model, send the multimedia information to the terminal 101 of the placement object, and the terminal 101 displays the multimedia information to the placement object.
[0115] In a possible implementation manner, an application provided by the server 102 is installed on the terminal 101. After the server 102 finishes training the placement prediction model, it predicts the placement object of the multimedia information through the placement prediction model and sends the multimedia information to the terminal 101 of the placement object. The terminal 101 can display the multimedia information on this application. Optionally, the application is an application in the operating system of the terminal 101 or an application provided by a third party. For example, the application is a multimedia application, and the multimedia application has the function of displaying multimedia information. Of course, the multimedia application can also have other functions, such as a review function, a shopping function, a navigation function, a game function, etc.
[0116] Among them, the object can be a user account, and the terminal 101 is used to log in to the application based on the object. After the server 102 finishes training the placement prediction model, it predicts the placement object of the multimedia information through the placement prediction model, and then sends the multimedia information to the terminal 101 that logs in to the placement object.
[0117] Figure 2 It is a flowchart of a model training method provided by an embodiment of the present application. This embodiment of the present application is executed by a computer device. Refer to Figure 2 , this method includes:
[0118] 201. The computer device obtains positive samples, negative samples, and unlabeled samples. The positive samples include the object characteristics of the placement object and positive sample labels. The negative samples include the object characteristics of the placement object and negative sample labels. The placement object refers to the object to which the multimedia information is placed. The positive sample label indicates that the placement object interacts with the multimedia information, and the negative sample label indicates that the placement object does not interact with the multimedia information. The unlabeled samples include the object characteristics of non-placement objects. The non-placement object refers to the object that is not placed with multimedia information.
[0119] In the scenario of information delivery, multimedia information can be delivered to some objects. The objects to which the multimedia information is delivered are called delivery objects, and the objects that are not delivered the multimedia information are called non-delivery objects. Among them, after being delivered the multimedia information, the delivery objects may interact with the multimedia information or may not interact with the multimedia information. Interacting with the multimedia information includes interaction behaviors such as clicking, liking, commenting, sharing, and collecting.
[0120] In the embodiments of the present application, positive samples, negative samples, and unlabeled samples are constructed according to historical delivery situations and historical interaction situations, and these samples are used to train a delivery prediction model for subsequent prediction of which objects will be used as the next batch of delivery objects. Among them, the positive samples include the object features and positive sample labels of the delivery objects that have interacted with the multimedia information, and the negative samples include the object features and negative sample labels of the delivery objects that have not interacted with the multimedia information. Since the non-delivery objects are not delivered the multimedia information, they cannot interact with the multimedia information, and thus it is impossible to determine whether the non-delivery objects will interact with the multimedia information when they are delivered the multimedia information. Therefore, the non-delivery objects do not have clear labels, and the unlabeled samples only include the object features of the non-delivery objects that have not been delivered the multimedia information.
[0121] Among them, the objects in the embodiments of the present application can be user accounts, etc., and the object features are used to reflect the characteristics of the objects. Optionally, the object feature is an embedding feature, which is a vector obtained by embedding the high-dimensional features of the object from a high-dimensional space into a low-dimensional space.
[0122] 202. The computer device trains a label prediction model based on the positive samples and negative samples. The label prediction model is used to predict the probability that any object will interact with the multimedia information under the condition that the multimedia information is delivered.
[0123] In the embodiments of the present application, the unlabeled samples do not have labels and cannot be directly used as training samples for the delivery prediction model. In order to be able to use the unlabeled samples when training the delivery prediction model, the pseudo-labels of the unlabeled samples are predicted first. Therefore, a label prediction model is trained using the positive samples and negative samples. Since the positive samples and negative samples can reflect the situation of whether the delivery objects interact with the multimedia information, the label prediction model can learn through the positive samples and negative samples which object features will interact with the multimedia information and which object features will not interact with the multimedia information under the condition that the multimedia information is delivered. Therefore, the label prediction model can predict the probability that an object will interact with the multimedia information under the condition that the multimedia information is delivered.
[0124] 203. The computer device uses the label prediction model obtained through training to predict the pseudo-label of the unlabeled sample. The pseudo-label represents the probability of a non-target object interacting with the multimedia information under the condition of being presented with the multimedia information.
[0125] After the computer device obtains the label prediction model through training, it uses this label prediction model to predict the probability of a non-target object in the unlabeled sample interacting with the multimedia information under the condition of being presented with the multimedia information. Furthermore, this probability can be used as the pseudo-label of the unlabeled sample. The pseudo-label does not represent the real probability but the predicted probability.
[0126] 204. The computer device trains a placement prediction model based on positive samples, negative samples, unlabeled samples, and the pseudo-labels of the unlabeled samples. The placement prediction model is used to predict the placement score of any object, and the placement score is used to determine whether to present the multimedia information to the object.
[0127] After predicting the pseudo-labels of the unlabeled samples, the positive samples, negative samples, and unlabeled samples all have their respective labels. Then, based on the positive samples and negative samples, the unlabeled samples and their pseudo-labels are integrated to train the placement prediction model, enabling the placement prediction model to not only learn based on the target objects but also learn based on non-target objects, thereby learning the features of the entire space, ensuring that the data space in the model training stage is consistent with the data space in the model usage stage, and effectively avoiding the problem of sample selection bias.
[0128] Among them, the placement prediction model is used to predict the placement score of any object, and then based on the placement scores of each object, it determines which objects are the target objects for placement, that is, it determines which objects to present the multimedia information to. It can be understood that the placement score of an object can reflect the benefit brought by presenting the multimedia information to this object. The higher the placement score of an object, the higher the benefit brought by presenting the multimedia information to this object, and the lower the placement score of an object, the lower the benefit brought by presenting the multimedia information to this object. Among them, the interaction between the object and the presented multimedia information can bring benefits. Therefore, it can also be understood that the placement score of this object can reflect the possibility of this object interacting with the multimedia information after presenting the multimedia information to this object. The higher the placement score of an object, the higher the possibility of this object interacting with the multimedia information, and the lower the placement score of an object, the lower the possibility of this object interacting with the multimedia information.
[0129] In the method provided by the embodiments of this application, the placement objects that interact with the multimedia information are used as positive samples, the placement objects that do not interact with the multimedia information are used as negative samples, and the non-placement objects that are not placed with multimedia information are used as unlabeled samples. The sample label indicates whether there is interaction with the multimedia information. Therefore, both positive samples and negative samples have their respective sample labels, while unlabeled samples do not have sample labels. Based on this, this application uses positive samples and negative samples to train a label prediction model, and uses the trained label prediction model to infer the probability that a non-placement object interacts with the multimedia information under the condition of being placed, so as to obtain the pseudo-labels of the unlabeled samples. Furthermore, using positive samples, negative samples, unlabeled samples, and the sample labels corresponding to each sample to train a placement prediction model enables the training process of the placement prediction model to globally learn the characteristics of placement objects and non-placement objects, which is beneficial to improving the accuracy of the placement prediction model and further improving the placement accuracy.
[0130] In the embodiments of this application, in order to implement Figure 2 the step of predicting the pseudo-labels of unlabeled samples by the label prediction model in the embodiments, it is necessary to model this label prediction model, and the modeling method is as follows.
[0131] In the scenario of information placement, there are three types of variables, namely feature variables, intervention variables, and outcome variables. Hereinafter, X represents the feature variable, T represents the intervention variable, and Y represents the outcome variable. Among them, X is the object feature of the object, T is used to indicate whether multimedia information is placed. For example, T is a binary variable. When T takes the value of 0, it means that the multimedia information is not placed, and when T takes the value of 1, it means that the multimedia information is placed. Y is used to indicate whether there is interaction with the multimedia information. For example, Y is a binary variable. When Y takes the value of 0, it means that there is no interaction with the multimedia information, and when Y takes the value of 1, it means that there is interaction with the multimedia information.
[0132] Based on the above three types of variables, model the probability of interacting with the multimedia information under the condition of being placed with multimedia information.
[0133] URT = p(Y = 1|X, do(T = 1)); Formula (1)
[0134] Among them, URT represents the probability that an object interacts with the multimedia information under the condition of being placed with multimedia information, Y = 1 represents interacting with the multimedia information, X represents the control variable (object feature), do represents the intervention operation, and T = 1 represents being placed with multimedia information.
[0135] In causal effect modeling, the causal effect of delivering multimedia information is equal to the difference between the probability of interacting with the multimedia information under the condition of delivering the multimedia information and the probability of interacting with the multimedia information when not delivering the multimedia information. Since the probability of interacting with the multimedia information when not delivering the multimedia information is 0, the causal effect is equivalent to the probability of interacting with the multimedia information under the condition of delivering the multimedia information. That is, the causal effect can be used as the probability of interacting with the multimedia information under the condition of delivering the multimedia information. This logic can be represented by the following formula.
[0136] ITE = p(Y = 1|X, do(T = 1)) - p(Y = 1|X, do(T = 0))
[0137] = UTR - 0 = UTR; Formula (2)
[0138] Where, ITE represents the causal effect, p(Y = 1|X, do(T = 1)) represents the probability of interacting with the multimedia information under the condition of delivering the multimedia information, and p(Y = 1|X, do(T = 0)) represents the probability of interacting with the multimedia information when not delivering the multimedia information.
[0139] Furthermore, by calculating the causal effect, the probability of interacting with the multimedia information under the condition of delivering the multimedia information can be obtained. The calculation method of this causal effect can be achieved through the DML (Double Machine Learning) model. The DML model decomposes the estimation problem of the causal effect into the following three sub-tasks: (1) Predicting the result variable based on the control variables, which can be achieved by the result variable prediction model; (2) Predicting the intervention variable based on the control variables, which can be achieved by the intervention variable prediction model; (3) Estimating the causal effect based on the first two prediction models, which can be achieved by the causal inference model.
[0140] In the DML model, the data generation process satisfies the following assumptions.
[0141] Y = θ(X)*T + g(X, W) + ∈, E[∈|X, W] = 0; Formula (3)
[0142] T = f(X, W) + η, E[η|X, W] = 0; Formula (4)
[0143] E[η*∈|X, W] = 0; Formula (5)
[0144] Where, θ(X) represents the causal effect, g(X, W) and f(X, W) represent functions related to the control variables, ∈ and η represent random errors, and E[∈|X, W], E[η|X, W] and E[η*∈|X, W] represent the expectations related to the random errors.
[0145] Through the above formula (3), it can be written as the following formula (6).
[0146] Y - E[Y|X, W] = θ(X) * (T - E[T|X, W]) + ∈; Formula (6)
[0147] Among them, E[Y|X, W] represents the expectation of the result variable, and E[T|X, W] represents the expectation of the intervention variable.
[0148] The residuals of the result variable and the intervention variable can be represented by the following formula (7) and formula (8).
[0149]
[0150] q(X, W) = E[Y|X, W], f(X, W) = E[T|X, W]; Formula (8)
[0151] Among them, represents the residual of the result variable, represents the residual of the intervention variable, and q(X, W) and f(X, W) represent functions related to the control variables.
[0152] According to the above formulas (6) to (8), the relationships among the residuals of the result variable, the residuals of the intervention variable, and the control variables can be represented by the following formula (9) and formula (10).
[0153]
[0154]
[0155] Among them, represents the residual of the result variable, represents the residual of the intervention variable, θ(X) represents the causal effect, represents the residual of the causal effect.
[0156] Therefore, after predicting the result variable and the intervention variable based on the result variable prediction model and the intervention variable prediction model, the causal inference model is modeled based on the above formulas (9) and (10). Subsequently, the causal effect can be predicted through the causal inference model, so as to obtain the probability of interacting with the multimedia information under the condition of delivering the multimedia information.
[0157] Based on the above modeling method of the label prediction model, the label prediction model includes an intervention variable prediction model, a result variable prediction model, and a causal inference model. The training processes of the label prediction model and the delivery prediction model are described in detail in the following Figure 3 embodiment. Figure 3It is a flowchart of another model training method provided by an embodiment of this application. The embodiment of this application is executed by a computer device. Refer to Figure 3 , the method includes:
[0158] 301. The computer device obtains positive samples, negative samples, and unlabeled samples. The positive samples include the object features of the target object and positive sample labels. The negative samples include the object features of the target object and negative sample labels. The target object refers to the object to which the multimedia information is to be delivered. The positive sample label indicates that the target object interacts with the multimedia information. The negative sample label indicates that the target object does not interact with the multimedia information. The unlabeled samples include the object features of non-target objects. The non-target object refers to the object that is not to be delivered with the multimedia information.
[0159] In a possible implementation manner, the multimedia information can be any form of multimedia information, such as text information, picture information, audio information, or card information, etc. Optionally, the multimedia information is the information provided by the server of the multimedia application, and the multimedia information is used to be delivered into the multimedia application. Optionally, the multimedia information is a game card, and the game card is used to promote the electronic game. For example, the game card includes the introduction information of the electronic game and the download entry of the electronic game, etc. Optionally, the multimedia application includes multiple sections, and a target section is included in the multiple sections. The multimedia information is used to be delivered into the target section of the multimedia application, that is, the multimedia application displays the multimedia information in the target section. For example, the multiple sections in the multimedia application include a recommendation section, a game section, a local section, a friend section, etc. If the multimedia information is a game card, then the target section is the game section.
[0160] In a possible implementation manner, the label prediction model and the delivery prediction model are trained every preset period. Then the process for the computer device to obtain positive samples, negative samples, and unlabeled samples includes: determining multiple exposed objects within a preset duration before the current period. The exposed object refers to the object that has browsed the target section, and the multimedia information is used to be delivered into the target section; when the exposed object belongs to the target object and the exposed object interacts with the multimedia information, determining the object features of the exposed object and the positive sample label as the positive sample; when the exposed object belongs to the target object and the exposed object does not interact with the multimedia information, determining the object features of the exposed object and the negative sample label as the negative sample; when the exposed object belongs to the non-target object, determining the object features of the exposed object as the unlabeled sample.
[0161] In the embodiments of the present application, a positive sample, a negative sample, and an unlabeled sample are obtained every preset period, and a model training process is performed based on the obtained samples. The preset period can be 1 day, 1 week, 72 hours, etc. Taking the current period as an example, a plurality of exposed objects within a preset duration before the current period are obtained. For example, if the preset period is 1 day and the preset duration is 1 week, and the current day is recorded as day T, then the plurality of exposed objects within the preset duration before the current period are the exposed objects within the period from [T - 7, T - 1].
[0162] Among them, the multimedia information is used to be placed in the target section, and the exposed object is the object that has browsed the target section. For example, the multimedia information is a game card provided by the server of the multimedia application, the target section is the game section in the multimedia application, and the game card is used to be placed in the game section of the multimedia application. Then, the exposed object refers to the object that has browsed the game section.
[0163] Among them, the exposed objects include placed objects and non - placed objects. The placed object is the object to which the multimedia information is placed, and the non - placed object is the object to which the multimedia information is not placed. The placed objects include interactive objects and non - interactive objects. The interactive object is the placed object that has interacted with the multimedia information, and the non - interactive object is the placed object that has not interacted with the multimedia information. Figure 4 is a schematic diagram of an exposed object provided by the embodiments of the present application. As Figure 4 shown, the exposed object includes placed objects. The part of the exposed object other than the placed objects is non - placed objects. The placed objects include interactive objects. The part of the placed objects other than the interactive objects is non - interactive objects. In the embodiments of the present application, the object features and positive sample labels of the interactive objects are used as positive samples, the object features and negative sample labels of the non - interactive objects are used as negative samples, and the object features of the non - placed objects are used as unlabeled samples.
[0164] In a possible implementation manner, after the computer device obtains a plurality of samples (positive samples, negative samples, unlabeled samples), the plurality of samples are divided into two groups according to a preset ratio. One group of samples is used as a training set for training the model, and the other group of samples is used as a test set for testing the model. For example, if the preset ratio is 4:1, then 80% of the samples are used as the training set, and 20% of the samples are used as the test set.
[0165] In a possible implementation manner, the number of samples obtained by the computer device is large, and usually the number of negative samples is significantly more than the number of positive samples, that is, there is a problem of sample imbalance. In order to relieve the storage pressure and calculation pressure, and at the same time avoid the learning process of the model being dominated by some samples, it is necessary to balance the samples. Therefore, the computer device performs down - sampling on the negative samples, and then uses the down - sampled negative samples and all the determined positive samples for model training.
[0166] 302. The computer device determines the preset label as the sample label of the unlabeled sample, and the preset label indicates that the non-target object has not interacted with the multimedia information.
[0167] In the embodiments of the present application, in order to ensure the comprehensiveness of the samples, when training the label prediction model, not only positive samples and negative samples can be considered, but also unlabeled samples can be considered. Since the non-target objects in the unlabeled samples have not been delivered with multimedia information, the non-target objects must not have interacted with the multimedia information. Therefore, a preset label can be set, and this preset label indicates that the non-target object has not interacted with the multimedia information, so as to determine the preset label as the sample label of the unlabeled sample.
[0168] 303. The computer device determines the positive samples, negative samples, unlabeled samples, and the preset labels of the unlabeled samples as the first sample set.
[0169] After the computer device determines the preset label of the unlabeled sample, it takes the positive samples, negative samples, unlabeled samples, and their preset labels as the first sample set, and then trains the label prediction model based on this first sample set. Among them, in the embodiments of the present application, the label prediction model includes an intervention variable prediction model, a result variable prediction model, and a causal inference model. The process of training the label prediction model based on this first sample set can be seen in the following steps 304-step 307.
[0170] 304. The computer device trains an intervention variable prediction model based on the first sample set. The intervention variable prediction model is used to predict an intervention variable based on the object features of any object, and the intervention variable indicates whether to deliver to the object.
[0171] The first sample set includes multiple samples, and each sample corresponds to object features and its respective label. The computer device trains the intervention variable prediction model based on the multiple samples in this first sample set. The intervention variable prediction model is used to predict an intervention variable based on a control variable, and this control variable is the object feature of the object. This intervention variable is used to indicate whether to deliver to the object. Therefore, the object feature in the sample is the control variable corresponding to this sample. By whether the object in the sample is delivered, the intervention variable corresponding to this sample can be determined. That is, each sample corresponds to a control variable and an intervention variable. Therefore, by learning from this first sample set, it is possible to learn how to predict the intervention variable (that is, whether to deliver) based on the control variable (that is, the object feature).
[0172] In a possible implementation, for any sample in the first sample set, a computer device determines a sample intervention variable of the sample, where the sample intervention variable is used to indicate whether the object in the sample is delivered with multimedia information; through an initial intervention variable prediction model, based on the object features in the sample, the computer device determines a first predicted intervention variable of the sample; and based on the difference between the first predicted intervention variable and the sample intervention variable, the computer device trains the initial intervention variable prediction model to obtain an intervention variable prediction model.
[0173] Among them, the sample intervention variables of positive samples and negative samples indicate that multimedia information is delivered, and the sample intervention variable of an unlabeled sample indicates that multimedia information is not delivered. The sample intervention variable is the true intervention variable of the sample, which can truly reflect whether the object in the sample is delivered. The first predicted intervention variable is the intervention variable obtained through model prediction. The smaller the difference between the first predicted intervention variable and the sample intervention variable, the more accurate the model. Therefore, the initial intervention variable prediction model can be trained based on the difference between the first predicted intervention variable and the sample intervention variable. The training objective is to reduce the difference between the first predicted intervention variable and the sample intervention variable. When the training end condition is reached, the trained intervention variable prediction model is obtained.
[0174] In a possible implementation, the computer device determines a first loss parameter based on the difference between the first predicted intervention variable and the sample intervention variable, and trains the initial intervention variable prediction model based on the first loss parameter. When the training end condition is met, the intervention variable prediction model is obtained. Among them, the computer device trains the initial intervention variable prediction model by minimizing the first loss parameter. The training end condition is that the loss parameter is less than a preset threshold, or the number of iterative trainings reaches a preset number, etc.
[0175] Optionally, if the computer device uses a cross-entropy loss function to determine the first loss parameter, the computer device determines the first loss parameter through the following formula (11).
[0176]
[0177] Among them, H1 represents the first loss parameter, t represents the sample intervention variable, represents the first predicted intervention variable.
[0178] In a possible implementation, the intervention variable prediction model is a LightGBM (Light Gradient-Boosting Machine) model. Optionally, the parameter initialization method of the LightGBM model is as follows: the number of weak learners n_estimators = 500, the learning rate learning_rate = 0.01, the maximum tree depth max_depth = 5, and the minimum number of samples in a leaf node min_data_in_leaf = 1000.
[0179] 305. The computer device trains a result variable prediction model based on the first sample set. The result variable prediction model is used to predict a result variable based on the object features of any object, and the result variable indicates whether the object interacts with multimedia information.
[0180] The first sample set includes multiple samples, and each sample corresponds to object features and its respective label. The computer device trains a result variable prediction model based on the multiple samples in the first sample set. The result variable prediction model is used to predict a result variable based on a control variable, and the control variable is the object feature of the object. The result variable is used to indicate whether it interacts with multimedia information. Therefore, the object features in the sample are the control variable corresponding to the sample, and the label corresponding to the sample is the result variable corresponding to the sample. That is, each sample corresponds to a control variable and a result variable. Therefore, by learning from the first sample set, it is possible to learn how to predict the result variable (that is, whether it interacts with multimedia information) based on the control variable (that is, the object feature).
[0181] In a possible implementation, for any sample in the first sample set, the computer device determines the label of the sample as the sample result variable of the sample. The sample result variable is used to indicate whether the object in the sample interacts with multimedia information; through the initial result variable prediction model, based on the object features in the sample, the computer device determines the first predicted result variable of the sample; based on the difference between the first predicted result variable and the sample result variable, the computer device trains the initial result variable prediction model to obtain the result variable prediction model.
[0182] Among them, the sample result variable of the positive sample indicates interaction with the multimedia information, and the sample result variables of the negative sample and the unlabeled sample indicate no interaction with the multimedia information. The sample result variable is the true result variable of the sample, which can truly reflect whether the object in the sample interacts with the multimedia information. The first predicted result variable is the result variable obtained by model prediction. The smaller the difference between the first predicted result variable and the sample result variable, the more accurate the model. Therefore, the initial result variable prediction model can be trained based on the difference between the first predicted result variable and the sample result variable. The training objective is to reduce the difference between the first predicted result variable and the sample result variable. When the training end condition is reached, the trained result variable prediction model is obtained.
[0183] In a possible implementation manner, the computer device determines a second loss parameter based on the difference between the first predicted result variable and the sample result variable, and trains the initial result variable prediction model based on the second loss parameter. When the training end condition is satisfied, the result variable prediction model is obtained. Among them, the computer device trains the initial result variable prediction model by minimizing the second loss parameter. The training end condition is that the loss parameter is less than a preset threshold, or the number of iterative trainings reaches a preset number, etc.
[0184] Optionally, if the computer device uses the cross-entropy loss function to determine the second loss parameter, the computer device determines the second loss parameter through the following formula (12).
[0185]
[0186] Among them, H2 represents the first loss parameter, y represents the sample result variable, represents the first predicted result variable.
[0187] In a possible implementation manner, the result variable prediction model is a LightGBM (Light Gradient-Boosting Machine) model. Optionally, the parameter initialization method of the LightGBM model is as follows: the number of weak learners n_estimators = 500, the learning rate learning_rate = 0.01, the maximum tree depth max_depth = 5, and the minimum number of samples in a leaf node min_data_in_leaf = 1000.
[0188] 306. The computer device determines the intervention variable residuals and the result variable residuals of multiple samples in the first sample set through the trained intervention variable prediction model and result variable prediction model respectively.
[0189] After the computer device trains the intervention variable prediction model and the outcome variable prediction model, it can determine the intervention variable residuals and outcome variable residuals of multiple samples in the first sample set according to the intervention variable prediction model and the outcome variable prediction model. Since it is considered that the trained intervention variable prediction model and outcome variable prediction model are accurate enough, it can be considered that the intervention variable residuals and outcome variable residuals determined by these two models are also accurate enough. Therefore, the intervention variable residuals and outcome variable residuals can be used as sample data to train the causal inference model subsequently.
[0190] In a possible implementation, for any sample in the first sample set, the computer device determines the sample intervention variable and the sample control variable of the sample; through the intervention variable prediction model, based on the object features in the sample, it determines the second predicted intervention variable, and determines the difference between the sample intervention variable and the second predicted intervention variable as the intervention variable residual of the sample; through the outcome variable prediction model, based on the object features in the sample, it determines the second predicted outcome variable, and determines the difference between the sample outcome variable and the second predicted outcome variable as the outcome variable residual of the sample.
[0191] Among them, the sample intervention variables of positive samples and negative samples represent being served with multimedia information, and the sample intervention variables of unlabeled samples represent not being served with multimedia information. The sample outcome variables of positive samples represent interacting with the multimedia information, and the sample outcome variables of negative samples and unlabeled samples represent not interacting with the multimedia information. For any sample, the computer device inputs the object features into the intervention variable prediction model, and the intervention variable prediction model outputs the second predicted intervention variable. It inputs the object features into the outcome variable prediction model, and the outcome variable prediction model outputs the second predicted outcome variable. The residual refers to the error between the true value and the predicted value. Therefore, the intervention variable residual of the sample is equal to the difference between the sample intervention variable and the second predicted intervention variable, and the outcome variable residual of the sample is equal to the difference between the sample outcome variable and the second predicted outcome variable.
[0192] 307. The computer device trains a causal inference model based on multiple samples in the first sample set and the intervention variable residuals and outcome variable residuals of multiple samples in the first sample set. The causal inference model is used to predict the probability that an object interacts with the multimedia information under the condition of being served with the multimedia information based on the object features, outcome variable residuals, and intervention variable residuals of any object.
[0193] After determining the intervention variable residuals and outcome variable residuals of each sample in the first sample set, the computer device trains a causal inference model based on the object features, outcome variable residuals, and intervention variable residuals corresponding to each sample in the first sample set. The input of the causal inference model is the object features, outcome variable residuals, and intervention variable residuals, and the output of the causal inference model is the causal effect, which is equal to the probability that the object interacts with the multimedia information under the condition that the multimedia information is delivered.
[0194] In a possible implementation, for any sample in the first sample set, the computer device determines the first causal effect of the sample through an initial causal inference model based on the object features, intervention variable residuals, and outcome variable residuals in the sample. The first causal effect represents the probability that the object in the sample interacts with the multimedia information under the condition that the multimedia information is delivered. Based on the first causal effect of the sample, the initial causal inference model is trained to obtain the causal inference model.
[0195] The computer device inputs the object features, intervention variable residuals, and outcome variable residuals in the sample into the initial causal inference model. The initial causal inference model outputs the first causal effect, and the initial causal inference model is trained based on the first causal effect. When the training end condition is reached, the trained causal inference model is obtained.
[0196] In a possible implementation, the computer device determines the target parameter based on the first causal effect of the sample, and trains the initial causal inference model based on the target parameter. When the training end condition is satisfied, the causal inference model is obtained. Among them, the computer device trains the initial causal inference model by maximizing the target parameter.
[0197] Optionally, the causal inference model is a causal forest model. The computer device uses a causal effect loss function to determine the target parameter, that is, by maximizing the heterogeneity score to split nodes. Then the computer device determines the target parameter through the following formula (13).
[0198]
[0199] Among them, Δ(C1,C2) represents the target parameter, P represents a parent node in the causal forest tree, and C1 and C2 represent two child nodes obtained by splitting the parent node. n C1 represents the number of samples in the C1 node, n C2 represents the number of samples in the C2 node, n P represents the number of samples in the P node, n P is equal to n C1 and n C2 sum. Represents the average causal effect of multiple samples in the C1 node, Represents the average causal effect of multiple samples in the C2 node.
[0200] In a possible implementation, the causal inference model is a causal forest model. Optionally, the parameter initialization method of the causal forest model is as follows: the number of weak learners n_estimators = 500, the learning criterion criterion = Δ(C1, C2), the maximum tree depth max_depth = 5, the minimum number of samples in a leaf node min_data_in_leaf = 1000, and the maximum sample ratio max_samples = 0.5.
[0201] It should be noted that the above embodiments are only described by taking the causal inference model as a causal forest model as an example. In addition, other types of causal inference models can also be used, such as a random forest model, etc. The embodiments of the present application do not limit this.
[0202] It should be noted that the above steps 303 - step 307 are only described by taking the label prediction model including an intervention variable prediction model, an outcome variable prediction model, and a causal inference model as an example for the training process of the label prediction model. In other embodiments, the computer device can also use other methods to train the label prediction model based on positive samples, negative samples, unlabeled samples, and the preset labels of unlabeled samples. The embodiments of the present application do not limit this.
[0203] 308. The computer device uses the trained label prediction model to predict the pseudo-label of the unlabeled sample. The pseudo-label represents the probability that a non-target object interacts with the multimedia information under the condition of being delivered with the multimedia information.
[0204] After the computer device trains the label prediction model, it can use the label prediction model to predict the pseudo-label of the unlabeled sample. The pseudo-label represents the probability that a non-target object in the unlabeled sample interacts with the multimedia information under the condition of being delivered with the multimedia information. In the embodiments of the present application, the unlabeled sample itself has no label. By training the label prediction model and using the label prediction model to predict the pseudo-label of the unlabeled sample, unbiased label imputation is achieved, so that subsequently, based on positive samples and negative samples, the unlabeled sample and its pseudo-label can be additionally combined to train the delivery prediction model, enabling the delivery prediction model to learn the features of the entire data space.
[0205] In a possible implementation, the computer device determines the intervention variable residual and the outcome variable residual of the unlabeled sample through an intervention variable prediction model and an outcome variable prediction model; through a causal inference model, based on the object features in the unlabeled sample, the intervention variable residual and the outcome variable residual of the unlabeled sample, determines the second causal effect of the unlabeled sample; and determines the second causal effect of the unlabeled sample as the pseudo-label of the unlabeled sample.
[0206] In the embodiments of this application, the label prediction model includes an intervention variable prediction model, an outcome variable prediction model, and a causal inference model. The object features of the non-target objects in the unlabeled sample are respectively input into the intervention variable prediction model and the outcome variable prediction model to obtain the second predicted intervention variable and the second predicted outcome variable of the unlabeled sample. The difference between the sample intervention variable and the second predicted intervention variable of the unlabeled sample is the intervention variable residual of the unlabeled sample, and the difference between the sample outcome variable and the second predicted outcome variable of the unlabeled sample is the outcome variable residual of the unlabeled sample. The object features, intervention variable residual, and outcome variable residual of the non-target object are input into the causal inference model, and the causal inference model outputs the second causal effect of the unlabeled sample. The second causal effect of the unlabeled sample is equivalent to the probability that the non-target object in the unlabeled sample interacts with the multimedia information under the condition of being presented with the multimedia information. Therefore, this second causal effect can be used as the pseudo-label of the unlabeled sample.
[0207] Optionally, the pseudo-label is a binary variable. When the second causal effect reaches a preset threshold, the pseudo-label of the unlabeled sample is set to indicate that the non-target object interacts with the multimedia information under the condition of being presented with the multimedia information. When the second causal effect does not reach the preset threshold, the pseudo-label of the unlabeled sample is set to indicate that the non-target object does not interact with the multimedia information under the condition of being presented with the multimedia information.
[0208] 309. The computer device trains a placement prediction model based on positive samples, negative samples, unlabeled samples, and the pseudo-labels of the unlabeled samples. The placement prediction model is used to predict the placement score of any object, and the placement score is used to determine whether to present multimedia information to the object.
[0209] After predicting the pseudo-label of the unlabeled sample, the positive samples, negative samples, and unlabeled samples all have their respective labels. Then, based on the positive samples and negative samples, the unlabeled sample and its pseudo-label are fused to train the placement prediction model, enabling the placement prediction model to not only learn based on the target objects but also learn based on the non-target objects, thereby learning the features of the entire space, making the data space in the model training stage consistent with the data space in the model usage stage, effectively avoiding the problem of sample selection bias, and improving the generalization ability of the placement prediction model.
[0210] The usage process of the placement prediction model is described in detail below Figure 7 and Figure 8 the embodiments of
[0211] In the method provided by the embodiment of this application, the placement object that interacts with the multimedia information is used as the positive sample, the placement object that does not interact with the multimedia information is used as the negative sample, and the non-placement object that has not been placed with the multimedia information is used as the unlabeled sample. The sample label indicates whether it interacts with the multimedia information. Therefore, both the positive sample and the negative sample have their respective sample labels, while the unlabeled sample does not have a sample label. Based on this, this application uses the positive sample and the negative sample to train the label prediction model, and uses the trained label prediction model to infer the probability that the non-placement object interacts with the multimedia information under the condition of being placed, so as to obtain the pseudo-label of the unlabeled sample. Furthermore, using the positive sample, the negative sample, the unlabeled sample, and the sample labels corresponding to each sample to train the placement prediction model enables the training process of the placement prediction model to globally learn the characteristics of the placement object and the non-placement object, which is beneficial to improving the accuracy of the placement prediction model and further improving the placement accuracy.
[0212] Moreover, since the non-placement object in the unlabeled sample has not been placed with the multimedia information, the non-placement object must not have interacted with the multimedia information. Therefore, a preset label can be set, and this preset label indicates that the non-placement object has not interacted with the multimedia information, so as to determine the preset label as the sample label of the unlabeled sample. Furthermore, when training the label prediction model, not only the positive sample and the negative sample can be considered, but also the unlabeled sample and its preset label can be considered, thus ensuring the comprehensiveness of the training samples of the label prediction model.
[0213] Based on the above embodiments, both the intervention variable prediction model and the result variable prediction model can be LightGBM models, and the LightGBM model can adopt the histogram algorithm, the GOSS (Gradient-based One-Side Sampling) algorithm, and the EFB (Exclusive Feature Bundling) algorithm to reduce the computational complexity. The detailed descriptions of the histogram algorithm, the GOSS algorithm, and the EFB algorithm are as follows.
[0214] Histogram algorithm: Discretize the continuous feature into k discrete values, construct a Histogram (histogram) with a width of k, and then traverse the training samples to determine the cumulative statistics of each discrete value in the histogram. When performing feature selection, according to the discrete values of the histogram, traverse to find the optimal split point. The histogram algorithm can significantly reduce memory consumption and improve the training speed.
[0215] GOSS algorithm: In a gradient boosting learner, the gradient of each sample reflects the learning degree of that sample. If the gradient of a sample is small, the training error of that sample is small and it has been learned relatively well. If the gradient of a sample is large, the training error of that sample is large and it has not been learned sufficiently. Based on this, the GOSS algorithm samples the samples according to the magnitude of the gradient, retains the samples with larger gradients, and randomly samples the samples with smaller gradients, so that the model pays more attention to the samples that have not been learned sufficiently.
[0216] EFB algorithm: High-dimensional features are usually sparse. In a sparse feature space, many different features are often mutually exclusive, that is, they do not take non-zero values simultaneously. Therefore, a group of mutually exclusive features can be bound into one feature, thereby achieving a reduction in the feature dimension. The features are sorted according to the number of non-zero values, the conflict ratio between different features is calculated, and each feature is traversed and the features are tried to be merged to minimize the conflict ratio.
[0217] It should be noted that only the intervention variable prediction model and the outcome variable prediction model are taken as examples of the LightGBM model in the above part. In addition, other types of intervention variable prediction models and outcome variable prediction models can also be used, and the embodiments of the present application do not limit this.
[0218] Based on the above embodiments, for the specific process of the computer device training the placement prediction model based on the positive samples, negative samples, unlabeled samples, and pseudo-labels of the unlabeled samples, reference can be made to the following Figure 5 embodiment. Figure 5 is a flowchart of a method for training a placement prediction model provided by an embodiment of the present application. The embodiment of the present application is executed by a computer device. Refer to Figure 5 , and the method includes:
[0219] 501. The computer device determines the positive samples, negative samples, unlabeled samples, and pseudo-labels of the unlabeled samples as the second sample set.
[0220] After predicting the pseudo-labels of the unlabeled samples, the positive samples, negative samples, and unlabeled samples all have their respective labels. Then, the positive samples, negative samples, unlabeled samples, and pseudo-labels of the unlabeled samples are determined as the second sample set. Since this second sample set includes the placement objects that interact with the multimedia information, the placement objects that do not interact with the multimedia information, and the non-placement objects that are not exposed to the multimedia information, training the placement prediction model using this second sample set can enable the placement prediction model to learn the features of the entire data space, which is beneficial to improving the generalization ability of the placement prediction model. For the detailed training process, reference can be made to the following steps 502-step 504.
[0221] 502. For any sample in the second sample set, the computer device determines the sample placement score of the sample based on the label of the sample.
[0222] In the embodiments of this application, the placement prediction model is used to predict the placement score of any object. This placement score can reflect the benefits brought by placing the multimedia information for this object, or it can be understood that the placement score of this object can reflect the possibility of the object interacting with the multimedia information after the multimedia information is placed for this object. And the label of the sample indicates whether the object in the sample interacts with the multimedia information. Therefore, the sample placement score of the sample can be determined based on the label of the sample.
[0223] In a possible implementation, when the label of the sample indicates that the object in the sample interacts with the multimedia information, the computer device uses the first value as the sample placement score of the sample; when the label of the sample indicates that the object in the sample does not interact with the multimedia information, the computer device uses the second value as the sample placement score of the sample, and the first value is greater than the second value.
[0224] For example, the value range of the placement score is data between 0 and 1, the first value is 1, and the second value is 0. That is, when the label of the sample indicates that the object in the sample interacts with the multimedia information, the sample placement score of the sample is equal to 1; when the label of the sample indicates that the object in the sample does not interact with the multimedia information, the sample placement score of the sample is equal to 0.
[0225] 503. The computer device determines the predicted placement score of the sample based on the object characteristics in the sample through the initial placement prediction model.
[0226] For any sample in the second sample set, the computer device inputs the object characteristics in the sample into the initial placement prediction model, and the initial placement prediction model outputs the predicted placement score of the sample.
[0227] 504. The computer device trains the initial placement prediction model based on the difference between the predicted placement score and the sample placement score to obtain the placement prediction model.
[0228] The sample placement score is the true placement score of the sample and can truly reflect the interaction between the object in the sample and the multimedia information. The predicted placement score is the placement score obtained through model prediction. The smaller the difference between the predicted placement score and the sample placement score, the more accurate the model. Therefore, the initial placement prediction model can be trained based on the difference between the predicted placement score and the sample placement score. The training objective is to reduce the difference between the predicted placement score and the sample placement score. When the training end condition is reached, the trained placement prediction model is obtained.
[0229] The method provided by the embodiments of this application uses positive samples, negative samples, unlabeled samples, and the sample labels corresponding to each sample to train a placement prediction model, enabling the training process of the placement prediction model to globally learn the characteristics of placement objects and non-placement objects, which is beneficial to improving the accuracy of the placement prediction model and thus improving the placement accuracy.
[0230] Based on the above embodiments, for the specific process of the computer device to determine the object characteristics in the positive samples, negative samples, and unlabeled samples, refer to the following Figure 6 embodiments. Figure 6 is a flowchart of a method for determining object characteristics provided by the embodiments of this application. The embodiments of this application are executed by a computer device. Refer to Figure 6 , and this method includes:
[0231] 601. For any object, the computer device obtains the attribute information and interaction information of the object.
[0232] It should be noted that the attribute information and interaction information of the object in the embodiments of this application are obtained under the condition that the user is fully informed and authorized.
[0233] Among them, the attribute information includes the information provided when registering the object, etc. The interaction information includes at least one of section interaction information, placement interaction information, or associated interaction information. The section interaction information represents the interaction situation between the object and the target section, and the multimedia information is to be placed in the target section. The placement interaction information represents the interaction situation between the object and the multimedia information. The associated content interaction information represents the interaction situation between the object and the associated information, and the associated information refers to other information associated with the multimedia information. For example, when the multimedia information is a game card, the associated information associated with the multimedia information refers to information related to the game.
[0234] In a possible implementation manner, the section interaction information includes information on behaviors such as browsing, clicking, playing, liking, commenting, and sharing performed by the object on the target section. In a possible implementation manner, the multimedia information is a game card, and the game card includes the introduction information and download entry of the electronic game. The placement interaction information includes information on behaviors such as browsing, clicking, downloading, and installing the object on the game card.
[0235] In a possible implementation manner, for any object, the computer device obtains the interaction information of the object within a preset time period before the current time point. For example, the preset time period is 90 days, 30 days, 1 week, or 1 day, etc. Optionally, the interaction information includes various different types of interaction information, such as click interaction information, play interaction information, and like interaction information, etc. For different types of interaction information, different statistical indicators can be used to convert the number of interaction behaviors into corresponding interaction information.
[0236] 602. The computer device generates object features of an object based on the attribute information and interaction information of the object.
[0237] The computer device extracts features from the attribute information and interaction information of the object to obtain the object features of the object.
[0238] In a possible implementation, the object information includes multiple types, the object features include multiple dimensions, and the features of each dimension belong to the features of one type of information. In the case where the interaction information of a certain type of an object is not obtained, the features on the dimension corresponding to this type will also be missing, and then the computer device fills in the missing features. For example, the proportion-based features and the count-based features are filled in with the constant 0, and the classification features are labeled with a separate value.
[0239] In a possible implementation, for the classification features, the computer device uses the one-hot encoding method to extract the classification features. In a possible implementation, for the high-dimensional classification features, considering that directly using the one-hot encoding method to extract will result in the features being too sparse, the classification statistical features are used instead, so as to generate low-dimensional dense features.
[0240] The method provided in the embodiments of the present application comprehensively considers the attribute information and interaction information of the object when extracting the object features of the object, making the amount of information contained in the object features large enough, improving the accuracy of extracting the object features, and thus being beneficial to improving the accuracy of subsequent placement prediction based on the object features.
[0241] After training the placement prediction model by using the methods of the above various embodiments, the placement scores of each candidate object can be predicted through the placement prediction model, and then based on the placement scores of each candidate object, it is determined which candidate objects are used as the placement objects. For the detailed process, refer to the following Figure 7 embodiments. Figure 7 is a flowchart of an information placement method provided by the embodiments of the present application. The embodiments of the present application are executed by a computer device. Refer to Figure 7 and the method includes:
[0242] 701. The computer device obtains the object features of multiple candidate objects.
[0243] In the embodiments of the present application, the computer device selects the placement objects among multiple candidate objects. The computer device obtains the object features of the multiple candidate objects, and the object features are used to reflect the characteristics of the object. Optionally, the object feature is an embedding feature, and the object feature is a vector obtained by embedding the high-dimensional features of the object from the high-dimensional space into the low-dimensional space.
[0244] 702. For any candidate object, the computer device determines the placement score of the candidate object based on the object characteristics of the candidate object through the placement prediction model corresponding to the multimedia information.
[0245] For any candidate object, the computer device inputs the object characteristics of the candidate object into the placement prediction model, and the placement prediction model outputs the placement score of the candidate object. Among them, the placement prediction model is trained based on the model training method provided in the above embodiments.
[0246] 703. Among multiple candidate objects, the computer device determines the candidate objects whose placement scores meet the placement conditions as placement objects.
[0247] After the computer device determines the placement scores of multiple candidate objects, it determines the candidate objects whose placement scores meet the placement conditions among the multiple candidate objects, and determines the candidate objects whose placement scores meet the placement conditions as placement objects. Among them, the placement scores of the candidate objects that meet the placement conditions are greater than the placement scores of the candidate objects that do not meet the placement conditions, and the placement objects can be flexibly set according to the actual situation.
[0248] 704. The computer device delivers multimedia information to the placement object.
[0249] After the computer device determines the placement object, it delivers the multimedia information to the placement object, that is, displays the multimedia information to the placement object.
[0250] In a possible implementation manner, the computer device is the server of the multimedia application, the multimedia information is provided by the server of the multimedia application, and the placement object is an account registered in the multimedia application. The computer device sends the multimedia information to the terminal of the placement object, and the terminal of the placement object refers to the terminal that logs in the placement object in the multimedia application. After receiving the multimedia information, the terminal displays the multimedia information in the multimedia application.
[0251] In the method provided in the embodiments of the present application, since the placement prediction model is trained using positive samples, negative samples, unlabeled samples, and the sample labels corresponding to each sample, the training process of the placement prediction model can globally learn the characteristics of placement objects and non-placement objects. Therefore, the accuracy of the placement prediction model is relatively high. Therefore, the placement scores determined by the placement prediction model are accurate enough, and using these placement scores to determine the placement objects can improve the placement accuracy.
[0252] In the above Figure 7 Based on the embodiments, for a more detailed implementation process of the information placement method, refer to the following Figure 8 embodiments. Figure 8It is a flowchart of another information delivery method provided by an embodiment of the present application. This embodiment of the present application is executed by a computer device. Refer to Figure 8 , the method includes:
[0253] 801. The computer device obtains the object features of multiple candidate objects.
[0254] In a possible implementation, the computer device determines multiple exposed objects within a preset duration before the current cycle, and determines the activity of these multiple exposed objects in the target section. Then, among the multiple exposed objects, it selects the exposed objects whose activity reaches a preset threshold, or selects a preset number of exposed objects ranked in the front, and uses the selected exposed objects as candidate objects. Among them, the multimedia information is used to be delivered to the target section.
[0255] In the embodiment of the present application, by selecting the exposed objects with higher activity as candidate objects, more objects with a higher possibility of interacting with the target section can be selected under a smaller scale of candidate objects, which is beneficial to improving the benefits obtained after delivering the multimedia information.
[0256] In a possible implementation, the computer device obtains the attribute information and interaction information of the candidate objects, and generates the object features of the candidate objects based on the attribute information and interaction information of the candidate objects. Among them, the interaction information includes at least one of section interaction information, delivery interaction information, or associated interaction information. The section interaction information represents the interaction situation between the object and the target section. The multimedia information is used to be delivered to the target section. The delivery interaction information represents the interaction situation between the candidate object and the multimedia information. The associated content interaction information represents the interaction situation between the candidate object and the associated information. The associated information refers to other information associated with the multimedia information. Among them, the process of generating the object features of the candidate objects is the same as that in the above Figure 6 embodiment and will not be elaborated here.
[0257] 802. The computer device determines the delivery prediction model trained in the current cycle.
[0258] In the embodiment of the present application, the delivery prediction model is trained every preset cycle. For the current cycle, the computer device determines the delivery prediction model trained in the current cycle, and this delivery prediction model is trained based on the model training method provided in the above embodiment.
[0259] 803. For any candidate object, the computer device determines the delivery score of the candidate object based on the object features of the candidate object through the delivery prediction model trained in the current cycle.
[0260] The computer device inputs the object features of the candidate objects into the placement prediction model trained in the current cycle, and the placement prediction model outputs the placement scores of the candidate objects.
[0261] 804. Among multiple candidate objects, the computer device determines the candidate objects whose placement scores meet the placement conditions as the placement objects.
[0262] In a possible implementation manner, the placement condition is that the placement score reaches a preset threshold, or the placement condition is that the placement scores are ranked within the top preset number in descending order.
[0263] 805. The computer device delivers multimedia information to the placement objects in the current cycle.
[0264] After determining the placement objects, the computer device delivers the multimedia information to the placement objects in the current cycle. In the next cycle, the computer device continues to train the placement prediction model using the above model training method, determines the placement objects in the next cycle based on the trained placement prediction model, and delivers multimedia information to the placement objects in the next cycle.
[0265] The method provided in the embodiments of this application uses positive samples, negative samples, unlabeled samples, and the sample labels corresponding to each sample to train the placement prediction model, enabling the training process of the placement prediction model to globally learn the features of placement objects and non-placement objects, which is beneficial to improving the accuracy of the placement prediction model and further improving the placement accuracy.
[0266] Figure 9 It is a schematic diagram of a model deployment method provided in the embodiments of this application. As Figure 9 shown, the computer device stores the data warehouse feature table and the data warehouse sample table generated in the current cycle offline. The data warehouse feature table includes the attribute information and interaction information of each object, and the object features of each object can be generated based on the data warehouse feature table. The data warehouse sample table can reflect whether each object is delivered multimedia information and whether it interacts with the multimedia information. According to the data warehouse feature table and the data warehouse sample table generated in the current cycle, the model is trained to obtain an updated placement prediction model. Based on the placement prediction model, the placement objects in the current cycle are determined, and multimedia information is delivered to the placement objects in the object service. In the next cycle, the data warehouse feature table and the data warehouse sample table are updated based on the data in the object service. The updated data warehouse feature table and data warehouse sample table can be used to train the model in the next cycle to determine the placement objects in the next cycle based on the updated placement prediction model in the next cycle.
[0267] In addition, as Figure 9As shown, the determined delivery targets can include multiple sections. When making a delivery, one can select the delivery targets of one section for delivery according to the actual situation. The delivery targets or the number of delivery targets of each section are different. For example, the delivery targets of Version 1 are the candidate objects whose delivery scores rank in the top 20%, and the delivery targets of Version 2 are the candidate objects whose delivery scores rank in the top 30%.
[0268] The solution provided by the embodiments of the present application can be applied to various scenarios. For example, in a multimedia application, if the multimedia information is a game card, the solution of the embodiments of the present application can be applied to the scenario of delivering game cards. Among them, as Figure 10 shown, the process of delivering game cards includes the following steps.
[0269] (1) Extract features from historical candidate objects and current candidate objects to obtain the object features of multiple historical candidate objects and the object features of current candidate objects, and determine the labels of historical candidate objects. The labels of historical candidate objects indicate whether historical candidate objects click on the game card.
[0270] (2) Construct a sample set based on the object features and labels of historical candidate objects for model training. The labels of historical candidate objects are the learning objectives of the model.
[0271] (3) Through the trained delivery prediction model, predict based on the object features of the current candidate objects to obtain the delivery scores of the current candidate objects, and sort the current candidate objects based on the delivery scores.
[0272] (4) Determine the delivery targets based on the sorting results. For example, take the top N current candidate objects as the delivery targets, where N is a positive integer set according to the actual situation.
[0273] (5) Display the game cards to the delivery targets.
[0274] To verify the effectiveness of the solution provided by the embodiments of the present application, an offline simulation experiment and an online control experiment are conducted on the delivery prediction model of the embodiments of the present application.
[0275] In the offline simulation experiment, it is divided into three groups of experiments. The first group of experiments is to train the delivery prediction model only based on the delivery targets. The second group of experiments is to train the delivery prediction model based on the delivery targets and non-delivery targets, and the non-delivery targets are directly used as negative samples, that is, it is defaulted that all non-delivery targets do not interact with the multimedia information. The third group of experiments uses the solution of the embodiments of the present application to train the delivery prediction model.
[0276] After obtaining the placement prediction models for three groups of experiments through training, they are tested on a test set with true labels. The Area Under Curve (AUC), a metric for evaluating the performance of a learner, is used to represent the test results of each placement prediction model, and the test results are shown in Table 1 below.
[0277] Table 1
[0278] Solution AUC Delivery prediction model for the first group of experiments 80.93% Delivery prediction model for the second group of experiments 85.99% Delivery prediction model of the embodiment of the present application 87.77%
[0279] As can be seen from Table 1, the placement prediction model provided by the embodiment of the present application has better effects.
[0280] To evaluate the offline learning effect of the placement prediction model provided by the embodiment of the present application, the AUCs of the placement prediction model on the training set and the test set are also calculated respectively. The training set is used to evaluate the fitting ability of the model for known data, and the test set is used to evaluate the generalization ability of the model for unknown data. The AUCs of the placement prediction model on the training set and the test set are shown in Table 2 below.
[0281] Table 2
[0282] Dataset AUC Training set 87.60% Test set 88.08%
[0283] As can be seen from Table 2, for the placement prediction model provided by the embodiment of the present application, the AUC is above 87% both in the training set and the test set.
[0284] In addition, considering unlabeled samples, the embodiment of the present application also uses AUUC (Area Under Uplift Curve, an evaluation metric for measuring the effect of causal effect estimation) and Decile Chart (an evaluation metric for judging whether the causal effect estimation meets the expectation) to measure the effect of the model in the complete sample space. Figure 11 is a schematic diagram of an AUUC metric provided by the embodiment of the present application. Figure 10 In it, the horizontal axis is the sample, the vertical axis is the cumulative causal effect of the top N samples arranged in descending order of the causal effect according to, and the area under the curve is AUUC. The larger the AUUC, the better the causal effect estimation. From Figure 10 it can be seen that in the case of random placement, AUUC is 0.5, and the AUUC estimated by the model reaches 0.8793. Therefore, the model can distinguish whether an object interacts with multimedia information. Figure 12 is a schematic diagram of a Decile Chart metric provided by the embodiment of the present application. From Figure 12 it can be seen that after arranging in descending order of the estimated causal effect, the actual causal effect basically decreases, indicating that the model can distinguish whether an object interacts with multimedia information.
[0285] To make a more fair comparison of different technical solutions, an online A / B test was also conducted. The deep learning solution without causal inference correction was used as the control group, and the solution provided in the embodiments of this application was used as the experimental group. Under the condition of controlling other conditions to be the same, the advantages and disadvantages of different technical solutions were compared. From the perspective of the impact on business indicators, the solution provided in the embodiments of this application has a significant improvement on business indicators, with an improvement rate of 44.3%. The solution provided in the embodiments of this application aims to optimize the user experience through precise placement. The experiment shows that the solution provided in the embodiments of this application not only brings optimization of traffic distribution, but also further brings an increase in the traffic of the business scenario, with an increase rate of 2.11%.
[0286] It can be seen that the solution provided in the embodiments of this application can improve the conversion rate of the placement object, accurately identify the objects with conversion tendency, and improve the interaction rate of multimedia information. Moreover, it expands the exposure resources of the business scenario. By optimizing the user experience through precise placement, while increasing the interaction rate, it reduces the occupation and waste of business scenario resources, and at the same time drives the growth of the overall exposure volume of the scenario, bringing an increase in the available resources of the business scenario.
[0287] Figure 13 It is a schematic structural diagram of a model training device provided in the embodiments of this application. Refer to Figure 13 , the device includes:
[0288] A sample acquisition module 1301, configured to acquire positive samples, negative samples, and unlabeled samples. The positive samples include the object features of the placement object and positive sample labels, the negative samples include the object features of the placement object and negative sample labels. The placement object refers to the object to which multimedia information is placed. The positive sample label indicates that the placement object interacts with the multimedia information, and the negative sample label indicates that the placement object does not interact with the multimedia information. The unlabeled samples include the object features of non-placement objects, and the non-placement object refers to the object to which multimedia information is not placed;
[0289] A first training module 1302, configured to train a label prediction model based on the positive samples and negative samples. The label prediction model is used to predict the probability that any object interacts with the multimedia information under the condition of being placed with the multimedia information;
[0290] A label determination module 1303, configured to predict the pseudo labels of the unlabeled samples through the trained label prediction model. The pseudo labels indicate the probability that non-placement objects interact with the multimedia information under the condition of being placed with the multimedia information;
[0291] A second training module 1304, configured to train a placement prediction model based on the positive samples, negative samples, unlabeled samples, and the pseudo labels of the unlabeled samples. The placement prediction model is used to predict the placement score of any object, and the placement score is used to determine whether to place multimedia information for the object.
[0292] The model training device provided by the embodiment of the present application uses the placement objects that interact with the multimedia information as positive samples, the placement objects that do not interact with the multimedia information as negative samples, and the non-placement objects that are not placed with the multimedia information as unlabeled samples. The sample label indicates whether to interact with the multimedia information. Therefore, both positive samples and negative samples have their respective sample labels, while unlabeled samples do not have sample labels. Based on this, the present application uses positive samples and negative samples to train a label prediction model, and uses the trained label prediction model to infer the probability that a non-placement object interacts with the multimedia information under the condition of being placed, so as to obtain the pseudo labels of unlabeled samples. Furthermore, using positive samples, negative samples, unlabeled samples, and the sample labels corresponding to each sample to train a placement prediction model enables the training process of the placement prediction model to globally learn the characteristics of placement objects and non-placement objects, which is beneficial to improving the accuracy of the placement prediction model and further improving the placement accuracy.
[0293] Optionally, the first training module 1302 is used for:
[0294] Determine a preset label as the sample label of the unlabeled sample, where the preset label indicates that the non-placement object does not interact with the multimedia information;
[0295] Train a label prediction model based on positive samples, negative samples, unlabeled samples, and the preset labels of unlabeled samples.
[0296] Optionally, the label prediction model includes an intervention variable prediction model, a result variable prediction model, and a causal inference model; the first training module 1302 is used for:
[0297] Determine positive samples, negative samples, unlabeled samples, and the preset labels of unlabeled samples as the first sample set;
[0298] Train an intervention variable prediction model based on the first sample set. The intervention variable prediction model is used to predict an intervention variable based on the object characteristics of any object, and the intervention variable indicates whether to place an object;
[0299] Train a result variable prediction model based on the first sample set. The result variable prediction model is used to predict a result variable based on the object characteristics of any object, and the result variable indicates whether the object interacts with the multimedia information;
[0300] Determine the intervention variable residuals and result variable residuals of multiple samples in the first sample set through the trained intervention variable prediction model and result variable prediction model respectively;
[0301] Train a causal inference model based on multiple samples in the first sample set, as well as the intervention variable residuals and outcome variable residuals of the multiple samples in the first sample set. The causal inference model is used to predict the probability that an object interacts with multimedia information under the condition that the multimedia information is delivered, based on the object characteristics, outcome variable residuals, and intervention variable residuals of any object.
[0302] Optionally, the first training module 1302 is configured to:
[0303] For any sample in the first sample set, determine the sample intervention variable of the sample, where the sample intervention variable is used to indicate whether the object in the sample is delivered with multimedia information;
[0304] Based on the object characteristics in the sample, determine the first predicted intervention variable of the sample through the initial intervention variable prediction model;
[0305] Train the initial intervention variable prediction model based on the difference between the first predicted intervention variable and the sample intervention variable, to obtain an intervention variable prediction model.
[0306] Optionally, the first training module 1302 is configured to:
[0307] For any sample in the first sample set, determine the label of the sample as the sample outcome variable of the sample, where the sample outcome variable is used to indicate whether the object in the sample interacts with multimedia information;
[0308] Based on the object characteristics in the sample, determine the first predicted outcome variable of the sample through the initial outcome variable prediction model;
[0309] Train the initial outcome variable prediction model based on the difference between the first predicted outcome variable and the sample outcome variable, to obtain an outcome variable prediction model.
[0310] Optionally, the first training module 1302 is configured to:
[0311] For any sample in the first sample set, determine the sample intervention variable and the sample control variable of the sample;
[0312] Based on the object characteristics in the sample, determine the second predicted intervention variable through the intervention variable prediction model, and determine the difference between the sample intervention variable and the second predicted intervention variable as the intervention variable residual of the sample;
[0313] Based on the object characteristics in the sample, determine the second predicted outcome variable through the outcome variable prediction model, and determine the difference between the sample outcome variable and the second predicted outcome variable as the outcome variable residual of the sample.
[0314] Optionally, the first training module 1302 is configured to:
[0315] For any sample in the first sample set, through the initial causal inference model, based on the object features in the sample, the intervention variable residual and the outcome variable residual of the sample, determine the first causal effect of the sample, where the first causal effect represents the probability that the object in the sample interacts with the multimedia information under the condition of being delivered with the multimedia information;
[0316] Based on the first causal effect of the sample, train the initial causal inference model to obtain a causal inference model.
[0317] Optionally, the label determination module 1303 is used for:
[0318] Determine the intervention variable residual and the outcome variable residual of the unlabeled sample through the intervention variable prediction model and the outcome variable prediction model;
[0319] Based on the object features in the unlabeled sample, the intervention variable residual and the outcome variable residual of the unlabeled sample, determine the second causal effect of the unlabeled sample through the causal inference model;
[0320] Determine the second causal effect of the unlabeled sample as the pseudo-label of the unlabeled sample.
[0321] Optionally, the second training module 1304 is used for:
[0322] Determine the positive samples, negative samples, unlabeled samples, and the pseudo-labels of the unlabeled samples as the second sample set;
[0323] For any sample in the second sample set, based on the label of the sample, determine the sample delivery score of the sample;
[0324] Through the initial delivery prediction model, based on the object features in the sample, determine the predicted delivery score of the sample;
[0325] Based on the difference between the predicted delivery score and the sample delivery score, train the initial delivery prediction model to obtain a delivery prediction model.
[0326] Optionally, the second training module 1304 is used for:
[0327] When the label of the sample indicates that the object in the sample interacts with the multimedia information, use the first value as the sample delivery score of the sample;
[0328] When the label of the sample indicates that the object in the sample does not interact with the multimedia information, use the second value as the sample delivery score of the sample, where the first value is greater than the second value.
[0329] Optionally, the label prediction model and the delivery prediction model are trained every preset period; the sample acquisition module 1301 is used for:
[0330] Determine multiple exposure objects within a preset duration before the current cycle. An exposure object refers to an object that has browsed the target section, and the multimedia information is to be delivered to the target section;
[0331] When the exposure object belongs to the delivery object and the exposure object interacts with the multimedia information, determine the object characteristics and positive sample labels of the exposure object as positive samples;
[0332] When the exposure object belongs to the delivery object and the exposure object does not interact with the multimedia information, determine the object characteristics and negative sample labels of the exposure object as negative samples;
[0333] When the exposure object belongs to a non-delivery object, determine the object characteristics of the exposure object as unlabeled samples.
[0334] Optionally, refer to Figure 14 , the device further includes a feature acquisition module 1305, which is used for:
[0335] For any object, obtain the attribute information and interaction information of the object. The interaction information includes at least one of section interaction information, delivery interaction information, or associated interaction information. The section interaction information represents the interaction situation between the object and the target section, the multimedia information is to be delivered to the target section, the delivery interaction information represents the interaction situation between the object and the multimedia information, and the associated content interaction information represents the interaction situation between the object and the associated information. The associated information refers to other information associated with the multimedia information;
[0336] Generate the object characteristics of the object based on the attribute information and interaction information of the object.
[0337] It should be noted that: The model training device provided in the above embodiments is only illustrated by dividing the above functional modules. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the computer device is divided into different functional modules to complete all or part of the functions described above. In addition, the model training device provided in the above embodiments and the model training method embodiments belong to the same concept. The specific implementation process can be found in the method embodiments and will not be repeated here.
[0338] Figure 15 is a schematic structural diagram of an information delivery device provided by an embodiment of the present application. Refer to Figure 15 , the device includes:
[0339] A feature acquisition module 1501, which is used to acquire the object characteristics of multiple candidate objects;
[0340] A score prediction module 1502, configured to, for any candidate object, determine a placement score of the candidate object based on the object features of the candidate object through a placement prediction model corresponding to the multimedia information;
[0341] An object determination module 1503, configured to determine, from multiple candidate objects, a candidate object whose placement score meets the placement condition as a placement object;
[0342] An information placement module 1504, configured to place multimedia information to the placement object;
[0343] Wherein, the placement prediction model is trained by using the model training method provided in the foregoing embodiment.
[0344] In the information placement device provided in the embodiment of the present application, a placement object that interacts with multimedia information is used as a positive sample, a placement object that does not interact with multimedia information is used as a negative sample, and a non-placement object that has not been placed with multimedia information is used as an unlabeled sample. The sample label indicates whether it interacts with multimedia information. Therefore, both the positive sample and the negative sample have their respective sample labels, while the unlabeled sample does not have a sample label. Based on this, the present application uses the positive sample and the negative sample to train a label prediction model, and uses the trained label prediction model to infer the probability that the non-placement object interacts with multimedia information under the condition of being placed, so as to obtain the pseudo label of the unlabeled sample. Furthermore, by using the positive sample, the negative sample, the unlabeled sample, and the sample labels corresponding to each sample, the placement prediction model is trained, so that the training process of the placement prediction model can globally learn the features of the placement object and the non-placement object, which is beneficial to improving the accuracy of the placement prediction model, and further improving the placement accuracy.
[0345] Optionally, the placement prediction model is trained once every preset period; the score prediction module 1502 is configured to:
[0346] Determine the placement prediction model obtained by training in the current period, and determine the placement score of the candidate object based on the object features of the candidate object through the placement prediction model obtained by training in the current period;
[0347] The information placement module 1504 is configured to:
[0348] Place multimedia information to the placement object in the current period.
[0349] Optionally, the feature acquisition module 1501 is configured to:
[0350] Obtain the attribute information and interaction information of the candidate object. The interaction information includes at least one of section interaction information, placement interaction information, or associated interaction information. The section interaction information represents the interaction situation between the object and the target section. The multimedia information is to be placed in the target section. The placement interaction information represents the interaction situation between the candidate object and the multimedia information. The associated content interaction information represents the interaction situation between the candidate object and the associated information. The associated information refers to other information associated with the multimedia information;
[0351] Generate the object features of the candidate object based on the attribute information and interaction information of the candidate object.
[0352] It should be noted that: The information placement device provided in the above embodiments is only illustrated by dividing the above functional modules. In practical applications, the above functions can be allocated to different functional modules as needed, that is, the internal structure of the computer device is divided into different functional modules to complete all or part of the functions described above. In addition, the information placement device provided in the above embodiments and the information placement method embodiments belong to the same concept. For the specific implementation process, please refer to the method embodiments, which will not be elaborated here.
[0353] An embodiment of the present application also provides a computer device, which includes a processor and a memory. At least one computer program is stored in the memory, and the at least one computer program is loaded and executed by the processor to implement the operations performed in the model training method of the above embodiments, or to implement the operations performed in the information placement method of the above embodiments.
[0354] Optionally, the computer device is provided as a terminal. Figure 16 The structural schematic diagram of a terminal 1600 provided by an exemplary embodiment of the present application is shown.
[0355] The terminal 1600 includes: a processor 1601 and a memory 1602.
[0356] The processor 1601 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor 1601 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field Programmable Gate Array), or PLA (Programmable Logic Array). The processor 1601 may also include a main processor and a coprocessor. The main processor is a processor used to process data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor 1601 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1601 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.
[0357] The memory 1602 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 1602 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1602 is used to store at least one computer program, and the at least one computer program is used to be possessed by the processor 1601 to implement the model training method or information delivery method provided in the method embodiments of the present application.
[0358] In some embodiments, the terminal 1600 may further optionally include: a peripheral device interface 1603 and at least one peripheral device. The processor 1601, the memory 1602, and the peripheral device interface 1603 may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 1603 through a bus, signal lines, or a circuit board. Optionally, the peripheral device includes at least one of a radio frequency circuit 1604, a display screen 1605, a camera component 1606, an audio circuit 1607, and a power supply 1608.
[0359] The peripheral device interface 1603 can be used to connect at least one I / O (Input / Output) related peripheral device to the processor 1601 and the memory 1602. In some embodiments, the processor 1601, the memory 1602, and the peripheral device interface 1603 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 1601, the memory 1602, and the peripheral device interface 1603 can be implemented on a separate chip or circuit board, and this embodiment does not limit this.
[0360] The radio frequency circuit 1604 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 1604 communicates with the communication network and other communication devices through electromagnetic signals. The radio frequency circuit 1604 converts an electrical signal into an electromagnetic signal for transmission, or converts the received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 1604 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and so on. The radio frequency circuit 1604 can communicate with other devices through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: metropolitan area network, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area network, and / or WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 1604 may further include a circuit related to NFC (Near Field Communication), and this application does not limit this.
[0361] The display screen 1605 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 1605 is a touch display screen, the display screen 1605 also has the ability to collect touch signals on or above the surface of the display screen 1605. The touch signal can be input to the processor 1601 as a control signal for processing. At this time, the display screen 1605 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 1605, which is disposed on the front panel of the terminal 1600; in other embodiments, there may be at least two display screens 1605, which are respectively disposed on different surfaces of the terminal 1600 or in a foldable design; in other embodiments, the display screen 1605 may be a flexible display screen, which is disposed on the curved surface or the folding surface of the terminal 1600. Even further, the display screen 1605 can also be set to an irregular non-rectangular shape, that is, a special-shaped screen. The display screen 1605 can be prepared using materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0362] The camera module 1606 is used to capture images or videos. Optionally, the camera module 1606 includes a front camera and a rear camera. The front camera is disposed on the front panel of the terminal 1600, and the rear camera is disposed on the back of the terminal 1600. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth-of-field camera, a wide-angle camera, and a telephoto camera, to implement functions such as background blurring by fusing the main camera and the depth-of-field camera, panoramic shooting by fusing the main camera and the wide-angle camera, and VR (Virtual Reality) shooting functions or other fusion shooting functions. In some embodiments, the camera module 1606 may also include a flash. The flash can be a single-color temperature flash or a two-color temperature flash. A two-color temperature flash refers to a combination of a warm light flash and a cold light flash, which can be used for light compensation under different color temperatures.
[0363] The audio circuit 1607 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 1601 for processing, or input to the radio frequency circuit 1604 to enable voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the terminal 1600. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signal from the processor 1601 or the radio frequency circuit 1604 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signal into sound waves audible to humans, but also convert the electrical signal into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 1607 may further include a headset jack.
[0364] The power supply 1608 is used to supply power to each component in the terminal 1600. The power supply 1608 may be alternating current, direct current, a disposable battery or a rechargeable battery. When the power supply 1608 includes a rechargeable battery, the rechargeable battery may support wired charging or wireless charging. The rechargeable battery may also be used to support fast charging technology.
[0365] Those skilled in the art can understand that Figure 16 the structure shown in does not limit the terminal 1600, and may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component layout.
[0366] Optionally, the computer device is provided as a server. Figure 17 FIG. is a schematic structural diagram of a server provided by an embodiment of the present application. The server 1700 may vary greatly due to different configurations or performances, and may include one or more processors (Central Processing Units, CPUs) 1701 and one or more memories 1702. Among them, at least one computer program is stored in the memory 1702, and the at least one computer program is loaded and executed by the processor 1701 to implement the methods provided by the above-mentioned various method embodiments. Of course, the server may also have components such as a wired or wireless network interface, a keyboard, and an input / output interface for input / output. The server may further include other components for implementing the functions of the device, which will not be elaborated here.
[0367] The embodiment of the present application further provides a computer-readable storage medium, in which at least one computer program is stored, and the at least one computer program is loaded and executed by a processor to implement the operations performed by the model training method in the above embodiment, or implement the operations performed by the information delivery method in the above embodiment.
[0368] The embodiments of the present application also provide a computer program product, including a computer program, which is loaded and executed by a processor to implement the operations performed by the model training method in the above embodiments or the operations performed by the information delivery method in the above embodiments.
[0369] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium, and the above-mentioned storage medium can be a read-only memory, a disk, an optical disc, etc.
[0370] The above are only optional embodiments of the embodiments of the present application, and are not intended to limit the embodiments of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the embodiments of the present application shall be included within the protection scope of the present application.
Claims
1. A model training method, characterized in that: The method comprises: Obtaining positive samples, negative samples, and unlabeled samples, wherein the positive samples include object features of a delivery object and a positive sample label, and the negative samples include object features of a delivery object and a negative sample label, wherein the delivery object refers to an object to which the multimedia information is delivered, the positive sample label indicates that the delivery object interacts with the multimedia information, and the negative sample label indicates that the delivery object does not interact with the multimedia information, and the unlabeled samples include object features of non-delivery objects, and the non-delivery objects indicate objects to which the multimedia information is not delivered; Based on the positive samples and the negative samples, a label prediction model is trained, wherein the label prediction model is used to predict the probability of any object interacting with the multimedia information under the condition that the multimedia information is presented; Predicting a pseudo label for the unlabeled sample using the trained label prediction model, where the pseudo label represents the probability that the non-delivery target will interact with the multimedia information under the condition that the multimedia information is delivered; Based on the positive samples, the negative samples, the unlabeled samples and the pseudo labels of the unlabeled samples, a delivery prediction model is trained, and the delivery prediction model is used to predict the delivery score of any object, and the delivery score is used to determine whether to deliver the multimedia information to the object.
2. The method according to claim 1, characterized in that The step of training a label prediction model based on the positive samples and the negative samples includes: Determining a preset label as the sample label of the unlabeled sample, wherein the preset label indicates that the non-delivery object has not interacted with the multimedia information; The label prediction model is trained based on the positive samples, the negative samples, the unlabeled samples, and the preset labels of the unlabeled samples.
3. The method according to claim 2, characterized in that The label prediction model includes an intervention variable prediction model, an outcome variable prediction model, and a causal inference model; the label prediction model is trained based on the positive sample, the negative sample, the unlabeled sample, and the preset label of the unlabeled sample, including: Determine the positive sample, the negative sample, the unlabeled sample, and the preset label of the unlabeled sample as a first sample set; Based on the first sample set, training the intervention variable prediction model, the intervention variable prediction model is used to predict the intervention variable based on the subject characteristics of any subject, the intervention variable indicating whether to administer the drug to the subject; Based on the first sample set, training the result variable prediction model, the result variable prediction model is used to predict a result variable based on an object feature of any object, the result variable indicating whether the object interacts with the multimedia information; Determining the intervention variable residuals and the outcome variable residuals of the plurality of samples in the first sample set respectively using the intervention variable prediction model and the outcome variable prediction model obtained through training; The causal inference model is trained based on multiple samples in the first sample set and the intervention variable residuals and outcome variable residuals of the multiple samples in the first sample set. The causal inference model is used to predict the probability of the object interacting with the multimedia information under the condition that the multimedia information is presented to the object based on the object characteristics, outcome variable residuals and intervention variable residuals of any object.
4. The method according to claim 3, characterized in that The step of training the intervention variable prediction model based on the first sample set includes: For any sample in the first sample set, determining a sample intervention variable of the sample, where the sample intervention variable is used to indicate whether an object in the sample is provided with the multimedia information; Determining a first predictive intervention variable for the sample based on characteristics of the subjects in the sample using an initial intervention variable prediction model; Based on the difference between the first prediction intervention variable and the sample intervention variable, the initial intervention variable prediction model is trained to obtain the intervention variable prediction model.
5. The method according to claim 3, characterized in that The step of training the outcome variable prediction model based on the first sample set includes: For any sample in the first sample set, determining a label of the sample as a sample result variable of the sample, where the sample result variable is used to indicate whether an object in the sample interacts with the multimedia information; Determining a first predicted outcome variable of the sample based on characteristics of the subjects in the sample using an initial outcome variable prediction model; Based on the difference between the first predicted outcome variable and the sample outcome variable, the initial outcome variable prediction model is trained to obtain the outcome variable prediction model.
6. The method according to claim 3, characterized in that The intervention variable prediction model and the outcome variable prediction model obtained through training respectively determine intervention variable residuals and outcome variable residuals of multiple samples in the first sample set, including: For any sample in the first sample set, determining a sample intervention variable and a sample control variable of the sample; Determining a second predictive intervention variable based on the subject characteristics in the sample using the intervention variable prediction model, and determining the difference between the sample intervention variable and the second predictive intervention variable as the intervention variable residual of the sample; A second predicted outcome variable is determined based on the object characteristics in the sample through the outcome variable prediction model, and the difference between the sample outcome variable and the second predicted outcome variable is determined as the sample outcome variable residual.
7. The method according to claim 3, characterized in that The training of the causal inference model based on the plurality of samples in the first sample set and the intervention variable residuals and the outcome variable residuals of the plurality of samples in the first sample set includes: For any sample in the first sample set, determining a first causal effect of the sample based on characteristics of the subject in the sample, residuals of the intervention variable, and residuals of the outcome variable of the sample using an initial causal inference model, where the first causal effect represents a probability that the subject in the sample will interact with the multimedia information under the condition that the multimedia information is presented to the subject; The initial causal inference model is trained based on the first causal effect of the sample to obtain the causal inference model.
8. The method according to claim 3, characterized in that The label prediction model obtained through training predicts the pseudo label of the unlabeled sample, including: Determining the intervention variable residual and the outcome variable residual of the unlabeled sample by using the intervention variable prediction model and the outcome variable prediction model; Determining, by the causal inference model, a second causal effect of the unlabeled sample based on the subject characteristics in the unlabeled sample, the intervention variable residuals and the outcome variable residuals of the unlabeled sample; The second causal effect of the unlabeled sample is determined as a pseudo label of the unlabeled sample.
9. The method according to any one of claims 1 to 8, characterized in that The training of the delivery prediction model based on the positive samples, the negative samples, the unlabeled samples, and the pseudo labels of the unlabeled samples includes: Determine the positive sample, the negative sample, the unlabeled sample, and the pseudo label of the unlabeled sample as a second sample set; For any sample in the second sample set, determining a sample delivery score of the sample based on the label of the sample; Determining a predicted delivery score for the sample based on the object characteristics in the sample using an initial delivery prediction model; Based on the difference between the predicted delivery score and the sample delivery score, the initial delivery prediction model is trained to obtain the delivery prediction model.
10. The method according to claim 9, characterized in that The determining of the sample delivery score of the sample based on the label of the sample includes: In a case where the label of the sample indicates that the object in the sample interacts with the multimedia information, using the first value as the sample delivery score of the sample; When the label of the sample indicates that the object in the sample has not interacted with the multimedia information, the second value is used as the sample delivery score of the sample, and the first value is greater than the second value.
11. The method according to any one of claims 1 to 8, characterized in that The label prediction model and the delivery prediction model are trained once every preset period; and the obtaining of positive samples, negative samples and unlabeled samples includes: Determine a plurality of exposure objects within a preset time period before a current cycle, wherein the exposure objects refer to objects that have browsed a target section, and the multimedia information is used to be delivered to the target section; In a case where the exposure object is a delivery object and the exposure object interacts with the multimedia information, determining the object features and the positive sample label of the exposure object as the positive sample; When the exposure object is a delivery object and the exposure object does not interact with the multimedia information, the object feature and the negative sample label of the exposure object are determined as the negative sample; In a case where the exposure object is a non-delivery object, the object feature of the exposure object is determined as the unlabeled sample.
12. The method according to any one of claims 1 to 8, characterized in that The method further comprises: For any object, obtaining attribute information and interaction information of the object, the interaction information including at least one of section interaction information, delivery interaction information, or associated interaction information, the section interaction information indicating the interaction between the object and a target section, the multimedia information being used for delivery to the target section, the delivery interaction information indicating the interaction between the object and the multimedia information, the associated content interaction information indicating the interaction between the object and associated information, and the associated information referring to other information associated with the multimedia information; An object feature of the object is generated based on the attribute information of the object and the interaction information.
13. An information delivery method, characterized in that: The method comprises: Obtaining object features of multiple candidate objects; For any candidate object, determine the delivery score of the candidate object based on the object features of the candidate object through the delivery prediction model corresponding to the multimedia information; Among the multiple candidate objects, determining the candidate object whose delivery score meets the delivery condition as the delivery object; delivering the multimedia information to the delivery object; The delivery prediction model is trained based on the method described in any one of claims 1-12.
14. The method according to claim 13, characterized in that The delivery prediction model is trained once every preset period; the delivery prediction model corresponding to the multimedia information determines the delivery score of the candidate object based on the object features of the candidate object, including: Determine a delivery prediction model obtained through training in the current cycle, and determine a delivery score of the candidate object based on the object features of the candidate object using the delivery prediction model obtained through training in the current cycle; The delivering the multimedia information to the delivery object includes: The multimedia information is delivered to the delivery object in a current cycle.
15. The method according to claim 13, characterized in that The obtaining of object features of multiple candidate objects includes: Acquiring attribute information and interaction information of the candidate object, the interaction information including at least one of section interaction information, delivery interaction information, or associated interaction information, the section interaction information indicating interaction between the object and a target section, the multimedia information being used for delivery to the target section, the delivery interaction information indicating interaction between the candidate object and the multimedia information, the associated content interaction information indicating interaction between the candidate object and associated information, and the associated information referring to other information associated with the multimedia information; An object feature of the candidate object is generated based on the attribute information of the candidate object and the interaction information.
16. A model training device, characterized in that: The device comprises: a sample acquisition module, configured to acquire positive samples, negative samples, and unlabeled samples, wherein the positive samples include object features of a delivery object and a positive sample label; the negative samples include object features of a delivery object and a negative sample label; the delivery object refers to an object to which multimedia information is delivered; the positive sample label indicates that the delivery object interacts with the multimedia information; the negative sample label indicates that the delivery object does not interact with the multimedia information; and the unlabeled samples include object features of non-delivery objects, wherein the non-delivery objects represent objects to which the multimedia information is not delivered; A first training module is configured to train a label prediction model based on the positive samples and the negative samples, wherein the label prediction model is configured to predict a probability of any object interacting with the multimedia information under a condition where the multimedia information is presented to the object; a label determination module, configured to predict a pseudo label of the unlabeled sample using a label prediction model obtained through training, wherein the pseudo label represents a probability that the non-delivery object will interact with the multimedia information under the condition that the multimedia information is delivered; The second training module is used to train a delivery prediction model based on the positive samples, the negative samples, the unlabeled samples and the pseudo labels of the unlabeled samples. The delivery prediction model is used to predict the delivery score of any object, and the delivery score is used to determine whether to deliver the multimedia information to the object.
17. An information delivery device, characterized in that: The device comprises: A feature acquisition module, used to acquire object features of multiple candidate objects; A score prediction module is used to determine, for any candidate object, a delivery score of the candidate object based on the object features of the candidate object using a delivery prediction model corresponding to the multimedia information; An object determination module, configured to determine, among the plurality of candidate objects, a candidate object whose delivery score meets the delivery condition as a delivery object; An information delivery module, configured to deliver the multimedia information to the delivery target; The delivery prediction model is trained based on the method described in any one of claims 1-12.
18. A computer device, characterized in that: The computer device includes a processor and a memory, wherein the memory stores at least one computer program, and the at least one computer program is loaded and executed by the processor to implement the operations performed by the model training method as described in any one of claims 1 to 12, or to perform the operations performed by the information delivery method as described in any one of claims 13 to 15.
19. A computer-readable storage medium, characterized in that The computer-readable storage medium stores at least one computer program, which is loaded and executed by a processor to implement the operations performed by the model training method as described in any one of claims 1 to 12, or to perform the operations performed by the information delivery method as described in any one of claims 13 to 15.
20. A computer program product comprising a computer program, characterized in that The computer program is loaded and executed by a processor to implement the operations performed by the model training method as described in any one of claims 1 to 12, or to perform the operations performed by the information delivery method as described in any one of claims 13 to 15.