Method, device and equipment for training interaction behavior prediction model and storage medium

CN117540073BActive Publication Date: 2026-09-29TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210896637.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-28
Publication Date
2026-09-29
Estimated Expiration
2042-07-28

AI Technical Summary

Technical Problem

[0004]然而,基于上述方式得到的模型参数的近似后验分布并不能很好地拟合模型参数的真实后验分布,导致最终将互动行为预测模型用于推荐系统时,该互动行为预测模型输出的预测结果准确度较低

Benefits of technology

[0033]在本申请实施例中,基于标准化流模型来获取互动行为预测模型的第一模型参数和第一模型参数的近似后验分布,在基于多个样本媒体资源和互动行为预测模型当前的模型参数获取到第一模型参数的真实后验分布之后,根据第一模型参数的近似后验分布与真实后验分布之间的相对熵,更新互动行为预测模型的模型参数。其中,由于标准化流模型能够将一个简单分布转换成一个复杂分布,因此基于该标准化流模型得到的模型参数的近似后验分布能够更接近模型参数的真实后验分布,从而有效提高了互动行为预测模型的准确度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117540073B_ABST
    Figure CN117540073B_ABST
Patent Text Reader

Abstract

The application provides a training method and device of an interactive behavior prediction model, equipment and a storage medium, belongs to the technical field of machine learning, and is applied to various scenes such as cloud technology, artificial intelligence, intelligent transportation and auxiliary driving. The method comprises the following steps: obtaining first model parameters of the interactive behavior prediction model and an approximate posterior distribution of the first model parameters based on a standardized flow model; and after obtaining a real posterior distribution of the first model parameters based on a plurality of sample media resources and current model parameters of the interactive behavior prediction model, updating the model parameters of the interactive behavior prediction model according to the relative entropy between the approximate posterior distribution of the first model parameters and the real posterior distribution. Since the standardized flow model can convert a simple distribution into a complex distribution, the approximate posterior distribution of the model parameters obtained based on the standardized flow model can be closer to the real posterior distribution of the model parameters, thereby effectively improving the accuracy of the interactive behavior prediction model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of machine learning technology, and in particular to a training method, apparatus, device, and storage medium for an interactive behavior prediction model. Background Technology

[0002] With the continuous development of network and computer technologies, media resource platforms often use recommendation systems to recommend media resources that an audience may be interested in. For example, the recommendation system trains an interactive behavior prediction model based on sample media resources to predict click-through rates. When recommending media resources to a specified audience, the interactive behavior prediction model is used to predict the probability that each media resource will be clicked by the specified audience, and the media resources recommended to the specified audience are selected based on the predicted click-through rates.

[0003] In related technologies, a Bayesian online learning approach is adopted. It is assumed that the model parameters of the interactive behavior prediction model are random variables and follow a normal distribution. The approximate posterior distribution of the model parameters is calculated using Gaussian mean field theory, and the model parameters of the interactive behavior prediction model are updated based on this approximate posterior distribution.

[0004] However, the approximate posterior distribution of the model parameters obtained in the above manner cannot fit the true posterior distribution of the model parameters well, resulting in low accuracy of the prediction results when the interaction behavior prediction model is finally used in the recommendation system. Summary of the Invention

[0005] This application provides a training method, apparatus, device, and storage medium for an interactive behavior prediction model, which can effectively improve the accuracy of the interactive behavior prediction model. The technical solution is as follows:

[0006] On the one hand, a training method for an interactive behavior prediction model is provided, the method comprising:

[0007] Based on the standardized flow model, the first vector obtained by sampling based on the target distribution is processed to obtain the first model parameters of the interactive behavior prediction model and the approximate posterior distribution of the first model parameters. The interactive behavior prediction model is used to predict the probability of performing interactive behavior on media resources.

[0008] Based on multiple sample media resources, the interactive behavior prediction model, and the parameters of the first model, the true posterior distribution of the parameters of the first model is obtained.

[0009] If the first relative entropy between the approximate posterior distribution and the true posterior distribution meets the target condition, the model parameters of the interaction behavior prediction model are updated to the first model parameters.

[0010] If the first relative entropy does not meet the target condition, the standardized flow model is updated. Based on the updated standardized flow model, the interaction behavior prediction model is trained again until the target relative entropy meets the target condition. Then, the model parameters of the interaction behavior prediction model are updated to the target model parameters corresponding to the target relative entropy.

[0011] On the one hand, a training device for an interactive behavior prediction model is provided, the device comprising:

[0012] The processing module is used to process the first vector obtained by sampling based on the target distribution based on the standardized flow model to obtain the first model parameters of the interactive behavior prediction model and the approximate posterior distribution of the first model parameters. The interactive behavior prediction model is used to predict the probability of performing interactive behavior on media resources.

[0013] The acquisition module is used to acquire the true posterior distribution of the first model parameters based on multiple sample media resources, the interactive behavior prediction model, and the first model parameters.

[0014] The first update module is used to update the model parameters of the interactive behavior prediction model to the first model parameters if the first relative entropy between the approximate posterior distribution and the true posterior distribution meets the target condition.

[0015] The second update module is used to update the standardized flow model if the first relative entropy does not meet the target condition, and continue to train the interaction behavior prediction model based on the updated standardized flow model until the obtained target relative entropy meets the target condition, and then update the model parameters of the interaction behavior prediction model to the target model parameters corresponding to the target relative entropy.

[0016] In some embodiments, the processing module includes:

[0017] The processing unit is configured to input the first vector into the normalized flow model, and process the first vector based on at least one bijective function in the normalized flow model to obtain the first model parameters.

[0018] The acquisition unit is used to acquire the approximate posterior distribution of the first model parameters based on the target distribution, the first vector, and the at least one bijective function.

[0019] In some embodiments, the processing unit is configured to:

[0020] Based on the bijective function with a flow depth of K, the intermediate model parameters obtained from the bijective function with a flow depth of K-1 are reversibly transformed to obtain the first model parameters. K indicates the flow depth of the standardized flow model, and K is a positive integer.

[0021] In some embodiments, the acquisition module is configured to:

[0022] Based on the distribution of the model parameters of the interactive behavior prediction model, the prior distribution of the first model parameters is obtained.

[0023] Based on the multiple sample media resources, the tag information of the multiple sample media resources, and the interactive behavior prediction model applying the first model parameters, the prediction results of the multiple sample media resources are obtained.

[0024] Based on the prior distribution of the first model parameters and the prediction results of the multiple sample media resources, the true posterior distribution of the first model parameters is obtained.

[0025] In some embodiments, the second update module is configured to:

[0026] If the first relative entropy does not meet the target condition, at least one bijective function in the normalized flow model is updated to obtain the updated normalized flow model.

[0027] In some embodiments, the apparatus further includes a sample media resource determination unit, configured to:

[0028] Obtain the prediction results and actual results of multiple candidate media resources predicted by the interactive behavior prediction model within the target time period;

[0029] If the difference between the predicted result and the actual result of the first candidate sample media resource is greater than or equal to the target threshold, the first candidate sample media resource is determined as a sample media resource. The first candidate sample media resource is any one of the multiple candidate sample media resources.

[0030] On one hand, a computer device is provided, which includes a processor and a memory for storing at least one computer program, which is loaded and executed by the processor to implement the training method of the interactive behavior prediction model in the embodiments of this application.

[0031] On the one hand, a computer-readable storage medium is provided, which stores at least one computer program, which is loaded and executed by a processor to implement the training method of the interactive behavior prediction model in the embodiments of this application.

[0032] On one hand, a computer program product or computer program is provided, which includes computer program code stored in a computer-readable storage medium. A processor of a computer device reads the computer program code from the computer-readable storage medium and executes the computer program code, causing the computer device to perform a training method for the interactive behavior prediction model described in the embodiments of this application.

[0033] In this embodiment, a standardized flow model is used to obtain the first model parameters and their approximate posterior distribution for the interactive behavior prediction model. After obtaining the true posterior distribution of the first model parameters based on multiple sample media resources and the current model parameters of the interactive behavior prediction model, the model parameters are updated according to the relative entropy between the approximate posterior distribution and the true posterior distribution of the first model parameters. Since the standardized flow model can transform a simple distribution into a complex distribution, the approximate posterior distribution of the model parameters obtained based on this standardized flow model is closer to the true posterior distribution of the model parameters, thereby effectively improving the accuracy of the interactive behavior prediction model. Attached Figure Description

[0034] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0035] Figure 1 This is a schematic diagram of the implementation environment for a training method for an interactive behavior prediction model provided in an embodiment of this application;

[0036] Figure 2 This is a schematic diagram of a recommendation system provided according to an embodiment of this application;

[0037] Figure 3 This is a schematic diagram of the structure of an interactive behavior prediction model provided according to an embodiment of this application;

[0038] Figure 4 This is a flowchart of a training method for an interactive behavior prediction model provided according to an embodiment of this application;

[0039] Figure 5 This is a flowchart of another training method for an interactive behavior prediction model provided according to an embodiment of this application;

[0040] Figure 6 This is a schematic diagram of a standardized flow model provided according to an embodiment of this application;

[0041] Figure 7 This is a schematic diagram of a model test result provided according to an embodiment of this application;

[0042] Figure 8 This is a schematic diagram of the structure of a training device for an interactive behavior prediction model according to an embodiment of this application;

[0043] Figure 9 This is a schematic diagram of the structure of a server according to an embodiment of this application. Detailed Implementation

[0044] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0045] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0046] In this application, the terms "first," "second," etc., are used to distinguish identical or similar items that have essentially the same function. It should be understood that there is no logical or temporal dependency between "first," "second," and "nth," nor does it limit the quantity or execution order. It should also be understood that although the following description uses the terms "first," "second," etc., to describe various elements, these elements should not be limited by the terms.

[0047] These terms are simply used to distinguish one element from another. For example, without departing from the various examples, a first function can be called a second function, and similarly, a second function can be called a first function. Both the first and second functions can be functions, and in some cases, they can be separate and distinct functions.

[0048] "At least one" refers to one or more functions. For example, at least one function can be one function, two functions, three functions, or any integer number of functions greater than or equal to one. "Multiple" refers to two or more functions. For example, multiple functions can be two functions, three functions, or any integer number of functions greater than or equal to two.

[0049] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the sample media resources involved in this application were obtained under full authorization. In some embodiments, this disclosure provides a permission inquiry page, which is used to inquire whether permission to obtain the above information is granted. The permission inquiry page displays an consent / authorization control and a denial / rejection control. When a trigger operation on the consent / authorization control is detected, the above information is obtained using the training method of the interactive behavior prediction model provided in this application.

[0050] The following describes the technologies that may be used in the embodiments of this application.

[0051] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.

[0052] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0053] Machine learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instruction-based learning.

[0054] Cloud technology refers to a managed technology that unifies a series of resources such as hardware, software, and networks within a wide area network or local area network to achieve data computing, storage, processing, and sharing.

[0055] Cloud computing is a computing model that distributes computing tasks across a resource pool consisting of a large number of computers, enabling various application systems to access computing power, storage space, and information services as needed. The network providing these resources is called the "cloud." From the user's perspective, the resources in the "cloud" are infinitely scalable, readily available, on-demand, and expandable, with payment based on usage.

[0056] The key terms or abbreviations involved in the embodiments of this application are described below.

[0057] Bayes' theorem is a theorem about the conditional or marginal probabilities of random events A and B. It can also be understood as describing the probability of certain events occurring given some known conditions. Its core statement is: p(A|B) = p(A)p(B|A) / p(B).

[0058] Normalizing Flow (NF) is a statistical method that uses the law of change of variable to transform a simple distribution into a complex distribution.

[0059] A bijective function is a function in mathematics that maps a set X to a set Y such that for every y in Y, there exists a unique x in X corresponding to it, and for every x in X, there exists a unique y in Y corresponding to it.

[0060] A determinant is a scalar calculated on a square matrix. It can be viewed as a generalization of the concepts of area or volume to general Euclidean space. In other words, in Euclidean space, the determinant describes the effect of a linear transformation on "volume".

[0061] A Jacobian matrix is ​​a matrix in which the first-order partial derivatives of a function are arranged in a certain way.

[0062] An affine transformation is a linear transformation performed on a vector space followed by a translation, transforming it into another vector space.

[0063] KL divergence (Kullback-Leibler Divergence), also known as relative entropy or information divergence, is a metric used to measure the divergence of a distribution.

[0064] The Hadamard product is a commonly used matrix operation in machine learning. It takes two matrices of the same shape as input and outputs a matrix of the same shape, where each element is equal to the product of the elements at the same position in the two input matrices.

[0065] Variational inference is an approximate inference method used when the posterior probability distribution of a variable has no analytical solution.

[0066] The Area Under the Curve (AUC), typically referring to the AUC based on the Receiver Operating Characteristic (ROC), is a metric for evaluating models. A larger AUC indicates that the model's output is closer to the true value, meaning higher model accuracy. For example, the AUC ranges from 0 to 1, with values ​​closer to 1 indicating higher model accuracy.

[0067] The implementation environment of the embodiments of this application is described below.

[0068] Figure 1 This is a schematic diagram illustrating the implementation environment of a training method for an interactive behavior prediction model according to an embodiment of this application. The implementation environment includes a terminal 101 and a server 102. The terminal 101 and server 102 can be directly or indirectly connected via a wired or wireless network, which is not limited herein.

[0069] Terminal 101 includes, but is not limited to, mobile phones, computers, smart voice interaction devices, smart home appliances, vehicle terminals, and aircraft. Indicatively, terminal 101 can install and run a target application that provides recommendation functions for media resources, such as video applications, audio applications, social applications, and conferencing applications, etc., without limitation. In some embodiments, terminal 101 can provide server 102 with information required for training an interactive behavior prediction model, such as training parameters, sample media resources, and an initial AI model.

[0070] In some embodiments, terminal 101 generally refers to one of a plurality of terminals; this embodiment uses terminal 101 as an example only. Those skilled in the art will understand that the number of terminals 101 can be greater. For example, there may be dozens or hundreds, or even more, terminals 101. In this case, the implementation environment of the training method for the interactive behavior prediction model may also include other terminals. This application does not limit the number of terminals or the type of device.

[0071] Server 102 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. The number of servers 102 can be more or less, and this application embodiment does not limit this. Of course, server 102 may also include other functional servers to provide more comprehensive and diversified services. In some embodiments, server 102 is used to execute the training method of the interactive behavior prediction model provided in this application embodiment, training the model based on information provided by terminal 101.

[0072] Schematic illustration: The server provides background services for the target application running on the terminal. The terminal can display media resources through this target application. Based on the terminal's resource acquisition request, the server retrieves multiple candidate media resources, predicts the probability of interactive behavior for each candidate media resource using an interactive behavior prediction model, determines the target media resource based on the prediction results, and sends the target media resource to the terminal. The terminal displays the target media resource and feeds back the interactive behavior for that target media resource to the server. Based on the interactive behavior for that target media resource, the server generates corresponding sample media resources. This process can also be understood as the server continuously updates the set of sample media resources as the prediction results of the interactive behavior prediction model increase within the online time period, providing sample support for the training of the interactive behavior prediction model.

[0073] In some embodiments, during model training, server 102 undertakes the main computational work and terminal 101 undertakes the secondary computational work; or, server 102 undertakes the secondary computational work and terminal 101 undertakes the main computational work; or, server 102 or terminal 101 can each undertake computational work independently.

[0074] In some embodiments, the wired or wireless networks described above use standard communication technologies and / or protocols. The network is typically the Internet, but can be any network, including but not limited to Local Area Networks (LANs), Metropolitan Area Networks (MANs), Wide Area Networks (WANs), mobile, wired or wireless networks, private networks, or any combination of virtual private networks. In some embodiments, technologies and / or formats including Hyper Text Markup Language (HTML), Extensible Markup Language (XML), etc., are used to represent data exchanged over the network. Furthermore, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Networks (VPNs), and Internet Protocol Security (IPsec) can be used to encrypt all or some links. In other embodiments, custom and / or dedicated data communication technologies can be used to replace or supplement the aforementioned data communication technologies.

[0075] The application scenarios of the embodiments of this application will be described below based on the above implementation environment.

[0076] This application provides a training method for an interactive behavior prediction model. This model can be applied to a recommendation system of a media resource platform to predict the probability of interactive behaviors performed on media resources. This allows the recommendation system to recommend media resources of interest to users based on the prediction results output by the interactive behavior prediction model. For example, media resources include videos, audio, images, and advertisements, and interactive behaviors performed on media resources include clicking, liking, and saving. Correspondingly, the interactive behavior prediction model can be a click-through rate prediction model, a like rate prediction model, and a save rate prediction model, etc., and this application does not limit the specific model used.

[0077] Indicatively, for reference Figure 2 , Figure 2 This is a schematic diagram of a recommendation system provided according to an embodiment of this application. For example... Figure 2As shown, taking media resources as advertisements and interactive behavior prediction models as click-through rate (CTR) prediction models as an example, the advertising platform receives an advertisement retrieval request, identifies multiple candidate advertisements from the advertisement library, predicts the CTR of each candidate advertisement using the CTR prediction model, and, based on the prediction results, selects at least one advertisement from the multiple candidate advertisements and returns it. All advertisement impressions, clicks, and conversions generated during this process are stored and used to generate input features and tag information for training the CTR prediction model through extraction, transformation, and loading. The trained CTR prediction model is then updated to the advertising platform for real-time online prediction of click-through rates for specific advertisements.

[0078] This application provides a method for training an interactive behavior prediction model in real time. By applying Bayes' theorem and normalization flow, the model parameters of the interactive behavior prediction model are updated in real time based on the continuously accumulating sample media resources, thereby improving the accuracy of the model's output. To more clearly illustrate the training method provided in this application, the structure of the interactive behavior prediction model provided in this application embodiment is first illustrated below.

[0079] Figure 3 This is a schematic diagram of the structure of an interactive behavior prediction model provided according to an embodiment of this application. Indicatively, as... Figure 3 As shown, this interactive behavior prediction model is built on a deep neural network (DNN). The model includes an input layer, an embedding layer, a DNN layer (composed of three fully connected layers (FC) and a ReLU activation function), and an output layer. The input to the interactive behavior prediction model is the feature data of the media resources, such as object features (including object identifier, type, interests, and age) and media resource features (such as media resource identifier, content, type, and optimization objectives). The output is the predicted probability of performing an interactive behavior on the media resource (e.g., click-through rate y).

[0080] Of course, the structure of the interactive behavior prediction model described above is merely illustrative. In some embodiments, the interactive behavior prediction model may also be a model with other structures. For example, the interactive behavior prediction model may be a model built based on a Convolutional Neural Network (CNN) or a Recurrent Neural Network (RNN). This application does not limit the structure of the interactive behavior prediction model.

[0081] The training method of the interactive behavior prediction model provided in the embodiments of this application will be described below.

[0082] Figure 4 This is a flowchart illustrating a training method for an interactive behavior prediction model provided in an embodiment of this application. For example... Figure 4 As shown, this embodiment of the application takes a server as an example for explanation. The method includes the following steps 401 to 404.

[0083] 401. The server processes the first vector obtained by sampling based on the target distribution using a standardized flow model to obtain the first model parameters of the interactive behavior prediction model and the approximate posterior distribution of the first model parameters.

[0084] In this embodiment, the interactive behavior prediction model is used to predict the probability of engaging in interactive behavior with media resources. The normalized flow model is used to transform a vector using at least one bijective function to obtain the model parameters of the interactive behavior prediction model. The target distribution is a simple distribution, such as a standard normal distribution, a Poisson distribution, or a uniform distribution, etc., without limitation. The first vector is a vector randomly sampled from the target distribution, and the first model parameters refer to all parameters of the interactive behavior prediction model, wherein the dimension of the first vector is the same as the dimension of the first model parameters. For example, taking an interactive behavior prediction model with 2000 parameters as an example, the dimension of the first vector is 2000, and correspondingly, the first model parameters indicate the 2000 parameters of the interactive behavior prediction model, without limitation.

[0085] It should be noted that the first model parameter is obtained by processing the first vector using a standardized flow model. This first model parameter can also be understood as a type of model parameter that can be applied to the interactive behavior prediction model, essentially a predicted model parameter. Based on this, the server determines whether to update the current model parameters of the interactive behavior prediction model to this first model parameter through the following steps to improve the accuracy of the interactive behavior prediction model.

[0086] 402. The server obtains the true posterior distribution of the first model parameters based on multiple sample media resources, the interactive behavior prediction model, and the first model parameters.

[0087] In this embodiment, the sample media resource refers to the media resource predicted by the interactive behavior prediction model within the online time period. Illustratively, the sample media resource indicates that it is based on multiple dimensions of features. For example, taking an advertisement as an example, these multiple dimensions of features include the profile information of the object initiating the resource acquisition request, the context information when the resource acquisition request is initiated, the type of advertisement, and the advertisement content, etc., without limitation.

[0088] In some embodiments, the sample media resource carries tag information indicating whether the sample object interacted with the sample media resource when it was recommended to the sample object. Illustratively, taking the sample media resource as an advertisement and the interaction behavior prediction model as an example to predict click-through rate, the tag information is divided into 1 and 0. A tag of 1 indicates that the sample media resource was clicked by the sample object, and a tag of 0 indicates that the sample media resource was not clicked by the sample object.

[0089] 403. If the first relative entropy between the approximate posterior distribution and the true posterior distribution meets the target condition, the server updates the model parameters of the interactive behavior prediction model to the first model parameters.

[0090] In this embodiment, the first relative entropy, also known as the first KL divergence, is used to measure the distance between the approximate posterior distribution and the true posterior distribution of the first model parameters. Indicatively, the smaller the distance (or the smaller the difference) between these two posterior distributions, the smaller the first relative entropy. Correspondingly, this indicates that the accuracy of the interaction behavior prediction model applying the first model parameters is higher than the accuracy of the current interaction behavior prediction model. Based on this, if the first relative entropy meets the target condition, the server updates the model parameters of the interaction behavior prediction model to the first model parameters. The updated interaction behavior prediction model then has higher accuracy, and when applied to the recommendation system, it can effectively improve the accuracy of the prediction results, thereby increasing the exposure and conversion rate of media resources.

[0091] 404. If the first relative entropy does not meet the target condition, the server updates the standardized flow model. Based on the updated standardized flow model, the interaction behavior prediction model continues to be trained until the obtained target relative entropy meets the target condition. Then, the model parameters of the interaction behavior prediction model are updated to the target model parameters corresponding to the target relative entropy.

[0092] In this embodiment, if the first relative entropy does not meet the target condition, it indicates that the accuracy of the interactive behavior prediction model using the first model parameters is lower than the accuracy of the current interactive behavior prediction model, or the difference is small. In other words, it indicates that the accuracy of the first model parameters output by the standardized flow model needs to be improved. Based on this, the server updates the standardized flow model. Based on the updated standardized flow model, the interactive behavior prediction model is trained again according to the same process as steps 401 and 402 above (or the model parameters that can be applied to the interactive behavior prediction model are predicted again, which can also be understood as performing the next iteration) until the target relative entropy meets the target condition. Then, the model parameters of the interactive behavior prediction model are updated to the target model parameters corresponding to the target relative entropy.

[0093] In this embodiment, a standardized flow model is used to obtain the first model parameters and their approximate posterior distribution for the interactive behavior prediction model. After obtaining the true posterior distribution of the first model parameters based on multiple sample media resources and the current model parameters of the interactive behavior prediction model, the model parameters are updated according to the relative entropy between the approximate posterior distribution and the true posterior distribution of the first model parameters. Since the standardized flow model can transform a simple distribution into a complex distribution, the approximate posterior distribution of the model parameters obtained based on this standardized flow model is closer to the true posterior distribution of the model parameters, thereby effectively improving the accuracy of the interactive behavior prediction model.

[0094] According to the above Figure 4 The illustrated embodiments briefly describe the training method of the interactive behavior prediction model provided in this application. The following is based on... Figure 5 The illustrated embodiment provides a detailed description of the training method for this interactive behavior prediction model.

[0095] Figure 5 This is a flowchart of another training method for an interactive behavior prediction model provided according to an embodiment of this application, such as... Figure 5 As shown, this application embodiment uses a server as an example for illustration. The method includes the following steps:

[0096] 501. The server acquires multiple sample media resources and interactive behavior prediction models.

[0097] In this embodiment, the server acquires multiple sample media resources and the current model parameters of the interactive behavior prediction model at preset time intervals to train the interactive behavior prediction model. For example, the preset time interval is 30 minutes, but this is not limited. In some embodiments, when the interactive behavior prediction model is already online, the server acquires multiple sample media resources and the interactive behavior prediction model at preset time intervals to achieve real-time training of the interactive behavior prediction model, but this is not limited.

[0098] In some embodiments, the plurality of sample media resources are all sample media resources stored in the server for training the interactive behavior prediction model. The interactive behavior prediction model trained in this way fully considers the global distribution of sample media resources, thereby improving the accuracy of the interactive behavior prediction model.

[0099] In other embodiments, the sample media resources are a subset of sample media resources stored on the server for training the interactive behavior prediction model. Illustratively, the server acquires the prediction results and actual results of multiple candidate sample media resources predicted by the interactive behavior prediction model within a target time period. Taking any one of these candidate sample media resources as the first candidate sample media resource, if the difference between the prediction result and the actual result of the first candidate sample media resource is greater than or equal to a target threshold, the first candidate sample media resource is determined as a sample media resource and used in the subsequent training process of the interactive behavior prediction model. Wherein, if the difference between the prediction result and the actual result of the first candidate sample media resource is greater than or equal to the target threshold, it indicates that there is a significant deviation between the prediction result and the actual result of the first candidate sample media resource. By incorporating such sample media resources into the subsequent training process of the interactive behavior prediction model, the learning of model parameters can be strengthened, thereby further improving the accuracy of the model. For example, the target time period is one week, and the target threshold is 0.5, which is not limited.

[0100] 502. The server processes the first vector obtained by sampling based on the target distribution using a standardized flow model to obtain the first model parameters of the interactive behavior prediction model and the approximate posterior distribution of the first model parameters.

[0101] In this embodiment, the server randomly samples based on the target distribution to obtain a first vector. The first vector is then processed using a standardized flow model to obtain first model parameters. An approximate posterior distribution of the first model parameters is then fitted based on at least one bijective function in the standardized flow model and the target distribution.

[0102] To facilitate understanding, the principles of the standardized flow model will be introduced below.

[0103] Schematic, the standardized flow model is mainly based on the Change of Variable Theorem, assuming that a random variable z follows a Poisson distribution, i.e., z ~ π(z), and at the same time constructing a bijective function g(·) and a new variable v = g(z). Based on the definition of probability distribution, the following differential equation can be obtained, i.e., formula (1):

[0104] ∫p(v)dv=∫π(z)dz (1)

[0105] Assuming that v and z are spatially continuous, then their integrals are the same, and the integrals of all distributions are 1. By transforming the differential equation above, we can obtain the following formula (2):

[0106] |p(v)dv|=|π(z)dz| (2)

[0107] Further transformation of the above formula (2) yields the following formula (3):

[0108]

[0109] In formula (3), det(·) represents the determinant, J g (·) denotes the Jacobian matrix.

[0110] Based on the variational theorem, g(·) makes π(z) approach p(v) by expanding or contracting the space. The absolute value of the Jacobian matrix quantifies the volume near the original vector z. Since the volume changes relatively after the g(·) transformation—that is, when an infinitesimal volume dz near z is mapped to an infinitesimal volume dv near v after g(·), then det(·) equals dv divided by dz, meaning the mapped volume is several times larger than the original. Because the probabilities contained in dz are equal to those contained in dv, if the volume of dz is enlarged, the probability density in dv should decrease. By establishing a long series of such invertible mappings (also called invertible transformations), the variational theorem can be repeatedly applied to obtain the final probability distribution. Illustratively, this process is referenced... Figure 6 , Figure 6 This is a schematic diagram of a standardized flow model provided according to an embodiment of this application, such as... Figure 6 As shown, by using multiple bijective functions in the normalized flow model, an invertible transformation is performed on the simple distribution z0~p0(z0), ultimately yielding the complex distribution z. k ~p k (z k ).

[0111] Based on the principles of the standardized flow model described above, the implementation method of step 502 is described below, including steps 1 and 2:

[0112] Step 1: Input the first vector into the normalized flow model, and process the first vector based on at least one bijective function in the normalized flow model to obtain the first model parameters.

[0113] In this standardized flow model, the flow depth is K, where K is a positive integer, and the number of at least one bijective function is also K. For example, K = 10, but this is not limited and can be set according to requirements.

[0114] Schematic, the server performs an invertible transformation on the intermediate model parameters obtained from the bijective function with a flow depth of K-1, based on the bijective function with a flow depth of K, to obtain the first model parameters. This process is described in the following formula (4):

[0115]

[0116] In the formula, θ represents the first model parameter, g(·) represents the bijective function, and z is the first vector. This represents a mapping multiplication, where g is the value of g. K The input is g K-1 The output.

[0117] The following example, using K=10, illustrates the principle of the invertible transformation of the first vector based on the bijective function g1(z) with a flow depth of 1. Illustratively, the first vector is z, and splitting the first vector z yields two sub-vectors zi. a and z b As shown in the following formula (5):

[0118] z a , z b =split(z) (5)

[0119] In the formula, split(·) means splitting the first vector z.

[0120] Next, the first intermediate model parameter θ1 is obtained based on the following formulas (6) to (10).

[0121] logs, t = NN(z) b (6)

[0122] s = exp(logs) (7)

[0123] θ a =s⊙z a +t (8)

[0124] θ b =z b (9)

[0125] θ1 = concat(θ) a θ b (10)

[0126] In formulas (6) to (10) above, NN(·) is a DNN structure, for example, the DNN structure is built based on three fully connected layers and a nonlinear activation function, which is not limited. s is the linear transformation coefficient, and t is the translation amount (which can also be understood as the intercept). ⊙ represents the Hadamard product. Schematically, formula (8) above can be understood as an affine transformation. concat(·) is the opposite of split(·), indicating that two vectors θ are concatenated. a and θ b They are spliced ​​together to form θ1.

[0127] After obtaining the first intermediate model parameter θ1, use θ1 as the input to the bijective function g2 with a flow depth of 2 for the next reversible transformation to obtain the second intermediate model parameter θ2, and so on, until the bijective function g2 with a flow depth of 10 is obtained. 10 The intermediate model parameters obtained based on the bijective function g9 with a flow depth of 9 are reversibly transformed to obtain the first model parameters.

[0128] Step 2: Based on the target distribution, the first vector, and the at least one bijective function, obtain the approximate posterior distribution of the first model parameters.

[0129] The server obtains the approximate posterior distribution of the first model parameter based on the following formula (11). This process can also be understood as fitting the posterior distribution of the first model parameter using a standardized flow model so that the approximate posterior distribution obtained by fitting is as close as possible to the true posterior distribution of the first model parameter.

[0130]

[0131] In the formula, x represents the sample media resource, y represents the tag information of the sample media resource, q(θ|x,y) represents the approximate posterior distribution of the first model parameters, and i is a positive integer.

[0132] 503. The server obtains the true posterior distribution of the first model parameters based on multiple sample media resources, the interactive behavior prediction model, and the first model parameters.

[0133] In this embodiment, the server obtains the true posterior distribution of the first model parameters based on the distribution of the current model parameters of the interaction behavior prediction model, according to Bayes' theorem. Illustratively, this process refers to the following formula (12):

[0134]

[0135] In the formula, x represents the sample media resource, y represents the tag information of the sample media resource, p(θ|x,y) represents the true posterior distribution of the first model parameter, p(θ) represents the prior distribution of the first model parameter, p(y|θ,x) represents the prediction result output by the interactive behavior prediction model applying the first model parameter, and p(y) represents the global distribution of multiple sample media resources and corresponding tag information.

[0136] Indicatively, step 503 includes the following steps:

[0137] Step 1: Based on the distribution of the model parameters of the interactive behavior prediction model, obtain the prior distribution of the first model parameters. This process involves substituting the first model parameters into the distribution of the current model parameters of the interactive behavior prediction model to obtain the prior distribution of the first model parameters.

[0138] Step 2: Based on the multiple sample media resources, their tag information, and the interactive behavior prediction model applying the first model parameters, obtain the prediction results for the multiple sample media resources. This process involves inputting the multiple sample media resources into the interactive behavior prediction model applying the first model parameters to obtain the corresponding prediction results.

[0139] Step 3: Based on the prior distribution of the first model parameters and the prediction results of the multiple sample media resources, obtain the true posterior distribution of the first model parameters.

[0140] 504. The server obtains the first relative entropy between the approximate posterior distribution and the true posterior distribution of the first model parameters.

[0141] In this embodiment of the application, after the server obtains the approximate posterior distribution and the true posterior distribution of the first model parameters through the above steps 502 to 503, the server calculates the first relative entropy between the approximate posterior distribution and the true posterior distribution, as shown in the following formula (13):

[0142]

[0143] In the formula, q(θ|x,y) represents the approximate posterior distribution of the first model parameters, and p(θ|x,y) represents the true posterior distribution of the first model parameters.

[0144] It should be understood that during the computation process, after obtaining the approximate posterior distribution of the first model parameters, the server can directly perform simplified computation based on the above formula (13) to obtain the first relative entropy, without having to perform the above step 503 to obtain the true posterior distribution of the first model parameters. The following is an introduction to this simplified computation process:

[0145] Substituting formula (12) into formula (13) and simplifying it, we can obtain the following formula (14):

[0146]

[0147] In the formula, To express the expectation, further substituting the approximate posterior distribution shown in the above formula (11) into formula (14), we can obtain the following formula (15):

[0148]

[0149] It should be noted that the meanings of the parameters in the above formulas (14) and (15) are the same as those in the aforementioned formulas (4) to (13), so they will not be repeated here.

[0150] 505. If the first relative entropy meets the target condition, the server updates the model parameters of the interactive behavior prediction model to the first model parameters.

[0151] In this embodiment of the application, this step is the same as the aforementioned Figure 4 Step 403 in the illustrated embodiment is similar. Furthermore, based on step 403 above, it can be seen that the smaller the first relative entropy, the higher the accuracy of the interaction behavior prediction model applying the first model parameters compared to the current interaction behavior prediction model. Therefore, the optimization objective of the interaction behavior prediction model is to minimize the relative entropy between the true posterior distribution and the approximate posterior distribution of the model parameters, as shown in the following formula (16):

[0152] minKL(q(θ|x,y)||p(θ|x,y)) (16)

[0153] In other words, theoretically, the optimal solution q learned after optimization by the interactive behavior prediction model is... * (θ|x, y) = p(θ|x, y).

[0154] Schematic, in some embodiments, the target condition indicates that the relative entropy is the minimum value in the most recent N iterations, where N is a positive integer. In other embodiments, the target condition indicates that the difference between the relative entropy and the relative entropy in the previous iteration is less than or equal to a preset threshold, which can also be understood as the relative entropy tending to level off. This application does not limit the specific setting of the above target condition, as long as the target condition can minimize the relative entropy between the true posterior distribution and the approximate posterior distribution of the model parameters.

[0155] 506. If the first relative entropy does not meet the target condition, the server updates at least one bijective function in the normalized flow model to obtain the updated normalized flow model.

[0156] In this embodiment of the application, if the first relative entropy does not meet the target condition, the server updates the parameters of at least one bijective function in the normalized flow model to obtain the updated normalized flow model.

[0157] 507. Based on the updated standardized flow model, the server continues to train the interaction behavior prediction model until the obtained target relative entropy meets the target condition, and then updates the model parameters of the interaction behavior prediction model to the target model parameters corresponding to the target relative entropy.

[0158] In this embodiment, the server continues to train the interactive behavior prediction model based on the updated standardized flow model, following a process similar to steps 502 to 505 above (or continues to predict model parameters that can be applied to the interactive behavior prediction model, which can also be understood as performing the next iteration), until the obtained target relative entropy meets the target condition, and then updates the model parameters of the interactive behavior prediction model to the target model parameters corresponding to the target relative entropy.

[0159] The following is an illustrative description of the prediction process based on the updated interactive behavior prediction model, referring to the following formula (17):

[0160]

[0161] In the formula, P(y * |x * (x, y) represents the prediction result for the media resource x* to be predicted. In addition, the meaning of each parameter in this formula (17) is the same as that of the aforementioned formulas (4) to (16), so it will not be repeated.

[0162] Following the training process for the interactive behavior prediction model as described in steps 501 to 507 above, a more accurate interactive behavior prediction model can be obtained. Applying this updated interactive behavior prediction model to a recommendation system can improve the exposure and conversion rates of media resources. (Illustratively, refer to...) Figure 7 , Figure 7 This is a schematic diagram of a model test result provided according to an embodiment of this application. For example... Figure 7 As shown, this application tested the model performance of related technologies and the interactive behavior prediction model trained by this application based on sample media resources. According to the AUC results, the AUC of the interactive behavior prediction model trained by this application is significantly higher than that of related technologies, indicating that the interactive behavior prediction model trained by this application has higher accuracy.

[0163] In summary, in this embodiment, the first model parameters and their approximate posterior distribution are obtained based on a standardized flow model. After obtaining the true posterior distribution of the first model parameters based on multiple sample media resources and the current model parameters of the interactive behavior prediction model, the model parameters are updated according to the relative entropy between the approximate posterior distribution and the true posterior distribution. Since the standardized flow model can transform a simple distribution into a complex distribution, the approximate posterior distribution of the model parameters obtained based on this standardized flow model is closer to the true posterior distribution of the model parameters, thereby effectively improving the accuracy of the interactive behavior prediction model.

[0164] Figure 8 This is a schematic diagram of a training device for an interactive behavior prediction model according to an embodiment of this application. The device is used to execute the training method for the aforementioned interactive behavior prediction model. See also... Figure 8 The device includes: a processing module 801, an acquisition module 802, a first update module 803, and a second update module 804.

[0165] The processing module 801 is used to process the first vector obtained by sampling based on the target distribution based on the standardized flow model to obtain the first model parameters of the interactive behavior prediction model and the approximate posterior distribution of the first model parameters. The interactive behavior prediction model is used to predict the probability of performing interactive behavior on media resources.

[0166] The acquisition module 802 is used to acquire the true posterior distribution of the first model parameters based on multiple sample media resources, the interactive behavior prediction model, and the first model parameters.

[0167] The first update module 803 is used to update the model parameters of the interactive behavior prediction model to the first model parameters if the first relative entropy between the approximate posterior distribution and the true posterior distribution meets the target condition.

[0168] The second update module 804 is used to update the standardized flow model if the first relative entropy does not meet the target condition, and continue to train the interaction behavior prediction model based on the updated standardized flow model until the obtained target relative entropy meets the target condition, and update the model parameters of the interaction behavior prediction model to the target model parameters corresponding to the target relative entropy.

[0169] In some embodiments, the processing module 801 includes:

[0170] The processing unit is configured to input the first vector into the normalized flow model, and process the first vector based on at least one bijective function in the normalized flow model to obtain the first model parameters.

[0171] The acquisition unit is used to acquire the approximate posterior distribution of the first model parameters based on the target distribution, the first vector, and the at least one bijective function.

[0172] In some embodiments, the processing unit is configured to:

[0173] Based on the bijective function with a flow depth of K, the intermediate model parameters obtained from the bijective function with a flow depth of K-1 are reversibly transformed to obtain the first model parameters. K indicates the flow depth of the standardized flow model, and K is a positive integer.

[0174] In some embodiments, the acquisition module 802 is configured to:

[0175] Based on the distribution of the model parameters of the interactive behavior prediction model, the prior distribution of the first model parameters is obtained.

[0176] Based on the multiple sample media resources, the tag information of the multiple sample media resources, and the interactive behavior prediction model applying the first model parameters, the prediction results of the multiple sample media resources are obtained.

[0177] Based on the prior distribution of the first model parameters and the prediction results of the multiple sample media resources, the true posterior distribution of the first model parameters is obtained.

[0178] In some embodiments, the second update module 804 is configured to:

[0179] If the first relative entropy does not meet the target condition, at least one bijective function in the normalized flow model is updated to obtain the updated normalized flow model.

[0180] In some embodiments, the apparatus further includes a sample media resource determination unit, configured to:

[0181] Obtain the prediction results and actual results of multiple candidate media resources predicted by the interactive behavior prediction model within the target time period;

[0182] If the difference between the predicted result and the actual result of the first candidate sample media resource is greater than or equal to the target threshold, the first candidate sample media resource is determined as a sample media resource. The first candidate sample media resource is any one of the multiple candidate sample media resources.

[0183] Using the aforementioned apparatus, the first model parameters and their approximate posterior distribution are obtained based on a standardized flow model for the interactive behavior prediction model. After obtaining the true posterior distribution of the first model parameters based on multiple sample media resources and the current model parameters of the interactive behavior prediction model, the model parameters are updated according to the relative entropy between the approximate posterior distribution and the true posterior distribution of the first model parameters. Since the standardized flow model can transform a simple distribution into a complex distribution, the approximate posterior distribution of the model parameters obtained based on this standardized flow model is closer to the true posterior distribution of the model parameters, thereby effectively improving the accuracy of the interactive behavior prediction model.

[0184] It should be noted that the training device for the interactive behavior prediction model provided in the above embodiments is only illustrated by the division of the above functional modules during the training process. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the training device for the interactive behavior prediction model and the training method embodiment for the interactive behavior prediction model provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiment, which will not be repeated here.

[0185] In an exemplary embodiment, a computer device is also provided, the computer device including a processor and a memory for storing at least one computer program, the at least one computer program being loaded and executed by the processor to implement the training method of the interactive behavior prediction model in the embodiments of this application.

[0186] Taking computer equipment as a server as an example, Figure 9 This is a schematic diagram of a server structure according to an embodiment of this application. The server 900 can vary considerably depending on its configuration or performance, and may include one or more Central Processing Units (CPUs) 901 and one or more memories 902. The memories 902 store at least one computer program, which is loaded and executed by the processor 901 to implement the training method for the interactive behavior prediction model provided in the various method embodiments described above. Of course, the server may also have wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. The server may also include other components for implementing device functions, which will not be elaborated upon here.

[0187] This application also provides a computer-readable storage medium applied to a computer device, wherein the computer-readable storage medium stores at least one computer program, which is loaded and executed by a processor to implement the training method of the interactive behavior prediction model in the above embodiments.

[0188] This application also provides a computer program product or computer program, which includes computer program code stored in a computer-readable storage medium. A processor of a computer device reads the computer program code from the computer-readable storage medium and executes the computer program code, causing the computer device to perform a training method for the interactive behavior prediction model described in the above embodiments.

[0189] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0190] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A training method for an interactive behavior prediction model, characterized in that, The method includes: Based on the K bijective functions in the standardized flow model, the first vector obtained by sampling from the target distribution is processed to obtain the first model parameters of the interactive behavior prediction model, where K is a positive integer and K is the flow depth of the standardized flow model. The bijective function at flow depth i is used to perform an invertible transformation on the intermediate model parameters output by the bijective function at flow depth i-1. The interactive behavior prediction model is used to predict the probability of engaging in interactive behavior with media resources. Based on the target distribution, the first vector, and the K bijective functions, obtain the approximate posterior distribution of the first model parameters; Based on multiple sample media resources, the interactive behavior prediction model, and the first model parameters, the true posterior distribution of the first model parameters is obtained. The multiple sample media resources are samples in the target time period whose difference between the prediction result and the true result is greater than or equal to the target threshold. If the first relative entropy between the approximate posterior distribution and the true posterior distribution meets the target condition, the model parameters of the interactive behavior prediction model are updated to the first model parameters. If the first relative entropy does not meet the target condition, the standardized flow model is updated. Based on the updated standardized flow model, the interaction behavior prediction model is trained until the obtained target relative entropy meets the target condition. Then, the model parameters of the interaction behavior prediction model are updated to the target model parameters corresponding to the target relative entropy.

2. The method according to claim 1, characterized in that, The step of obtaining the true posterior distribution of the first model parameters based on multiple sample media resources, the interactive behavior prediction model, and the first model parameters includes: Based on the distribution of the model parameters of the interactive behavior prediction model, the prior distribution of the first model parameters is obtained. Based on the multiple sample media resources, the tag information of the multiple sample media resources, and the interactive behavior prediction model applying the first model parameters, the prediction results of the multiple sample media resources are obtained. Based on the prior distribution of the first model parameters and the prediction results of the multiple sample media resources, the true posterior distribution of the first model parameters is obtained.

3. The method according to claim 1, characterized in that, If the first relative entropy does not meet the target condition, updating the standardized flow model includes: If the first relative entropy does not meet the target condition, at least one bijective function in the standardized flow model is updated to obtain the updated standardized flow model.

4. The method according to claim 1, characterized in that, The method further includes: Obtain the prediction results and actual results of multiple candidate media resources predicted by the interactive behavior prediction model within the target time period; If the difference between the predicted result and the actual result of the first candidate sample media resource is greater than or equal to the target threshold, the first candidate sample media resource is determined as a sample media resource, and the first candidate sample media resource is any one of the plurality of candidate sample media resources.

5. A training device for an interactive behavior prediction model, characterized in that, The device includes: The processing module is used to process the first vector obtained by sampling based on the target distribution, based on K bijective functions in the standardized flow model, to obtain the first model parameters of the interactive behavior prediction model, where K is a positive integer and K is the flow depth of the standardized flow model. The bijective function at flow depth i is used to perform an invertible transformation on the intermediate model parameters output by the bijective function at flow depth i-1. The interactive behavior prediction model is used to predict the probability of performing interactive behavior on media resources; based on the target distribution, the first vector, and the K bijective functions, the approximate posterior distribution of the first model parameters is obtained; The acquisition module is used to acquire the true posterior distribution of the first model parameters based on multiple sample media resources, the interactive behavior prediction model, and the first model parameters. The multiple sample media resources are samples in the target time period whose difference between the prediction result and the true result is greater than or equal to a target threshold. The first update module is used to update the model parameters of the interactive behavior prediction model to the first model parameters if the first relative entropy between the approximate posterior distribution and the true posterior distribution meets the target condition. The second update module is used to update the standardized flow model if the first relative entropy does not meet the target condition, and continue to train the interaction behavior prediction model based on the updated standardized flow model until the obtained target relative entropy meets the target condition, and update the model parameters of the interaction behavior prediction model to the target model parameters corresponding to the target relative entropy.

6. A computer device, characterized in that, The computer device includes a processor and a memory, the memory being used to store at least one computer program, the at least one computer program being loaded by the processor and executed as a training method for an interactive behavior prediction model as described in any one of claims 1 to 4.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one computer program, which is loaded and executed by a processor to implement the training method of the interactive behavior prediction model as described in any one of claims 1 to 4.

8. A computer program product, characterized in that, The computer program product includes at least one computer program, which is loaded and executed by a processor to implement the training method of the interactive behavior prediction model as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Recommendation model training method, medium, electronic equipment and recommendation model

    CN112184391A