Model processing method and device, information pushing method and device, computer equipment and storage medium

By introducing multiple predictive sub-models into the interaction probability prediction model, using the historical interaction results and description information of sample users, gradually predicting the probability of each interactive task, solving the problem of low model accuracy during the user's multi-step conversion process, and improving the accuracy of information push.

CN120386914APending Publication Date: 2025-07-29TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410124343.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-29
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing interaction probability prediction model has low accuracy during the user's multi-step conversion process, resulting in poor information push effect.

Method used

By introducing multiple predictive sub-models into the interaction probability prediction model, the historical interaction results and description information of sample users are used to gradually predict the probability of each interactive task, and train it in combination with sampling information, making full use of the sequence dependence between tasks.

Benefits of technology

The accuracy of the interaction probability prediction model is improved, thereby improving the accuracy and effectiveness of information push.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120386914A_ABST
    Figure CN120386914A_ABST
Patent Text Reader

Abstract

The invention relates to a model processing method and device, an information pushing method and device, computer equipment and a storage medium. The method comprises the following steps: in an ith prediction sub-model in an interaction probability prediction model, predicting to obtain a second interaction probability based on user description information of a sample user under an ith interaction task and first task information determined based on a historical interaction result under an (i-1) th interaction task; in the ith prediction sub-model, predicting a first interaction probability corresponding to the ith interaction task based on user description information of the sample user under the ith interaction task and second task information determined based on the sampling information; and training an interaction probability prediction model based on the first interaction probability and the second interaction probability corresponding to the ith interaction task. The method can improve the accuracy of the interaction probability prediction model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technologies, and in particular, to a model processing method, apparatus, computer device, storage medium, and computer program product, as well as an information push method, apparatus, computer device, storage medium, and computer program product. Background Art

[0002] With the development of computer technologies, information push technologies have emerged. Information push refers to pushing information to the terminals of target users so that the target users can perform interactive operations on the pushed information, thereby achieving the purpose of targeted customer acquisition. Here, the target users usually need to predict the interaction probability of candidate users through an interaction probability prediction model, so as to recall the target users from the candidate users for information push.

[0003] The interaction probability prediction model needs to be trained with samples. However, targeted customer acquisition is usually a process of multi-step conversion of users, that is, users need to perform multiple interaction tasks. For example, in financial advertisements, such as credit card services, user conversion is usually exposure → click → application → card approval. Due to the longer path dependence of multi-step user conversion and the gradual sparsity of positive samples, the accuracy of the trained interaction probability prediction model is low. Summary of the Invention

[0004] Based on this, in view of the above technical problems, it is necessary to provide a model processing method, apparatus, computer device, computer-readable storage medium, and computer program product that can improve the accuracy of the interaction probability prediction model. In addition, an information push method, apparatus, computer device, and storage medium that can improve the accuracy of information push are also provided.

[0005] In a first aspect, the present application provides a model processing method. The method includes:

[0006] Obtain user description information of a sample user under each interaction task in n interaction tasks, where n is a positive integer;

[0007] In the first prediction sub-model of the interaction probability prediction model, determine a first interaction probability corresponding to the first interaction task based on the user description information of the sample user under the first interaction task;

[0008] In the i-th prediction sub-model of the interaction probability prediction model, predict a second interaction probability corresponding to the i-th interaction task based on the user description information of the sample user under the i-th interaction task and the first task information of the (i - 1)-th interaction task determined by the (i - 1)-th prediction sub-model based on the historical interaction result under the (i - 1)-th interaction task, where i is a positive integer greater than or equal to 2 and not greater than n;

[0009] In the i-th prediction sub-model, based on the user description information of the sample user under the i-th interaction task and the second task information of the (i - 1)-th interaction task determined by the (i - 1)-th prediction sub-model based on the sampling information, the first interaction probability corresponding to the i-th interaction task is predicted; the sampling information is sampled by the (i - 1)-th prediction sub-model from the historical interaction results of the sample user under the (i - 1)-th interaction task and the first interaction probability corresponding to the (i - 1)-th interaction task;

[0010] Based on the first interaction probability corresponding to the i-th interaction task and the second interaction probability, the interaction probability prediction model is trained.

[0011] In a second aspect, the present application also provides a model processing device. The device includes:

[0012] A user description information acquisition module, configured to acquire the user description information of the sample user under each interaction task in n interaction tasks, where n is a positive integer;

[0013] A first prediction module, configured to determine, in the first prediction sub-model of the interaction probability prediction model, the first interaction probability corresponding to the first interaction task based on the user description information of the sample user under the first interaction task;

[0014] A second prediction module, configured to, in the i-th prediction sub-model of the interaction probability prediction model, predict the second interaction probability corresponding to the i-th interaction task based on the user description information of the sample user under the i-th interaction task and the first task information of the (i - 1)-th interaction task determined by the (i - 1)-th prediction sub-model based on the historical interaction results under the (i - 1)-th interaction task, where i is a positive integer greater than or equal to 2 and not greater than n;

[0015] The second prediction module is further configured to, in the i-th prediction sub-model, predict the first interaction probability corresponding to the i-th interaction task based on the user description information of the sample user under the i-th interaction task and the second task information of the (i - 1)-th interaction task determined by the (i - 1)-th prediction sub-model based on the sampling information; the sampling information is sampled by the (i - 1)-th prediction sub-model from the historical interaction results of the sample user under the (i - 1)-th interaction task and the first interaction probability corresponding to the (i - 1)-th interaction task;

[0016] A model training module, configured to train the interaction probability prediction model based on the first interaction probability corresponding to the i-th interaction task and the second interaction probability.

[0017] In a third aspect, the present application also provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0018] Obtain user description information of a sample user for each interaction task in n interaction tasks, where n is a positive integer;

[0019] In the first prediction sub-model of the interaction probability prediction model, based on the user description information of the sample user for the first interaction task, determine the first interaction probability corresponding to the first interaction task;

[0020] In the i-th prediction sub-model of the interaction probability prediction model, based on the user description information of the sample user for the i-th interaction task and the first task information of the (i - 1)-th interaction task determined by the (i - 1)-th prediction sub-model based on the historical interaction result for the (i - 1)-th interaction task, predict the second interaction probability corresponding to the i-th interaction task, where i is a positive integer greater than or equal to 2 and not greater than n;

[0021] In the i-th prediction sub-model, based on the user description information of the sample user for the i-th interaction task and the second task information of the (i - 1)-th interaction task determined by the (i - 1)-th prediction sub-model based on sampling information, predict the first interaction probability corresponding to the i-th interaction task; the sampling information is sampled by the (i - 1)-th prediction sub-model from the historical interaction result of the sample user for the (i - 1)-th interaction task and the first interaction probability corresponding to the (i - 1)-th interaction task;

[0022] Train the interaction probability prediction model based on the first interaction probability and the second interaction probability corresponding to the i-th interaction task.

[0023] In a fourth aspect, the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, and when the computer program is executed by a processor, the following steps are implemented:

[0024] Obtain user description information of a sample user for each interaction task in n interaction tasks, where n is a positive integer;

[0025] In the first prediction sub-model of the interaction probability prediction model, based on the user description information of the sample user for the first interaction task, determine the first interaction probability corresponding to the first interaction task;

[0026] In the i-th prediction sub-model of the interaction probability prediction model, based on the user description information of the sample user under the i-th interaction task, and the first task information of the (i - 1)-th interaction task determined by the (i - 1)-th prediction sub-model based on the historical interaction results under the (i - 1)-th interaction task, the second interaction probability corresponding to the i-th interaction task is predicted, where i is a positive integer greater than or equal to 2 and not greater than n;

[0027] In the i-th prediction sub-model, based on the user description information of the sample user under the i-th interaction task, and the second task information of the (i - 1)-th interaction task determined by the (i - 1)-th prediction sub-model based on the sampling information, the first interaction probability corresponding to the i-th interaction task is predicted; the sampling information is sampled by the (i - 1)-th prediction sub-model from the historical interaction results of the sample user under the (i - 1)-th interaction task and the first interaction probability corresponding to the (i - 1)-th interaction task;

[0028] Based on the first interaction probability and the second interaction probability corresponding to the i-th interaction task, the interaction probability prediction model is trained.

[0029] In a fifth aspect, the present application also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the following steps are implemented:

[0030] Obtain the user description information of the sample user under each interaction task in n interaction tasks, where n is a positive integer;

[0031] In the first prediction sub-model of the interaction probability prediction model, based on the user description information of the sample user under the first interaction task, determine the first interaction probability corresponding to the first interaction task;

[0032] In the i-th prediction sub-model of the interaction probability prediction model, based on the user description information of the sample user under the i-th interaction task, and the first task information of the (i - 1)-th interaction task determined by the (i - 1)-th prediction sub-model based on the historical interaction results under the (i - 1)-th interaction task, the second interaction probability corresponding to the i-th interaction task is predicted, where i is a positive integer greater than or equal to 2 and not greater than n;

[0033] In the i-th prediction sub-model, based on the user description information of the sample user in the i-th interaction task and the second task information of the (i - 1)-th interaction task determined by the (i - 1)-th prediction sub-model based on the sampling information, the first interaction probability corresponding to the i-th interaction task is predicted; the sampling information is sampled by the (i - 1)-th prediction sub-model from the historical interaction results of the sample user in the (i - 1)-th interaction task and the first interaction probability corresponding to the (i - 1)-th interaction task.

[0034] Train the interaction probability prediction model based on the first interaction probability corresponding to the i-th interaction task and the second interaction probability.

[0035] In the above model processing method, device, computer device, storage medium, and computer program product, since in the i-th prediction sub-model, based on the user description information of the sample user in the i-th interaction task and the second task information of the (i - 1)-th interaction task determined by the (i - 1)-th prediction sub-model based on the sampling information, the first interaction probability corresponding to the i-th interaction task is predicted, and the sampling information is sampled by the (i - 1)-th prediction sub-model from the historical interaction results of the sample user in the (i - 1)-th interaction task and the first interaction probability corresponding to the (i - 1)-th interaction task, the downstream task can obtain relatively accurate pre-task information during the entire training process. And since in the i-th prediction sub-model of the interaction probability prediction model, based on the user description information of the sample user in the i-th interaction task and the first task information of the (i - 1)-th interaction task determined by the (i - 1)-th prediction sub-model based on the historical interaction results in the (i - 1)-th interaction task, the second interaction probability corresponding to the i-th interaction task is predicted, thus the sequential dependence relationship between tasks can be utilized more fully, so that for the i-th task, the entire training process can utilize the label information of the (i - 1)-th task to obtain a more accurate prediction value. By training the interaction probability prediction model based on the first interaction probability and the second interaction probability corresponding to the i-th interaction task, a more accurate interaction probability prediction model can be trained.

[0036] In a sixth aspect, the present application provides an information push method. The method includes:

[0037] Determine a candidate user set for the information to be pushed;

[0038] For each candidate user in the candidate user set, determine the user description information of the candidate user in each interaction task among n interaction tasks, where n is a positive integer;

[0039] Input the user description information under each interaction task into each prediction sub - model of the interaction probability prediction model respectively to obtain the interaction probability of the candidate user for the information to be pushed; wherein, the interaction probability prediction model is trained by using the above - mentioned model processing method;

[0040] Based on the interaction probability of each candidate user, determine the target push user of the information to be pushed from the candidate user set;

[0041] Push the information to be pushed to the terminal corresponding to the target push user.

[0042] In a seventh aspect, the present application further provides an information push device. The device includes:

[0043] A candidate user determination module, configured to determine a candidate user set of the information to be pushed;

[0044] A user description information determination module, configured to determine, for each candidate user in the candidate user set, the user description information of the candidate user under each interaction task in n interaction tasks, where n is a positive integer;

[0045] An interaction probability prediction module, configured to input the user description information under each interaction task into each prediction sub - model of the interaction probability prediction model respectively to obtain the interaction probability of the candidate user for the information to be pushed; wherein, the interaction probability prediction model is trained by using the method described in any one of claims 1 to 11;

[0046] A target push user determination module, configured to determine the target push user of the information to be pushed from the candidate user set based on the interaction probability of each candidate user;

[0047] An information push module, configured to push the information to be pushed to the terminal corresponding to the target push user.

[0048] In an eighth aspect, the present application further provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0049] Determine a candidate user set of the information to be pushed;

[0050] For each candidate user in the candidate user set, determine the user description information of the candidate user under each interaction task in n interaction tasks, where n is a positive integer;

[0051] Input the user description information under each interaction task into each prediction sub-model of the interaction probability prediction model respectively to obtain the interaction probability of the candidate user for the information to be pushed; wherein, the interaction probability prediction model is trained by using the above model processing method;

[0052] Based on the interaction probability of each candidate user, determine the target push users of the information to be pushed from the candidate user set;

[0053] Push the information to be pushed to the terminal corresponding to the target push user.

[0054] In a ninth aspect, the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, and when the computer program is executed by a processor, the following steps are implemented:

[0055] Determine a candidate user set of the information to be pushed;

[0056] For each candidate user in the candidate user set, determine the user description information of the candidate user under each interaction task in n interaction tasks, where n is a positive integer;

[0057] Input the user description information under each interaction task into each prediction sub-model of the interaction probability prediction model respectively to obtain the interaction probability of the candidate user for the information to be pushed; wherein, the interaction probability prediction model is trained by using the above model processing method;

[0058] Based on the interaction probability of each candidate user, determine the target push users of the information to be pushed from the candidate user set;

[0059] Push the information to be pushed to the terminal corresponding to the target push user.

[0060] In a tenth aspect, the present application also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the following steps are implemented:

[0061] For the above information pushing method, device, computer device, storage medium and computer program product, since the interaction probability prediction model is trained by using the above model processing method, a more accurate interaction probability can be predicted by using the interaction probability prediction model. Therefore, based on the predicted interaction probability, the target users can be accurately determined, and when the information is pushed to the terminals of these target users, the possibility of obtaining positive feedback is greater, thereby improving the accuracy of information pushing. Description of the Drawings

[0062] Figure 1It is an application environment diagram of the model processing method and the information pushing method in some embodiments;

[0063] Figure 2 It is a schematic flowchart of the model processing method in some embodiments;

[0064] Figure 3 It is a schematic diagram of the connection relationship of the prediction sub-model in some embodiments;

[0065] Figure 4 It is a schematic flowchart of the steps for training the interaction probability prediction model in some other embodiments;

[0066] Figure 5 It is an example schematic diagram for determining the first ranking loss in some embodiments;

[0067] Figure 6 It is an example schematic diagram for determining the first ranking loss in some embodiments;

[0068] Figure 7 It is a schematic flowchart of the information pushing method in some embodiments;

[0069] Figure 8A It is a schematic diagram of the exposure interface in the advertising audience expansion scenario in some embodiments;

[0070] Figure 8B It is a schematic diagram of the click interface in the advertising audience expansion scenario in some embodiments;

[0071] Figure 8C It is a schematic diagram of the application interface in the advertising audience expansion scenario in some embodiments;

[0072] Figure 8D It is a schematic diagram of the card nuclear interface in the advertising audience expansion scenario in some embodiments;

[0073] Figure 9 It is a schematic diagram of the structure of the shared base model in some embodiments;

[0074] Figure 10 It is a schematic diagram of the structure of the PIKD model in some embodiments;

[0075] Figure 11 It is an example schematic diagram of the application of the PIT module in some embodiments;

[0076] Figure 12 It is a structural block diagram of the model processing device in some embodiments;

[0077] Figure 13 It is a structural block diagram of the information pushing device in some embodiments;

[0078] Figure 14 It is an internal structure diagram of a computer device in some embodiments;

[0079] Figure 15 It is the internal structure diagram of a computer device in some other embodiments. Specific implementation manners

[0080] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0081] The model processing method and information pushing method provided by the embodiments of the present application relate to technologies such as machine learning (ML) and reinforcement learning (RL) in artificial intelligence, where:

[0082] Artificial intelligence (AI) is a theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence is also the study of the design principles and implementation methods of various intelligent machines, enabling machines to have the functions of perception, reasoning and decision-making.

[0083] Artificial intelligence technology is a comprehensive discipline, involving a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include, for example, sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, pre-trained models, also known as large models and basic models, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0084] Machine Learning (ML) is an interdisciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning. Pre-trained models are the latest development results of deep learning, integrating the above technologies.

[0085] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, digital twins, virtual humans, robots, artificial intelligence-generated content (AIGC), conversational interaction, smart healthcare, smart customer service, game AI, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0086] The model processing method and information pushing method provided by the embodiments of this application can be applied to an application environment as Figure 1 shown. Among them, the terminal 102 and the server 104 can communicate with each other through a network, such as a wired or wireless network. The data storage system can store the data that the server 104 needs to process. The data storage system can be set up separately, integrated on the server 104, or placed in the cloud or on other servers. The terminal 102 can be, but is not limited to, various desktop computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc.

[0087] For the model processing method and information pushing method provided by the embodiments of this application, the execution entity of each step can be a computer device, which refers to an electronic device with data computing, processing, and storage capabilities. Taking Figure 1 the application environment shown as an example, the model processing method can be executed independently by the terminal 102, or independently by the server 104, or the terminal 102 and the server 104 can interact and cooperate to execute the model processing method. This application does not make any restrictions on this.

[0088] In some embodiments, as Figure 2As shown, a model processing method is provided. This method is executed by a computer device, which can be Figure 1 the server 104 or the terminal 102 in Figure 1 . In the embodiments of the present application, taking the case where this method is applied to

[0089] the server 104 in

[0090] as an example for illustration, the method includes the following steps:

[0091] Specifically, the server can obtain the description data of the user, input the description data of the user into the feature representation model, and output the user description information of the sample user under each interaction task in the n interaction tasks through the feature representation model. Optionally, the user description information under each interaction task can be the feature vector representation under each interaction task.

[0092] Optionally, the feature representation model can be a module of the interaction probability prediction model, and its parameters are adjusted through the process of training the interaction probability prediction model. Alternatively, the feature representation model can be a trained model structure, and its parameters can remain unchanged during the process of training the interaction probability prediction model.

[0093] In some embodiments, the feature representation model may include a shared representation layer corresponding to each interaction task and a feature representation layer unique to each interaction task. After the server inputs the description data of the sample user into the shared representation layer, the shared representation layer outputs the shared feature representations of the n interaction tasks. After the shared feature representations are respectively input into the feature representation layers unique to each interaction task, each feature representation layer outputs the user description information corresponding to its respective interaction task, and the user description information of the sample user under each interaction task in the n interaction tasks is obtained.

[0094] In some other embodiments, the feature representation model may also be implemented using structures such as a multi-gate expert network.

[0095] Step 204, in the first prediction sub-model of the interaction probability prediction model, based on the user description information of the sample user under the first interaction task, determine the first interaction probability corresponding to the first interaction task.

[0096] Among them, the interaction probability prediction model contains multiple prediction sub-models, and these prediction sub-models correspond one-to-one with the interaction tasks in the n interaction tasks required for obtaining customers through the to-be-pushed information. Moreover, the prediction sub-models in these prediction sub-models are connected in the task order of the corresponding interaction tasks. Thus, the first prediction sub-model in these prediction sub-models corresponds to the first interaction task in the n interaction tasks, the i-th prediction sub-model corresponds to the i-th interaction task in the n interaction tasks, where i is a positive integer greater than or equal to 2 and not greater than n. In other words, the first prediction sub-model is used to predict the interaction probability of the user under the first interaction task, and the i-th prediction sub-model is used to predict the interaction probability of the user under the i-th interaction task. For example, referring to Figure 3 , it is a schematic diagram of the connection relationship of the prediction sub-models in some embodiments. Assume n = 3, that is, after the to-be-pushed information is pushed to the corresponding terminal of the user, the user needs to perform interaction operations under 3 interaction tasks. Then the interaction probability prediction model contains 3 prediction sub-models. Among them, the first prediction sub-model corresponds to the first interaction task, the second prediction sub-model corresponds to the second interaction task, and the third prediction sub-model corresponds to the third interaction task.

[0097] Specifically, the server inputs the user description information under the first interaction task into the first prediction sub-model, and the first prediction sub-model performs interaction probability prediction based on the user description information to obtain the predicted interaction probability of the sample user under the first interaction task, and this predicted interaction probability is the first interaction probability.

[0098] Step 206: In the i-th prediction sub-model of the interaction probability prediction model, based on the user description information of the sample user in the i-th interaction task and the first task information of the (i - 1)-th interaction task determined by the (i - 1)-th prediction sub-model based on the historical interaction results in the (i - 1)-th interaction task, predict the second interaction probability corresponding to the i-th interaction task, where i is a positive integer greater than or equal to 2 and not greater than n.

[0099] Among them, for the i-th prediction sub-model in the interaction probability prediction model, its input includes the first task information of the previous task, that is, the first task information of the (i - 1)-th interaction task. This first task information is output by the prediction sub-model corresponding to the (i - 1)-th interaction task, that is, the (i - 1)-th prediction sub-model. In the (i - 1)-th prediction sub-model, this first task information can be determined based on the historical interaction results in the (i - 1)-th interaction task. Here, the historical interaction results are the training labels of the (i - 1)-th interaction task, which are used to represent whether the user has an interaction operation in the (i - 1)-th interaction task. Since the first task information is determined by the historical interaction results in the (i - 1)-th interaction task, the (i - 1)-th prediction sub-model can accurately transmit the task information to the i-th prediction sub-model.

[0100] Specifically, for each prediction sub-model starting from the second prediction sub-model in the interaction probability prediction model (hereinafter referred to as the i-th prediction sub-model), the server inputs the first task information output by the previous prediction sub-model corresponding to the i-th prediction sub-model, that is, the (i - 1)-th prediction sub-model, into the i-th prediction sub-model. Through the i-th prediction sub-model, predict the interaction probability of the sample user in the i-th interaction task to obtain the predicted interaction probability of the sample user in the i-th interaction task. This predicted interaction probability is the second interaction probability. That is to say, the second interaction probability predicted by each prediction sub-model starting from the second prediction sub-model is the result obtained by predicting the interaction probability using the training labels of the previous task.

[0101] Optionally, the specific method for the (i - 1)-th prediction sub-model to determine the first task information based on the historical interaction results in the (i - 1)-th interaction task can be: In the (i - 1)-th prediction sub-model, the server can map the historical interaction results in the (i - 1)-th interaction task through a neural network to obtain an embedding, and determine the first task information based on this embedding.

[0102] In a specific application, in the (i - 1)-th prediction sub-model, the server can also fuse the embedding obtained by mapping the historical interaction results under the (i - 1)-th interaction task with the implicit representation of the (i - 1)-th interaction task to obtain the first task information. Here, the embedding is obtained based on the historical interaction results and can be regarded as the explicit representation of the (i - 1)-th interaction task. That is to say, the first task information is obtained based on the explicit and implicit representations of the (i - 1)-th interaction task and can accurately represent the (i - 1)-th interaction task. The implicit representation of the (i - 1)-th interaction task is as follows: when i = 2, the (i - 1)-th interaction task is the first interaction task, and the implicit representation of the first interaction task is generated in the first prediction sub-model. In the first prediction sub-model, the server determines the input features of the prediction layer in the (i - 1)-th interaction task based on the user description information under the first interaction task, and after transforming the input features through a fully connected layer, the implicit representation of the first interaction task is obtained; when i > 2, the implicit representation of the (i - 1)-th interaction task is generated in the (i - 1)-th prediction sub-model. In the (i - 1)-th prediction sub-model, the server determines the input features of the prediction layer in the (i - 1)-th prediction sub-model based on the task information passed from the previous interaction task, and after transforming the input features through a fully connected layer, the implicit representation of the (i - 1)-th interaction task is obtained.

[0103] Furthermore, in the i-th prediction sub-model, the server can determine the input features of the prediction layer in the i-th prediction sub-model based on the task information passed from the (i - 1)-th interaction task, and after transforming the input features through a fully connected layer, the implicit representation of the i-th interaction task is obtained. The server can continue to pass this implicit representation downward. For example, when the i-th prediction sub-model is not the last prediction sub-model, this implicit representation continues to be passed to the (i + 1)-th prediction sub-model.

[0104] Step 208, in the i-th prediction sub-model, based on the user description information of the sample user under the i-th interaction task and the second task information of the (i - 1)-th interaction task determined by the (i - 1)-th prediction sub-model based on the sampling information, predict the first interaction probability corresponding to the i-th interaction task.

[0105] Among them, the sampling information is sampled by the (i - 1)-th prediction sub-model from the historical interaction results of the sample user under the (i - 1)-th interaction task and the first interaction probability corresponding to the (i - 1)-th interaction task. The historical interaction results under the (i - 1)-th interaction task are the training labels corresponding to the (i - 1)-th interaction task and are used to represent whether the sample user has an interaction operation under the (i - 1)-th interaction task. The first interaction probability corresponding to the (i - 1)-th interaction task is the result obtained by predicting the interaction probability using the sampling information in the (i - 1)-th prediction sub-model.

[0106] For the i-th prediction sub-model in the interaction probability prediction model, its input further includes the second task information of the previous task, that is, the second task information of the (i - 1)-th interaction task, which is output by the prediction sub-model corresponding to the (i - 1)-th interaction task, that is, the (i - 1)-th prediction sub-model. In the (i - 1)-th prediction sub-model, the second task information can be determined based on the sampling information under the (i - 1)-th interaction task.

[0107] Specifically, for each prediction sub-model starting from the second prediction sub-model in the interaction probability prediction model (hereinafter referred to as the i-th prediction sub-model), the server inputs the second task information output by the previous prediction sub-model corresponding to the i-th prediction sub-model, that is, the (i - 1)-th prediction sub-model, into the i-th prediction sub-model. The i-th prediction sub-model predicts the interaction probability of the sample user under the i-th interaction task, and obtains the predicted interaction probability of the sample user under the i-th interaction task. This predicted interaction probability is the second interaction probability. That is to say, the second interaction probability predicted by each prediction sub-model starting from the second prediction sub-model is the result obtained by predicting the interaction probability using the sampling information corresponding to the previous task.

[0108] In a specific application, sampling from the historical interaction results of the sample user under the (i - 1)-th interaction task and the first interaction probability corresponding to the (i - 1)-th interaction task is randomly performed with probabilities of P and P - 1. Among them, the sampling probability p is dynamically changed. In the initial stage of training, the network parameters of the (i - 1)-th prediction sub-model are not sufficiently learned, and the prediction value error is relatively large. Therefore, the p value is set to be relatively large, and the model is more inclined to pass the actual label information to the downstream task. As the number of training rounds increases, the p value gradually decreases, and the model is inclined to pass the predicted value as explicit upstream task information to the downstream task.

[0109] Optionally, the (i - 1)-th prediction sub-model determines the second task information based on the sampling information under the (i - 1)-th interaction task. Specifically, in the (i - 1)-th prediction sub-model, the server can map the sampling information under the (i - 1)-th interaction task through a neural network to obtain an embedding, and determine the second task information based on this embedding.

[0110] In a specific application, in the (i - 1)-th prediction sub-model, the server can also fuse the embedding obtained by mapping the sampling information under the (i - 1)-th interaction task with the implicit representation of the (i - 1)-th interaction task to obtain the first task information. Here, the embedding is obtained based on the sampling information and can also be regarded as the explicit representation of the (i - 1)-th interaction task. That is to say, the second task information is obtained based on the explicit representation and the implicit representation of the (i - 1)-th interaction task, and can accurately represent the (i - 1)-th interaction task. The way to obtain the implicit representation of the (i - 1)-th interaction task can refer to the above embodiments.

[0111] Step 210: Train an interaction probability prediction model based on the first interaction probability and the second interaction probability corresponding to the i-th interaction task.

[0112] Specifically, the server can determine the loss information of the i-th prediction sub-model based on the first interaction probability and the second interaction probability corresponding to the i-th interaction task. After obtaining the loss information of each prediction sub-model, the server can statistically process these loss information to obtain the final loss information of the interaction probability prediction model, and train the interaction probability prediction model based on this final loss information.

[0113] When the training process meets the training stop condition, a trained interaction probability prediction model is obtained. This trained interaction probability prediction model can be used to predict the interaction probability for the information to be pushed, so as to achieve accurate pushing of the information to be pushed. Here, the training stop condition can be any one of the loss reaching the minimum value, the training iteration times reaching the maximum number, the training duration reaching the preset duration, etc.

[0114] Optionally, for each prediction sub-model, during the process of calculating the loss information of the prediction sub-model, the training label can also be further considered. That is, for each prediction sub-model, the server can calculate the cross-entropy loss based on at least one of the first interaction probability or the second interaction probability output by the prediction sub-model and in combination with the training label, and then jointly determine the loss information of the prediction sub-model based on the cross-entropy loss and the loss determined based on the first interaction probability and the second interaction probability output by the prediction sub-model. <A

[0115] Optionally, in the process of determining the loss based on the first interaction probability and the second interaction probability output by the prediction sub-model, the second interaction probability can be used as the soft label of the first interaction probability to guide the model to use the second interaction probability as the "teacher signal" for model training through knowledge distillation. In some specific embodiments, the server can specifically calculate the loss based on the difference between the first interaction probability and the second interaction probability. Optionally, the KL divergence can be calculated based on the first interaction probability and the second interaction probability to obtain the loss, or the MES (Mean Square Error) loss can be calculated.

[0116] It can be understood that for the first prediction sub-model, since only the first interaction probability is predicted, the server can calculate the cross-entropy loss only based on the first interaction probability and the training label under the first interaction task, and use the cross-entropy loss as the loss information of the first prediction sub-model.

[0117] In the above model processing method, in the i-th prediction sub-model, based on the user description information of the sample user under the i-th interaction task and the second task information of the (i - 1)-th interaction task determined by the (i - 1)-th prediction sub-model based on the sampling information, the first interaction probability corresponding to the i-th interaction task is predicted. The sampling information is sampled by the (i - 1)-th prediction sub-model from the historical interaction results of the sample user under the (i - 1)-th interaction task and the first interaction probability corresponding to the (i - 1)-th interaction task. Therefore, the downstream task can obtain relatively accurate pre-task information throughout the training process. And because in the i-th prediction sub-model of the interaction probability prediction model, based on the user description information of the sample user under the i-th interaction task and the first task information of the (i - 1)-th interaction task determined by the (i - 1)-th prediction sub-model based on the historical interaction results under the (i - 1)-th interaction task, the second interaction probability corresponding to the i-th interaction task is predicted. Thus, the sequential dependence relationship between tasks can be utilized more fully, so that for the i-th task, the entire training process can utilize the label information of the (i - 1)-th task to obtain a more accurate prediction value. Based on the first interaction probability and the second interaction probability corresponding to the i-th interaction task, training the interaction probability prediction model can train a more accurate interaction probability prediction model.

[0118] The following embodiments will specifically introduce optional implementation manners of training the interaction probability prediction model based on the first interaction probability and the second interaction probability.

[0119] In some embodiments, the sample user is the user corresponding to the training sample in the training sample set, such as Figure 4 As shown, training the interaction probability prediction model based on the first interaction probability and the second interaction probability corresponding to the i-th interaction task includes:

[0120] Step 402: Determine the ranking loss corresponding to the i-th prediction sub-model based on the first interaction probability and the second interaction probability corresponding to the i-th interaction task.

[0121] Among them, the ranking loss is positively correlated with the difference between the first ranking information and the second ranking information corresponding to the training samples. The first ranking information is the ranking information of the training samples under the first interaction probability corresponding to the i-th interaction task, and the second ranking information is the ranking information of the training samples under the second interaction probability.

[0122] Specifically, the server can rank the training samples respectively based on the first interaction probability corresponding to the i-th interaction task to obtain the first ranking information, and rank the training samples based on the second interaction probability corresponding to the i-th interaction task to obtain the first ranking information. Furthermore, the ranking loss corresponding to the i-th prediction sub-model can be determined based on the difference between the first ranking information and the second ranking information.

[0123] Step 404: Determine the cross-entropy loss corresponding to the i-th prediction sub-model based on the first interaction probability corresponding to the i-th interaction task and the historical interaction results under the i-th interaction task.

[0124] Specifically, the historical interaction results under the i-th interaction task can be used as the training labels of the i-th prediction sub-model. The server can calculate the cross-entropy based on the first interaction probability corresponding to the i-th interaction task and the historical interaction results under the i-th interaction task, so as to obtain the cross-entropy loss corresponding to the i-th prediction sub-model.

[0125] Step 406: Train the interaction probability prediction model based on the ranking loss and the cross-entropy loss.

[0126] Specifically, the server can determine the loss information corresponding to the i-th prediction sub-model based on the ranking loss and the cross-entropy loss corresponding to the i-th prediction sub-model. Finally, based on the loss information corresponding to each prediction sub-model respectively, determine the final loss of the interaction probability prediction model, and then train the interaction probability prediction model based on the final loss.

[0127] In the above embodiments, based on the first interaction probability and the second interaction probability corresponding to the i-th interaction task, determine the ranking loss corresponding to the i-th prediction sub-model, based on the first interaction probability corresponding to the i-th interaction task and the historical interaction results under the i-th interaction task, determine the cross-entropy loss corresponding to the i-th prediction sub-model, and train the interaction probability prediction model based on the ranking loss and the cross-entropy loss. It can make full use of the dependency relationship between tasks to generate the second interaction probability as teacher knowledge and guide model training in the form of soft labels, which can effectively alleviate the problems of sample imbalance and scarcity of positive samples in downstream tasks, and improve the prediction accuracy and ranking ability of the model.

[0128] In some embodiments, determining the ranking loss corresponding to the i-th prediction sub-model based on the first interaction probability and the second interaction probability corresponding to the i-th interaction task includes: determining a first ranking score corresponding to the training sample based on the first interaction probability corresponding to the i-th interaction task, where the first ranking score is positively correlated with the probability that the training sample ranks among the top in the training sample set; determining a second ranking score corresponding to the training sample based on the second interaction probability, where the first ranking score is positively correlated with the probability that the training sample ranks among the top in the training sample set; and determining the first ranking loss corresponding to the i-th prediction sub-model based on the difference between the first ranking score and the second ranking score.

[0129] Among them, the ranking score is positively correlated with the probability that the training sample ranks among the top in the training sample set. The probability that the training sample ranks among the top in the training sample set is the probability that the training sample is most likely to obtain positive feedback under the i-th interaction task. That is to say, the larger the ranking score, the greater the probability that the training sample is most likely to obtain positive feedback in the training sample set. In other words, after pushing the information to be pushed to the terminals corresponding to the sample users of each training sample in the training sample set, the sample user corresponding to the training sample is more likely to interact under the i-th interaction task.

[0130] Specifically, for each training sample in the training sample set, the server can respectively determine the first ranking score and the second ranking score corresponding to the training sample based on the first interaction probability and the second interaction probability corresponding to the i-th interaction task predicted by using the training sample through the interaction probability prediction model. Then, based on the difference between the first ranking score and the second ranking score corresponding to the training sample, determine the first ranking loss corresponding to the training sample. Then, statistically analyze the first ranking losses of all training samples to obtain the first ranking loss corresponding to the i-th prediction sub-model. Here, the statistics can be any one of summation, averaging, median, etc. The first ranking loss is positively correlated with the difference between the first ranking score and the second ranking score.

[0131] Illustrate by way of example, refer to Figure 5, assuming that the training sample set contains n training samples, the first ranking score corresponding to training sample 1 can be determined based on the first interaction probability corresponding to training sample 1, the second ranking score corresponding to training sample 1 can be determined based on the second interaction probability corresponding to training sample 1, and then the first ranking loss corresponding to training sample 1 can be determined based on the difference between the first ranking score and the second ranking score corresponding to training sample 1. The first ranking score corresponding to training sample 2 can be determined based on the first interaction probability corresponding to training sample 2, the second ranking score corresponding to training sample 2 can be determined based on the second interaction probability corresponding to training sample 2, and then the first ranking loss corresponding to training sample 2 can be determined based on the difference between the first ranking score and the second ranking score corresponding to training sample 2, ……, the first ranking score corresponding to training sample n can be determined based on the first interaction probability corresponding to training sample n, the second ranking score corresponding to training sample n can be determined based on the second interaction probability corresponding to training sample n, and then the first ranking loss corresponding to training sample n can be determined based on the difference between the first ranking score and the second ranking score corresponding to training sample n. Finally, the first ranking losses corresponding to all training samples are statistically calculated to obtain the first ranking loss corresponding to the i-th prediction sub-model.

[0132] In the above embodiment, the server determines the first ranking score corresponding to the training sample based on the first interaction probability corresponding to the i-th interaction task, determines the second ranking score corresponding to the training sample based on the second interaction probability, and determines the first ranking loss corresponding to the i-th prediction sub-model based on the difference between the first ranking score and the second ranking score, which can enhance the ranking ability of the interaction probability prediction model.

[0133] In some embodiments, determining the ranking loss corresponding to the i-th prediction sub-model based on the first interaction probability and the second interaction probability corresponding to the i-th interaction task includes: forming training sample pairs with the training sample and other training samples in the training sample set, where m is a positive integer; determining the second ranking loss corresponding to the training sample based on the first probability magnitude relationship and the second probability magnitude relationship of the training sample pair, where the first probability magnitude relationship is the magnitude relationship of the training sample pair under the first interaction probability, and the first probability magnitude relationship is the magnitude relationship of the training sample pair under the second interaction probability; determining the ranking loss corresponding to the i-th prediction sub-model based on the second ranking loss corresponding to the training sample pair.

[0134] Among them, the size relationship of training sample pairs under the first interaction probability refers to the size relationship between two first interaction probabilities obtained by the interaction probability prediction model using the two training samples in the training sample pair for the same interaction task. The size relationship of training sample pairs under the second interaction probability refers to the size relationship between two second interaction probabilities obtained by the interaction probability prediction model using the two training samples in the training sample pair for the same interaction task.

[0135] Specifically, the server can pair up the training samples in the training sample set in pairs to obtain multiple training sample pairs. For each training sample pair, the server can determine the size relationship between the two training samples in the training sample pair under the first interaction probability, and determine the size relationship between the two training samples in the training sample pair under the second interaction probability. Then, based on the first probability size relationship and the second probability size relationship of the training sample pair, the server determines the second ranking loss corresponding to the training sample pair. Finally, the server statistically analyzes the second ranking losses corresponding to each training sample pair to obtain the second ranking loss corresponding to the i-th prediction sub-model.

[0136] For example, referring to Figure 6 , taking the training sample set containing three training samples x, y, and z as an example, the three training samples can form a total of three training sample pairs, namely training sample pair 1, training sample pair 2, and training sample pair 3. For training sample pair 1, compare the size relationship between the first interaction probability of training sample x and the first interaction probability of training sample y to obtain the first probability size relationship of training sample pair 1, and compare the size relationship between the second interaction probability of training sample x and the second interaction probability of training sample y to obtain the second probability size relationship of training sample pair 1; for training sample pair 2, compare the size relationship between the first interaction probability of training sample x and the first interaction probability of training sample z to obtain the first probability size relationship of training sample pair 2, and compare the size relationship between the second interaction probability of training sample x and the second interaction probability of training sample z to obtain the second probability size relationship of training sample pair 2; for training sample pair 3, compare the size relationship between the first interaction probability of training sample y and the first interaction probability of training sample z to obtain the first probability size relationship of training sample pair 3, and compare the size relationship between the second interaction probability of training sample y and the second interaction probability of training sample z to obtain the second probability size relationship of training sample pair 3. Finally, based on the first probability size relationship and the second probability size relationship of each training sample pair, determine the second ranking loss corresponding to each training sample pair, and further determine the second ranking loss corresponding to the i-th prediction sub-model based on each second ranking loss.

[0137] In the above embodiments, by forming training sample pairs from the training samples and determining the second ranking loss based on the first probability magnitude relationship and the second probability magnitude relationship of the training sample pairs, the ranking ability of the interaction probability prediction model can be enhanced.

[0138] In some embodiments, determining the second ranking loss corresponding to a training sample based on the first probability magnitude relationship and the second probability magnitude relationship of the training sample pairs includes: when the first probability magnitude relationship and the second probability magnitude relationship are inconsistent, calculating the first absolute difference of the two training samples included in the training sample pair at the first interaction probability and the second absolute difference at the first interaction probability; determining the second ranking loss corresponding to the training sample pair based on the absolute difference between the first absolute difference and the second absolute difference, and the second ranking loss is positively correlated with the absolute difference; obtaining the second ranking loss corresponding to the i-th prediction sub-model based on the second ranking loss corresponding to the training sample pair.

[0139] Specifically, for each training sample pair, when the first probability magnitude relationship and the second probability magnitude relationship are inconsistent, for example, assuming that the training sample pair includes training sample i and training sample j in the training sample set, if the first interaction probability of training sample i is greater than the first interaction probability of training sample j and the second interaction probability of training sample i is less than the second interaction probability of training sample j, it means that the first probability magnitude relationship and the second probability magnitude relationship of this training sample pair are inconsistent. The server can calculate the first absolute difference of the first interaction probabilities of the two training samples included in this training sample pair, and calculate the second absolute difference of the second interaction probabilities of the two training samples included in this training sample pair, and further calculate the absolute difference between the first absolute difference and the second absolute difference, and determine the second ranking loss of this training sample pair based on the calculated absolute difference. The server can further count the second ranking losses of each training sample pair to obtain the second ranking loss corresponding to the i-th prediction sub-model.

[0140] In the above embodiments, when the first probability magnitude relationship and the second probability magnitude relationship are inconsistent, calculating the first absolute difference of the two training samples included in the training sample pair at the first interaction probability and the second absolute difference at the second interaction probability, and determining the second ranking loss corresponding to the training sample pair based on the absolute difference between the first absolute difference and the second absolute difference can make the ranking at the first interaction probability and the ranking at the second interaction probability as consistent as possible during the training process of the model, thereby improving the ranking ability of the model.

[0141] In some embodiments, determining the second ranking loss corresponding to a training sample based on the first probability magnitude relationship and the second probability magnitude relationship of a pair of training samples includes: when the first probability magnitude relationship and the second probability magnitude relationship are consistent, determining that the second ranking loss corresponding to the pair of training samples is a preset value, and the preset value is less than the second ranking loss corresponding to the pair of training samples when the first probability magnitude relationship and the second probability magnitude relationship are inconsistent.

[0142] Specifically, for each pair of training samples, when the first probability magnitude relationship and the second probability magnitude relationship are inconsistent. For example, assume that the pair of training samples includes training sample i and training sample j in the training sample set. If the first interaction probability of training sample i is greater than the first interaction probability of training sample j and the second interaction probability of training sample i is greater than the second interaction probability of training sample j, it indicates that the first probability magnitude relationship and the second probability magnitude relationship of this pair of training samples are inconsistent. Then the server can determine that the second ranking loss corresponding to this pair of training samples is a preset value, and this value is less than the second ranking loss corresponding to the pair of training samples when the first probability magnitude relationship and the second probability magnitude relationship are inconsistent. In a specific application, the preset value can be set to 0.

[0143] In some embodiments, predicting the first interaction probability corresponding to the i-th interaction task based on the user description information of the sample user in the i-th interaction task and the second task information of the (i - 1)-th interaction task determined by the (i - 1)-th prediction submodel based on sampling information includes: fusing the user description information of the sample user in the i-th interaction task and the second task information of the (i - 1)-th interaction task determined by the (i - 1)-th prediction submodel based on sampling information to obtain a first fused feature; predicting the interaction probability based on the first fused feature to obtain the first interaction probability corresponding to the i-th interaction task.

[0144] Specifically, in the i-th prediction submodel, the server can fuse the user description information of the sample user in the i-th interaction task and the second task information of the (i - 1)-th interaction task determined by the (i - 1)-th prediction submodel based on sampling information to obtain a first fused feature. The server can use this first fused feature as the input feature of the prediction layer in the i-th prediction submodel and input it into the prediction layer, and through the prediction layer, predict the interaction probability to obtain the first interaction probability corresponding to the i-th interaction task.

[0145] In the above embodiments, by fusing the user description information of the sample user under the i-th interaction task and the second task information of the (i - 1)-th interaction task determined by the (i - 1)-th prediction sub-model based on the sampling information, a first fusion feature is obtained. The obtained first fusion feature contains both user representation information and second task information, which can better predict the interaction probability, thereby obtaining a more accurate prediction result and further improving the accuracy of the interaction probability prediction model.

[0146] In some embodiments, fusing the user description information of the sample user under the i-th interaction task and the second task information of the (i - 1)-th interaction task determined by the (i - 1)-th prediction sub-model based on the sampling information to obtain a first fusion feature includes: calculating a first self-attention weight for the user description information of the sample user under the i-th interaction task, and calculating a second self-attention weight for the second task information of the (i - 1)-th interaction task determined by the (i - 1)-th prediction sub-model based on the sampling information; performing weighted processing on the user description information of the sample user under the i-th interaction task based on the first self-attention weight to obtain a first attention feature; performing weighted processing on the second task information of the (i - 1)-th interaction task based on the first self-attention weight to obtain a second attention feature; and fusing the first attention feature, the second attention feature, and the user description information of the sample user under the i-th interaction task to obtain a first fusion feature.

[0147] Specifically, in the i-th prediction sub-model, the server can calculate self-attention weights for the user description information of the sample user under the i-th interaction task and the second task information of the (i - 1)-th interaction task respectively, obtaining a first self-attention weight corresponding to the user description information and a second self-attention weight corresponding to the second task information. Furthermore, weighted processing can be performed on the user description information of the sample user under the i-th interaction task based on the first self-attention weight to obtain a first attention feature, and weighted processing can be performed on the second task information of the (i - 1)-th interaction task based on the first self-attention weight to obtain a second attention feature. After adding the first attention feature and the second attention feature, they are added to the user description information of the sample user under the i-th interaction task to obtain a first fusion feature.

[0148] In the above embodiments, due to the self-attention processing of the user description information and the second task information, better feature fusion can be performed to obtain a more accurate first fusion feature.

[0149] In some embodiments, based on the user description information of a sample user under the i-th interaction task and the first task information of the (i - 1)-th interaction task determined by the (i - 1)-th prediction sub-model based on the historical interaction results under the (i - 1)-th interaction task, predicting the second interaction probability corresponding to the i-th interaction task includes: fusing the user description information of the sample user under the i-th interaction task and the first task information of the (i - 1)-th interaction task determined by the (i - 1)-th prediction sub-model based on the historical interaction results under the (i - 1)-th interaction task to obtain a second fused feature; performing an interaction probability prediction based on the second fused feature to obtain the second interaction probability corresponding to the i-th interaction task.

[0150] Specifically, in the i-th prediction sub-model, the server can fuse the user description information of the sample user under the i-th interaction task and the first task information of the (i - 1)-th interaction task determined by the (i - 1)-th prediction sub-model based on the historical interaction results under the (i - 1)-th interaction task to obtain a second fused feature. The server can use this second fused feature as the input feature of the prediction layer in the i-th prediction sub-model and input it into the prediction layer, and perform an interaction probability prediction through the prediction layer to obtain the second interaction probability corresponding to the i-th interaction task.

[0151] In the above embodiments, by fusing the user description information of the sample user under the i-th interaction task and the first task information of the (i - 1)-th interaction task determined by the (i - 1)-th prediction sub-model based on the historical interaction results under the (i - 1)-th interaction task to obtain a second fused feature, the obtained second fused feature contains both user representation information and first task information, which can better perform the interaction probability prediction, thereby obtaining a more accurate prediction result and further improving the accuracy of the interaction probability prediction model.

[0152] In some embodiments, fusing the user description information of the sample user under the i-th interaction task and the first task information of the (i - 1)-th interaction task determined by the (i - 1)-th prediction sub-model based on the historical interaction results under the (i - 1)-th interaction task to obtain a second fused feature includes: calculating a third self-attention weight for the user description information of the sample user under the i-th interaction task and calculating a fourth self-attention weight for the first task information of the (i - 1)-th interaction task determined by the (i - 1)-th prediction sub-model based on the historical interaction results under the (i - 1)-th interaction task; performing a weighted processing on the user description information of the sample user under the i-th interaction task based on the third self-attention weight to obtain a third attention feature; performing a weighted processing on the first task information of the (i - 1)-th interaction task based on the fourth self-attention weight to obtain a fourth attention feature; fusing the third attention feature, the fourth attention feature, and the user description information of the sample user under the i-th interaction task to obtain a second fused feature.

[0153] Specifically, in the i-th prediction sub-model, the server can calculate the self-attention weights for the user description information of the sample user in the i-th interaction task and the first task information of the (i - 1)-th interaction task respectively, obtaining the third self-attention weight corresponding to the user description information and the fourth self-attention weight corresponding to the first task information. Furthermore, the server can perform weighted processing on the user description information of the sample user in the i-th interaction task based on the third self-attention weight to obtain the third attention feature, perform weighted processing on the first task information of the (i - 1)-th interaction task based on the fourth self-attention weight to obtain the fourth attention feature. After adding the third attention feature and the fourth attention feature, the server adds the result to the user description information of the sample user in the i-th interaction task to obtain the second fusion feature.

[0154] In the above embodiments, since self-attention processing is performed on the user description information and the first task information, better feature fusion can be achieved, and a more accurate first fusion feature can be obtained.

[0155] In some embodiments, as Figure 7 shown, an information push method is provided. This method is executed by a computer device, which can be Figure 1 the server 104 or the terminal 102 in Figure 1 . In the embodiments of the present application, taking the method being applied to

[0156] the server 104 in

[0157] as an example for illustration, the method includes the following steps:

[0158] Optionally, the interaction probability prediction model may include a feature representation layer. For each candidate user in the candidate user set, the server can input the description data of the candidate user into the feature representation layer to obtain the user description information of the candidate user in each interaction task among the n interaction tasks. Optionally, the user description information of the candidate user in each interaction task among the n interaction tasks may be the feature vector representation in each interaction task.

[0159] Step 702, determining a candidate user set for the information to be pushed.

[0160] Step 704, for each candidate user in the candidate user set, determining the user description information of the candidate user in each interaction task among the n interaction tasks, where n is a positive integer.

[0161] Specifically, for each candidate user in the candidate user set, the server can determine the user description information of the candidate user under each interaction task among the n interaction tasks, and then input the user description information under each interaction task into the corresponding prediction sub-model of each interaction task in the interaction probability prediction model, and determine the interaction probability output by the last prediction sub-model as the interaction probability of the candidate user for the information to be pushed.

[0162] Step 708: Based on the interaction probabilities of each candidate user, determine the target push users of the information to be pushed from the candidate user set.

[0163] Step 710: Push the information to be pushed to the terminals corresponding to the target push users.

[0164] Specifically, after obtaining the interaction probabilities of each candidate user, the server can determine the candidate users with higher interaction probabilities as the target push users of the information to be pushed, and the server further pushes the information to be pushed to the terminals corresponding to each target push user.

[0165] In some embodiments, an interaction probability threshold can be preset. After the server obtains the interaction probabilities of each candidate user, it filters out the candidate users with interaction probabilities greater than the interaction probability threshold from the candidate user set as the target push users of the information to be pushed.

[0166] For the above information push method, since the interaction probability prediction model is trained using the above model processing method, more accurate interaction probabilities can be predicted using this interaction probability prediction model. Thus, based on the predicted interaction probabilities, the target users can be accurately determined, and when pushing information to the terminals of these target users, the possibility of obtaining positive feedback is greater, thereby improving the accuracy of information push.

[0167] In some embodiments, determining the target push users of the information to be pushed from the candidate user set based on the interaction probabilities of each candidate user includes: determining the sorting information of each candidate user for the information to be pushed based on the interaction probabilities of each candidate user; and determining the target push users of the information to be pushed from the candidate user set based on the sorting information of each candidate user.

[0168] Among them, the sorting information of the candidate user for the information to be pushed is used to represent the ranking of the interaction probability of the candidate user. For example, the interaction probabilities of each candidate user can be sorted from largest to smallest, and the rankings of the interaction probabilities of each candidate user can be determined as the sorting information of each candidate user respectively. For example, assuming that the interaction probability of a certain candidate user is ranked third, then the sorting information of this candidate user is "third".

[0169] Specifically, the server may determine the sorting information of each candidate user for the information to be pushed based on the respective interaction probabilities of the candidate users, and determine, based on the respective sorting information of the candidate users, a candidate user with a relatively high interaction probability from the candidate user set as the target push user.

[0170] Optionally, when the sorting information is a ranking representing the interaction probabilities sorted from high to low, the target push user is a candidate user with an interaction probability in the TopN.

[0171] In the above embodiments, since the target push user of the information to be pushed can be determined from the candidate user set based on the respective sorting information of the candidate users, the target push user can be determined more accurately, thereby further improving the accuracy of information push.

[0172] In some specific embodiments, the present application further provides an application scenario in which the model processing method and information push method of the present application are used to expand the advertising audience in the credit card placement business, so as to improve the advertising placement effect of multiple channels, find the target users of the bank, and improve the target conversion efficiency. In this application scenario, the placement mainly includes the following steps:

[0173] 1. Exposure: Refer to Figure 8A , after the advertisement is placed to the user, the user sees the relevant advertisement.

[0174] 2. Click: Refer to Figure 8B , if the user is interested in the advertisement, the user will click on this advertisement and enter the application form page.

[0175] 3. Submission: Refer to Figure 8C , after entering the application form page, the user completes the filling of the credit card application, which is the submission process.

[0176] 4. Card approval: Card approval is also called credit granting, which means that the user has good credit, passes the application and is granted a certain credit card limit by the bank. This goal is usually the most important goal for marketing credit card issuance marketing activities. The schematic diagram of the card approval interface can be referred to Figure 8D .

[0177] It can be seen from Figures 8A to 8D that in this application scenario, multiple interaction tasks need to be passed through to complete the user conversion process. This kind of multi-task conversion has a long path dependence, the positive samples are gradually sparse, and the class imbalance of the downstream tasks becomes more and more serious. The positive sample data of the core task is usually very scarce. In related technologies, to solve this problem of sample sparsity, multi-task learning is usually used, which can associate the data of other tasks to achieve the effect of expanding the data set, thereby alleviating the problem of sample sparsity. Refer to Figure 9, which is a schematic structural diagram of a multi-task model implemented through a shared Bottom model in the related art. In this model, multiple tasks learn the commonalities between different tasks through a shared layers, and learn the differences between different tasks through task-specific layers corresponding to each task. Although this model can improve the quality and efficiency of modeling to a certain extent, since this model only shares multi-task information through the underlying shared layers, it is difficult to effectively utilize the information between different tasks, and there is still a problem of low accuracy when the obtained model is used for interactive probability prediction under multiple tasks.

[0178] Based on this, in this application scenario, PIKD (Prior Information Knowledge Distillation) is used as an interactive probability prediction model for interactive probability prediction. This model can combine multi-task sequence relationship information and knowledge distillation to expand the advertising audience. The specific structure of PIKD can be referred to Figure 10 , and it includes the following parts:

[0179] 1. Shared representation layer: The input user description data is transformed into a dense vector representation through an embedding operation, and the embeddings of all features are concatenated to form a shared representation.

[0180] 2. Task-specific tower: For the t-th task, the shared representation is transformed through a neural network to obtain user description information for task t .

[0181] 3. PIT (Prior Information Transfer) module: For all the t-th tasks except the first task, the PIT module first generates an embedding containing explicit prior task information according to the sampling probability based on the actual label or predicted value of task t, and then combines the embedding with the implicit information representation of task t passed over to obtain . In addition, the PIT module will completely transfer the actual label embedding of task t for the entire batch of samples during the training phase, that is, Figure 10 in . The following is a specific introduction to this process:

[0182] Considering the following characteristics of the application scenario of the multi-task estimation problem in the long conversion link: for any sample, only when the label of task t-1 is positive, this sample is likely to be a positive sample for task t, that is is a necessary but not sufficient condition; as the task path deepens, the number of positive samples gradually thins out, and the positive-negative sample ratio becomes more imbalanced; important tasks are often located at the end of the conversion path.

[0183] Based on the above characteristics, the PIT module is proposed. The PIT module can use the curriculum learning paradigm for information transfer. Specifically, the PIT module randomly selects the true label of task t-1 with probabilities p and 1-p

[0184] or the predicted value . Among them, the sampling probability p is dynamically changing. At the initial stage of training, the network parameters of task t-1 are not well learned, and the error of the predicted value is large. Therefore, the p value is set to be large, that is, it is more inclined to transfer the actual label information to the downstream task. As the number of training rounds increases, the p value gradually decreases, and the PIM tends to transfer the predicted value as explicit upstream task information to task t. Through this learning method from easy to difficult, each downstream task can obtain relatively accurate upstream task information throughout the training process, thereby reducing the difficulty of model learning.

[0185] On the other hand, in order to make more full use of the sequential dependency between tasks, so that for task t, the label information of task t-1 can be utilized throughout the training process to obtain a more accurate predicted value as the "teacher signal" to guide the model training, the PIT module will also transfer the label information of task t-1. Therefore, the information transfer in PIT includes two parts:

[0186] 1) Randomly sample from with probabilities p and 1-p to obtain the upstream task information , and after non-linear transformation, obtain a d-dimensional vector , where is the activation function, is the trainable parameter. Further fuse the implicit representation transferred from the upstream task and add them to obtain .

[0187] 2) For all samples in the batch (training batch), generate according to

[0188] , further fuse the implicit representation transferred from the upstream task and add them to obtain .

[0189] For example Figure 11 , take one that contains " Taking the case of "two-step conversion" as an example, the conversion behavior of the user will only occur after the user clicks, that is, the click behavior is a necessary but not sufficient condition for conversion. During the training process of PIKD, the user representation is input into the click prediction network and the conversion prediction network respectively. At the same time, the conversion prediction network will receive information from the click prediction network through the PIT module:

[0190] 1) : For all samples in the batch, the actual label indicating whether a click occurs is mapped through a neural network to obtain an embedding. For example, if a user clicks, then is multiplied by a 1*d-dimensional vector W to obtain a d-dimensional embedding containing label information. Finally, this embedding and the implicit representation learned by the click network are added together to obtain .

[0191] 2) : The d-dimensional embedding that mixes the actual label and click prediction value information in a batch and are added together. For example, in the early stage of training, assuming the label sampling probability is p = 0.2, then 20% of the samples in the batch use the actual label , and 80% of the samples use the click task prediction value

[0192] (for example, the prediction value 0.7 represents a click probability of 70%). Multiply (or ) by W to obtain a d-dimensional embedding, and then add it to to obtain the of this batch of samples. Note that in the later stage of training, the label sampling probability gradually decreases, only contains click prediction value information.

[0193] In the conversion task, the predicted user conversion rate prediction value obtained by using , and the predicted user conversion rate prediction value obtained by using is denoted as . Without loss of generality, if the model knows the information about whether the user clicks, then the prediction of the conversion task will be more accurate. Therefore, The accuracy ofis not lower than , and can be used as a sample sorting basis and soft label to guide the model training during the training process.

[0194] 4. Attention layer: For task t, the passed by the upstream task PIT module and

[0195] , respectively, and Integrate using the self-attention mechanism to obtain (i.e., the first fusion feature in the above text) and (i.e., the second fusion feature in the above text), and respectively obtain (i.e., the first interaction probability in the above text) and (i.e., the second interaction probability in the above text) through the output layer. At the same time will be transformed through the fully connected layer to obtain , which is passed as the implicit representation of task t to the (t + 1)-th task. It will be used as more accurate prediction information to supplement the ranking relationship between samples and as a soft label to assist model training, thereby alleviating the problem of sample imbalance in downstream tasks.

[0196] It can be understood that Figure 10 In the model structure shown, the prediction sub-model in the above text can be a model structure composed of an attention layer, an FC layer, a PIT module, and an output layer (not shown in the figure). The output by the PIT module is the first task information in the above text, and the output by the PIT module is the second task information in the above text, and the

[0197] The following combines Figure 10 the model structure shown to introduce the model training process indicated by the model processing method of the present application.

[0198] In the related art, for the multi-task sequence dependence modeling method, the information of the previous task only assists the learning of the downstream task in the form of a representation. Although it can alleviate the training difficulty to a certain extent, it still does not solve the problems of label imbalance and extremely sparse positive samples in the downstream task. The PIKD model provided by the embodiments of the present application, on the one hand, introduces curriculum learning, and on the other hand, introduces a sample pairing ranking loss and a list-wise ranking loss function within a batch, and makes full use of the dependence relationship between tasks to generate oracle prediction as teacher knowledge to guide model training in the form of soft labels. For the unique tower output

[0199] of task t, separately and 、 are fused in a self-attention manner combined with residual connections, such as combining

[0200] Generating the input of the prediction layer (i.e., the first fusion feature in the above text) can be referred to the following formula:

[0201] (1)

[0202] Where are the trainable parameters that map the d-dimensional vector a to another d-dimensional vector through a linear transformation.

[0203] represents the weight (i.e., the self-attention weight in the above text), and its calculation method can be referred to the following formula:

[0204] (2)

[0205] (3)

[0206] Among them, Q, K, and V are similar and are linear transformation parameters. After passing through the output layer network, the predicted value of task t can be obtained . Similarly, replace the part involving in formulas (1)-(3) with

[0207] , and through the same operation, generate the oracle prediction of task t, that is .

[0208] During the training process, the model inputs a batch of samples in each round of training. For task t, the samples can be sorted according to to obtain the oracle ranking. Due to the logical dependency relationships between tasks, it can be assumed that has an accuracy not lower than . Based on , for each sample, the list ranking score can be calculated according to the following formula:

[0209] (4)

[0210] Its value reflects the probability that sample i ranks at the forefront of this batch of samples, that is, the probability that it is most likely to obtain positive feedback in task t among this batch of samples.

[0211] Similarly, based on can generate the ranking according to . By applying the Kullback-Leibler divergence calculation to , the loss is obtained, which enhances the model's ranking ability at the list-wise level. Specifically, it can be referred to the following formula:

[0212] (5)

[0213] In addition, k samples in a batch are paired pairwise. For a sample pair <i, j>, the size relationship between and the sorting should be as consistent as possible. Therefore, the pairwise sorting loss is added. Specifically, it can be referred to the following formula:

[0214] (6)

[0215] By (i.e., the first sorting loss in the above text) and (i.e., the second sorting loss in the above text), soft labels containing sorting information are added, alleviating the problems of sample imbalance and scarcity of positive samples in downstream tasks, and improving the prediction accuracy and sorting ability of the model.

[0216] Finally, for task t (t > 0), the loss function of its corresponding model sub-model can be referred to the following formula:

[0217] (7)

[0218] where is the commonly used cross-entropy loss function.

[0219] At the beginning of training, p is usually set to 1 and gradually decreases to 0 as the training iterations progress. Therefore and are basically the same in the early stage of model training. The model is mainly trained based on the cross-entropy loss at this stage. In the later stage of training, and begin to play a role, enabling the model to learn the sorting information of samples from the teacher signal.

[0220] In the scenario of credit card advertising audience expansion, the commonly used AUC (Area Under ROC) and recall@k% metrics in this type of scenario are used for offline evaluation. Recall@k% represents the proportion of positive samples among all positive samples in the top k% users after the model sorts the predicted scores of the target task from high to low. The online evaluation uses the ad delivery card approval rate (AR, Approval Rate) metric, which is usually presented in the form of a percentage. Theoretically, the more accurate the population we expand, the higher the card approval rate. AR can be calculated using the following formula:

[0221] AR = Number of users with approved cards / Number of users with ad delivery × 100%

[0222] The online test results are shown in Table 1. The data in Table 1 reflect the relative improvement ratios of various indicators of the model compared to the standard model:

[0223] Table 1

[0224]

[0225] The baseline methods involved in the table are as follows:

[0226] 1. Standard model: Transfer the output probability of the previous task and multiply it by the estimated conditional probability to model sequence dependence.

[0227] 2. Model 1: The MMOE (Multi-gate Mixture-of-Experts) model. This model introduces a gating network for each task to learn the influence of multiple experts on the weights of the current task, thus largely alleviating the problem of reduced accuracy caused by low correlation of multi-objective tasks. By adopting multiple groups of expert networks, this model can adaptively learn knowledge at different levels. In addition, each task learns a gating network to determine the weights of each expert network.

[0228] 3. Model 2: The ATIM (Adaptive Information Transfer Multi-task) model. This model uses the attention mechanism to adaptively transfer implicit representations from previous tasks.

[0229] 4. Model 3: The MNCM (Multi-level Network Cascades Model). By sharing features among multiple related tasks, tasks can be associated at different layers, thus realizing cross-task knowledge transfer.

[0230] It can be seen that compared with other models, the PIKD model of this application has greatly improved in various indicators during the offline test and online test processes, that is, the PIKD of this application has higher accuracy.

[0231] It should be understood that although the steps in the flowcharts involved in the above embodiments are sequentially shown according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least some of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least some of the steps or stages in other steps or other steps.

[0232] Based on the same inventive concept, the embodiments of the present application also provide a model processing device for implementing the model processing method involved above and an information pushing device for implementing the information pushing method involved above. The implementation solutions provided by the device for solving problems are similar to the implementation solutions recorded in the above method. Therefore, the specific limitations in one or more of the following model processing device embodiments can refer to the limitations on the model processing method in the above text, and will not be repeated here.

[0233] In some embodiments, as Figure 12 shown, a model processing device 1200 is provided, including:

[0234] A user description information acquisition module 1202, configured to acquire the user description information of a sample user for each interaction task in n interaction tasks, where n is a positive integer;

[0235] A first prediction module 1204, configured to determine a first interaction probability corresponding to the first interaction task based on the user description information of the sample user in the first interaction task in the first prediction sub-model of the interaction probability prediction model;

[0236] A second prediction module 1206, configured to predict a second interaction probability corresponding to the i-th interaction task based on the user description information of the sample user in the i-th interaction task and the first task information of the (i - 1)-th interaction task determined by the (i - 1)-th prediction sub-model based on the historical interaction result in the (i - 1)-th interaction task, where i is a positive integer greater than or equal to 2 and not greater than n;

[0237] The second prediction module is further configured to, in the i-th prediction sub-model, based on the user description information of the sample user in the i-th interaction task and the second task information of the (i - 1)-th interaction task determined by the (i - 1)-th prediction sub-model based on the sampling information, predict the first interaction probability corresponding to the i-th interaction task; the sampling information is sampled by the (i - 1)-th prediction sub-model from the historical interaction results of the sample user in the (i - 1)-th interaction task and the first interaction probability corresponding to the (i - 1)-th interaction task.

[0238] The model training module 1208 is configured to train the interaction probability prediction model based on the first interaction probability and the second interaction probability corresponding to the i-th interaction task.

[0239] In the above model processing device, since in the i-th prediction sub-model, based on the user description information of the sample user in the i-th interaction task and the second task information of the (i - 1)-th interaction task determined by the (i - 1)-th prediction sub-model based on the sampling information, the first interaction probability corresponding to the i-th interaction task is predicted, and the sampling information is sampled by the (i - 1)-th prediction sub-model from the historical interaction results of the sample user in the (i - 1)-th interaction task and the first interaction probability corresponding to the (i - 1)-th interaction task, the downstream task can obtain relatively accurate pre-task information during the entire training process. And since in the i-th prediction sub-model of the interaction probability prediction model, based on the user description information of the sample user in the i-th interaction task and the first task information of the (i - 1)-th interaction task determined by the (i - 1)-th prediction sub-model based on the historical interaction results in the (i - 1)-th interaction task, the second interaction probability corresponding to the i-th interaction task is predicted, thus the sequential dependence relationship between tasks can be utilized more fully, so that for the i-th task, the entire training process can utilize the label information of the (i - 1)-th task to obtain a more accurate prediction value. Based on the first interaction probability and the second interaction probability corresponding to the i-th interaction task, training the interaction probability prediction model can train a more accurate interaction probability prediction model.

[0240] In some embodiments, the sample user is the user corresponding to the training sample in the training sample set. The model training module is further configured to determine the ranking loss corresponding to the i-th prediction sub-model based on the first interaction probability and the second interaction probability corresponding to the i-th interaction task, and the ranking loss is positively correlated with the difference between the first ranking information and the second ranking information corresponding to the training sample. The first ranking information is the ranking information of the training sample under the first interaction probability corresponding to the i-th interaction task, and the second ranking information is the ranking information of the training sample under the second interaction probability; determine the cross-entropy loss corresponding to the i-th prediction sub-model based on the first interaction probability corresponding to the i-th interaction task and the historical interaction results in the i-th interaction task; train the interaction probability prediction model based on the ranking loss and the cross-entropy loss.

[0241] In some embodiments, the model training module is further configured to determine a first ranking score corresponding to a training sample based on the first interaction probability corresponding to the i-th interaction task, where the first ranking score is positively correlated with the probability that the training sample ranks among the top in the training sample set; determine a second ranking score corresponding to the training sample based on the second interaction probability, where the first ranking score is positively correlated with the probability that the training sample ranks among the top in the training sample set; and determine a first ranking loss corresponding to the i-th prediction sub-model based on the difference between the first ranking score and the second ranking score.

[0242] In some embodiments, the model training module is further configured to form a training sample pair by using the training sample and other training samples in the training sample set, where m is a positive integer; determine a second ranking loss corresponding to the training sample pair based on the first probability magnitude relationship and the second probability magnitude relationship of the training sample pair, where the first probability magnitude relationship is the magnitude relationship of the training sample pair under the first interaction probability, and the first probability magnitude relationship is the magnitude relationship of the training sample pair under the second interaction probability; and determine the ranking loss corresponding to the i-th prediction sub-model based on the second ranking loss corresponding to the training sample pair.

[0243] In some embodiments, the model training module is further configured to, when the first probability magnitude relationship and the second probability magnitude relationship are inconsistent, calculate a first absolute difference between the two training samples included in the training sample pair under the first interaction probability and a second absolute difference between the two training samples included in the training sample pair under the second interaction probability; determine a second ranking loss corresponding to the training sample pair based on the absolute difference between the first absolute difference and the second absolute difference, where the second ranking loss is positively correlated with the absolute difference; and obtain the second ranking loss corresponding to the i-th prediction sub-model based on the second ranking loss corresponding to the training sample pair.

[0244] In some embodiments, the model training module is further configured to, when the first probability magnitude relationship and the second probability magnitude relationship are consistent, determine that the second ranking loss corresponding to the training sample pair is a preset value, where the preset value is less than the second ranking loss corresponding to the training sample pair when the first probability magnitude relationship and the second probability magnitude relationship are inconsistent.

[0245] In some embodiments, the first prediction module is configured to fuse the user description information of the sample user under the i-th interaction task and the second task information of the (i - 1)-th interaction task determined by the (i - 1)-th prediction sub-model based on the sampling information to obtain a first fusion feature; and predict an interaction probability based on the first fusion feature to obtain a first interaction probability corresponding to the i-th interaction task.

[0246] In some embodiments, a first prediction module is configured to calculate a first self-attention weight for the user description information of a sample user under the i-th interaction task, and calculate a second self-attention weight for the second task information of the (i - 1)-th interaction task determined by the (i - 1)-th prediction sub-model based on sampling information; perform weighted processing on the user description information of the sample user under the i-th interaction task based on the first self-attention weight to obtain a first attention feature; perform weighted processing on the second task information of the (i - 1)-th interaction task based on the first self-attention weight to obtain a second attention feature; fuse the first attention feature, the second attention feature, and the user description information of the sample user under the i-th interaction task to obtain a first fusion feature.

[0247] In some embodiments, a second prediction module is configured to fuse the user description information of the sample user under the i-th interaction task and the first task information of the (i - 1)-th interaction task determined by the (i - 1)-th prediction sub-model based on the historical interaction results under the (i - 1)-th interaction task to obtain a second fusion feature; perform interaction probability prediction based on the second fusion feature to obtain a second interaction probability corresponding to the i-th interaction task.

[0248] In some embodiments, a second prediction module is configured to calculate a third self-attention weight for the user description information of the sample user under the i-th interaction task, and calculate a fourth self-attention weight for the first task information of the (i - 1)-th interaction task determined by the (i - 1)-th prediction sub-model based on the historical interaction results under the (i - 1)-th interaction task; perform weighted processing on the user description information of the sample user under the i-th interaction task based on the third self-attention weight to obtain a third attention feature; perform weighted processing on the first task information of the (i - 1)-th interaction task based on the fourth self-attention weight to obtain a fourth attention feature; fuse the third attention feature, the fourth attention feature, and the user description information of the sample user under the i-th interaction task to obtain a second fusion feature.

[0249] In some embodiments, as Figure 13 shown, an information push device 1300 is provided, including:

[0250] A candidate user determination module 1302, configured to determine a set of candidate users for the information to be pushed;

[0251] A user description information determination module 1304, configured to determine, for each candidate user in the set of candidate users, the user description information of the candidate user under each interaction task in n interaction tasks, where n is a positive integer;

[0252] The interaction probability prediction module 1306 is configured to input the user description information under each interaction task into each prediction sub-model of the interaction probability prediction model respectively, so as to obtain the interaction probability of the candidate user for the information to be pushed; wherein, the interaction probability prediction model is trained by using the method according to any one of claims 1 to 10.

[0253] The target push user determination module 1308 is configured to determine the target push user of the information to be pushed from the candidate user set based on the respective interaction probabilities of the candidate users.

[0254] The information push module 1310 is configured to push the information to be pushed to the terminal corresponding to the target push user.

[0255] In the above information push device, since the interaction probability prediction model is trained by using the above model processing method, the interaction probability prediction model can be used to predict a more accurate interaction probability. Thus, based on the predicted interaction probability, the target users can be accurately determined, and when the information is pushed to the terminals of these target users, the possibility of obtaining positive feedback is greater, thereby improving the accuracy of information push.

[0256] In some embodiments, the target push user determination module is further configured to determine the sorting information of each candidate user for the information to be pushed based on the respective interaction probabilities of the candidate users; and determine the target push user of the information to be pushed from the candidate user set based on the respective sorting information of the candidate users.

[0257] Each module in the above model processing device and information push device can be implemented in whole or in part by software, hardware and their combination. The above modules can be embedded in the processor in the computer device in the form of hardware or be independent of the processor, or can be stored in the memory in the computer device in the form of software, so as to facilitate the processor to call and execute the operations corresponding to the above modules.

[0258] In some embodiments, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 14As shown in the figure. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store training sample data. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it is used to implement the model processing method or the information pushing method of the present application.

[0259] In some embodiments, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as Figure 15 shown in the figure. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner. The wireless manner can be achieved through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it is used to implement the model processing method or the information pushing method of the present application. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.

[0260] Those skilled in the art can understand that Figure 14 、 Figure 15The structure shown is only a block diagram of some of the structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. A specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.

[0261] In some embodiments, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps of the model processing method or information push in any of the above embodiments are implemented.

[0262] In some embodiments, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the model processing method or information push in any of the above embodiments are implemented.

[0263] In some embodiments, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps of the model processing method or information push in any of the above embodiments are implemented.

[0264] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.

[0265] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.

[0266] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0267] The above-described embodiments only represent several implementation manners of the present application. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A model processing method, characterized in that, The method includes: Obtaining user description information of a sample user for each interaction task in n interaction tasks, where n is a positive integer; In the first prediction sub-model of the interaction probability prediction model, determining a first interaction probability corresponding to the first interaction task based on the user description information of the sample user for the first interaction task; In the i-th prediction sub-model of the interaction probability prediction model, predicting a second interaction probability corresponding to the i-th interaction task based on the user description information of the sample user for the i-th interaction task and the first task information of the (i - 1)-th interaction task determined by the (i - 1)-th prediction sub-model based on the historical interaction result for the (i - 1)-th interaction task, where i is a positive integer greater than or equal to 2 and not greater than n; In the i-th prediction sub-model, predicting a first interaction probability corresponding to the i-th interaction task based on the user description information of the sample user for the i-th interaction task and the second task information of the (i - 1)-th interaction task determined by the (i - 1)-th prediction sub-model based on sampling information; the sampling information is sampled by the (i - 1)-th prediction sub-model from the historical interaction result of the sample user for the (i - 1)-th interaction task and the first interaction probability corresponding to the (i - 1)-th interaction task; Training the interaction probability prediction model based on the first interaction probability and the second interaction probability corresponding to the i-th interaction task.

2. The method according to claim 1, characterized in that, The sample user is the user corresponding to the training sample in the training sample set; the training the interaction probability prediction model based on the first interaction probability and the second interaction probability corresponding to the i-th interaction task includes: Determining a ranking loss corresponding to the i-th prediction sub-model based on the first interaction probability and the second interaction probability corresponding to the i-th interaction task, where the ranking loss is positively correlated with the difference between the first ranking information and the second ranking information corresponding to the training sample; the first ranking information is the ranking information of the training sample under the first interaction probability corresponding to the i-th interaction task, and the second ranking information is the ranking information of the training sample under the second interaction probability; Determining a cross-entropy loss corresponding to the i-th prediction sub-model based on the first interaction probability corresponding to the i-th interaction task and the historical interaction result for the i-th interaction task; Training the interaction probability prediction model based on the ranking loss and the cross-entropy loss.

3. The method according to claim 2, wherein The determining the ranking loss corresponding to the i-th prediction sub-model based on the first interaction probability and the second interaction probability corresponding to the i-th interaction task includes: Determining a first ranking score corresponding to the training sample based on the first interaction probability corresponding to the i-th interaction task, where the first ranking score is positively correlated with the probability that the training sample is ranked ahead in the training sample set; Determining a second ranking score corresponding to the training sample based on the second interaction probability, where the first ranking score is positively correlated with the probability that the training sample is ranked ahead in the training sample set; Based on the difference between the first sorting score and the second sorting score, determine the first sorting loss corresponding to the i-th prediction sub-model.

4. The method according to claim 2, wherein The determining the sorting loss corresponding to the i-th prediction sub-model based on the first interaction probability and the second interaction probability corresponding to the i-th interaction task includes: Form training sample pairs by using the training sample and other training samples in the training sample set, where m is a positive integer; Based on the first probability magnitude relationship and the second probability magnitude relationship of the training sample pair, determine the second sorting loss corresponding to the training sample pair, where the first probability magnitude relationship is the magnitude relationship of the training sample pair under the first interaction probability, and the first probability magnitude relationship is the magnitude relationship of the training sample pair under the second interaction probability; Based on the second sorting loss corresponding to the training sample pair, determine the sorting loss corresponding to the i-th prediction sub-model.

5. The method according to claim 4, wherein The determining the second sorting loss corresponding to the training sample based on the first probability magnitude relationship and the second probability magnitude relationship of the training sample pair includes: When the first probability magnitude relationship and the second probability magnitude relationship are inconsistent, calculate the first absolute difference of the two training samples included in the training sample pair under the first interaction probability and the second absolute difference under the second interaction probability; Based on the absolute difference between the first absolute difference and the second absolute difference, determine the second sorting loss corresponding to the training sample pair, and the second sorting loss is positively correlated with the absolute difference; Based on the second sorting loss corresponding to the training sample pair, obtain the second sorting loss corresponding to the i-th prediction sub-model.

6. The method according to claim 4, wherein The determining the second sorting loss corresponding to the training sample based on the first probability magnitude relationship and the second probability magnitude relationship of the training sample pair includes: When the first probability magnitude relationship and the second probability magnitude relationship are consistent, determine that the second sorting loss corresponding to the training sample pair is a preset value, and the preset value is less than the second sorting loss corresponding to the training sample pair when the first probability magnitude relationship and the second probability magnitude relationship are inconsistent.

7. The method according to claim 1, wherein The predicting the first interaction probability corresponding to the i-th interaction task based on the user description information of the sample user under the i-th interaction task and the second task information of the (i - 1)-th interaction task determined by the (i - 1)-th prediction sub-model based on sampling information includes: Fuse the user description information of the sample user under the i-th interaction task and the second task information of the (i - 1)-th interaction task determined by the (i - 1)-th prediction sub-model based on sampling information to obtain a first fusion feature; Based on the first fusion feature, perform an interaction probability prediction to obtain the first interaction probability corresponding to the i-th interaction task.

8. The method according to claim 7, characterized in that, The fusing the user description information of the sample user under the i-th interaction task and the second task information of the (i - 1)-th interaction task determined by the (i - 1)-th prediction sub-model based on sampling information to obtain a first fusion feature includes: Calculate the first self-attention weight for the user description information of the sample user under the i-th interaction task, and calculate the second self-attention weight for the second task information of the (i - 1)-th interaction task determined by the (i - 1)-th prediction submodel based on the sampling information; Perform weighted processing on the user description information of the sample user under the i-th interaction task based on the first self-attention weight to obtain the first attention feature; Perform weighted processing on the second task information of the (i - 1)-th interaction task based on the first self-attention weight to obtain the second attention feature; Fuse the first attention feature, the second attention feature, and the user description information of the sample user under the i-th interaction task to obtain the first fusion feature.

9. The method according to claim 1, wherein Predicting the second interaction probability corresponding to the i-th interaction task based on the user description information of the sample user under the i-th interaction task and the first task information of the (i - 1)-th interaction task determined by the (i - 1)-th prediction submodel based on the historical interaction results under the (i - 1)-th interaction task includes: Fuse the user description information of the sample user under the i-th interaction task and the first task information of the (i - 1)-th interaction task determined by the (i - 1)-th prediction submodel based on the historical interaction results under the (i - 1)-th interaction task to obtain the second fusion feature; Perform interaction probability prediction based on the second fusion feature to obtain the second interaction probability corresponding to the i-th interaction task.

10. The method according to claim 9, characterized in that, The fusing the user description information of the sample user under the i-th interaction task and the first task information of the (i - 1)-th interaction task determined by the (i - 1)-th prediction submodel based on the historical interaction results under the (i - 1)-th interaction task to obtain the second fusion feature includes: Calculate the third self-attention weight for the user description information of the sample user under the i-th interaction task, and calculate the fourth self-attention weight for the first task information of the (i - 1)-th interaction task determined by the (i - 1)-th prediction submodel based on the historical interaction results under the (i - 1)-th interaction task; Perform weighted processing on the user description information of the sample user under the i-th interaction task based on the third self-attention weight to obtain the third attention feature; Perform weighted processing on the first task information of the (i - 1)-th interaction task based on the fourth self-attention weight to obtain the fourth attention feature; Fuse the third attention feature, the fourth attention feature, and the user description information of the sample user under the i-th interaction task to obtain the second fusion feature.

11. An information push method, characterized in that, The method includes: Determine a set of candidate users for the information to be pushed; For each candidate user in the set of candidate users, determine the user description information of the candidate user under each interaction task in n interaction tasks, where n is a positive integer; Input the user description information under each interaction task into each prediction submodel of the interaction probability prediction model respectively to obtain the interaction probability of the candidate user for the information to be pushed; wherein, the interaction probability prediction model is trained by the method according to any one of claims 1 to 10. Determine the target push user of the information to be pushed from the set of candidate users based on the respective interaction probabilities of each of the candidate users; Push the information to be pushed to the terminal corresponding to the target push user.

12. The method according to claim 11, wherein The determining the target push user of the information to be pushed from the set of candidate users based on the respective interaction probabilities of each of the candidate users includes: Determine the sorting information of each candidate user for the information to be pushed based on the respective interaction probabilities of each of the candidate users; Determine the target push user of the information to be pushed from the set of candidate users based on the respective sorting information of each candidate user.

13. A model processing device, characterized in that, The device includes: A user description information acquisition module, configured to acquire the user description information of a sample user for each interaction task in n interaction tasks, where n is a positive integer; A first prediction module, configured to determine a first interaction probability corresponding to the first interaction task in the first prediction sub-model of the interaction probability prediction model based on the user description information of the sample user in the first interaction task; A second prediction module, configured to, in the i-th prediction sub-model of the interaction probability prediction model, predict a second interaction probability corresponding to the i-th interaction task based on the user description information of the sample user in the i-th interaction task and the first task information of the (i - 1)-th interaction task determined by the (i - 1)-th prediction sub-model based on the historical interaction result in the (i - 1)-th interaction task, where i is a positive integer greater than or equal to 2 and not greater than n; The second prediction module is further configured to, in the i-th prediction sub-model, predict a first interaction probability corresponding to the i-th interaction task based on the user description information of the sample user in the i-th interaction task and the second task information of the (i - 1)-th interaction task determined by the (i - 1)-th prediction sub-model based on the sampling information; the sampling information is sampled by the (i - 1)-th prediction sub-model from the historical interaction result of the sample user in the (i - 1)-th interaction task and the first interaction probability corresponding to the (i - 1)-th interaction task; A model training module, configured to train the interaction probability prediction model based on the first interaction probability and the second interaction probability corresponding to the i-th interaction task.

14. An information push device, characterized in that, The device includes: A candidate user determination module, configured to determine a set of candidate users for the information to be pushed; A user description information determination module, configured to determine, for each candidate user in the set of candidate users, the user description information of the candidate user for each interaction task in n interaction tasks, where n is a positive integer; An interaction probability prediction module, configured to input the user description information for each interaction task into the respective prediction sub-models of the interaction probability prediction model to obtain the interaction probabilities of the candidate users for the information to be pushed; wherein, the interaction probability prediction model is trained by the method described in any one of claims 1 to 10; A target push user determination module, configured to determine the target push user of the information to be pushed from the set of candidate users based on the respective interaction probabilities of each of the candidate users; An information push module, configured to push the information to be pushed to the terminal corresponding to the target push user.

15. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 10 or 11 to 12 are implemented.

16. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 10 or 11 to 12 are implemented.

17. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 10 or 11 to 12 are implemented.