Push processing method, related device and medium
By using the probability prediction model and the fusion model in the push processing, combined with the first loss function training, the problem of out-of-synchronization of task score changes and sorting changes is solved, and the sorting accuracy and accuracy of push is improved.
Patent Information
- Application Number
- CN202311585188.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-23
- Publication Date
- 2025-05-23
AI Technical Summary
In the existing push processing technology, the change of task scores is not synchronized with the sorting changes, resulting in low sorting accuracy and insufficient pushing accuracy.
By obtaining the target push basic characteristics, input a probability prediction model to obtain multiple first probabilities, input a fusion model to obtain the first overall probability, and push sorting is performed based on this. The fusion model is trained based on the first loss function, and uses the ranking information of the push basic feature samples to improve the sorting accuracy.
It improves the sorting accuracy in push processing, thereby improving push accuracy, and overcoming the problem of small actual sorting differences in the case of large differences in score sorting.
Smart Images

Figure CN120034576A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of big data, and in particular to a push processing method, related devices and media. Background Art
[0002] In the current push processing, the push processing model often predicts the score or probability of the target object completing multiple tasks (for example, likes, comments, collections, etc.) for the pushed content (such as short videos, etc.), and then determines the push order for the target object based on these scores or probabilities (such as the order of pushing short videos). When determining the push order based on these scores or probabilities, the prior art can use formula fusion or model fusion methods. Formula fusion generally determines different weights for each score or probability, and obtains the final weighted score or probability through the formula. Model fusion is to input each score or probability into the fusion model, and obtain the fusion score or probability from the fusion model, so as to sort based on the fusion score or probability. Both methods calculate the scores corresponding to each task (such as weighting or model processing) and sort according to the final results.
[0003] However, the score changes corresponding to each task may not be synchronized with the ranking changes. For example, for the same task, the same score at different times may represent completely different rankings, and scores with large differences may not have much difference in ranking. Moreover, these processing methods will smooth out the distribution differences of scores between multiple tasks, which deviates from actual applications, resulting in low ranking accuracy and thus low push accuracy. Therefore, a technology that can improve the sorting accuracy in push processing and thus improve push accuracy is desired. Summary of the invention
[0004] The embodiments of the present disclosure provide a push processing method, a related device and a medium, which can improve the sorting accuracy in push processing, thereby improving the push accuracy.
[0005] According to one aspect of the present disclosure, a push processing method is provided, including:
[0006] Acquire a plurality of target push basic features, wherein the plurality of target push basic features include target object features and to-be-pushed content features;
[0007] Inputting the plurality of target push basic features into a probability prediction model to obtain a plurality of first probabilities of the target object completing a plurality of tasks for the content to be pushed;
[0008] Multiple first probabilities are cascaded and input into a fusion model to obtain a first overall probability that the target object completes multiple tasks for the content to be pushed, and the content to be pushed is pushed to the target object based on the first overall probability, wherein the fusion model is trained based on a first loss function, and the first loss function is calculated based on the first ranking of the push basic feature samples in the push basic feature sample set in the first sorting and the second ranking in the second sorting, the first sorting is for each of the tasks, after multiple push basic feature samples in the push basic feature sample set are input into the probability prediction model, the output of the probability prediction model corresponding to the task, and the second sorting is the sorting of the output of the fusion model after multiple push basic feature samples in the push basic feature sample set are input into the probability prediction model.
[0009] According to one aspect of the present disclosure, a push processing device is provided, comprising:
[0010] An acquisition unit, configured to acquire a plurality of target push basic features, wherein the plurality of target push basic features include target object features and to-be-pushed content features;
[0011] An input unit, configured to input the plurality of target push basic features into a probability prediction model to obtain a plurality of first probabilities of the target object completing a plurality of tasks for the content to be pushed;
[0012] A pushing unit is used to cascade multiple first probabilities and input them into a fusion model to obtain a first overall probability that the target object completes multiple tasks for the content to be pushed, and pushes the content to be pushed to the target object based on the first overall probability, wherein the fusion model is trained based on a first loss function, and the first loss function is calculated based on the first ranking of the push basic feature samples in the push basic feature sample set in the first sorting and the second ranking in the second sorting, the first sorting is for each of the tasks, after the multiple push basic feature samples in the push basic feature sample set are input into the probability prediction model, the output of the probability prediction model corresponding to the task, and the second sorting is the sorting of the output of the fusion model after the multiple push basic feature samples in the push basic feature sample set are input into the probability prediction model.
[0013] Optionally, the fusion model is trained by the following process:
[0014] Acquire a push basic feature sample set, wherein each push basic feature sample in the push basic feature sample set includes a sample object feature and a sample content feature;
[0015] Inputting the pushed basic feature sample set into the probability prediction model, and cascading the output of the probability prediction model into the fusion model;
[0016] For each of the tasks, obtaining, output by the probability prediction model, a probability ranking of the sample objects of each of the pushed basic feature samples in terms of completing the task for the sample content, as the first ranking;
[0017] Obtaining the overall probability ranking of the sample objects of each of the pushed basic feature samples output by the fusion model for completing the plurality of tasks for the sample content as the second ranking;
[0018] Calculating a first loss function based on a first ranking of the pushed basic feature sample in the first ranking corresponding to the task and a second ranking in the second ranking;
[0019] Based on the first loss function, the fusion model is trained.
[0020] Optionally, the calculating a first loss function based on a first ranking of the pushed basic feature sample in the first ranking corresponding to the task and a second ranking in the second ranking includes:
[0021] Calculating a first sub-loss function corresponding to the task based on a first ranking of the pushed basic feature sample in the first ranking corresponding to the task and a second ranking in the second ranking;
[0022] Determining a first weight corresponding to the task based on the first ranking and the second ranking;
[0023] Based on the first weights corresponding to the multiple tasks, a weighted sum of the first sub-loss functions corresponding to the tasks is calculated to obtain the first loss function.
[0024] Optionally, the calculating a first sub-loss function corresponding to the task based on a first ranking of the pushed basic feature sample in a first ranking corresponding to the task and a second ranking in the second ranking includes:
[0025] Setting a first label corresponding to the task, wherein if the first ranking in the first ranking corresponding to the task is before the second ranking, the first label is a first value; otherwise, the first label is a second value;
[0026] Obtaining the overall probability output by the fusion model;
[0027] Based on the first label and the overall probability, the first sub-loss function corresponding to the task is calculated.
[0028] Optionally, determining a first weight corresponding to the task based on the first ranking and the second ranking includes:
[0029] Based on the first ranking, determining a first normalized cumulative loss gain corresponding to the task;
[0030] Based on the second ranking, determining a second normalized discounted cumulative gain;
[0031] An absolute value of a difference between a first normalized discounted cumulative gain corresponding to the task and the second normalized discounted cumulative gain is used as the first weight corresponding to the task.
[0032] Optionally, determining a first normalized cumulative loss gain corresponding to the task based on the first ranking includes:
[0033] Get the distribution control coefficient;
[0034] Taking the logarithm of the sum of the product of the distribution control coefficient and the first rank and a predetermined constant to obtain a first intermediate value;
[0035] Based on the logarithm of the predetermined constant and the first intermediate value, the first normalized discounted cumulative gain corresponding to the task is obtained.
[0036] Optionally, the training the fusion model based on the first loss function includes:
[0037] Calculating a second loss function of the probability prediction model for the pushed basic feature sample;
[0038] Determine a single sample loss function based on the first loss function and the second loss function;
[0039] Based on the single sample loss function, the fusion model is trained.
[0040] Optionally, the training the fusion model based on the single sample loss function includes:
[0041] Obtaining a sample weight of each of the pushed basic feature samples;
[0042] Using the sample weights, a weighted sum of the single sample loss functions of each of the pushed basic feature samples is calculated to obtain a multi-sample loss function;
[0043] The fusion model is trained using the multi-sample loss function.
[0044] Optionally, the push basic feature sample has a plurality of second tags corresponding to a plurality of the tasks, and the second tags indicate whether the sample object completes the task for the sample content;
[0045] The calculating a second loss function of the probability prediction model for the pushed basic feature sample includes:
[0046] Acquire a plurality of third probabilities output by the probability prediction model, that the sample object in the push basic feature sample completes a plurality of the tasks for the sample content;
[0047] The second loss function is calculated based on a plurality of second labels corresponding to a plurality of the tasks and a plurality of the third probabilities.
[0048] Optionally, the probability prediction model includes multiple first expert sub-models, multiple first gating nodes corresponding to multiple tasks, and multiple first prediction sub-models, each of the first gating nodes is connected to multiple first expert sub-models, the first gating node corresponding to the task is connected to the first prediction sub-model corresponding to the task, and the first prediction sub-model outputs the first probability corresponding to the task;
[0049] The input unit is specifically used for:
[0050] Based on the plurality of target push basic features, obtaining a first cascade vector;
[0051] Inputting the first cascade vector into a plurality of the first expert sub-models respectively to obtain a plurality of first vectors;
[0052] By using the first gating weights of the first expert sub-models and the first gating node corresponding to the task, a weighted sum of the first vectors is obtained to obtain a second vector corresponding to the task;
[0053] The second vector is input into the first prediction sub-model corresponding to the task to obtain the first probability corresponding to the task.
[0054] Optionally, the input unit is further specifically used for:
[0055] Quantizing the plurality of target push basic feature vectors into a plurality of target push basic feature vectors;
[0056] A plurality of the target pushed basic feature vectors are concatenated into the first concatenated vector.
[0057] Optionally, after quantizing the plurality of target push basic feature vectors into a plurality of target push basic feature vectors, the push processing device further comprises:
[0058] A determining unit, configured to determine whether a dimension of the target push basic feature vector is greater than a first threshold;
[0059] A pooling unit is used to perform pooling processing on the target pushed basic feature vector so that the dimension of the pooled target pushed basic feature vector is equal to the first threshold.
[0060] Optionally, the first expert sub-model comprises a first weight matrix, the number of columns of the first weight matrix is equal to the dimension of the first cascade vector, and the number of rows of the first weight matrix is equal to the dimension of the first vector;
[0061] The input unit is also specifically used for:
[0062] The first weight matrices in the plurality of the first expert sub-models are multiplied by the transpose of the first cascade vector to obtain a plurality of the first vectors.
[0063] Optionally, the first gating node comprises a second weight matrix, the number of columns of the second weight matrix is equal to the dimension of the first cascade vector, and the number of rows of the second weight matrix is equal to the number of the first expert sub-models;
[0064] The input unit is also specifically used for:
[0065] Determine, by means of the first gating node, a first product vector of the second weight matrix and the transpose of the first cascade vector;
[0066] performing exponential normalization on the first product vector to obtain a first gating weight vector, wherein the first gating weight vector includes the first gating weights of a plurality of the first expert sub-models;
[0067] Using the first gating weights of the plurality of the first expert sub-models in the first gating weight vector, a weighted sum is calculated for the plurality of the first vectors to obtain a second vector corresponding to the task.
[0068] Optionally, the input unit is further specifically used for:
[0069] Inputting the second vector into the first prediction sub-model corresponding to the task to obtain a first result value corresponding to the task;
[0070] Apply a nonlinear activation function to the first result value to obtain the first probability.
[0071] Optionally, the fusion model includes a fusion weight vector, and the dimension of the fusion weight vector is equal to the number of the tasks;
[0072] The push unit is specifically used for:
[0073] Cascading a plurality of the first probabilities into a first probability vector, wherein the dimension of the first probability vector is equal to the number of the tasks;
[0074] Multiplying the transpose of the first probability vector by the fusion weight vector to obtain a second result value;
[0075] Apply a nonlinear activation function to the second result value to obtain the first overall probability.
[0076] According to one aspect of the present disclosure, an electronic device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor implements the push processing method as described above when executing the computer program.
[0077] According to one aspect of the present disclosure, a computer-readable storage medium is provided, wherein the storage medium stores a computer program, and when the computer program is executed by a processor, the push processing method as described above is implemented.
[0078] According to one aspect of the present disclosure, a computer program product is provided. The computer program product includes a computer program. The computer program is read and executed by a processor of a computer device, so that the computer device executes the push processing method as described above.
[0079] In the disclosed embodiment, during model training, for a sample, instead of processing the scores corresponding to each task and sorting them according to the processing results, the ranking of the sample in all samples in each task is processed and sorted according to the processing results. Since the ranking is more direct than the score, it can overcome the situation that the same score may have a large difference in ranking, and the score with a large difference may have a small difference in actual ranking, so as to improve the accuracy of sorting. In the present invention, a fusion model is also cascaded behind the probability prediction model. Not only can multiple first rankings of a sample for multiple tasks predicted by the probability prediction model be obtained, but also the overall second ranking of the sample predicted by the fusion model can be obtained. Therefore, by constructing the first loss function through the first ranking and the second ranking, and training the fusion model based on the first loss function, it can better reflect the difference and discrimination between the ranking ranking of a sample for a specific task and the ranking ranking for all tasks as a whole, and improve the sorting accuracy of the overall model including the probability prediction model and the trained fusion model, thereby improving the push accuracy.
[0080] Other features and advantages of the present disclosure will be described in the following description, and partly become apparent from the description, or understood by practicing the present disclosure. The purpose and other advantages of the present disclosure can be realized and obtained by the structures particularly pointed out in the description, claims and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0081] The accompanying drawings are used to provide further understanding of the technical solution of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the technical solution of the present disclosure and do not constitute a limitation on the technical solution of the present disclosure.
[0082] Figure 1 is a system architecture diagram of a system to which a push processing method according to an embodiment of the present disclosure is applied;
[0083] Figure 2A-2C It is a schematic diagram of an interface of an embodiment of the present disclosure applied in a video content push scenario;
[0084] Figure 3 is a flowchart of a push processing method according to an embodiment of the present disclosure;
[0085] Figure 4A yes Figure 3 Step 310 in FIG. 1 is a schematic diagram showing that the basic characteristics of target push include characteristics of the target object and characteristics of the content to be pushed;
[0086] Figure 4B yes Figure 3 Step 310 in the diagram is a schematic diagram of target push basic features including target object features, to-be-pushed content features, and target scene features;
[0087] Figure 5 yes Figure 3 Step 320 in FIG. 1 is a schematic diagram of a model structure of a probability prediction model;
[0088] Figure 6 yes Figure 3 Step 320 in the flowchart is a flowchart of obtaining multiple first probabilities based on multiple target push basic features and probability prediction models;
[0089] Figure 7 yes Figure 6 Step 610 in the flowchart of acquiring a first cascade vector based on multiple target push basic features;
[0090] Fig. 8A yes Figure 4A A schematic diagram of a target push basic feature vector corresponding to the target push basic feature in ;
[0091] Figure 8B yes Figure 4B A schematic diagram of a target push basic feature vector corresponding to the target push basic feature in ;
[0092] Fig. 9 is a flow chart of performing pooling processing on a target push basic feature vector based on the dimension of the target push basic feature vector;
[0093] Fig. 10A yes Fig. 9A schematic diagram of a specific implementation process of performing average value pooling on the target push basic feature vector in step 920;
[0094] Fig. 10B yes Fig. 9 A schematic diagram of a specific implementation process of performing maximum pooling on the target push basic feature vector in step 920;
[0095] Fig.11 yes Figure 7 A schematic diagram of a specific implementation process of step 720 in FIG. 7 for cascading multiple target push basic feature vectors into a first cascade vector;
[0096] Fig.12 yes Figure 6 Step 630 in FIG. 1 is a schematic diagram of obtaining a second vector through a first gating node corresponding to a task;
[0097] Fig.13 yes Figure 6 Step 640 in the flowchart is about inputting the second vector into the first prediction sub-model corresponding to the task to obtain the first probability;
[0098] Fig.14 yes Figure 3 Step 330 in FIG. 1 is a schematic diagram of a model structure of a cascade probability prediction model and a fusion model;
[0099] Fig.15 yes Figure 3 Step 330 in the flowchart is about cascading multiple first probabilities and inputting them into the fusion model to obtain a first overall probability;
[0100] Fig.16A yes Figure 3 A schematic diagram of a specific implementation process of pushing the content to be pushed to the target object based on the first overall probability sorting in step 330;
[0101] Fig. 16B is a schematic diagram of a specific implementation process of pushing the content to be pushed to the target object based on the first overall probability ranking and the push threshold;
[0102] Fig.17 yes Figure 3 Step 330 in FIG. 1 is a flowchart of a training process for a fusion model;
[0103] Fig.18 yes Fig.17 Step 1730 in FIG. 1 is a schematic diagram of obtaining a first sorting;
[0104] Fig.19 yes Fig.17 Step 1740 in FIG. 1 is a schematic diagram of obtaining a second sorting;
[0105] Fig. 20 yes Fig.17 Step 1750 in the flowchart of calculating a first loss function based on the first rank and the second rank;
[0106] Fig.21 yes Fig. 20 Step 2010 in the flowchart is based on the first ranking and the second ranking calculation task corresponding to the first sub-loss function;
[0107] Fig. 22 yes Fig. 20 Step 220 in the flowchart is a flowchart for determining a first weight based on the first ranking and the second ranking;
[0108] Fig.23 yes Fig. 22 Step 2210 in the flowchart is a flowchart for determining a first normalized discounted cumulative gain based on the first ranking;
[0109] Fig.24 yes Fig.17 Step 1760 in the flowchart of training the fusion model based on the first loss function;
[0110] Fig.25 yes Fig.24 Step 2410 in the flowchart of calculating the second loss function;
[0111] Fig.26 yes Fig.24 Step 2430 in the flowchart of training the fusion model based on the single sample loss function;
[0112] Fig. 27 It is an implementation detail diagram of the push processing method of the embodiment of the present disclosure.
[0113] Fig.28 is a module diagram of a push processing device according to an embodiment of the present disclosure;
[0114] Fig.29 According to the embodiment of the present disclosure, Figure 3 The terminal structure diagram of the push processing method shown;
[0115] Fig.30 According to the embodiment of the present disclosure, Figure 3 The server structure diagram of the push processing method shown. DETAILED DESCRIPTION
[0116] In order to make the purpose, technical solution and advantages of the present disclosure more clear, the present disclosure is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present disclosure and are not used to limit the present disclosure.
[0117] Before further describing the embodiments of the present disclosure in detail, the nouns and terms involved in the embodiments of the present disclosure are described. The nouns and terms involved in the embodiments of the present disclosure are subject to the following interpretations:
[0118] Deep Neural Networks (DNN): is a multi-layer unsupervised neural network. The deep neural network uses the output features of the previous layer as the input of the next layer for feature learning. After layer-by-layer feature mapping, the features of the existing spatial samples are mapped to another feature space, so as to learn to have better feature expression for the existing input. The deep neural network has multiple nonlinear mapping feature transformations and can fit highly complex functions.
[0119] Push: refers to the method by which a system or platform proactively sends information, content or services to users without explicit requests or searches by users. Push processing methods are usually targeted based on personal information such as user interests, preferences, historical behavior, etc. to provide personalized content push.
[0120] Nonlinear activation function (sigmoid function): also known as logistic function or S-shaped function. The sigmoid function maps the input value to an output value between 0 and 1, that is, the output value ranges from 0 to 1. When the input value approaches negative infinity, the output value is close to 0; when the input value approaches positive infinity, the output value is close to 1.
[0121] Normalized exponential (softmax function): compresses a K-dimensional vector containing arbitrary real numbers into another K-dimensional vector so that each element ranges between (0,1) and the sum of all elements is 1.
[0122] System architecture and scenario description of the application of the embodiments of the present disclosure
[0123] Figure 1 The system architecture diagram of the push processing method according to the embodiment of the present disclosure includes: a target terminal 110 , the Internet 120 , a gateway 130 , and a push processing server 140 .
[0124] The object terminal 110 is a device used by the object to view the reminder message corresponding to the pushed content. It includes various forms such as desktop computers, laptops, PDAs (personal digital assistants), mobile phones, car terminals, home theater terminals, and dedicated terminals. In addition, it can be a single device or a collection of multiple devices. For example, multiple devices are connected through a local area network and work together using a common display device to form a terminal. The object terminal 110 can also communicate with the Internet 120 in a wired or wireless manner to exchange data.
[0125] The gateway 130 is also called an internetwork connector or a protocol converter. The gateway 130 realizes network interconnection at the transport layer and is a computer system or device that acts as a converter. The gateway 130 is a translator between two systems that use different communication protocols, data formats or languages, or even have completely different architectures. At the same time, the gateway 130 can also provide filtering and security functions. The message sent by the object terminal 110 to the push processing server 140 is sent to the corresponding push processing server 140 through the gateway 130. The message sent by the push processing server 140 to the object terminal 110 is also sent to the corresponding object terminal 110 through the gateway 130.
[0126] The push processing server 140 refers to a computer system that can provide content message push services to the target terminal 110. Compared with the target terminal 110, the push processing server 140 has higher requirements in terms of stability, security, performance, etc. The push processing server 140 can be a high-performance computer in the network platform, a cluster of multiple high-performance computers, a part of a high-performance computer (such as a virtual machine), a combination of parts of multiple high-performance computers (such as virtual machines), etc. The push processing server 140 can also communicate with the Internet 120 in a wired or wireless manner to exchange data.
[0127] The embodiments of the present disclosure can be applied in various scenarios, such as Figures 2A-2C The scenario shown is of viewing video content push messages in an instant short messaging application, etc.
[0128] like Figure 2A As shown, in the interface of the instant short message application in the target terminal 110, the instant short message application can be used to communicate with multiple contacts and can also be used to watch videos. After triggering the "Message" option in the function bar below, a short message list is displayed. The short message list contains message bars with multiple contacts. At 10:30 am, in the "Short Video" option in the function bar below, there is a content prompt ( Figure 2A The content prompt is used to remind the object that new content is pushed. The object can watch the recommended content by triggering the "Short Video" option in the function bar below.
[0129] like Figure 2B As shown, at 12:00 am, the subject sees the "small black dot" in the interface of the instant short message application and chooses to trigger the "short video" in the function bar below.
[0130] like Figure 2CAs shown in the figure, after the "short video" option is triggered, the instant short message application opens the content push interface and plays the pushed content, which is a video V. The content push interface also shows the number of likes (176,000), favorites (2,392), and comments (2,965) that the video received before 12:00 am. The subject can also like, favorite, and comment on the video in the interface. The video content pushed to the subject is determined by the push processing server 140.
[0131] In the current push processing, the push processing model often predicts the scores or probabilities of the target object completing multiple tasks (e.g., likes, comments, collections, etc.) for the pushed content (e.g., short videos, etc.), and then determines the push order (e.g., the order of pushing short videos) for the target object based on these scores or probabilities. However, the score changes corresponding to each task may not be synchronized with the ranking changes. For example, during the off-duty time period, the frequency of user use of the terminal device increases, and the scores for each task are generally high. During the working time period, the frequency of user use of the terminal device is low, and the scores for each task are generally low. Therefore, for an estimated task, the same score at different times may represent completely different rankings, and the scores with large differences may not differ much in the ranking. Since the existing processing method will smooth out the distribution differences of scores between multiple tasks, there is a deviation from the actual application, resulting in low sorting accuracy, thereby resulting in low push accuracy. When pushing, the embodiment of the present disclosure processes the ranking of the sample in all samples in each task and sorts them according to the processing results. Since ranking is more direct than scores, it can overcome the situation where the same score may have very different rankings, or scores with large differences may have very small actual ranking differences, so as to improve the accuracy of ranking.
[0132] General description of the disclosed embodiments
[0133] According to an embodiment of the present disclosure, a push processing method is provided.
[0134] The push processing method is a process of determining the target content for the target object and pushing the reminder message of the target content to the target object's object terminal 110. The target object refers to the object that wants to view the target content. The target content can be a variety of types such as pictures, videos, articles, applications, etc.
[0135] In the current push processing, the push processing model often predicts the scores or probabilities of the target object completing multiple tasks (e.g., likes, comments, collections, etc.) for the pushed content (e.g., short videos, etc.), and then determines the push order (e.g., the order of pushing short videos) for the target object based on these scores or probabilities. When determining the push order based on these scores or probabilities, the prior art can use formula fusion or model fusion methods. Formula fusion generally determines different weights for each score or probability, and obtains the final weighted score or probability through the formula. Model fusion is to input each score or probability into the fusion model, and obtain the fusion score or probability by the fusion model, so as to sort based on the fusion score or probability. Both methods calculate the scores corresponding to each task (such as weighting or model processing) and sort according to the final results. However, the changes in the scores corresponding to each task may not be synchronized with the changes in the sorting. For example, for the same task, the same score at different times may represent completely different sortings, and scores with large differences may not differ much in the sorting. The existing processing method will smooth out the distribution differences of scores between multiple tasks, which deviates from the actual application, resulting in low sorting accuracy, and thus low push accuracy. In addition, if there is a mutually exclusive relationship between the two tasks, for example, Task A is the dwell time and Task B is the 5-second like rate, then there is a certain contradiction between the two tasks. At this time, if a simple addition fusion is used (i.e., the score of Task A + the score of Task B), then since Task A and Task B are mutually exclusive, their summation result may remain constant, which may easily result in a situation where the fused score does not make much difference in the score of the head sample, which in turn leads to a disadvantageous effect of sorting based on the fused score for both Task A and Task B.
[0136] In the disclosed embodiment, during model training, for a sample, instead of processing the scores corresponding to each task and sorting them according to the processing results, the ranking of the sample in all samples in each task is processed and sorted according to the processing results. Since the ranking is more direct than the score, it can overcome the situation that the same score may have a large difference in ranking, and the score with a large difference may have a small difference in actual ranking, so as to improve the accuracy of sorting. In the present invention, a fusion model is also cascaded behind the probability prediction model. Not only can multiple first rankings of a sample for multiple tasks predicted by the probability prediction model be obtained, but also the overall second ranking of the sample predicted by the fusion model can be obtained. Therefore, by constructing the first loss function through the first ranking and the second ranking, and training the fusion model based on the first loss function, it can better reflect the difference and discrimination between the ranking ranking of a sample for a specific task and the ranking ranking for all tasks as a whole, and improve the sorting accuracy of the overall model including the probability prediction model and the trained fusion model, thereby improving the push accuracy.
[0137] It should be noted that the specific model structure and training process of the probability prediction model and the fusion model will be described in detail in the subsequent embodiments and will not be described in detail here.
[0138] The push processing method of the embodiment of the present disclosure is executed on the server 140. After determining the target content to be pushed, the reminder message corresponding to the target content is pushed to the target terminal 110 through the network and the Internet 120.
[0139] like Figure 3 As shown, according to one embodiment of the present disclosure, the push processing method includes:
[0140] Step 310: Acquire multiple target push basic features, where the multiple target push basic features include target object features and to-be-pushed content features;
[0141] Step 320: Input multiple target push basic features into the probability prediction model to obtain multiple first probabilities of the target object completing multiple tasks for the content to be pushed;
[0142] Step 330: Concatenate multiple first probabilities and input them into a fusion model to obtain a first overall probability that the target object completes multiple tasks for the content to be pushed, and push the content to be pushed to the target object based on the first overall probability.
[0143] The advantage of the embodiment of step 310-step 330 is that, during model training, for a sample, instead of processing the scores corresponding to each task and sorting them according to the processing results, the ranking of the sample in each task among all samples is processed and sorted according to the processing results. Since the ranking is more direct than the score, it can overcome the situation that the same score may have a large difference in ranking, and the score with a large difference may have a small difference in actual ranking, so as to improve the accuracy of sorting. In the present invention, a fusion model is also cascaded behind the probability prediction model. Not only can multiple first rankings of a sample for multiple tasks predicted by the probability prediction model be obtained, but also the overall second ranking of the sample predicted by the fusion model can be obtained. Therefore, by constructing the first loss function through the first and second rankings, and training the fusion model based on the first loss function, it can better reflect the difference and discrimination between the ranking of a sample for a specific task and the ranking of all tasks as a whole, and improve the sorting accuracy of the overall model including the probability prediction model and the trained fusion model, thereby improving the push accuracy.
[0144] The above steps 310 to 330 are described in detail below.
[0145] Detailed description of step 310
[0146] In step 310, multiple target push basic features are obtained, wherein the multiple target push basic features include target object features and to-be-pushed content features. The target object features and to-be-pushed content features are features used to predict different aspects of the target content to be pushed to the target object to watch.
[0147] The target object feature is the feature of the target object itself. Usually, target objects with a common target object feature may show certain commonalities in viewing target content. Therefore, using the target object feature for target content prediction helps improve the effect of push processing. Figure 4A As shown, the target object characteristics include education level characteristics, working years characteristics, work characteristics, hobby characteristics, historical viewing content characteristics, etc.
[0148] In terms of education level, people with higher education levels usually receive more comprehensive and in-depth professional education, and are often more interested in in-depth and professional content. They may want to obtain more knowledge and information in professional fields, so the content pushed may tend to provide in-depth and professional content. For example, lawyers tend to watch videos that analyze relevant legal knowledge in depth.
[0149] As for years of work, the increase in years of work means the accumulation of experience in a certain industry or field. People may be more interested in the professional knowledge and information in that field, and the pushed content may be more inclined to provide practical information related to their work, as well as trends and changes in industry development. For example, for people who have worked for 15 years, it is more likely that the push content will be more inclined to provide industry trends, cutting-edge technologies and other related content.
[0150] In terms of job characteristics, people who work in the same job are more likely to choose the same content that they often need to watch in their job. For example, students tend to watch educational or growth-related video content.
[0151] Subjects with the same hobbies may tend to watch the same content. For example, subjects who like mathematics tend to watch mathematics teaching content; for subjects who like painting, painting teaching content is more likely to be recommended and displayed to them.
[0152] For the historical viewing content feature, the subject may like to watch content with the same type of historical viewing content feature. For example, the subject likes to watch video content with a content length of less than 100 seconds, or the subject likes to watch video content with more than 1,200 likes.
[0153] In one embodiment, the target object features of the target object may be obtained by obtaining the target object features selected by the target object in advance, or by obtaining the target object features statistically obtained from the target object's viewing content log.
[0154] Note that when obtaining the target object characteristics of the target object, the target object's consent must be obtained in advance. In addition, the collection, use and processing of these object characteristics will comply with relevant laws, regulations and standards. When obtaining the target object's consent, the target object's separate permission or separate consent can be obtained through pop-ups or jumps to a confirmation page.
[0155] The characteristics of the content to be pushed are the characteristics of the content to be pushed. Usually, the content to be pushed with the same characteristics is easy to be viewed by the same target object. Therefore, using the characteristics of the content to be pushed for target content prediction helps to improve the effect of push processing. Figure 4A As shown, the features of the content to be pushed include content type features, content length features, like quantity features, comment quantity features, collection quantity features, etc.
[0156] It should be noted that a target object feature is used as a target push basic feature, and a feature of the content to be pushed is also used as a target push basic feature. Figure 4A As shown, there are 5 target object features and 5 to-be-pushed content features, so there are 10 target push basic features in total.
[0157] In another embodiment, if Figure 4B As shown, multiple target push basic features also include target scene features. Target scene features are features of the current scene in which the target object is located. Usually the target object may watch similar content in the same scene. For example, in a work scene, the target object may want to view push content related to work more; when at home, he may want to view push content related to his own preferences more. Therefore, using target scene features for target content prediction helps to improve the effect of push processing. Target scene features include geographical area features, current input day type features (weekdays, weekends, or holidays), weather features, etc.
[0158] When obtaining the target scene characteristics of the target object, the consent of the target object must be obtained in advance, which is similar to the above and will not be repeated here.
[0159] The characteristics of the current geographical area refer to the characteristics of the geographical area where the target object is currently located. This geographical area can be an administrative geographical area, such as City A, City B, etc., or it can be a geographical area divided by entities on the map, such as the XX shopping mall area, the XX hospital area, etc. For administrative geographical areas, the push content that the target object is interested in has an impact. For example, if the target object is in City A, the push content can be tourist attractions or food culture related to City A. For geographical areas divided by entities, the push content that the target object is interested in also has an impact. For example, if the target object is in the XX shopping mall area, the push content can be activity information or special projects in the mall; if the target object is in the XX hospital area, the push content can be content related to health and wellness.
[0160] The current input day type feature refers to the feature of the type of day when the target object views the push content. The current input day type is divided into weekdays, weekends, and holidays. The type of push content that the target object wants to view may be different on different types of days. For example, on weekdays (Monday to Friday), the target object is often at work, and the push content he wants to view may be work-related; on weekends (Saturday and Sunday), when the target object does not need to work, the push content he wants to view may be related to his hobbies; on holidays, the target object may want to view push content related to holiday culture.
[0161] Weather characteristics refer to the characteristics of the weather in the current geographical area where the target object is viewing the pushed content. This is divided into sunny, rainy, cloudy, etc. The push content that the target object wants to view may be different in different weather conditions. For example, if the target object's current weather is sunny, the push content may be related tourist attractions that can be visited nearby on sunny days; if the target object's current weather is rainy, the push content may be indoor places that can be visited nearby on rainy days, etc.
[0162] The advantage of the above embodiment is that by using the target object characteristics, the characteristics of the content to be pushed, and the target scene characteristics to predict the target content to be pushed for the target object, the accuracy of the push processing is improved.
[0163] Detailed description of step 320
[0164] In step 320, multiple target push basic features are input into the probability prediction model to obtain multiple first probabilities of the target object completing multiple tasks for the content to be pushed. The target push basic features have been described in detail in the above step 310 and will not be repeated here.
[0165] The probability prediction model refers to a model structure that considers multiple target push basic features to predict the target object's completion of multiple tasks for the content to be pushed. The probability prediction model is a model structure built on a multi-gate expert mixture model network. Specifically, it means that multiple gated expert networks are used to combine and select features to improve the performance of the model. In this model, different expert sub-models can learn different feature representations, and then dynamically select and combine these features through a gating mechanism to adapt to different tasks and data distributions. This model is better suited for processing complex nonlinear relationships and multi-task learning.
[0166] like Figure 5 As shown, the probability prediction model 510 of the disclosed embodiment includes multiple first expert sub-models, multiple first gating nodes corresponding to multiple tasks, and multiple first prediction sub-models, each first gating node is connected to multiple first expert sub-models, the first gating node corresponding to the task is connected to the first prediction sub-model corresponding to the task, and the first prediction sub-model outputs the first probability corresponding to the task. For example, three target push basic features are target push basic feature T1, target push basic feature T2, and target push basic feature T3. After the three target push basic features are input into the probability prediction model 510, they are first vectorized and then into three target push basic feature vectors, namely target push basic feature vector TV1, target push basic feature vector TV2, and target push basic feature vector TV3. The three target push basic feature vectors are cascaded to obtain a first cascade vector. The first cascade vector passes through the first expert sub-model, the second gating node, and the first prediction sub-model in the probability prediction model 510 in sequence. If the tasks to be completed by the target object for the content to be pushed include task A and task B, the first prediction sub-model corresponding to task A finally outputs the first probability of task A, and the first prediction sub-model corresponding to task B finally outputs the first probability of task B.
[0167] Tasks refer to operations such as likes and comments on pushed content. The first probability refers to the prediction of the target object completing the corresponding task for the pushed content, such as 0.8, 0.95, etc. The larger the first probability, the higher the possibility that the target object will complete the corresponding task for the pushed content.
[0168] In one embodiment, Figure 6 As shown, step 320 includes:
[0169] Step 610: Push basic features based on multiple targets to obtain a first cascade vector;
[0170] Step 620: input the first cascade vector into a plurality of first expert sub-models respectively to obtain a plurality of first vectors;
[0171] Step 630: Using the first gating weights of the first expert sub-models and the first gating node corresponding to the task, calculate the weighted sum of the multiple first vectors to obtain the second vector corresponding to the task;
[0172] Step 640: Input the second vector into the first prediction sub-model corresponding to the task to obtain a first probability corresponding to the task.
[0173] The above steps 610 to 640 are described in detail below.
[0174] In step 610, the target push basic features have been described in detail in the above step 310, and will not be repeated here. Cascading refers to the process of connecting multiple vectors in a certain order to form a longer vector. The first cascade vector refers to a longer vector formed by connecting the vectors of multiple target push basic features in a certain order. The order of connecting the vectors of multiple target push basic features is not specifically limited here.
[0175] In one embodiment, Figure 7 As shown, step 610 includes:
[0176] Step 710: quantize multiple target push basic feature vectors into multiple target push basic feature vectors;
[0177] Step 720: Cascade multiple target push basic feature vectors into a first cascade vector.
[0178] In step 710, one target push basic feature corresponds to one target push basic feature vector. In one embodiment, the target push basic feature vector corresponding to the target push basic feature can be determined by looking up a table.
[0179] In another embodiment, if the target push basic feature is a character feature, the character feature is input into the embedding layer to obtain a target push basic feature vector corresponding to the character feature; if the target push basic feature is a numerical feature, the numerical feature is numerically mapped to obtain a target push basic feature vector corresponding to the numerical feature.
[0180] In this embodiment, the advantage of using different vectorization methods for different types of target push basic features is that the features can be vectorized according to the characteristics of different types of target push basic features, thereby ensuring the accuracy of the feature vectorization results.
[0181] A vector is an array of values in different dimensions, and is a point in a multidimensional space. The line segment between the point and the origin in the multidimensional coordinate system has a magnitude and a direction, which are the magnitude and direction of the vector. Each value in the vector is the point value projected on the corresponding coordinate axis in the multidimensional coordinate system, i.e., a vector element. A vector element can be a numerical value or a symbol. The vector elements of the target push basic feature vector in the disclosed embodiment are numerical values.
[0182] In one embodiment, the multiple target push basic features include target object features and to-be-pushed content features. Figure 4A As shown in , there are 5 target object features and 5 content features to be pushed, so there are 10 target push basic features. Fig. 8A As shown, 10 target push basic features correspond to 10 target push basic feature vectors. Specifically, after quantizing multiple target push basic feature vectors, the target push basic feature vector with the education level feature of "college" is [1,2,2,0,2,1,0,0], the target push basic feature vector with the working years feature of "15" is [5,8,6,0,2,1,0,0], and the target push basic feature vector with the working years feature of "lawyer" is "[10,20,30,40,68,50,60,10]". The target push basic feature vector with the hobby feature of "movie" is "[10,40,50,30,20,76,90,100]". The target push basic feature vector with the historical viewing content feature of "viewing content sequence for more than 10 seconds" is "[58,54,42,30,44,54,64,68]". The target push basic feature vector for content type feature "movie" is "[10,40,50,30,20,80,90,104]". The target push basic feature vector for content length feature "70 seconds" is "[5,6,8,9,10,13,5,32]". The target push basic feature vector for like quantity feature "1300" is "[505,76,68,97,10,13,55,320]". The target push basic feature vector for comment quantity feature "500" is "[203,33,53,66,9,9,31,84]". The target push basic feature vector for collection quantity feature "100" is "[99,23,24,18,6,8,20,54]".
[0183] In another embodiment, the multiple target push basic features also include target scene features. Figure 4B As shown in , there are 5 target object features, 5 content features to be pushed, and 3 target scene features, so there are 13 target push basic features. Figure 8BAs shown, the 13 target push basic features correspond to 13 target push basic feature vectors. Specifically, after quantizing multiple target push basic feature vectors, the target push basic feature vector with the education level feature of "college" is [1,2,2,0,2,1,0,0], the target push basic feature vector with the working years feature of "15" is [5,8,6,0,2,1,0,0], and the target push basic feature vector with the working years feature of "lawyer" is "[10,20,30,40,68,50,60,10]". The target push basic feature vector with the hobby feature of "movie" is "[10,40,50,30,20,76,90,100]". The target push basic feature vector with the historical viewing content feature of "viewing content sequence for more than 10 seconds" is "[58,54,42,30,44,54,64,68]". The target push basic feature vector for the geographical area feature "A place B city" is "[23,1,5,62,18,99,10,30]". The target push basic feature vector for the current input day type feature "weekday" is "[11,23,22,17,49,45,61,19]". The target push basic feature vector for the weather feature "sunny day" is "[10,20,43,55,18,9,1,32]". The target push basic feature vector for the content type feature "movie" is "[10,40,50,30,20,80,90,104]". The target push basic feature vector for the content length feature "70 seconds" is "[5,6,8,9,10,13,5,32]". The target push basic feature vector with the like quantity feature of "1300" is "[505,76,68,97,10,13,55,320]". The target push basic feature vector with the comment quantity feature of "500" is "[203,33,53,66,9,9,31,84]". The target push basic feature vector with the collection quantity feature of "100" is "[99,23,24,18,6,8,20,54]".
[0184] In one embodiment, among the multiple target push basic features, the target push basic feature vector of each target push basic feature is not necessarily of a predetermined dimension. For example, for the historical viewing content feature of "content sequence viewed for more than 10 seconds", the length of the content sequence is not fixed, and the dimension of the corresponding target push basic feature vector is uncertain. In order to alleviate the unstable effect caused by different dimensions, dimension adjustment is required.
[0185] To adjust the dimensions, such as Fig. 9 As shown, after step 710, according to an embodiment of the present disclosure, the push processing method further includes:
[0186] Step 910: Determine whether the dimension of the target push basic feature vector is greater than a first threshold;
[0187] Step 920: Perform pooling processing on the target push basic feature vector so that the dimension of the pooled target push basic feature vector is equal to the first threshold.
[0188] In step 910, the first threshold refers to a value that determines whether the target push basic feature vector needs to be pooled based on the dimension of the target push basic feature vector. If the dimension of the target push basic feature vector is greater than the first threshold, the target push basic feature vector needs to be pooled; if the dimension of the target push basic feature vector is less than or equal to the first threshold, the target push basic feature vector does not need to be pooled. For example, assume that the target push basic feature is a target object feature, and the target object feature is "a sequence of content viewed for more than 10 seconds". After vectorizing the "sequence of content viewed for more than 10 seconds", the target push basic feature vector is obtained as "[70,90,80,10,40,50,30,20,80,90,100,50]" with a dimension of 12. If the first threshold set in advance is 8, the target push basic feature vector with a dimension of 12 is determined to be a target push basic feature vector with a dimension greater than the first threshold.
[0189] In step 920, the first threshold also refers to the standard value of the dimension of the target push basic feature vector. The target push basic feature vector is pooled so that the dimension of the pooled target push basic feature vector is equal to the first threshold. For example, if the first threshold set in advance is 8, the target push basic feature vector with a dimension of 12 is determined to be the target push basic feature vector with a dimension greater than the first threshold. Therefore, the target push basic feature vector with a dimension of 12 is pooled into a target push basic feature vector with a dimension of 8.
[0190] Pooling, also known as pooling, is actually sampling in essence. Pooling pushes basic features to the input target, and can choose a certain way to reduce the dimension and compress it to speed up the operation. Pooling includes average pooling and maximum pooling.
[0191] In one example, if Fig. 10AAs shown in the figure, assuming the size of the pooling kernel is 5×1 and the stride is 1, we can get {[70,90,80,10,40], [90,80,10,40,50], [80,10,40,50], [10,40,50,30], [40,50,30,20,80], [50,30,20,80,90], [30,20,80,90,100], [20,80,90,100,50]} from [70,90,80,10,40,50]. There are a total of 8 pooling intermediate vectors.
[0192] If average pooling is used, the average value of these 8 pooled intermediate vectors needs to be calculated. Specifically, (70+90+80+10+40) / 5=58. (90+80+10+40+50) / 5=54. (80+10+40+50+30) / 5=42. (10+40+50+30+20) / 5=30. (40+50+30+20+80) / 5=44. (50+30+20+80+90) / 5=54. (30+20+80+90+100) / 5=64. (20+80+90+100+50) / 5]=68. Then the target push basic feature vector is obtained as "[58,54,42,30,44,54,64,68]".
[0193] In another example, if Fig. 10B As shown, Fig. 10A Different average pooling methods are used. Fig. 10B The maximum value pooling is used. Specifically, the maximum values are found from {[70,90,80,10,40], [90,80,10,40,50], [80,10,40,50,30], [10,40,50,30,20], [40,50,30,20,80], [50,30,20,80,90], [30,20,80,90,100], [20,80,90,100,50]} in turn, and 90, 90, 80, 50, 80, 90, 100, 100 are obtained. Then the target push basic feature vector is "[90,90,80,50,80,90,100,100]".
[0194] The advantage of the above-mentioned embodiment of step 910-step 920 is that the dimension of the target push basic feature vector is adjusted by using the pooling operation so that each target push basic feature vector has the same dimension, thereby being able to improve the accuracy of sorting the completion status of multiple tasks when considering the overall feature information of multiple target push basic features based on the cascade results of the target push basic feature vectors.
[0195] In step 720, a plurality of target push basic feature vectors are concatenated into a first concatenated vector.
[0196] The first cascade vector refers to the vector after cascading multiple target push basic feature vectors. Fig.11 As shown in the figure, there are 6 target push basic features, corresponding to which there are 6 target push basic feature vectors. The 6 target push basic feature vectors are {[1,2,2], [5,8,6], [10,20,30], [10,40,50], [58,54,42], [10,40,50]}, and the second cascade vector is [1,2,2,5,8,6,10,20,30,10,40,50,58,54,42,10,40,50].
[0197] The advantage of the embodiment of the above steps 710-720 is that after the multiple target push basic feature vectors are vectorized into multiple target push basic feature vectors, the feature independence of each target can be retained, and at the same time, their dimensions can be ensured to be consistent, which is convenient for subsequent processing and analysis. By cascading multiple target push basic feature vectors into a first cascade vector, the feature representations of multiple targets can be unified into one vector to obtain a richer feature representation, which helps to improve the performance and generalization ability of the model. At the same time, it can also better capture the association and interaction information between different targets, thereby improving the ranking accuracy of the probability prediction model.
[0198] In step 620, the first cascade vector is input into multiple first expert sub-models respectively to obtain multiple first vectors. Each first expert sub-model is a model built based on a deep neural network. Multiple first expert sub-models are used to analyze and calculate each dimension of the first cascade vector in order to balance the sharing and mutual exclusion between multiple tasks. After obtaining the first cascade vector, the first cascade vector is input into multiple first expert sub-models respectively. Each first expert sub-model will obtain a corresponding first vector based on the input first cascade vector. The number of first expert sub-models can be flexibly set according to actual needs, which is not specifically limited here and will not be repeated.
[0199] It should be noted that the first expert sub-model includes a first weight matrix, the number of columns of the first weight matrix is equal to the dimension of the first cascade vector, and the number of rows of the first weight matrix is equal to the dimension of the first vector. For example, if the dimension of the first cascade vector is 6 and the dimension of the first vector is 3, then the first weight matrix is a 3×6 matrix.
[0200] In one embodiment, the first cascade vector is input into a plurality of first expert sub-models to obtain a plurality of first vectors, including: multiplying the first weight matrix in the plurality of first expert sub-models by the transpose of the first cascade vector to obtain a plurality of first vectors. For example, the first weight matrix is It can be seen that the dimension of the first cascade vector is 6, and the dimension of the first vector is 3. If the first cascade vector is [1,6,1,2,3,5], then the first vector obtained by multiplying the first weight matrix and the transpose of the first cascade vector is:
[0201] In another embodiment, a target push basic feature is input into multiple first expert sub-models. In this case, the target push basic feature does not need to be cascaded. After passing through the first gating node and the first prediction sub-model, a first probability associated with task A and a first probability associated with task B for the content to be pushed are obtained.
[0202] In step 630, a weighted sum of multiple first vectors is calculated using the first gating weights of multiple first expert sub-models through the first gating node corresponding to the task to obtain a second vector corresponding to the task.
[0203] The first gating node refers to a node in the probability prediction model that is used to receive the first vector output by the corresponding first expert sub-model. Each task corresponds to a first gating node, and the first gating node is used to integrate the first vector output by the corresponding first expert sub-model into a high-order vector for the task. Therefore, after the first vectors output by multiple first expert models are input into the first gating node corresponding to the task, in each first gating node, the first gating weights of multiple first expert sub-models are used to calculate the weighted sum of multiple first vectors to obtain the second vector corresponding to the task.
[0204] The first gating weight refers to the weight used to control each first expert sub-model for input data in the probability prediction model. Since each first expert sub-model will learn different feature representations, the first gating weight is used to control the proportion of each first expert sub-model in processing input data. The role of the first gating weight is to dynamically select and allocate experts between different input data and tasks to improve the performance and generalization ability of the model. In a probability prediction model, the sum of the weights of multiple first gating weights contained is 1.
[0205] In one embodiment, a first gating node corresponding to the task is used to calculate a weighted sum of multiple first vectors using first gating weights of multiple first expert sub-models to obtain a second vector corresponding to the task, including:
[0206] Determine, through the first gating node, a first product vector of a second weight matrix and a transpose of the first cascade vector;
[0207] Exponentially normalizing the first product vector to obtain a first gating weight vector, where the first gating weight vector includes first gating weights of multiple first expert sub-models;
[0208] Use the first gating weights of multiple first expert sub-models in the first gating weight vector to calculate the weighted sum of multiple first vectors, and obtain the second vector corresponding to the task.
[0209] The first gating node includes a second weight matrix, which is a trainable parameter matrix. The number of columns of the second weight matrix is equal to the dimension of the first concatenated vector, and the number of rows of the second weight matrix is equal to the number of first expert sub-models. The dimension of the first concatenated vector is the sum of the dimensions of the vectorized results of all target push base features after concatenation. For example, if the number of first expert sub-models is 3 and the dimension of the first concatenated vector is 10, then the number of rows of the second weight matrix is 3 and the number of columns is 10, that is, the second weight matrix is a 3×10 matrix.
[0210] Therefore, the process of calculating the dimension of the first concatenated vector is shown in Formula 1:
[0211]
[0212] In Formula 1, d X represents the dimension of the first concatenated vector, X represents the first concatenated vector, and d n represents the dimension of the nth target push base feature vector, and N represents the number of target push base features concatenated in the first concatenated vector.
[0213] The first product vector refers to the vector result obtained by multiplying the second weight matrix by the transpose of the first concatenated vector. At this time, the first product vector is equivalent to a matrix vector of the number of first expert sub-models × the dimension of the first concatenated vector (i.e., the second weight matrix) multiplied by a column vector of the dimension of the first concatenated vector × 1 (i.e., the transpose of the first concatenated vector), so the first product vector is a column vector with a dimension of the number of first expert sub-models × 1. For example, as Fig.12 shown, the second weight matrix is It can be seen that the number of first expert sub-models is 3 and the dimension of the first concatenated vector is 6. If the first concatenated vector is [1, 6, 10, 2, 3, 5], then the first product vector obtained by multiplying the second weight matrix by the transpose of the first concatenated vector is:
[0214] The exponential normalization process is to limit each value in the result of the multiplication operation within a specific range, so that the values of the first gating weights corresponding to different first expert sub-models are mapped within the same range, making them comparable.
[0215] In one embodiment, the exponential normalization process uses a softmax function, which inputs the second weight matrix and the first product vector of the transpose of the first concatenated vector through vector multiplication, and outputs a first gating weight vector.
[0216] The first product vector includes multiple first vector elements, and a first vector element is a parameter used to determine the weight of a first expert sub-model. The softmax function exponentially normalizes each first vector element in the first product vector, and numerically concatenates the values of multiple first vector elements after exponential normalization to obtain a first gating weight vector. The first gating weight vector contains the first gating weights of multiple first expert sub-models. At this time, the values of multiple first gating weights are mapped in the same range, that is, the first gating weight is a value in the range of [0,1], and the sum of all first gating weights in the first gating weight vector is 1.
[0217] Therefore, the process of calculating the first gating weight vector in the probability prediction model is shown in Formula 2:
[0218] G t (X) = softmax(W t ·X T ) (Formula 2).
[0219] In Formula 2, W t Refers to the second weight matrix in the first gating node corresponding to the t-th task, t is a positive integer, t∈[1,T1], T1 represents the total number of tasks, W t The dimension M×d X , M refers to the number of the first expert sub-model, d X refers to the dimension of the first cascade vector, X refers to the first cascade vector, X T is the transposed vector of the first cascade vector, G t (X) refers to the first gating weight vector. Each row of the first gating weight vector represents the first gating weight of a first expert sub-model. Fig.12 As shown, the first product vector is determined as After that, the first product vector is subjected to exponential normalization, and the first gating weight vector obtained is In other words, the first gating weights of the three first expert sub-models included in the probability prediction model are 0.63, 0.17, and 0.20, respectively.
[0220] The second vector refers to the vector output by the first gating node corresponding to each task. Therefore, the process of calculating the second vector is shown in Formula 3:
[0221]
[0222] In Formula 3, X refers to the first cascade vector, G t (X) refers to the first gating weight vector of the first gating node corresponding to task t, [G t (X)] m represents the first gating weight corresponding to the mth first expert sub-model in the first gating weight vector, e m (X) refers to the first vector output by the mth first expert sub-model, Refers to the second vector output by the first gating node corresponding to task t. Therefore, the first vector of the first expert sub-model is multiplied by the corresponding first gating weight, and then the results of the multiplication of M (the number of first expert sub-models) vectors are added together to obtain the second vector corresponding to task t. Fig.12 As shown in FIG. 1 , if the probability prediction model includes three first expert sub-models, the first expert sub-model M1, the first expert sub-model M2, and the first expert sub-model M3, then the first gating weight vector is Then the first gating weight corresponding to the first expert sub-model M1 is 0.63, the first gating weight corresponding to the first expert sub-model M2 is 0.17, and the first gating weight corresponding to the first expert sub-model M3 is 0.20. If the first vector of the first expert sub-model M1 is (20, 45, 23) T , the first vector of the first expert sub-model M2 is (10,32,52) T , the first vector of the first expert sub-model M3 is (16,22,18) T At this time, the first gating weights of the multiple first expert sub-models are obtained by weighting the multiple first vectors: The second vector obtained is
[0223] The advantage of the above embodiment is that the first gating weights of each first expert sub-model are mapped within the same numerical range through exponential normalization processing, and the first vector of the first expert sub-model is weighted and summed based on the first gating weight vector, which reflects the different impacts of different first expert sub-models on the task, improves the accuracy of determining the second vector corresponding to the task, and thus improves the accuracy of ranking a specific task by the probability prediction model.
[0224] Step 640: Input the second vector into the first prediction sub-model corresponding to the task to obtain a first probability corresponding to the task.
[0225] The first prediction sub-model refers to a model structure that predicts the target object to complete a corresponding task for the content to be pushed based on the second vector. The first prediction sub-model can be a deep neural network model.
[0226] The first probability refers to the probability value of the target object completing the corresponding task for the pushed content determined based on the probability prediction model, such as 0.8, 0.95, etc. The larger the first probability, the greater the possibility that the target object will complete the corresponding task for the pushed content predicted based on the probability prediction model.
[0227] In one embodiment, Fig.13 As shown, step 640 includes:
[0228] Step 1310: input the second vector into the first prediction sub-model corresponding to the task to obtain a first result value corresponding to the task;
[0229] Step 1320: Apply a nonlinear activation function to the first result value to obtain a first probability.
[0230] In step 1310, the first result value refers to the value output after the second vector is input into the first prediction sub-model corresponding to the task. The first result value is a real scalar. A real scalar refers to a real number, which is a single value without a unit or direction. Real numbers are the set of all real numbers including integers, decimals, and irrational numbers. This means that a real scalar can be any real number.
[0231] Therefore, the process of calculating the first result value corresponding to the task is shown in Formula 4:
[0232]
[0233] In formula 4, It refers to the second vector output by the first gating node corresponding to task t, h t (·) refers to the function of the first prediction sub-model corresponding to task t to obtain the first result value, logit t Refers to the first result value corresponding to task t. For example, the second vector is: [0.8, 0.3, 230, 200]. After calculation by formula 4, the first result value corresponding to the task is 18.
[0234] In step 1320, after obtaining the first result value, the first result value is processed by a nonlinear activation function and can be converted into a first probability. The first result value is mapped to an output value between 0 and 1 by the nonlinear activation function, that is, the numerical range of the first probability is between 0 and 1, so that different first probabilities are comparable.
[0235] Therefore, if the sigmoid function is used to process the first result value with a nonlinear activation function, then the process of calculating the first probability corresponding to the task output by the first prediction sub-model is as shown in Formula 5:
[0236]
[0237] In formula 5, p t refers to the first probability corresponding to task t, refers to the second vector output by the first gating node corresponding to task t, θ t (·) refers to the function of the first prediction sub-model corresponding to task t to obtain the first probability, logit t refers to the first result value corresponding to task t, and sigmoid(·) refers to the nonlinear activation function.
[0238] The advantage of the embodiment of the above-mentioned steps 1310-1320 is that the first result value in the first prediction sub-model is processed by a nonlinear activation function to obtain the first probability, and the first result value corresponding to the task can be mapped to the same numerical range. The preference of the target object for the content to be pushed is analyzed through multiple tasks, so that the prediction result is more comprehensive, and the accuracy of determining the target content based on the first probability is improved, thereby improving the push accuracy.
[0239] The advantage of the embodiment of the above steps 610-640 is that by vectorizing multiple target push basic feature vectors into a first cascade vector and inputting them into multiple first expert sub-models respectively, a diversified feature representation and expert processing can be obtained, which helps to improve the performance and generalization ability of the model. Through the first gating node and the first gating weight, multiple first vectors can be dynamically selected and combined, so that the probability prediction model can dynamically adjust the contribution of experts according to different tasks, so as to better adapt to different input data and tasks. By predicting the second vector through the first prediction sub-model to obtain the first probability corresponding to the task, the prediction accuracy of the probability prediction model based on the first probability can be improved, thereby improving the push accuracy.
[0240] Detailed description of step 330
[0241] In step 330, multiple first probabilities are cascaded and input into a fusion model to obtain a first overall probability that the target object completes multiple tasks for the content to be pushed, and the content to be pushed is pushed to the target object based on the first overall probability.
[0242] The first overall probability refers to the probability value of the fusion model outputting the content to be pushed to the target object. The larger the first overall probability, the greater the possibility of pushing the content to be pushed to the target object.
[0243] The fusion model refers to a model structure that predicts the completion of a task by considering the ranking of a sample for a specific task and its ranking for all tasks as a whole. Since the ranking is more direct than the score, it can overcome the situation where the same score may have very different rankings, and the scores with large differences may have very small actual ranking differences. Therefore, after multiple first probabilities are cascaded and input into the fusion model, the first overall probability of the task obtained based on the fusion model can better reflect the difference and distinction between the ranking of a content to be pushed for a specific task and the ranking for all tasks as a whole, so as to improve the accuracy of pushing for the target object.
[0244] The disclosed embodiment also cascades a fusion model behind the probability prediction model, that is, the probability prediction model and the fusion model constitute an overall model for pushing target content to the target object. The specific training process of the probability prediction model and the fusion model in the overall model will be described in detail in the subsequent embodiments, and will not be described in detail here.
[0245] The fusion model is a model structure built based on a deep neural network. In one embodiment, Fig.14 As shown, the three target push basic features are target push basic feature T1, target push basic feature T2, and target push basic feature T3. After the three target push basic features are input into the probability prediction model 510, they are first vectorized and then into three target push basic feature vectors, namely target push basic feature vector TV1, target push basic feature vector TV2, and target push basic feature vector TV3. The three target push basic feature vectors are cascaded to obtain a first cascade vector. The first cascade vector passes through the first expert sub-model, the second gating node, and the first prediction sub-model in the probability prediction model 510 in sequence. If the tasks completed by the target object for the content to be pushed include task A and task B, the first prediction sub-model corresponding to task A finally outputs the first probability of task A, and the first prediction sub-model corresponding to task B finally outputs the first probability of task B. The first probability of task A and the first probability of task B are cascaded and input into the fusion model to obtain the first overall probability of the target object completing multiple tasks for the content to be pushed.
[0246] In one embodiment, Fig.15 As shown, in step 330, multiple first probabilities are cascaded and input into the fusion model to obtain a first overall probability that the target object completes multiple tasks for the content to be pushed, including:
[0247] Step 1510: cascade multiple first probabilities into a first probability vector, where the dimension of the first probability vector is equal to the number of tasks;
[0248] Step 1520: multiply the transpose of the first probability vector by the fusion weight vector to obtain a second result value;
[0249] Step 1530: Apply a nonlinear activation function to the second result value to obtain a first overall probability.
[0250] The above steps 1510 to 1530 are described in detail below.
[0251] In step 1510, the first probability vector refers to a vector obtained by concatenating the first probabilities corresponding to multiple tasks. If the first probability is a one-dimensional value, the dimension of the first probability vector is equal to the number of tasks. Fig.14 As shown in the figure, the first prediction sub-model corresponding to task A finally outputs the first probability of task A as 0.95, and the first prediction sub-model corresponding to task B finally outputs the first probability of task B as 0.8. Then, the first probability of task A 0.95 and the first probability of task B 0.8 are concatenated to obtain the first probability vector [0.95, 0.8], and the dimension of the first probability vector is 2.
[0252] In step 1520, the fusion model includes a fusion weight vector, and the dimension of the fusion weight vector is equal to the number of tasks. The fusion weight vector refers to a vector of weights that control the influence of the first probability corresponding to different tasks in the fusion model. The fusion weight vector is a matrix vector of the number of tasks × 1. Each row of the fusion weight vector represents the fusion weight corresponding to a task. For example, the fusion model includes task A and task B, and the fusion weight vector is Then, the fusion weight corresponding to task A is 0.7, and the fusion weight corresponding to task B is 0.3.
[0253] The second result value refers to the predicted values corresponding to the multiple tasks output after the first probability vector is input into the fusion model. The second result value is a real scalar. A real scalar is a real number, which is a single value without a unit or direction. Real numbers are the set of all real numbers including integers, decimals, and irrational numbers. This means that a real scalar can be any real number.
[0254] Therefore, the process of calculating the second result value corresponding to the task is shown in Formula 6:
[0255]
[0256] In Formula 6, It refers to the first probability vector cascaded from multiple first probabilities. is the transposed vector of the first probability vector, W p refers to the fusion weight vector, logit pRefers to the second result value corresponding to the first probability vector. For example, the first probability vector is: [0.95, 0.8]. After calculation by formula 6, the second result value corresponding to the overall number of tasks is 20.
[0257] In step 1530, after obtaining the second result value, the second result value is processed by a nonlinear activation function and can be converted into a first overall probability. Therefore, the second result value is mapped to an output value between 0 and 1 through the nonlinear activation function, that is, the numerical range of the first overall probability is between 0 and 1, so that different first overall probabilities are comparable.
[0258] Therefore, if the sigmoid function is used to process the second result value with a nonlinear activation function, then the process of calculating the first overall probability corresponding to the multiple tasks output by the fusion model is as shown in Formula 7:
[0259]
[0260] In Formula 7, p refers to the first overall probability of multiple tasks, refers to the first probability vector cascaded from multiple first probabilities, θ(·) refers to the function of the fusion model to obtain the first overall probability, logit p It refers to the second result value corresponding to the first probability vector, and sigmoid(·) refers to the nonlinear activation function.
[0261] The advantage of the embodiment of the above-mentioned steps 1510 to 1530 is that by processing the second result value in the fusion model through a nonlinear activation function to obtain the first overall probability, the second result values corresponding to multiple tasks can be mapped to the same numerical range, and the target object's preference for the content to be pushed can be analyzed through multiple tasks, so that the prediction result is more comprehensive, and the accuracy of determining the target content based on the first overall probability is improved, thereby improving the push accuracy.
[0262] In one embodiment, pushing the content to be pushed to the target object based on the first overall probability in step 330 includes:
[0263] sorting the plurality of contents to be pushed based on the first overall probability;
[0264] Based on the sorting, push the content to be pushed to the target object.
[0265] In one embodiment, the plurality of contents to be pushed are sorted based on the first overall probability, which means that the contents to be pushed are sorted in descending order based on the first overall probability. At this time, if the number of pushed contents to the target object is not limited, the contents to be pushed are pushed to the target object based on the sorting, that is, the target contents are pushed to the target object from high to low according to the numerical value of the first overall probability. Fig.16AAs shown, the target object features of the target object include: "university", "15 years of work experience", "lawyer", and "movie". The content features of the content R1 to be pushed in the content set to be pushed include: "picture", "1000 likes", and the corresponding first overall probability is 50%. The content features of the content R2 to be pushed in the content set to be pushed include: "movie", "length 70 seconds", "1300 likes", "500 comments", "100 collections", and the corresponding first overall probability is 90%. The content features of the content R3 to be pushed in the content set to be pushed include: "movie", "length 3600 seconds", "30 likes", and the corresponding first overall probability is 78%. The content features of the content R4 to be pushed in the content set to be pushed include: "picture", "30 likes", "10 comments", and the corresponding first overall probability is 20%. The content features of the content R5 to be pushed in the content set to be pushed include: "movie", "length 120 seconds", "3000 likes", "5000 comments", "2000 collections", and the corresponding first overall probability is 95%. In this way, the plurality of contents to be pushed are sorted based on the first overall probability, from high to low: content to be pushed R5> content to be pushed R2> content to be pushed R3> content to be pushed R1> content to be pushed R4. Therefore, the content to be pushed R5 can be used as the target content T1, the content to be pushed R2 can be used as the target content T2, the content to be pushed R3 can be used as the target content T3, the content to be pushed R1 can be used as the target content T4, and the content to be pushed R4 can be used as the target content T5, and the target contents in the target content set can be pushed to the target object in sequence.
[0266] In another embodiment, pushing the content to be pushed to the target object based on the ranking further includes: pushing the content to be pushed to the target object based on the ranking and a preset push quantity threshold. The preset push quantity threshold refers to the maximum value of the amount of content to be pushed to the target object. Fig. 16BAs shown, the target object features of the target object include: "university", "15 years of work experience", "lawyer", and "movie". The content features of the content R1 to be pushed in the content set to be pushed include: "picture", "1000 likes", and the corresponding first overall probability is 50%. The content features of the content R2 to be pushed in the content set to be pushed include: "movie", "length 70 seconds", "1300 likes", "500 comments", "100 collections", and the corresponding first overall probability is 90%. The content features of the content R3 to be pushed in the content set to be pushed include: "movie", "length 3600 seconds", "30 likes", and the corresponding first overall probability is 78%. The content features of the content R4 to be pushed in the content set to be pushed include: "picture", "30 likes", "10 comments", and the corresponding first overall probability is 20%. The content features of the content R5 to be pushed in the content set to be pushed include: "movie", "length 120 seconds", "3000 likes", "5000 comments", "2000 collections", and the corresponding first overall probability is 95%. In this way, the multiple first overall probabilities are arranged in descending order, and the result is 95%>90%>78%>66%>50%. If the preset push quantity threshold is 3, then the to-be-pushed content R5 corresponding to the first overall probability of 95% can be used as the target content T1, the to-be-pushed content R2 corresponding to the first overall probability of 90% can be used as the target content T2, and the to-be-pushed content R3 corresponding to the first overall probability of 78% can be used as the target content T1. Therefore, the target content T1, the target content T2, and the target content T3 form a target content set, and the target content in the target content set is pushed to the target object.
[0267] In another embodiment, pushing the content to be pushed to the target object based on the first overall probability also includes: pushing the content to be pushed to the target object based on the first overall probability and a preset probability threshold. The preset probability threshold refers to the minimum value of the first overall probability of the content to be pushed pushed to the target object. In this way, if the first overall probability of the content to be pushed is greater than or equal to the preset probability threshold, the content to be pushed is determined to be the target content. Then, based on the first overall probability, the target content in the target content set is pushed to the target object in order from high to low.
[0268] The advantage of the above embodiment is that the first overall probability is the push probability determined by the ranking of the corresponding content to be pushed in all pushed content, which can overcome the situation that the same score may have a large difference in ranking, and the score with a large difference may have a small difference in actual ranking. Therefore, pushing the content to be pushed for the target object based on the first overall probability can better reflect the difference and distinction between the ranking of a sample for a specific task and the ranking of all tasks as a whole, improve the ranking accuracy of the overall model including the probability prediction model and the trained fusion model, and thus improve the push accuracy.
[0269] The training process of the fusion model in the embodiment of the present disclosure
[0270] In one embodiment, Fig.14 As shown in FIG. 5 , the probability prediction model and the fusion model constitute an overall model for pushing target content to a target object. The model structure of the probability prediction model 510 is shown in FIG. Figure 5 The model structure of the fusion model is as follows: Fig.14 The specific training process of the probability prediction model in the overall model is described in detail in the subsequent embodiments, and will not be described in detail here.
[0271] In one embodiment, Fig.17 As shown, the fusion model is trained in the following way:
[0272] Step 1710: Acquire a push basic feature sample set, where each push basic feature sample in the push basic feature sample set includes a sample object feature and a sample content feature;
[0273] Step 1720: input the pushed basic feature sample set into the probability prediction model, and cascade the output of the probability prediction model into the fusion model;
[0274] Step 1730: For each task, obtain the probability ranking of each sample object that pushes the basic feature sample to complete the task for the sample content, which is output by the probability prediction model, as the first ranking;
[0275] Step 1740: Obtain the overall probability ranking of the sample objects of each pushed basic feature sample output by the fusion model for completing multiple tasks for the sample content as the second ranking;
[0276] Step 1750: Calculate a first loss function based on the first ranking of the pushed basic feature sample in the first ranking corresponding to the task and the second ranking in the second ranking;
[0277] Step 1760: Train the fusion model based on the first loss function.
[0278] The above steps 1710 to 1760 are described in detail below.
[0279] In step 1710, the push basic feature sample set refers to a set for storing multiple push basic feature samples. The push basic feature samples are the same as the target push basic features, except that they are all used as training samples. In order to save space, they are not described here.
[0280] The basic feature samples pushed also include sample scene features, which are similar to the target scene features mentioned above, except that they are all used as training samples. In order to save space, they will not be repeated here. Therefore, obtaining the scene features of the sample objects can improve the ranking accuracy of the overall model including the probability prediction model and the trained fusion model, thereby improving the push accuracy.
[0281] In step 1720, the process of pushing the basic feature sample set into the probability prediction model is the same as the process of step 320 above, except that the basic feature sample set is pushed into the probability prediction model here as a training for the model. In order to save space, it is not repeated here. The process of cascading the output of the probability prediction model and inputting it into the fusion model is the same as the process of step 330 above, except that the output result after pushing the basic feature sample set into the probability prediction model is cascaded and input into the fusion model here as a training for the model. In order to save space, it is not repeated here.
[0282] In step 1730, the probability output by the probability prediction model of each sample object that pushes the basic feature sample completing the task for the sample content is the same as the first probability in the above step 320, except that they are all used as training samples here, and will not be repeated here to save space.
[0283] The first ranking refers to the ranking result of multiple probabilities of completing a task for multiple pushed basic feature samples predicted by the probability prediction model. The number of first rankings is equal to the number of tasks. At this time, a pushed basic feature sample has a first ranking in the first ranking corresponding to a task. Therefore, the embodiment of the present disclosure can obtain multiple first rankings of a sample for multiple tasks predicted by the probability prediction model, that is, it can determine multiple first rankings, so as to better reflect the ranking of a sample for a specific task.
[0284] Take a task A as an example. Fig.18As shown, the sample object features of the sample object include: "university", "10 years of work experience", "lawyer", and "music". The content features of the sample content S1 to be pushed in the sample content set to be pushed include: "picture", "1000 likes", and the corresponding probability of completing task A is 52%. The content features of the sample content S2 to be pushed in the sample content set to be pushed include: "movie", "length 70 seconds", "1300 likes", "500 comments", "100 collections", and the corresponding probability of completing task A is 89%. The content features of the sample content S3 to be pushed in the sample content set to be pushed include: "movie", "length 3600 seconds", "30 likes", and the corresponding probability of completing task A is 75%. The content features of the sample content S4 to be pushed in the sample content set to be pushed include: "picture", "30 likes", "10 comments", and the corresponding probability of completing task A is 25%. The content features of the sample content S5 to be pushed in the sample content set to be pushed include: "movie", "length 120 seconds", "3000 likes", "5000 comments", "2000 collections", and the corresponding probability of completing task A is 98%. Because 98%>89%>75%>52%>25%, then the ranking content in the first sorting corresponding to task A is from high to low: sample content S5 to be pushed (first ranking is 1), sample content S2 to be pushed (first ranking is 2), sample content S3 to be pushed (first ranking is 3), sample content S1 to be pushed (first ranking is 4), sample content S4 to be pushed (first ranking is 5). At this time, sample content S5 to be pushed (first ranking is 1) has the highest ranking in the first sorting of task A.
[0285] In step 1740, the overall probability of the sample objects of each pushed basic feature sample completing multiple tasks for the sample content output by the fusion model is the same as the first overall probability in the above step 330, except that they are all used as training samples here, and will not be repeated here to save space.
[0286] The second ranking refers to the ranking result of multiple overall probabilities of completing multiple tasks for multiple pushed basic feature samples predicted by the fusion model. For one training, the number of second rankings is 1. At this time, a pushed basic feature sample has a second ranking in the second rankings corresponding to multiple tasks. Therefore, the embodiment of the present disclosure can obtain a second ranking of a sample for multiple tasks predicted by the fusion model, that is, a second ranking can be determined, so as to better reflect the ranking of a sample for all tasks as a whole.
[0287] like Fig.19As shown, the sample object features of the sample object include: "university", "10 years of work experience", "lawyer", and "music". The content features of the sample content S1 to be pushed in the sample content set to be pushed include: "picture", "1000 likes", and the corresponding overall probability of completing multiple tasks is 40%. The content features of the sample content S2 to be pushed in the sample content set to be pushed include: "movie", "length 70 seconds", "1300 likes", "500 comments", "100 collections", and the corresponding overall probability of completing multiple tasks is 921%. The content features of the sample content S3 to be pushed in the sample content set to be pushed include: "movie", "length 3600 seconds", "30 likes", and the corresponding overall probability of completing multiple tasks is 78%. The content features of the sample content S4 to be pushed in the sample content set to be pushed include: "picture", "30 likes", "10 comments", and the corresponding overall probability of completing multiple tasks is 45%. The content features of the sample content S5 to be pushed in the sample content set to be pushed include: "movie", "length 120 seconds", "3000 likes", "5000 comments", "2000 collections", and the corresponding overall probability of completing multiple tasks is 95%. Because 95%>92%>78%>45%>40%, then the rankings in the second sorting are from high to low: sample content S5 to be pushed (second ranking is 1), sample content S2 to be pushed (second ranking is 2), sample content S3 to be pushed (second ranking is 3), sample content S4 to be pushed (second ranking is 4), sample content S1 to be pushed (second ranking is 5). At this time, the sample content S5 to be pushed (second ranking is 1) has the highest ranking in the second sorting.
[0288] In the embodiment of the present disclosure, when training the model, for a sample, instead of processing the scores corresponding to each task and sorting them according to the processing results, the ranking of the sample in each task among all samples is processed and sorted according to the processing results. Since the ranking is more direct than the score, it can overcome the situation that the same score may have a large difference in ranking, and the scores with a large difference may have a small difference in actual ranking, so as to improve the accuracy of sorting.
[0289] In step 1750 , the first loss function refers to a loss function of multiple tasks determined based on the first ranking of the pushed basic feature sample in the first ranking corresponding to the task and the second ranking in the second ranking.
[0290] The disclosed embodiment constructs a first loss function through the first rank and the second rank, and trains the fusion model based on the first loss function, which can better reflect the difference and distinction between the ranking rank of a sample for a specific task and the ranking rank for all tasks as a whole, and improves the ranking accuracy of the overall model including the probability prediction model and the trained fusion model, thereby improving the push accuracy.
[0291] In one embodiment, Fig. 20 As shown, step 1750 includes:
[0292] Step 2010: Calculate a first sub-loss function corresponding to the task based on the first ranking of the pushed basic feature sample in the first ranking corresponding to the task and the second ranking in the second ranking;
[0293] Step 2020: Determine a first weight corresponding to the task based on the first ranking and the second ranking;
[0294] Step 2030: Based on the first weights corresponding to the multiple tasks, a weighted sum is calculated for the first sub-loss functions corresponding to the tasks to obtain a first loss function.
[0295] In step 2010, if the pushed basic feature sample has a first ranking in the first ranking corresponding to a task, the number of the first rankings corresponding to the pushed basic feature samples is the same as the number of tasks. The first sub-loss function refers to the loss function of a single task determined based on the first ranking in the first ranking corresponding to the task and the second ranking in the second ranking.
[0296] In one embodiment, Fig.21 As shown, step 2010 includes:
[0297] Step 2110: Set a first label corresponding to the task, wherein if the first rank in the first ranking corresponding to the task is before the second rank, the first label is the first value; otherwise, the first label is the second value;
[0298] Step 2120, obtaining the overall probability of the fusion model output;
[0299] Step 2130: Calculate a first sub-loss function corresponding to the task based on the first label and the overall probability.
[0300] In step 2110, the first label refers to the label value corresponding to the result of comparing the first rank and the second rank in the first ranking corresponding to the task. The first value refers to the numerical value set corresponding to the first label when the first rank in the first ranking corresponding to the task is before the second rank, such as 1. The first value refers to the numerical value set corresponding to the first label when the first rank in the first ranking corresponding to the task is after the second rank (including the first rank and the second rank in the first ranking are equal), such as 0. The first value and the second value are not equal. The first value and the second value are pre-set values and are not specifically limited here.
[0301] Therefore, the first rank in the first sorting corresponding to the task and the second rank in the second sorting, the process of determining the first label corresponding to the task is shown in Formula 8:
[0302]
[0303] In formula 8, represents the first label corresponding to task t, r t represents the first rank in the first ranking corresponding to task t, and r represents the second rank in the second ranking. t -r)<0 means that the first task in the first ranking is before the second task, then the first value corresponding to the first label is 1; (r t -r)≥0 means that the first ranking in the first ranking corresponding to the task is after the second ranking or the first ranking in the first ranking corresponding to the task is the same as the second ranking, then the second value corresponding to the first label is 0. For example, if the first ranking of a push basic feature sample in the first ranking corresponding to the "Like" task is 2nd, and the first ranking of the push basic feature sample in the first ranking corresponding to the entire push basic feature sample set is 3rd, then, combined with the above formula 8, the first ranking in the first ranking corresponding to the "Like" task is before the second ranking, then the first label corresponding to the "Like" task is 1.
[0304] In step 2120, the overall probability output by the fusion model is the same as the first overall probability in the above step 330, except that they are all used as training samples here, and will not be repeated here to save space.
[0305] In step 2130, after obtaining the overall probability of the fusion model output, the process of calculating the first sub-loss function corresponding to the task is shown in Formula 9:
[0306]
[0307] In Formula 9, represents the first label corresponding to task t, represents the first sub-loss function corresponding to task t, and p represents the overall probability of the fusion model output. For example, if a push basic feature sample ranks second in the first ranking corresponding to the "Like" task, and the push basic feature sample ranks third in the first ranking corresponding to the entire push basic feature sample set, then, combined with the above formula 8, the first ranking in the first ranking corresponding to the "Like" task is before the second ranking, then the first label corresponding to the "Like" task is 1, that is, If the overall probability of the push basic feature sample in the fusion model output is 0.8, combined with Formula 9, the first sub-loss function corresponding to the "Like" task = -log0.8≈0.097.
[0308] The advantage of the embodiment of the above-mentioned steps 2110 to 2130 is that the first label is determined based on the ranking, and the first loss function corresponding to the task is calculated based on the first label and the overall probability, which can better reflect the difference and discrimination between the ranking of a sample for a specific task and the ranking of all tasks as a whole, thereby improving the ranking accuracy of the trained overall model.
[0309] In step 220, the first weight refers to the weight value corresponding to the task in the first loss function determined based on the result of comparing the first ranking and the second ranking in the first ranking corresponding to the task. The sum of the first weight values of all tasks is 1.
[0310] In one embodiment, Fig. 22 As shown, step 2020 includes:
[0311] Step 2210: Based on the first ranking, determine the first normalized cumulative loss gain corresponding to the task;
[0312] Step 2220: Determine a second normalized discounted cumulative gain based on the second ranking;
[0313] Step 2230: The absolute value of the difference between the first normalized discounted cumulative gain and the second normalized discounted cumulative gain corresponding to the task is used as the first weight corresponding to the task.
[0314] In step 2210, the first normalized discounted cumulative gain refers to the gain value determined in the indicator function corresponding to the normalized discounted cumulative gain based on the first rank. Normalized Discounted Cumulative Gain (NDCG) is an indicator used to evaluate ranking quality in the field of information retrieval. It combines relevance and ranking performance to measure the quality of the returned ranked result list. NDCG can normalize the results of the discounted cumulative gain (DCG) for comparison across different queries or different data sets. The value of NDCG is between 0 and 1, where 1 represents the best ranking and 0 represents the worst ranking. Therefore, in information retrieval and recommendation systems, the use of NDCG can better evaluate the performance of the ranking algorithm in order to optimize the quality of search results or recommendation lists.
[0315] In one embodiment, Fig.23 As shown, step 2210 includes:
[0316] Step 2310, obtaining a distribution control coefficient;
[0317] Step 2320: Take the logarithm of the sum of the product of the distribution control coefficient and the first rank and the predetermined constant to obtain a first intermediate value;
[0318] Step 2330: Based on the logarithm of the predetermined constant and the first intermediate value, obtain a first normalized discounted cumulative gain corresponding to the task.
[0319] The distribution control coefficient refers to a value used to control the distribution of NDCG. The distribution control coefficient is a given global parameter greater than zero, such as 0.5, 2, etc., and is not specifically limited here.
[0320] The predetermined constant is a value preset for calculating the first intermediate value. The first constant can be 2, 3, etc., and is not specifically limited here.
[0321] Therefore, based on the first place, the process of determining the first normalized cumulative loss gain corresponding to the task is shown in formula 10:
[0322]
[0323] In formula 10, NDCG(·) represents the indicator function corresponding to the normalized discounted cumulative gain, r t represents the first rank in the first ranking corresponding to task t, λ represents the distribution control coefficient, 2 represents a predetermined constant (the predetermined constant can be adjusted according to actual needs), log(2+λ·r t ) represents the first intermediate value obtained, NDCG(r t) represents the first normalized discounted cumulative gain corresponding to task t.
[0324] In step 2220, the second normalized depreciation cumulative gain refers to the gain value determined in the indicator function corresponding to the normalized depreciation cumulative gain based on the second rank. The calculation process of the second normalized depreciation cumulative gain is the same as the process of calculating the first normalized depreciation cumulative gain in step 2210 above, except that the first rank is replaced by the second rank. In order to save space, it will not be repeated here.
[0325] In step 2230, the absolute value of the difference between the first normalized discounted cumulative gain and the second normalized discounted cumulative gain corresponding to the task is used as the first weight corresponding to the task.
[0326] Therefore, the process of calculating the first weight corresponding to the task is shown in formula 11:
[0327]
[0328] In formula 11, represents the first weight corresponding to task t, NDCG(r t ) represents the first normalized discounted cumulative gain corresponding to task t, and NDCG(r) represents the second normalized discounted cumulative gain.
[0329] The advantage of the embodiment of the above-mentioned steps 2210-2230 is that the first normalized cumulative loss gain and the second normalized cumulative loss gain are determined based on the ranking, and the absolute value of the difference between the first normalized cumulative loss gain and the second normalized cumulative loss gain is used as the first weight corresponding to the task. In other words, the first weight is also determined based on the ranking. The present application, combined with the NDCG method, can focus more on aligning the content with a high ranking, and further increase the distinction between samples with a high ranking and samples with a low ranking, that is, by focusing on the influence of samples with a high ranking during model training, the training accuracy of the model can be effectively improved to improve the push accuracy. Therefore, the present application can better reflect the difference and distinction between the ranking ranking of a sample for a specific task and the ranking ranking for all tasks as a whole, improve the ranking accuracy of the overall model including the probability prediction model and the trained fusion model, and thus improve the push accuracy.
[0330] In one embodiment, in step 2030, based on the first weights corresponding to the multiple tasks, a weighted sum is calculated for the first sub-loss functions corresponding to the tasks to obtain a first loss function.
[0331] Therefore, the process of calculating the first loss function is shown in Formula 12:
[0332]
[0333] In formula 12, represents the first loss function of a single sample, represents the first weight corresponding to task t, represents the first sub-loss function corresponding to task t, and T1 represents the total number of tasks.
[0334] In another embodiment, a first loss function is calculated based on the first ranking of the pushed basic feature sample in the first ranking corresponding to the task, and the second ranking in the second ranking, including: calculating a first sub-loss function corresponding to the task based on the first ranking of the pushed basic feature sample in the first ranking corresponding to the task, and the second ranking in the second ranking; taking the weighted sum of the first sub-loss functions corresponding to the task based on a pre-set task weight to obtain a first loss function.
[0335] The advantage of the embodiment of the above-mentioned steps 2010 to 2030 is that, based on multiple first rankings and second rankings of multiple tasks, the first weight and the first sub-loss function in the first loss function are respectively constructed, which can deeply reflect the difference and distinction between the ranking ranking of a sample for a specific task and the ranking ranking for all tasks as a whole, thereby improving the ranking accuracy of the overall model including the probability prediction model and the trained fusion model, thereby improving the push accuracy.
[0336] In step 1760, in order to prevent the interference of the fusion model training on the probability prediction model training and to increase the influence of the top ranked samples on the model training, the present application combines the stop gradient method in the back propagation training of the fusion model, that is, based on the first loss function, only trains the fusion model. In deep learning, the stop gradient method refers to stopping the propagation of the gradient during the back propagation process, so that the gradient of some parts will not be updated to control the propagation path of the gradient.
[0337] In one embodiment, Fig.24 As shown, step 1760 includes:
[0338] Step 2410, calculating a second loss function of the probability prediction model for the pushed basic feature sample;
[0339] Step 2420: determine a single sample loss function based on the first loss function and the second loss function;
[0340] Step 2430: train the fusion model based on the single sample loss function.
[0341] In step 2410, the second loss function refers to a loss function calculated based on the probability output of the probability prediction model and the probability that the sample object that pushes the basic feature sample completes the corresponding task for the sample content.
[0342] In one embodiment, Fig.25 As shown, step 2410 includes:
[0343] Step 2510: Obtain multiple third probabilities output by the probability prediction model, that the sample object in the push basic feature sample completes multiple tasks for the sample content;
[0344] Step 2520: Calculate a second loss function based on multiple second labels corresponding to multiple tasks and multiple third probabilities.
[0345] In step 2510, the third probability is the same as the first probability in step 320, except that they are both used as training samples. To save space, they will not be described here.
[0346] In step 2520, the push basic feature sample has multiple second tags corresponding to multiple tasks. That is, for one push basic feature sample, one task corresponds to one second tag. The second tag indicates whether the sample object has completed the task for the sample content. For example, in the task of liking the sample, the second tag corresponding to the task includes {Content A: 1, Content B; 0}, where the second tag is 1 for task completion, and the second tag is 0 for task incomplete. That is, in this example, the sample object has completed the task of liking content A, but has not completed the task of liking content B.
[0347] In one embodiment, a second loss function is calculated based on multiple second labels corresponding to multiple tasks and multiple third probabilities, including: calculating a second sub-loss function based on the second labels corresponding to the tasks and the third probabilities corresponding to the tasks; calculating a second loss function based on the second sub-loss function.
[0348] Therefore, the process of calculating the second sub-loss function is shown in Formula 13:
[0349] L t =-y t log p t -(1-y t )log(1-p t ) (Formula 13).
[0350] In Formula 13, L t is the second sub-loss function of task t, y t is the second label of the basic feature sample pushed under task t, p tIt refers to the third probability of task t. If the second label indicates that the task is completed, then the second label is 1, otherwise, the second label is 0.
[0351] In one embodiment, based on the second sub-loss function, calculating the second loss function includes: performing sum calculation on the second sub-loss function, and the result of the sum calculation is the second loss function. At this time, the process of calculating the second loss function is shown in Formula 14:
[0352]
[0353] In Formula 14, L represents the second loss function, L t is the second sub-loss function of task t, and T1 represents the total number of tasks.
[0354] In another embodiment, based on the second sub-loss function, calculating the second loss function includes: performing a weighted sum calculation on the second sub-loss function and the weight of the second sub-loss function, and the result of the weighted sum calculation is the second loss function. In this case, the process of calculating the second loss function is shown in Formula 15:
[0355]
[0356] In Formula 15, L represents the second loss function, L t is the second sub-loss function of task t, T1 represents the total number of tasks, and w′ t represents the weight of the second sub-loss function. The weight of the second sub-loss function can be flexibly set according to actual needs and is not specifically limited here. The sum of the weights of all second sub-loss functions of a sample is 1.
[0357] The advantage of the above embodiment is that by constructing the second sub-loss function of the task to determine the second loss functions of multiple tasks, and training the fusion model based on the second loss function and the first loss function, it can better reflect the difference and distinction between the ranking ranking of a sample for a specific task and the ranking ranking for all tasks as a whole, thereby improving the ranking accuracy of the overall model including the probability prediction model and the trained fusion model, thereby improving the push accuracy.
[0358] In step 2420, the single sample loss function refers to a function determined based on a first loss function and a second loss function corresponding to a pushed basic feature sample.
[0359] Therefore, based on the first loss function and the second loss function, the process of calculating the single sample loss function is shown in Formula 16:
[0360]
[0361] In formula 16, refers to the single sample loss function, refers to the first loss function, and L refers to the second loss function.
[0362] It should be noted that the above formula 16 is not unique. For example, the first loss function and the second loss function in formula 16 may be weighted to obtain the final single sample loss function. This embodiment does not specifically limit this.
[0363] In step 2430, after obtaining the single sample loss function, the fusion model is trained based on the single sample loss function, that is, the model parameters of the fusion model are adjusted.
[0364] In one embodiment, based on the single sample loss function, training the fusion model includes: summing each single sample loss function to determine the multi-sample loss function; and training the fusion model based on the multi-sample loss function. The purpose of the training is to reduce the multi-sample loss function. The reduction of the multi-sample loss function means that the ranking accuracy of the overall model including the probability prediction model and the trained fusion model is higher.
[0365] In the disclosed embodiment, the advantage of determining the multi-sample loss function based on the sum of each single-sample loss function is that the single-sample loss function corresponding to each pushed basic feature sample makes the same contribution to determining the multi-sample loss function of the pushed basic feature sample set, thereby improving fairness when training the model.
[0366] In another embodiment, Fig.26 As shown, step 2430 includes:
[0367] Step 2610: Obtain the sample weight of each pushed basic feature sample;
[0368] Step 2620: Calculate the weighted sum of the single sample loss function of each pushed basic feature sample using the sample weight to obtain a multi-sample loss function;
[0369] Step 2630: Use the multi-sample loss function to train the fusion model.
[0370] Sample weight refers to the importance of each push basic feature sample in the push basic feature sample set, that is, the contribution of the corresponding sample to the fusion model training. Sample weight can be adjusted according to the specific algorithm or task setting. If the push basic feature sample is more important, the value of the sample weight setting corresponding to the push basic feature sample is larger. For example, the push basic feature samples in the push basic feature sample set include training samples A to training samples E, the loss function of training sample A is 0.21, the loss function of training sample B is 0.05, the loss function of training sample C is 0.10, the loss function of training sample D is 0.23, and the loss function of training sample E is 0.21. If the sample weight corresponding to training sample A is 0.4, the sample weight corresponding to training sample B is 0.1, the sample weight corresponding to training sample C is 0.05, the sample weight corresponding to training sample D is 0.15, and the sample weight corresponding to training sample E is 0.3. Then, using the sample weights, we calculate the weighted sum of the single sample loss functions of each pushed basic feature sample, and the resulting multi-sample loss function is: 0.4*0.21+0.1*0.05+0.05*0.1+0.15*0.23+0.3*0.21≈0.19.
[0371] In the disclosed embodiment, a fusion model is trained using a multi-sample loss function. By updating the model parameters in the fusion model, the multi-sample loss function determined by the sample weights obtained based on the updated fusion model is reduced. Specifically, a loss threshold can be set in advance. The loss threshold refers to the maximum critical value at which the multi-sample loss function meets the accuracy requirement. Therefore, if the multi-sample loss function is less than the loss threshold, the training process ends. When the multi-sample loss function is greater than or equal to the loss threshold, the model parameters of the fusion model are adjusted until the multi-sample loss function is less than the loss threshold.
[0372] The advantage of the above embodiment is that by weighting each single sample loss function and determining a multi-sample loss function, different sample weights can be set for different training samples according to the actual application scenario, thereby improving the flexibility of the training model. Therefore, by training the fusion model with the constructed diversity loss function, it is possible to better reflect the difference and discrimination between the ranking of a sample for a specific task and the ranking of all tasks as a whole, thereby improving the ranking accuracy of the trained overall model.
[0373] It should be noted that after determining the second loss function, the probability prediction model can be trained based on the second loss function to improve the probability prediction model's prediction accuracy of the first probability of the target object completing the task for the content to be pushed, thereby improving the push accuracy.
[0374] The advantage of the embodiment of the above steps 1710 to 1760 is that a fusion model is cascaded behind the probability prediction model. Not only can multiple first rankings of a sample for multiple tasks predicted by the probability prediction model be obtained, but also the overall second ranking of the sample predicted by the fusion model can be obtained. Therefore, by constructing the first loss function through the first ranking and the second ranking, and training the fusion model based on the first loss function, it is possible to better reflect the difference and distinction between the ranking ranking of a sample for a specific task and the ranking ranking for all tasks as a whole, thereby improving the ranking accuracy of the overall model including the probability prediction model and the trained fusion model, thereby improving the push accuracy.
[0375] Detailed diagram of the implementation of the push processing method of the disclosed embodiment
[0376] Refer to the following Fig. 27 , the implementation details of the push processing method of the embodiment of the present disclosure are explained in detail and by way of example.
[0377] In step 2710, a plurality of target push basic features are obtained, wherein the plurality of target push basic features include target object features and features of content to be pushed.
[0378] In step 2720, multiple target push basic feature vectors are quantized into multiple target push basic feature vectors; it is determined that the dimension of the target push basic feature vector is greater than the first threshold; the target push basic feature vector is pooled so that the dimension of the pooled target push basic feature vector is equal to the first threshold; the multiple target push basic feature vectors are cascaded into a first cascade vector; the first weight matrix in the multiple first expert sub-models is multiplied by the transpose of the first cascade vector to obtain multiple first vectors; through the first gating node, the second weight matrix and the first product vector of the transpose of the first cascade vector are determined; the first product vector is exponentially normalized to obtain a first gating weight vector, and the first gating weight vector includes the first gating weights of the multiple first expert sub-models; the first gating weights of the multiple first expert sub-models in the first gating weight vector are used to calculate the weighted sum of the multiple first vectors to obtain the second vector corresponding to the task; the second vector is input into the first prediction sub-model corresponding to the task to obtain the first result value corresponding to the task; the first result value is exponentially normalized to obtain the first probability.
[0379] In step 2730, multiple first probabilities are cascaded into a first probability vector, and the dimension of the first probability vector is equal to the number of tasks; the first probability vector is multiplied by the transpose of the fusion weight vector to obtain a second result value; the second result value is exponentially normalized to obtain a first overall probability, and the content to be pushed is pushed to the target object based on the first overall probability, wherein the fusion model is trained based on a first loss function, and the first loss function is calculated based on the first ranking of the push basic feature samples in the push basic feature sample set in the first sorting and the second ranking in the second sorting, the first sorting is for each task, after multiple push basic feature samples of the push basic feature sample set are input into the probability prediction model, the output of the probability prediction model corresponding to the task is sorted, and the second sorting is the sorting of the output of the fusion model after multiple push basic feature samples of the push basic feature sample set are input into the probability prediction model.
[0380] Description of the apparatus and device of the present disclosure
[0381] It is to be understood that, although the steps in the above-mentioned flowcharts are sequentially displayed according to the characterization of arrows, these steps are not necessarily executed in sequence according to the order of arrow characterization. Unless there is a clear description in the present embodiment, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the above-mentioned flowcharts can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of the steps or stages in other steps.
[0382] It should be noted that in each specific implementation of the present application, when it comes to the need to perform relevant processing based on data related to the characteristics of the target content such as target content attribute information or attribute information sets, the permission or consent of the target content will be obtained first, and the collection, use and processing of these data will comply with relevant laws, regulations and standards. In addition, when the embodiment of the present application needs to obtain the attribute information of the target content, it will obtain the separate permission or separate consent of the target content through a pop-up window or jump to a confirmation page. After clearly obtaining the separate permission or separate consent of the target content, the necessary target content-related data used to enable the normal operation of the embodiment of the present application will be obtained.
[0383] Fig.28 The structure diagram of the push processing device 2800 provided in the embodiment of the present disclosure. The push processing device 2800 includes:
[0384] The acquisition unit 2810 is used to acquire a plurality of target push basic features, wherein the plurality of target push basic features include target object features and to-be-pushed content features;
[0385] An input unit 2820 is used to input a plurality of target push basic features into a probability prediction model to obtain a plurality of first probabilities of the target object completing a plurality of tasks for the content to be pushed;
[0386] The push unit 2830 is used to input the cascaded multiple first probabilities into the fusion model to obtain the first overall probability that the target object completes multiple tasks for the content to be pushed, and pushes the content to be pushed to the target object based on the first overall probability, wherein the fusion model is trained based on a first loss function, and the first loss function is calculated based on the first ranking of the push basic feature samples in the push basic feature sample set in the first sorting and the second ranking in the second sorting. The first sorting is for each task, after multiple push basic feature samples of the push basic feature sample set are input into the probability prediction model, the output of the probability prediction model corresponding to the task is sorted, and the second sorting is the sorting of the output of the fusion model after multiple push basic feature samples of the push basic feature sample set are input into the probability prediction model.
[0387] Optionally, the ensemble model is trained by the following process:
[0388] Acquire a push basic feature sample set, where each push basic feature sample in the push basic feature sample set includes a sample object feature and a sample content feature;
[0389] The push basic feature sample set is input into the probability prediction model, and the output of the probability prediction model is cascaded and input into the fusion model;
[0390] For each task, the probability ranking of the sample objects of each pushed basic feature sample in terms of completing the task on the sample content output by the probability prediction model is obtained as the first ranking;
[0391] Obtaining the overall probability ranking of the sample objects of each pushed basic feature sample output by the fusion model for completing multiple tasks for the sample content as the second ranking;
[0392] Calculate a first loss function based on a first ranking of the pushed basic feature sample in the first ranking corresponding to the task and a second ranking in the second ranking;
[0393] Based on the first loss function, train the fusion model.
[0394] Optionally, calculating a first loss function based on a first ranking of the pushed basic feature sample in a first ranking corresponding to the task and a second ranking in a second ranking includes:
[0395] Calculate a first sub-loss function corresponding to the task based on the first ranking of the pushed basic feature sample in the first ranking corresponding to the task and the second ranking in the second ranking;
[0396] Determining a first weight corresponding to the task based on the first ranking and the second ranking;
[0397] Based on the first weights corresponding to the multiple tasks, a weighted sum is calculated for the first sub-loss functions corresponding to the tasks to obtain a first loss function.
[0398] Optionally, based on the first ranking of the pushed basic feature sample in the first ranking corresponding to the task and the second ranking in the second ranking, calculating the first sub-loss function corresponding to the task includes:
[0399] Set a first label corresponding to the task, wherein if the first rank in the first ranking corresponding to the task is before the second rank, the first label is the first value; otherwise, the first label is the second value;
[0400] Get the overall probability of the fusion model output;
[0401] Based on the first label and the overall probability, calculate the first sub-loss function corresponding to the task.
[0402] Optionally, determining a first weight corresponding to the task based on the first ranking and the second ranking includes:
[0403] Based on the first place, determine the first normalized cumulative loss gain corresponding to the task;
[0404] Based on the second rank, determining a second normalized discounted cumulative gain;
[0405] The absolute value of the difference between the first normalized discounted cumulative gain and the second normalized discounted cumulative gain corresponding to the task is used as the first weight corresponding to the task.
[0406] Optionally, based on the first ranking, determining a first normalized cumulative loss gain corresponding to the task includes:
[0407] Get the distribution control coefficient;
[0408] Taking the logarithm of the sum of the product of the distribution control coefficient and the first rank and the predetermined constant to obtain a first intermediate value;
[0409] Based on the logarithm of the predetermined constant and the first intermediate value, a first normalized discounted cumulative gain corresponding to the task is obtained.
[0410] Optionally, based on the first loss function, training the fusion model includes:
[0411] Calculate the second loss function of the probability prediction model for the pushed basic feature sample;
[0412] Determine a single sample loss function based on the first loss function and the second loss function;
[0413] Based on the single sample loss function, the fusion model is trained.
[0414] Optionally, based on a single sample loss function, a fusion model is trained, including:
[0415] Get the sample weight of each pushed basic feature sample;
[0416] Using the sample weights, calculate the weighted sum of the single sample loss function of each pushed basic feature sample to obtain the multi-sample loss function;
[0417] Use multi-sample loss function to train the fusion model.
[0418] Optionally, the push basic feature sample has a plurality of second tags corresponding to a plurality of tasks, the second tags indicating whether the sample object completes the task for the sample content;
[0419] Calculate the second loss function of the probability prediction model for the push basic feature sample, including:
[0420] Obtain multiple third probabilities output by the probability prediction model, that the sample object in the push basic feature sample completes multiple tasks for the sample content;
[0421] A second loss function is calculated based on a plurality of second labels corresponding to a plurality of tasks and a plurality of third probabilities.
[0422] Optionally, the probability prediction model includes multiple first expert sub-models, multiple first gating nodes corresponding to multiple tasks, and multiple first prediction sub-models, each first gating node is connected to the multiple first expert sub-models, the first gating node corresponding to the task is connected to the first prediction sub-model corresponding to the task, and the first prediction sub-model outputs the first probability corresponding to the task;
[0423] The input unit 2820 is specifically used for:
[0424] Push basic features based on multiple targets to obtain a first cascade vector;
[0425] Inputting the first cascade vector into a plurality of first expert sub-models respectively to obtain a plurality of first vectors;
[0426] By using the first gating weights of the plurality of first expert sub-models through the first gating node corresponding to the task, a weighted sum of the plurality of first vectors is obtained to obtain a second vector corresponding to the task;
[0427] The second vector is input into the first prediction sub-model corresponding to the task to obtain a first probability corresponding to the task.
[0428] Optionally, the input unit 2820 is further specifically configured to:
[0429] Quantizing a plurality of target push basic feature vectors into a plurality of target push basic feature vectors;
[0430] Multiple target push basic feature vectors are concatenated into a first concatenated vector.
[0431] Optionally, after vectorizing the multiple target push basic feature vectors into multiple target push basic feature vectors, the push processing device further includes:
[0432] a determining unit (not shown), configured to determine whether a dimension of a target push basic feature vector is greater than a first threshold;
[0433] A pooling unit (not shown) is used to perform pooling processing on the target pushed basic feature vector so that the dimension of the pooled target pushed basic feature vector is equal to the first threshold.
[0434] Optionally, the first expert sub-model comprises a first weight matrix, the number of columns of the first weight matrix is equal to the dimension of the first concatenated vector, and the number of rows of the first weight matrix is equal to the dimension of the first vector;
[0435] The input unit 2820 is further specifically used for:
[0436] The first weight matrices in the plurality of first expert sub-models are multiplied by the transpose of the first cascade vector to obtain a plurality of first vectors.
[0437] Optionally, the first gating node comprises a second weight matrix, the number of columns of the second weight matrix is equal to the dimension of the first cascade vector, and the number of rows of the second weight matrix is equal to the number of the first expert sub-models;
[0438] The input unit 2820 is further specifically used for:
[0439] Determine, through the first gating node, a first product vector of a second weight matrix and a transpose of the first cascade vector;
[0440] Exponentially normalizing the first product vector to obtain a first gating weight vector, where the first gating weight vector includes first gating weights of multiple first expert sub-models;
[0441] The first gating weights of the multiple first expert sub-models in the first gating weight vector are used to calculate the weighted sum of the multiple first vectors to obtain a second vector corresponding to the task.
[0442] Optionally, the input unit 2820 is further specifically configured to:
[0443] Inputting the second vector into a first prediction sub-model corresponding to the task to obtain a first result value corresponding to the task;
[0444] A nonlinear activation function is applied to the first result value to obtain a first probability.
[0445] Optionally, the fusion model includes a fusion weight vector, and the dimension of the fusion weight vector is equal to the number of tasks;
[0446] The push unit 2830 is specifically used for:
[0447] Cascading multiple first probabilities into a first probability vector, where the dimension of the first probability vector is equal to the number of tasks;
[0448] Multiplying the transpose of the first probability vector by the fusion weight vector to obtain a second result value;
[0449] A nonlinear activation function is applied to the second result value to obtain a first overall probability.
[0450] Reference Fig.29 , Fig.29 The structural block diagram of the object terminal 110 for implementing the push processing method of the embodiment of the present disclosure includes: a radio frequency (RF) circuit 2910, a memory 2915, an input unit 2930, a display unit 2940, a sensor 2950, an audio circuit 2960, a wireless fidelity (WiFi) module 2970, a processor 2980, and a power supply 2990. Those skilled in the art can understand that Fig.29 The structure of the object terminal 110 shown does not constitute a limitation on a mobile phone or a computer, and may include more or less components than shown in the figure, or a combination of certain components, or a different arrangement of components.
[0451] RF circuit 2910 can be used for receiving and sending signals during information transmission or calls. In particular, after receiving the downlink information from the base station, it is sent to processor 2980 for processing; in addition, the designed uplink data is sent to the base station.
[0452] The memory 2915 may be used to store software programs and modules. The processor 2980 executes various functional applications and data processing of the content terminal by running the software programs and modules stored in the memory 2915 .
[0453] The input unit 2930 may be used to receive input digital or character information and generate key signal input related to the setting and function control of the content terminal. Specifically, the input unit 2930 may include a touch panel 2931 and other input devices 2932.
[0454] The display unit 2940 may be used to display input information or provided information and various menus of the content terminal. The display unit 2940 may include a display panel 2941.
[0455] The audio circuit 2960, the speaker 2961, and the microphone 2962 may provide an audio interface.
[0456] In this embodiment, the processor 2980 included in the target terminal 110 can execute the push processing method of the previous embodiment.
[0457] The target terminal 110 of the embodiment of the present disclosure includes but is not limited to mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle terminals, aircraft, etc. The embodiment of the present invention can be applied to various scenarios, including but not limited to content recommendation, data screening, etc.
[0458] Fig.30 A block diagram of the structure of a portion of a push processing server 140 for implementing the push processing method of an embodiment of the present disclosure. The push processing server 140 may have relatively large differences due to different configurations or performances, and may include one or more central processing units (CPUs) 3022 (e.g., one or more processors) and a memory 3032, and one or more storage media 3030 (e.g., one or more mass storage devices) storing application programs 3042 or data 3044. Among them, the memory 3032 and the storage medium 3030 may be short-term storage or permanent storage. The program stored in the storage medium 3030 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Furthermore, the central processing unit 3022 may be configured to communicate with the storage medium 3030 and execute a series of instruction operations in the storage medium 3030 on the server.
[0459] The push processing server 140 may also include one or more power supplies 3026, one or more wired or wireless network interfaces 3050, one or more input and output interfaces 3058, and / or, one or more operating systems 3041, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.
[0460] The central processor 3022 in the push processing server 140 may be used to execute the push processing method of the embodiment of the present disclosure.
[0461] The embodiment of the present disclosure also provides a computer-readable storage medium, which is used to store program codes, and the program codes are used to execute the push processing methods of the aforementioned embodiments.
[0462] The embodiment of the present disclosure also provides a computer program product, which includes a computer program. A processor of a computer device reads and executes the computer program, so that the computer device executes and implements the above-mentioned push processing method.
[0463] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present disclosure and the above-mentioned drawings are used to distinguish similar contents, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present disclosure described herein can, for example, be implemented in an order other than those illustrated or described herein. In addition, the terms "comprises" and "comprising" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0464] It should be understood that in the present disclosure, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated content, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the associated content before and after is in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0465] It should be understood that in the description of the embodiments of the present disclosure, the meaning of multiple (or multiple items) is more than two, greater than, less than, exceed, etc. are understood to not include the number, and above, below, within, etc. are understood to include the number.
[0466] In the several embodiments provided in the present disclosure, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0467] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0468] In addition, each functional unit in each embodiment of the present disclosure may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0469] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present disclosure is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the various embodiments of the present disclosure. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, referred to as ROM), random access memory (Random Access Memory, referred to as RAM), disk or optical disk and other media that can store program codes.
[0470] It should also be understood that the various implementations provided in the embodiments of the present disclosure can be combined arbitrarily to achieve different technical effects.
[0471] The above is a specific description of the implementation methods of the present disclosure, but the present disclosure is not limited to the above implementation methods. Technical personnel familiar with the art can also make various equivalent modifications or substitutions without violating the spirit of the present disclosure. These equivalent modifications or substitutions are all included in the scope defined by the claims of the present disclosure.
Claims
1. A push processing method, It is characterized in that include: Acquire a plurality of target push basic features, wherein the plurality of target push basic features include target object features and to-be-pushed content features; Inputting the plurality of target push basic features into a probability prediction model to obtain a plurality of first probabilities of the target object completing a plurality of tasks for the content to be pushed; Multiple first probabilities are cascaded and input into a fusion model to obtain a first overall probability that the target object completes multiple tasks for the content to be pushed, and the content to be pushed is pushed to the target object based on the first overall probability, wherein the fusion model is trained based on a first loss function, and the first loss function is calculated based on the first ranking of the push basic feature samples in the push basic feature sample set in the first sorting and the second ranking in the second sorting, the first sorting is for each of the tasks, after multiple push basic feature samples in the push basic feature sample set are input into the probability prediction model, the output of the probability prediction model corresponding to the task, and the second sorting is the sorting of the output of the fusion model after multiple push basic feature samples in the push basic feature sample set are input into the probability prediction model.
2. The push processing method according to claim 1, It is characterized in that The fusion model is trained through the following process: Acquire a push basic feature sample set, wherein each push basic feature sample in the push basic feature sample set includes a sample object feature and a sample content feature; Inputting the pushed basic feature sample set into the probability prediction model, and cascading the output of the probability prediction model into the fusion model; For each of the tasks, obtaining, output by the probability prediction model, a probability ranking of the sample objects of each of the pushed basic feature samples in terms of completing the task for the sample content, as the first ranking; Obtaining the overall probability ranking of the sample objects of each of the pushed basic feature samples output by the fusion model for completing the plurality of tasks for the sample content as the second ranking; Calculating a first loss function based on a first ranking of the pushed basic feature sample in the first ranking corresponding to the task and a second ranking in the second ranking; Based on the first loss function, the fusion model is trained.
3. The push processing method according to claim 2, It is characterized in that The calculating a first loss function based on the first ranking of the pushed basic feature sample in the first ranking corresponding to the task and the second ranking in the second ranking includes: Calculating a first sub-loss function corresponding to the task based on a first ranking of the pushed basic feature sample in the first ranking corresponding to the task and a second ranking in the second ranking; Determining a first weight corresponding to the task based on the first ranking and the second ranking; Based on the first weights corresponding to the multiple tasks, a weighted sum of the first sub-loss functions corresponding to the tasks is calculated to obtain the first loss function.
4. The push processing method according to claim 3, It is characterized in that The calculating a first sub-loss function corresponding to the task based on the first ranking of the pushed basic feature sample in the first ranking corresponding to the task and the second ranking in the second ranking includes: Setting a first label corresponding to the task, wherein if the first ranking in the first ranking corresponding to the task is before the second ranking, the first label is a first value; otherwise, the first label is a second value; Obtaining the overall probability output by the fusion model; Based on the first label and the overall probability, the first sub-loss function corresponding to the task is calculated.
5. The push processing method according to claim 3, It is characterized in that The determining a first weight corresponding to the task based on the first ranking and the second ranking includes: Based on the first ranking, determining a first normalized cumulative loss gain corresponding to the task; Based on the second ranking, determining a second normalized discounted cumulative gain; An absolute value of a difference between a first normalized discounted cumulative gain corresponding to the task and the second normalized discounted cumulative gain is used as the first weight corresponding to the task.
6. The push processing method according to claim 5, It is characterized in that The determining, based on the first ranking, a first normalized cumulative loss gain corresponding to the task includes: Get the distribution control coefficient; Taking the logarithm of the sum of the product of the distribution control coefficient and the first rank and a predetermined constant to obtain a first intermediate value; Based on the logarithm of the predetermined constant and the first intermediate value, the first normalized discounted cumulative gain corresponding to the task is obtained.
7. The push processing method according to claim 2, It is characterized in that The step of training the fusion model based on the first loss function includes: Calculating a second loss function of the probability prediction model for the pushed basic feature sample; Determine a single sample loss function based on the first loss function and the second loss function; Based on the single sample loss function, the fusion model is trained.
8. The push processing method according to claim 7, It is characterized in that The step of training the fusion model based on the single sample loss function includes: Obtaining a sample weight of each of the pushed basic feature samples; Using the sample weights, a weighted sum of the single sample loss functions of each of the pushed basic feature samples is calculated to obtain a multi-sample loss function; The fusion model is trained using the multi-sample loss function.
9. The push processing method according to claim 7, It is characterized in that The push basic feature sample has a plurality of second tags corresponding to the plurality of tasks, the second tags indicating whether the sample object completes the task for the sample content; The calculating a second loss function of the probability prediction model for the pushed basic feature sample includes: Acquire a plurality of third probabilities output by the probability prediction model, that the sample object in the push basic feature sample completes a plurality of the tasks for the sample content; The second loss function is calculated based on a plurality of second labels corresponding to a plurality of the tasks and a plurality of the third probabilities.
10. The push processing method according to claim 1, It is characterized in that The probability prediction model includes a plurality of first expert sub-models, a plurality of first gating nodes corresponding to a plurality of the tasks, and a plurality of first prediction sub-models, each of the first gating nodes is connected to a plurality of the first expert sub-models, the first gating nodes corresponding to the tasks are connected to the first prediction sub-models corresponding to the tasks, and the first prediction sub-models output the first probability corresponding to the tasks; The step of inputting the plurality of target push basic features into a probability prediction model to obtain a plurality of first probabilities of the target object completing a plurality of tasks for the content to be pushed includes: Based on the plurality of target push basic features, obtaining a first cascade vector; Inputting the first cascade vector into a plurality of the first expert sub-models respectively to obtain a plurality of first vectors; By using the first gating weights of the first expert sub-models and the first gating node corresponding to the task, a weighted sum of the first vectors is obtained to obtain a second vector corresponding to the task; The second vector is input into the first prediction sub-model corresponding to the task to obtain the first probability corresponding to the task.
11. The push processing method according to claim 10, It is characterized in that The acquiring a first cascade vector based on the plurality of target push basic features includes: Quantizing the plurality of target push basic feature vectors into a plurality of target push basic feature vectors; A plurality of the target pushed basic feature vectors are concatenated into the first concatenated vector.
12. The push processing method according to claim 11, It is characterized in that After the step of quantizing the plurality of target push basic feature vectors into a plurality of target push basic feature vectors, the push processing method further includes: Determining that a dimension of the target push basic feature vector is greater than a first threshold; Pooling is performed on the target pushing basic feature vector so that the dimension of the pooled target pushing basic feature vector is equal to the first threshold.
13. The push processing method according to claim 10, It is characterized in that The first expert sub-model comprises a first weight matrix, the number of columns of the first weight matrix is equal to the dimension of the first cascade vector, and the number of rows of the first weight matrix is equal to the dimension of the first vector; The step of inputting the first cascade vector into a plurality of the first expert sub-models respectively to obtain a plurality of first vectors includes: multiplying the first weight matrix in a plurality of the first expert sub-models by the transpose of the first cascade vector to obtain a plurality of the first vectors.
14. The push processing method according to claim 10, It is characterized in that The first gating node comprises a second weight matrix, the number of columns of the second weight matrix is equal to the dimension of the first cascade vector, and the number of rows of the second weight matrix is equal to the number of the first expert sub-models; The step of obtaining a second vector corresponding to the task by performing weighted sum calculation on a plurality of the first vectors using the first gating weights of a plurality of the first expert sub-models through the first gating node corresponding to the task comprises: Determine, by means of the first gating node, a first product vector of the second weight matrix and the transpose of the first cascade vector; performing exponential normalization on the first product vector to obtain a first gating weight vector, wherein the first gating weight vector includes the first gating weights of a plurality of the first expert sub-models; Using the first gating weights of the plurality of the first expert sub-models in the first gating weight vector, a weighted sum is calculated for the plurality of the first vectors to obtain a second vector corresponding to the task.
15. The push processing method according to claim 10, It is characterized in that The step of inputting the second vector into the first prediction sub-model corresponding to the task to obtain the first probability corresponding to the task includes: Inputting the second vector into the first prediction sub-model corresponding to the task to obtain a first result value corresponding to the task; Apply a nonlinear activation function to the first result value to obtain the first probability.
16. The push processing method according to claim 1, It is characterized in that The fusion model includes a fusion weight vector, and the dimension of the fusion weight vector is equal to the number of the tasks; The step of cascading the plurality of the first probabilities and inputting them into a fusion model to obtain a first overall probability that the target object completes the plurality of the tasks for the content to be pushed includes: Cascading a plurality of the first probabilities into a first probability vector, wherein the dimension of the first probability vector is equal to the number of the tasks; Multiplying the transpose of the first probability vector by the fusion weight vector to obtain a second result value; Apply a nonlinear activation function to the second result value to obtain the first overall probability.
17. A push processing device, It is characterized in that include: An acquisition unit, configured to acquire a plurality of target push basic features, wherein the plurality of target push basic features include target object features and to-be-pushed content features; An input unit, configured to input the plurality of target push basic features into a probability prediction model to obtain a plurality of first probabilities of the target object completing a plurality of tasks for the content to be pushed; A pushing unit is used to cascade multiple first probabilities and input them into a fusion model to obtain a first overall probability that the target object completes multiple tasks for the content to be pushed, and pushes the content to be pushed to the target object based on the first overall probability, wherein the fusion model is trained based on a first loss function, and the first loss function is calculated based on the first ranking of the push basic feature samples in the push basic feature sample set in the first sorting and the second ranking in the second sorting, the first sorting is for each of the tasks, after the multiple push basic feature samples in the push basic feature sample set are input into the probability prediction model, the output of the probability prediction model corresponding to the task, and the second sorting is the sorting of the output of the fusion model after the multiple push basic feature samples in the push basic feature sample set are input into the probability prediction model.
18. An electronic device comprising a memory and a processor, wherein the memory stores a computer program. It is characterized in that When the processor executes the computer program, the push processing method according to any one of claims 1 to 16 is implemented.
19. A computer-readable storage medium storing a computer program. It is characterized in that When the computer program is executed by a processor, the push processing method according to any one of claims 1 to 16 is implemented.
20. A computer program product, comprising a computer program, wherein the computer program is read and executed by a processor of a computer device, so that the computer device executes the push processing method according to any one of claims 1 to 16.