Push processing method, related device and medium
By using the probability prediction model on the content push platform, the push of long-tail content is optimized based on the total loss function training and weighted long-tail task loss function, the balance of long-tail content push and popular content push in the existing technology is solved, and long-tail content push with high interest correlation is achieved.
Patent Information
- Application Number
- CN202311603729.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-27
- Publication Date
- 2025-05-27
AI Technical Summary
When the prior art improves the object interest correlation of long-tail content pushes, it is easy to sacrifice the push efficiency of popular content, or the pushed long-tail content has a low correlation with object interest.
By obtaining the basic characteristics of the target push, including the target object characteristics and the content to be pushed, input a probability prediction model to obtain the first probability that the content to be pushed belongs to the long-tail content of the target object's interest. This probability prediction model is trained based on the total loss function. The total loss function is weighted by the long-tail task weight of the long-tail task loss function to determine the long-tail task weight of the push basic feature sample, and preferentially recommend long-tail content with high interest correlation of the object.
Without sacrificing the push efficiency of popular content, the object interest relevance of pushed long-tail content is improved, thereby improving the quantity and quality of pushed long-tail content.
Smart Images

Figure CN120045771A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence, and in particular, to a push processing method, related device, and medium. Background Art
[0002] On a content push platform, a push processing model can be used to recommend content that an object may be interested in based on object features and content features. The push processing model will collect effective information required for pushing according to the situation of the object's completion of various tasks (click, like, comment, etc.) for the content predicted, so as to improve the accuracy of pushing matching content for the object. Since popular content has better completion of tasks such as clicks and likes, it will receive more and more pushes, while the number of pushes for unpopular content is getting less and less. Unpopular content is also called long-tail content. The emergence of long-tail content has led to a decline in the personalized push ability of the push processing model for different objects.
[0003] Existing technologies also have solutions to increase the push of long-tail content, but either sacrifice the push effect of popular content or the long-tail content pushed has a low relevance to the object's interests. Summary of the Invention
[0004] Embodiments of the present disclosure provide a push processing method, related device, and medium, which can improve the object interest relevance of the pushed long-tail content without sacrificing the push efficiency of popular content, thereby increasing the quantity and quality of the pushed long-tail content.
[0005] According to an aspect of the present disclosure, there is provided a push processing method, including:
[0006] Obtain a plurality of target push basic features, where the plurality of target push basic features include target object features and content features to be pushed;
[0007] Input the plurality of target push basic features into a probability prediction model to obtain a first probability that the content to be pushed belongs to long-tail content that the target object is interested in, where the probability prediction model is trained based on a total loss function, and the total loss function is a weighted sum of the long-tail task loss functions of the push basic feature samples in the push basic feature sample set weighted by the long-tail task weights of the push basic feature samples. If the sample content corresponding to the push basic feature sample meets the object interest condition and the sample content meets the long-tail sample condition, the long-tail task weight of the push basic feature sample is determined as a first value, otherwise, the long-tail task weight of the push basic feature sample is determined as a second value, where the first value is greater than the second value;
[0008] Push the content to be pushed for the target object based on the first probability.
[0009] According to one aspect of the present disclosure, a push processing device is provided, including:
[0010] An acquisition unit configured to acquire a plurality of target push basic features, where the plurality of target push basic features include target object features and content-to-be-pushed features;
[0011] An input unit configured to input the plurality of target push basic features into a probability prediction model to obtain a first probability that the content-to-be-pushed belongs to long-tail content of interest to the target object, where the probability prediction model is trained based on a total loss function, and the total loss function is a weighted sum of long-tail task loss functions of the push basic feature samples in the push basic feature sample set using the long-tail task weights of the respective push basic feature samples, where if the sample content corresponding to the push basic feature sample meets the object interest condition and the sample content meets the long-tail sample condition, the long-tail task weight of the push basic feature sample is determined to be a first value, otherwise, the long-tail task weight of the push basic feature sample is determined to be a second value, where the first value is greater than the second value;
[0012] A push unit configured to push the content-to-be-pushed for the target object based on the first probability.
[0013] Optionally, the probability prediction model is trained in the following manner:
[0014] Acquire a push basic feature sample set, where each push basic feature sample in the push basic feature sample set includes sample object features and sample content features;
[0015] Input the push basic feature sample into the probability prediction model to obtain the long-tail task loss function for pushing the sample content for the sample object;
[0016] If it is determined that the sample content corresponding to the push basic feature sample meets the object interest condition and it is determined that the sample content meets the long-tail sample condition, the long-tail task weight of the push basic feature sample is determined to be the first value, otherwise, the long-tail task weight of the push basic feature sample is determined to be the second value, where the first value is greater than the second value;
[0017] Generate a total loss function based on the weighted sum of the long-tail task loss functions of the push basic feature samples using the long-tail task weights of the push basic feature samples;
[0018] Train the probability prediction model based on the total loss function.
[0019] Optionally, determining that the sample content corresponding to the push basic feature sample meets the object interest condition includes:
[0020] Determining that the sample object has clicked on the sample content;
[0021] Obtaining the playback duration of the sample content by the sample object;
[0022] Determining that the playback duration is greater than a first threshold.
[0023] Optionally, determining that the sample content meets the long-tail sample condition includes:
[0024] Determining the first click volume of the sample content;
[0025] Determining the proportion of the first exposure volume of the sample content in the push basic feature sample set whose click volume is not greater than the first click volume;
[0026] If the proportion of the first exposure volume is less than a second threshold, determining that the sample content meets the long-tail sample condition.
[0027] Optionally, determining the proportion of the first exposure volume of the sample content in the push basic feature sample set whose click volume is not greater than the first click volume includes:
[0028] Determining the first sum of the first exposure volumes of the sample content in the push basic feature sample set whose click volume is not greater than the first click volume;
[0029] Determining the second sum of the first exposure volumes of each sample content in the push basic feature sample set;
[0030] Based on the first sum and the second sum, determining the proportion of the first exposure volume.
[0031] Optionally, determining the proportion of the first exposure volume of the sample content in the push basic feature sample set whose click volume is not greater than the first click volume includes:
[0032] Substituting the first click volume into a first function to obtain the proportion of the first exposure volume, where the parameters in the first function are fitted by the following method:
[0033] Setting a plurality of critical proportions, where the plurality of critical proportions are arranged in an arithmetic progression;
[0034] Obtaining the second click volume of each fitting content sample in the fitting content sample set;
[0035] Sort the fitting content samples in ascending order of the second click volume, accumulate the second exposure volume ratios of the fitting content samples in the order from front to back of the sorting, and record the second click volume when the second exposure volume ratio reaches each of the critical ratios as the critical click volume corresponding to the critical ratio;
[0036] Based on the critical ratio and the critical click volume corresponding to the critical ratio, fit the parameters in the first function.
[0037] Optionally, the fitting the parameters in the first function based on the critical ratio and the critical click volume corresponding to the critical ratio includes:
[0038] Substitute the critical click volume into the first function to obtain the actual ratio corresponding to the critical click volume;
[0039] Based on the first difference between the actual ratio corresponding to the critical click volume and the critical ratio, determine the parameters in the first function.
[0040] Optionally, the determining the parameters in the first function based on the first difference between the actual ratio corresponding to the critical click volume and the critical ratio includes:
[0041] Determine the sum of squares of the first differences corresponding to each of the critical click volumes;
[0042] Determine the parameters that minimize the sum of squares as the parameters in the first function.
[0043] Optionally, the parameters include a first parameter, a second parameter, and a third parameter;
[0044] The substituting the critical click volume into the first function to obtain the actual ratio corresponding to the critical click volume includes:
[0045] Calculate the first power of the critical click volume as a first intermediate value;
[0046] Based on the first intermediate value and the second parameter, calculate a second intermediate value;
[0047] Calculate the second difference between a first constant and the second intermediate value power of a predetermined base;
[0048] Take the third power of the second difference as the actual ratio.
[0049] Optionally, the inputting the push base feature sample into the probability prediction model to obtain the long-tail task loss function for pushing the sample content to the sample object includes:
[0050] Input the push basic feature sample into the probability prediction model to obtain the long-tail task loss function and other task loss functions for pushing the sample content to the sample object;
[0051] Generating a total loss function based on the weighted sum of the long-tail task loss function of the push basic feature sample using the long-tail task weight of the push basic feature sample, including:
[0052] Perform a weighted sum of the long-tail task loss function of the push basic feature sample using the long-tail task weight of the push basic feature sample to obtain a first result;
[0053] Perform a weighted sum of the other task loss function of the push basic feature sample using the other task weight of the push basic feature sample to obtain a second result;
[0054] Generate the total loss function based on the first result and the second result.
[0055] Optionally, determining the long-tail task weight of the push basic feature sample as the first value includes:
[0056] Determine a first candidate value based on the first exposure ratio;
[0057] Determine the second value as the second candidate value;
[0058] Determine the larger of the first candidate value and the second candidate value as the first value, and determine the long-tail task weight of the push basic feature sample as the first value.
[0059] Optionally, determining the first candidate value based on the first exposure ratio includes:
[0060] Obtain the second threshold and the third threshold;
[0061] Determine a first coefficient based on the second threshold and the third threshold;
[0062] Add the result of multiplying the first coefficient by the first exposure ratio to the third threshold to obtain the first candidate value.
[0063] Optionally, determining the first coefficient based on the second threshold and the third threshold includes:
[0064] Calculate the third difference between the second constant and the third threshold, where the second constant is less than the third threshold;
[0065] Divide the third difference by the second threshold to obtain the first coefficient.
[0066] Optionally, inputting the push basic feature sample into the probability prediction model to obtain the long-tail task loss function for pushing the sample content to the sample object includes:
[0067] Inputting the push basic feature sample into the probability prediction model to obtain a second probability that the sample content belongs to the long-tail content of interest to the sample object;
[0068] Obtaining a sample label of the push basic feature sample, where the sample label indicates whether the sample object clicks on the sample content;
[0069] Based on the sample label and the second probability, obtaining the long-tail task loss function for pushing the sample content to the sample object.
[0070] Optionally, the input unit is specifically configured to:
[0071] Inputting multiple target push basic features into a probability prediction model to obtain a first probability that the content to be pushed belongs to the long-tail content of interest to the target object and a third probability that the target object completes multiple other tasks for the content to be pushed;
[0072] The pushing the content to be pushed to the target object based on the first probability includes:
[0073] Pushing the content to be pushed to the target object based on the first probability and the third probability of each of the other tasks.
[0074] Optionally, the input unit is specifically configured to:
[0075] Determining a first task weight of the long-tail task corresponding to the first probability and second task weights of the multiple other tasks;
[0076] Obtaining a multi-task score according to the first probability and the first task weight, the third probability of each of the other tasks and the second task weight;
[0077] Pushing the content to be pushed to the target object according to the multi-task score.
[0078] According to one aspect of the present disclosure, there is provided an electronic device including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the above-mentioned push processing method is implemented.
[0079] According to one aspect of the present disclosure, there is provided a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the above-mentioned push processing method is implemented.
[0080] According to one aspect of the present disclosure, there is provided a computer program product, which includes a computer program that is read and executed by a processor of a computer device, so that the computer device executes the push processing method as described above.
[0081] In the embodiments of the present disclosure, when training a probability prediction model, a total loss function is used, and the total loss function is obtained by weighted summation of the long-tail task loss functions of each push basic feature sample in the push basic feature sample set using the long-tail task weights of the push basic feature samples. If the sample content corresponding to a push basic feature sample meets the object interest condition, that is, the object interest relevance is high, and the sample content meets the long-tail sample condition, the long-tail task weight of the push basic feature sample is determined to be higher than the long-tail task weights of other push basic feature samples, that is, the first value. That is to say, the long-tail task weights of those long-tail contents that meet the object interest are increased, so that when actually pushing, those long-tail contents with high object interest relevance are recommended first. And the embodiments of the present disclosure do not reduce the weights of those popular contents, and their weights still remain at a second value lower than the first value. Therefore, without sacrificing the push efficiency of popular contents, the embodiments of the present disclosure improve the object interest relevance of the pushed long-tail contents, thereby improving the quantity and quality of the pushed long-tail contents.
[0082] Other features and advantages of the present disclosure will be described in the following specification, and part of them will become obvious from the specification, or will be understood by implementing the present disclosure. The objectives and other advantages of the present disclosure can be achieved and obtained through the structures specifically pointed out in the specification, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0083] The drawings are used to provide a further understanding of the technical solutions of the present disclosure, and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the technical solutions of the present disclosure, and do not constitute a limitation to the technical solutions of the present disclosure.
[0084] Figure 1 is an architecture diagram of the push processing method according to an embodiment of the present disclosure;
[0085] Figures 2A - 2B is a schematic interface diagram of an embodiment of the present disclosure applied to the content push scenario in a content push application;
[0086] Figure 3 is a flowchart of the push processing method according to an embodiment of the present disclosure;
[0087] Figure 4 is a schematic structural diagram of a probability prediction model for single-task prediction according to an embodiment of the present disclosure;
[0088] Figure 5 It is a schematic structural diagram of a probability prediction model for multi-task prediction according to an embodiment of the present disclosure;
[0089] Figure 6 It is a flowchart for training a probability prediction model according to an embodiment of the present disclosure;
[0090] Figure 7 It is Figure 6 A flowchart of step 620 in
[0091] Figure 8 It is Figure 6 A flowchart of step 630 in
[0092] Figure 9 It is Figure 6 Another flowchart of step 630 in
[0093] Figure 10 It is Figure 9 A flowchart of step 920 in
[0094] Figure 11 It is a flowchart for fitting a first function according to an embodiment of the present disclosure;
[0095] Figure 12 It is a schematic diagram for determining a critical click volume according to an embodiment of the present disclosure;
[0096] Figure 13 It is a curve graph of the mapping relationship between the critical click volume and the critical proportion according to an embodiment of the present disclosure;
[0097] Figure 14 It is Figure 11 A flowchart of step 1140 in
[0098] Figure 15 It is Figure 14 A flowchart of step 1410 in
[0099] Figure 16 It is Figure 14 A flowchart of step 1420 in
[0100] Figure 17 It is a function curve graph of a first function after fitting according to an embodiment of the present disclosure;
[0101] Figure 18 It is a flowchart for determining a second threshold according to an embodiment of the present disclosure;
[0102] Figure 19 It is a curve graph of the mapping relationship between the third click volume ranking and the third exposure proportion according to an embodiment of the present disclosure;
[0103] Figure 20 is Figure 6 A flowchart of step 630 in
[0104] Figure 21 is Figure 20 A flowchart of step 2010 in
[0105] Figure 22 is Figure 21 A flowchart of step 2120 in
[0106] Figure 23 A curve graph of the mapping relationship between the first click volume and the long-tail task weight according to an embodiment of the present disclosure;
[0107] Figure 24 is Figure 6 A flowchart of step 640 in
[0108] Figure 25 A flowchart of pushing the content to be pushed to the target object based on the first probability and the third probability of other tasks according to an embodiment of the present disclosure;
[0109] Figure 26 An implementation detail diagram of the pushing processing method according to an embodiment of the present disclosure;
[0110] Figure 27A A comparison graph of the click content volume before and after using the pushing processing method according to an embodiment of the present disclosure by a content pushing application;
[0111] Figure 27B A comparison graph of the total content playing duration before and after using the pushing processing method according to an embodiment of the present disclosure by a content pushing application;
[0112] Figure 28 A module diagram of the pushing processing device according to an embodiment of the present disclosure;
[0113] Figure 29 is according to an embodiment of the present disclosure Figure 3 The terminal structure diagram of the pushing processing method shown in
[0114] Figure 30 is according to an embodiment of the present disclosure Figure 3 The server structure diagram of the pushing processing method shown in Detailed implementation manners
[0115] In order to make the objectives, technical solutions and advantages of the present disclosure clearer, the present disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present disclosure, and are not used to limit the present disclosure.
[0116] Before further elaborating on the embodiments of the present disclosure, the nouns and terms involved in the embodiments of the present disclosure are described. The nouns and terms involved in the embodiments of the present disclosure are applicable to the following explanations:
[0117] Deep neural network: It is a multi-layer unsupervised neural network. The deep neural network uses the output features of the previous layer as the input of the next layer for feature learning. After layer-by-layer feature mapping, the features of the existing space samples are mapped to another feature space, so as to learn better feature expressions for the existing inputs. The deep neural network has multiple non-linear mapping feature transformations and can fit highly complex functions.
[0118] Gradient descent algorithm: A method for minimizing the loss function. It calculates the gradient of the loss function at the current position and updates the parameters in the direction of the gradient descent to find the optimal solution to minimize the value of the loss function.
[0119] Cross-entropy: It measures the degree of difference between two different probability distributions in the same random variable, and in machine learning, it represents the difference between the true probability distribution and the predicted probability distribution. Given two probability distributions p and q, where p represents the true distribution and q represents the distribution predicted by the model, the definition of cross-entropy is as follows:
[0120]
[0121] p i represents the probability of the occurrence of the i-th event in the true distribution, and q i represents the probability of the occurrence of the i-th event in the distribution predicted by the model. The smaller the value of the cross-entropy, the better the prediction effect of the model. Therefore, cross-entropy is often used as a loss function to guide the optimization of the model.
[0122] System architecture and scenario description applied in the embodiments of the present disclosure
[0123] Figure 1 It is a system architecture diagram applied to the push processing method according to the embodiments of the present disclosure. It includes: server 110, gateway 120, Internet 130, and terminal 140.
[0124] Server 110 refers to a computer system that determines the content to be pushed to the target object. Compared with the terminal 140, higher requirements are imposed on the stability, security, performance, etc. of the server 110. The server 110 can be a high-performance computer in the network platform, a combination of a part (such as a virtual machine) drawn from multiple high-performance computers, etc. The server 110 can also communicate with the Internet 130 through wired or wireless means to exchange data.
[0125] The gateway 120 is also known as an internetwork connector or protocol converter. The gateway 120 realizes network interconnection at the transport layer and is a computer system or device that acts as a converter. Between two systems using different communication protocols, data formats, or languages, and even with completely different architectures, the gateway 120 is a translator. At the same time, the gateway 120 can also provide filtering and security functions. The messages sent by the terminal 140 to the server 110 need to be sent to the corresponding server 110 through the gateway 120. The messages sent by the server 110 to the terminal 140 also need to be sent to the corresponding terminal 140 through the gateway 120.
[0126] The terminal 140 is a device used to push content for an object to view the pushed content. It includes various forms such as desktop computers, laptops, PDAs (Personal Digital Assistants), mobile phones, in-vehicle terminals, home theater terminals, dedicated terminals, etc. Additionally, it can be a single device or a collection of multiple devices. For example, multiple devices are connected through a local area network and share a display device to work collaboratively, jointly forming a terminal 140. The terminal 140 can also communicate with the Internet 130 in a wired or wireless manner to exchange data.
[0127] Embodiments of the present disclosure can be applied in various scenarios, such as Figures 2A - 2B the scenario of pushing video content to an object in a content push application as shown, etc.
[0128] As Figure 2A shown in the interface of the content push application in the terminal 140, when the object opens the content push application, the content push application will push video content that the object may be interested in. The cover and content title of the pushed video content can be displayed in a two-column information flow form in the interface, and the object can select the content they want to view based on the cover content or content title. Each content also shows the number of likes for this content, and when the object views a content, they can choose to like the content they like. The more the number of likes for a content, the higher the content popularity, and this content will be pushed to the object more. Using the push processing method of the embodiments of the present disclosure, the content push application will preferentially push long-tail content that the object may be interested in. In Figure 2A it, the number of likes for content A is 42, and the content popularity is low, but through the prediction of the content push application, content A may be long-tail content that the object is interested in. Therefore, content A is pushed to the object and displayed in the first position on the interface. At the same time, popular content that the object is interested in will also be pushed, such as content D.
[0129] As Figure 2BAs shown, Content A has been successfully pushed. The object clicks to view Content A and likes Content A, which represents a high degree of interest association between Content A and the object. Thus, it can be seen that the push processing method of the embodiments of the present disclosure can improve the association degree between the pushed long-tail content and the object's interest, thereby improving the quantity and quality of the pushed long-tail content in the content push application. At the same time, it will not sacrifice the push efficiency of popular content.
[0130] General description of the embodiments of the present disclosure
[0131] According to an embodiment of the present disclosure, a push processing method is provided.
[0132] The push processing method is a process of pushing content that the target object is interested in to the terminal 140 of the target object. The target object is the object that wants to view the content to be pushed. The embodiments of the present disclosure can be applied to the push of different types of content, and the content to be pushed can be various types such as pictures, videos, articles, applications, etc.
[0133] When determining the content to be pushed to the target object, it is necessary to consider the long-tail content in the content to be pushed. Long-tail content refers to the content that has been pushed very few times after a period of push processing by the push processing model. Since the push processing model will collect effective information from each push event to improve the push accuracy, it is difficult for the push processing model to obtain effective information from long-tail content. This results in fewer and fewer pushes of long-tail content and more and more pushes of popular content, thereby leading to a decline in the personalized push ability of the push processing model for different objects. Thus, it can be seen that improving the quantity and quality of pushed long-tail content can improve the push accuracy of the push processing model. Improving the quantity of pushed long-tail content means that the number of pushes of long-tail content can be increased without sacrificing the push efficiency of popular content. Improving the quality of pushed long-tail content means that the interest association degree between the pushed long-tail content and the target object can be increased.
[0134] The push processing method of the embodiments of the present disclosure can be executed by the terminal 140 or by the server 110. When executed by the server 110, after execution, the determined content to be pushed is transmitted to the terminal 140 through the Internet 130, and the terminal 140 displays the content to be pushed to the target object.
[0135] As Figure 3 shown, according to an embodiment of the present disclosure, the push processing method includes:
[0136] Step 310, obtain multiple target push basic features;
[0137] Step 320: Input multiple target push basic features into a probability prediction model to obtain a first probability that the content to be pushed belongs to the long-tail content of interest to the target object. The probability prediction model is trained based on a total loss function, which is the weighted sum of the long-tail task loss functions of the push basic feature samples in the push basic feature sample set using the long-tail task weights of each push basic feature sample. If the sample content corresponding to the push basic feature sample meets the object interest condition and the sample content meets the long-tail sample condition, the long-tail task weight of the push basic feature sample is determined as a first value; otherwise, the long-tail task weight of the push basic feature sample is determined as a second value, where the first value is greater than the second value.
[0138] Step 330: Push the content to be pushed to the target object based on the first probability.
[0139] The above steps 310 - 330 are described in detail below.
[0140] Detailed description of step 310
[0141] In step 310, the multiple target push basic features include target object features and content-to-be-pushed features.
[0142] Target object features refer to features related to the target object. Target object features can be features of the target object itself. For example, features such as the educational level of the target object. For the same type of feature, if the feature values are different, the target object's preference for the pushed content may be different. For example, objects in different industries may like different types of pushed content. Taking video content as an example, objects in the education industry may be more concerned about educational popular science-related videos, and objects in the medical industry may be more concerned about health care-related videos.
[0143] Note that when obtaining the target object features of the target object, the consent of the target object should be obtained in advance. Moreover, the collection, use, and processing of these object features will comply with relevant laws, regulations, and standards. When obtaining the consent of the target object, the target object's separate permission or separate consent can be obtained through methods such as pop-up windows or redirecting to a confirmation page.
[0144] The target object feature can also be the object features of the members in the target object group (such as family, work group) where the target object is located. This is because, in some cases, the preference of the target object for the pushed content is affected by the target object group it belongs to. For example, the target object itself is interested in relevant video content on health care, but since there is an object engaged in the news industry in the family group of the target object, the target object may also be interested in relevant video content on hot events. Considering the object features of the members in the target object group where the target object is located can improve the comprehensiveness of obtaining the target object features, determine the preferences of the target object from multiple aspects, and thus improve the accuracy of pushing the content to be pushed to the target object. Similarly to the above, when obtaining the object features of the target object group (such as family, work group) where the target object is located, the consent of the object group members also needs to be obtained in advance, so it will not be elaborated here.
[0145] In one embodiment, the target object feature can be obtained through registration information. Here, the registration information can be the registration information of the target object on the terminal 140, or the registration information of the target object in the content push application, etc. When the target object first uses the terminal 140, the terminal 140 may require the target object to register, and relevant object information needs to be filled in as the registration information during registration. In this way, the target object feature can be obtained from the terminal 140 used by the target object. When the target object first uses the content push application, registration is also required. The content push application may require the target object to fill in relevant object information as the registration information. Therefore, the target object feature can also be obtained by using the registration information of the target object in the content push application. The advantage of obtaining the target object feature through registration information is: convenient and fast, with high acquisition efficiency. When obtaining the registration information of the target object, as mentioned above, the consent of the target object also needs to be obtained in advance, which will not be elaborated here.
[0146] In another embodiment, the target object features are obtained based on the object actions of the target object in the content push application. Object actions refer to dynamic events dominated by the target object in the content push application. For example, the target object interacts with other objects, or the target object searches for content in the content push application, etc. By counting the interaction events between the target object and other objects, the interaction frequency can be used as a target object feature; by querying the search records of the target object in the content push application, the object search records can also be used as a target object feature. Object actions can directly or indirectly reflect the needs of the target object when using the content push application. For example, a target object with a high interaction frequency has a high social need when using the content push application and may be more interested in content that can provide interaction and communication. Another example is that if the search records of the target object contain a large amount of content related to medical and health, then such content can be pushed to the target object. Thus, obtaining the target object features based on the object actions of the target object in the content push application can more comprehensively understand the object's needs and improve the accuracy of pushing the content to be pushed to the target object. When obtaining the object actions of the target object, as described above, the consent of the target object also needs to be obtained in advance, which will not be elaborated here.
[0147] The features of the content to be pushed are used to represent the characteristics of multiple aspects of the content to be pushed. For example, content type, content length, content release time, content keywords, etc. For the same feature of the content to be pushed, if the feature values are different, the preferences of the target object for the content to be pushed may also be different. For example, different objects may have different preferences for the content length. Taking video content as an example, some objects may like to watch short video content of less than 3 minutes, while some objects may like to watch long video content of more than 10 minutes.
[0148] Note that when obtaining the features of the content to be pushed, the consent of the content publisher of the content to be pushed needs to be obtained in advance. The specific method is the same as the method of obtaining the consent of the target object, which will not be elaborated here. Moreover, the collection, use, and processing of these features of the content to be pushed will comply with relevant laws, regulations, and standards.
[0149] In one embodiment, the multiple target push basic features also include target scenario features. Target scenario features refer to the features of the current scenario where the target object views the content to be pushed. The same object may want to view different content in different scenarios. For example, when the object is in a work scenario, it may want to view content related to work; when the object is at home, it may be more likely to view content related to its own preferences. Target scenario features include: geographical region features, time period features, input day type features (weekday, weekend, or holiday), etc. when the target object views the content to be pushed.
[0150] When obtaining the target scene features of the target object, the consent of the target object must be obtained in advance. Similar to the above, it will not be elaborated here.
[0151] The geographical area feature refers to the features of the geographical area where the target object is located. This geographical area can be an administrative geographical area, such as City A, City B, etc.; it can also be a geographical area divided by entities on the map, such as the XX School area, the XX Shopping Mall area, etc. The administrative geographical area has an impact on the content that the object wants to view. For example, when the object is in City A, it may want to view content related to tourist attractions or food culture in City A. The geographical area divided by entities also has an impact on the content that the object wants to view. For example, when the object is in the XX Shopping Mall area, it may want to view the activity information or special projects in the mall.
[0152] The geographical area feature where the target object is located can be obtained from the positioning information of the target object's terminal 140. The terminal 140 of the target object has positioning devices such as GPS and Beidou, and this positioning device can continuously obtain the positioning information where the terminal 140 is located. The geographical area where the target object is located when viewing the push content can be determined according to this positioning information as the geographical area feature. For example, according to the GPS positioning information obtained by the positioning device and mapping it to an electronic map, it is found that it belongs to City A, then the geographical area feature when viewing the push content is obtained as "City A".
[0153] The time period feature refers to the time period within a day when the target object views the content to be pushed. For example, 9:00 - 11:00, 11:00 - 13:00, 13:00 - 18:00, 18:00 - 22:00, etc. Since the activities of the object are different in different time periods of a day, and the purpose of viewing content may also be different, the push content will be affected by the time period feature. For example, from 9:00 to 11:00, the object may be at work, and during this time period, the object is very likely to want to view content related to work; from 18:00 to 22:00, the object is usually resting, and during this time period, the object is very likely to want to view content related to its own interests.
[0154] The input day type feature refers to the feature of the type of day when the target object views the content to be pushed. It is divided into working days, weekends, and festivals, etc. The types of content that the object wants to view may be different on different types of days. For example, on working days (Monday to Friday), the object is usually at work and may want to view content related to work; on weekends (Saturday and Sunday), when the object does not need to work, it may want to view content related to going out for fun; on festivals, the object may want to view content related to festival culture.
[0155] The time period feature and the input date type feature can be obtained from the system time of the terminal 140 of the target object. The system time generally includes information such as year, month, day, day of the week, hour, minute, and second. The input time period feature can be determined from the information of hour, minute, and second. For example, if the current time is 13:52:26, it is determined that it belongs to 13:00 - 18:00, that is, the input time period feature is 13:00 - 18:00. If the date is Tuesday, June 20, 2023, it is determined that the current input date type feature is a working day.
[0156] In addition to the target object feature and the content to be pushed feature, obtaining the target scenario feature takes into account the influence of the scenario factor on the target object's preference, which is beneficial to improving the accuracy of pushing the content to be pushed to the target object.
[0157] Detailed description of step 320
[0158] In step 320, multiple target push basic features are input into the probability prediction model to obtain the first probability that the content to be pushed belongs to the long-tail content that the target object is interested in.
[0159] Based on the foregoing embodiments, the target scenario feature is included in the multiple target push basic features. Therefore, multiple target push basic features are input into the probability prediction model to obtain the first probability that the content to be pushed belongs to the long-tail content that the target object is interested in in the current scenario.
[0160] The probability prediction model can be a single-task prediction model only used to predict whether the content to be pushed belongs to the long-tail content that the target object is interested in. Based on this, the structure of the probability prediction model can be as Figure 4 shown, specifically including: an embedding layer, a concatenation layer, and a long-tail content prediction sub-model.
[0161] The role of the embedding layer is to convert the target object feature and the content to be pushed feature into a vector form that can be processed by the model. Since the data of the target object feature and the content to be pushed feature have various types, such as character type or character sequence, etc. Therefore, after inputting the target push basic features, it is necessary to uniformly convert the features of different data types into vector forms.
[0162] The role of the concatenation layer is to concatenate multiple feature vectors obtained according to multiple target push basic features into a single vector to generate a concatenated vector. The concatenated vector can be obtained by concatenating multiple feature vectors end to end. For example, multiple feature vectors include: [1,0,1,1,0,1,1,0], [0,0,1,0,1,0,1,0], and [0,0,0,1,1,0,1,1]. The concatenated vector obtained by concatenating the feature vectors is: [1,0,1,1,0,1,1,0,0,0,1,0,1,0,1,0,0,0,0,1,1,0,1,1].
[0163] The long-tail content prediction sub-model is used to predict the first probability that the content to be pushed belongs to the long-tail content of interest to the target object according to the cascaded vector.
[0164] The probability prediction model can also be a multi-task prediction model that can be used to predict the task completion probabilities of multiple tasks. One of the tasks is to predict whether the content to be pushed belongs to the long-tail content of interest to the target object, and the other tasks refer to the operations performed by the target object on the content to be pushed, such as: click, like, favorite, comment, etc. Therefore, in one embodiment, inputting multiple target push basic features into the probability prediction model to obtain the first probability that the content to be pushed belongs to the long-tail content of interest to the target object includes:
[0165] Inputting multiple target push basic features into the probability prediction model to obtain the first probability that the content to be pushed belongs to the long-tail content of interest to the target object and the third probability of the target object completing multiple other tasks for the content to be pushed.
[0166] Based on this, the structure of the probability prediction model can be as Figure 5 shown, specifically including: an embedding layer, a cascading layer, multiple expert sub-models, a gating node corresponding to each task, and a prediction sub-model corresponding to each task.
[0167] The functions of the embedding layer and the cascading layer are the same as those in the single-task prediction model in the foregoing embodiment, and will not be elaborated here.
[0168] Multiple expert sub-models are used to analyze and calculate each dimension of the cascaded vector in order to balance the sharing and mutual exclusion between multiple tasks. After each expert sub-model analyzes the cascaded vector, it outputs and inputs it into each gating node. Each task in the probability prediction model corresponds to a gating node. The gating node receives the outputs from multiple expert sub-models, and uses the weight vector in the gating node to perform a weighted sum on the outputs of multiple expert sub-models to obtain an intermediate vector corresponding to the task. For example Figure 5 as shown, the long-tail content gating node corresponding to the long-tail content prediction task uses the weight vector in the long-tail content gating node to perform a weighted sum on the outputs of expert sub-model A, expert sub-model B, and expert sub-model C to obtain intermediate vector A. The weight vectors in the gating nodes corresponding to different tasks may be different. This is because for different tasks, the contributions of the same expert sub-model may be different.
[0169] After the intermediate vector is output by the gating node corresponding to the task, the intermediate vector is input into the prediction sub-model corresponding to the task. The prediction sub-model will output the target object's probability of completing the task for the content to be pushed according to the task. For the long-tail content prediction task, the corresponding prediction sub-model will output the first probability that the content to be pushed belongs to the long-tail content that the target object is interested in; for other tasks, such as task A, the corresponding prediction sub-model will output the third probability that the target object completes task A for the content to be pushed; similarly, the prediction sub-model corresponding to task B will output the third probability that the target object completes task B for the content to be pushed.
[0170] Using the multi-task prediction model as the probability prediction model, in addition to the first probability, the third probability of the target object completing other tasks can also be obtained. The third probability can be used to evaluate the degree of interest of the target object in the content to be pushed. Therefore, the multi-task prediction model can comprehensively consider the long-tail probability of the content to be pushed and the degree of interest of the target object in the content to be pushed to determine the content to be pushed to the target object. On the basis of improving the quantity and quality of long-tail content pushed, the accuracy of pushing the content to be pushed to the target object is further improved.
[0171] In order to ensure that the probability prediction model can obtain accurate prediction results, in one embodiment, as Figure 6 shown, the probability prediction model is trained in the following way:
[0172] Step 610: Obtain a set of push basic feature samples. Each push basic feature sample in the set of push basic feature samples includes sample object features and sample content features;
[0173] Step 620: Input the push basic feature samples into the probability prediction model to obtain the long-tail task loss function for pushing the sample content to the sample object;
[0174] Step 630: If it is determined that the sample content corresponding to the push basic feature sample meets the object interest condition and the sample content meets the long-tail sample condition, determine the long-tail task weight of the push basic feature sample as the first value; otherwise, determine the long-tail task weight of the push basic feature sample as the second value, where the first value is greater than the second value;
[0175] Step 640: Generate a total loss function based on the weighted sum of the long-tail task loss functions of the push basic feature samples using the long-tail task weights of the push basic feature samples;
[0176] Step 650: Train the probability prediction model based on the total loss function.
[0177] In step 610, a push basic feature sample set is obtained. The push basic feature sample set includes multiple push basic feature samples. A push basic feature sample indicates an event of pushing sample content to a sample object. Each push basic feature sample includes a sample object feature and a sample content feature. For example, if content A is pushed to object A, then the push basic feature sample contains the sample object feature of object A and the sample content feature of content A. The sample object feature is the same as the target object feature in the foregoing embodiment and will not be elaborated here; the sample content feature is the same as the content feature to be pushed in the foregoing embodiment and will not be elaborated here.
[0178] In one embodiment, the push basic feature sample further includes a sample scenario feature, and the sample scenario feature refers to the scenario feature of the scenario where the sample object views the sample content. The sample scenario feature is the same as the target scenario feature in the foregoing embodiment and will not be elaborated here.
[0179] In one embodiment, obtaining the push basic feature sample set includes: obtaining historical push records in a content push application within a predetermined time period before the current time point; obtaining the push basic feature sample set from the historical push records.
[0180] The historical push record refers to multiple push event records of pushing content to an object within a predetermined time period before the current time point in a content push application. The predetermined time period can be one hour, one day, one week, one month, etc. Each record in the historical push record may include the object feature of the object and the content feature of the pushed content. For example, Table 1 shows the historical push records in a content push application within one day before the current time point.
[0181]
[0182] Table 1
[0183] Through the historical push records in the content push application within one day before the current time point shown in Table 1, multiple push events can be obtained. The object feature in each push event can be used as the sample object feature in a push basic feature sample, and the pushed content feature can be used as the sample content feature. For example, taking the first push event as a push basic feature sample, the sample object features include: sports, undergraduate, music, etc., and the sample content features include: movie, 15 minutes, etc.
[0184] Multiple push basic feature samples obtained from multiple push events in the historical push records constitute a push basic feature sample set. The advantage of using the historical push records to obtain the push basic feature sample set is that the real push events that occur in the content push application are used as training samples, making the trained probability prediction model more suitable for the content push scenario in the content push application and improving the prediction accuracy of the probability prediction model.
[0185] In step 620, the push basic feature sample is input into the probability prediction model to obtain a long-tail task loss function. The long-tail task loss function represents the probability accuracy of the probability prediction model predicting that the sample content belongs to the long-tail content that the sample object is interested in.
[0186] In one embodiment, as Figure 7 shown, inputting the push basic feature sample into the probability prediction model to obtain a long-tail task loss function for pushing the sample content to the sample object includes:
[0187] Step 710: Input the push basic feature sample into the probability prediction model to obtain a second probability that the sample content belongs to the long-tail content that the sample object is interested in;
[0188] Step 720: Obtain the sample label of the push basic feature sample, and the sample label indicates whether the sample object clicks on the sample content;
[0189] Step 730: Based on the sample label and the second probability, obtain a long-tail task loss function for pushing the sample content to the sample object.
[0190] In this embodiment, when the push basic feature sample is input into the probability prediction model, a second probability that the sample content belongs to the long-tail content that the sample is interested in will be obtained. The process of obtaining the second probability is the same as the process of obtaining the first probability in the foregoing embodiment and will not be elaborated here.
[0191] Each push basic feature sample corresponds to a sample label, and the sample label indicates whether the sample object clicks on the sample content. If the sample object clicks on the sample content, the sample label can be 1; otherwise, the sample label can be 0.
[0192] If the second probability is closer to the sample label, it means that the probability prediction model predicts the push basic feature sample more accurately. The long-tail task loss function can be calculated based on the sample label and the second probability using the mean square error, as shown in formula 1 specifically:
[0193] L i =(y i -p i ) 2 (Formula 1).
[0194] In formula 1, Li represents the long-tail task loss function of the i-th push base feature sample in the push base feature sample set; y i represents the sample label; p i represents the second probability. For example, if the sample label of a push base feature sample is 1 and the second probability is 0.8, then the long-tail task loss function of this push base feature sample is 0.04.
[0195] Calculating the long-tail task loss function based on the sample label and the second probability can also be done using cross-entropy, as shown in formula 2 specifically:
[0196] L i = -(y i log(p i ) + (1 - y i )log(1 - p i )) (Formula 2).
[0197] The meanings of each parameter in formula 2 are the same as those in formula 1, and will not be elaborated here. If the sample label of a push base feature sample is 1 and the second probability is 0.8, then the long-tail task loss function of this push base feature sample is -(1 * log(0.8) + 0 * log(0.2)) = 0.097.
[0198] In this embodiment, the long-tail task loss function is calculated using the sample label indicating whether the sample object clicks on the sample object and the second probability, improving the model training efficiency.
[0199] Based on the foregoing embodiment, since the probability prediction model may also be a multi-task prediction model, therefore, in another embodiment, the push base feature sample is input into the probability prediction model to obtain the long-tail task loss function for pushing the sample content to the sample object, including: inputting the push base feature sample into the probability prediction model to obtain the long-tail task loss function for pushing the sample content to the sample object and other task loss functions.
[0200] The long-tail task loss function is the same as that in the above embodiment and will not be elaborated here. The other task loss function is used to evaluate whether the task completion probability of the probability prediction model predicting other tasks is accurate. The calculation process of the other task loss function is similar to the calculation process of calculating the long-tail task loss function in the foregoing embodiment. Using the sample task label corresponding to the other task and the task completion probability of the other task obtained by the probability prediction model based on the push base feature sample, the other task loss function for the sample object to complete the corresponding task for the sample content can be obtained. The sample task label corresponding to the other task indicates whether the sample object completes the corresponding task for the sample content. If completed, the sample task label can be 1, otherwise, the sample task label can be 0.
[0201] The advantage of this embodiment is that for the probability prediction model of predictable multi-tasks, the corresponding task loss function is calculated by using the labels of each task respectively, so that the training process of the probability prediction model can comprehensively consider the prediction accuracy of each task, and improve the accuracy of training the multi-task probability prediction model.
[0202] In step 630, it is necessary to determine the long-tail task weight corresponding to the push basic feature sample. The long-tail task weight indicates the ability of the push basic feature sample to be used for training the probability prediction model to push the long-tail content that the target object is interested in. Therefore, if the sample content in the push basic feature sample is of interest to the sample object and the sample content is long-tail content, then the long-tail sample weight of this push basic feature sample is relatively high. Based on this, if it is determined that the sample content corresponding to the push basic feature sample meets both the object interest condition and the long-tail sample condition, the long-tail sample weight of this push basic feature sample is determined to be the first value, otherwise the long-tail sample weight is determined to be the second value, where the first value is greater than the second value.
[0203] In one embodiment, determining that the sample content corresponding to the push basic feature sample meets the object interest condition includes:
[0204] Determining whether the sample object has completed multiple tasks for the sample content;
[0205] Determining that the sample object has completed a predetermined number of tasks for the sample content.
[0206] In this embodiment, if the sample object has completed a predetermined number of tasks for the sample content, it is determined that the sample content meets the object interest condition. For example, the multiple tasks include: liking, commenting, collecting, full play, etc. The predetermined number is 2, then as long as the sample object has completed any two tasks for the sample content, it can be indicated that the sample content corresponding to the push basic feature sample meets the object interest condition.
[0207] Determining whether the sample object has completed multiple tasks for the sample content can be determined through the sample task label in the push basic feature sample. Therefore, the judgment efficiency of determining whether the sample content meets the object interest condition by determining whether the sample object has completed a predetermined number of tasks for the sample content is relatively high.
[0208] Different objects may have different task completion habits. Some objects may habitually like all the content they have viewed, or may also comment on the content they don't like. Therefore, the task completion situation of the sample object for the sample content may not accurately indicate whether the sample object is interested in the sample content. Therefore, in another embodiment, as Figure 8 shown, determining that the sample content corresponding to the push basic feature sample meets the object interest condition includes:
[0209] Step 810: Determine that the sample object has clicked on the sample content;
[0210] Step 820: Obtain the playback duration of the sample object for the sample content;
[0211] Step 830: Determine that the playback duration is greater than the first threshold.
[0212] In this embodiment, if in the push basic feature sample, the sample object has clicked on the sample content and the playback duration of the sample content is greater than the first threshold, it is determined that the sample content meets the object interest condition.
[0213] Determining that the sample object has clicked on the sample content can be determined by using a sample label indicating whether the sample object has clicked on the sample content. The sample content feature in the push basic feature sample can also include the playback duration of the sample content. Therefore, the playback duration can be directly extracted from the push basic feature sample.
[0214] The first threshold can be determined based on the content duration of all the content in the content push application. In one implementation, after calculating the average content duration of all the content, the duration obtained by multiplying the average content duration by a first predetermined ratio can be used as the first threshold. For example, in a content push application, the average content duration of all the content is 2 minutes and 20 seconds (140 seconds), and the first predetermined ratio is 0.3. Then the first threshold corresponding to this content push application is 2 minutes and 20 seconds (140 seconds) * 0.3 = 42 seconds. That is to say, if the sample object clicks on the sample content and the playback duration of the sample content is greater than 42 seconds, then it can be determined that the sample content in the push basic feature sample meets the object interest condition. The advantage of determining the first threshold based on the content duration of all the content in the content push application is that the same object interest condition is adopted for different content in the content push application, improving the fairness of determining whether the sample content meets the object interest condition.
[0215] The first threshold can also be determined based on the content duration of each sample content, and different sample contents correspond to different first thresholds. In one implementation, the push basic feature sample contains the content duration feature of the sample content. The first threshold corresponding to the sample content can be the content duration multiplied by a second predetermined ratio. For example, the content duration of a sample content is 55 seconds, and the second predetermined ratio is 0.4. Then the first threshold corresponding to this sample content is 55 * 0.4 = 22 seconds. That is to say, for a sample content with a duration of 55 seconds, if the sample object clicks on this content and the playback duration is not less than 22 seconds, then this sample content meets the object interest condition. The advantage of calculating different first thresholds for different sample contents is that different object interest conditions are determined for different sample contents, improving the flexibility and accuracy of determining whether the sample content meets the object interest condition.
[0216] After an object clicks on content it is interested in, there is a very high probability that it will watch for longer than a predetermined duration. Therefore, using the condition that the sample object clicks on the sample content and the playback duration of the sample content is greater than the first threshold as the object interest condition can improve the accuracy of determining whether the sample content in the push basic feature sample meets the object interest condition.
[0217] After determining that the sample content in the push basic feature sample meets the object interest condition, it is necessary to determine whether it meets the long-tail sample condition. The long-tail sample condition is the condition for determining whether the sample content is long-tail content.
[0218] In one embodiment, as Figure 9 shown, determining that the sample content meets the long-tail sample condition includes:
[0219] Step 910: Determine the first click volume of the sample content;
[0220] Step 920: Determine the first exposure volume ratio of the sample content in the push basic feature sample set whose click volume is not greater than the first click volume;
[0221] Step 930: If the first exposure volume ratio is less than the second threshold, determine that the sample content meets the long-tail sample condition.
[0222] The first click volume of the sample content refers to the number of times the sample content is clicked and viewed by the object in the content push application within a predetermined time. The first exposure volume ratio refers to the ratio of the sum of the exposure volumes of the sample content whose click volume is not greater than the first click volume to the sum of the exposure volumes of all the sample content in the push basic feature sample set. The exposure volume refers to the number of times the sample content is pushed to the object. The click volume and exposure volume of the sample content can be obtained as sample content features.
[0223] In one embodiment, as Figure 10 shown, determining the first exposure volume ratio of the sample content in the push basic feature sample set whose click volume is not greater than the first click volume includes:
[0224] Step 1010: Determine the first sum of the first exposure volumes of the sample content in the push basic feature sample set whose click volume is not greater than the first click volume;
[0225] Step 1020: Determine the second sum of the first exposure volumes of each sample content in the push basic feature sample set;
[0226] Step 1030: Based on the first sum and the second sum, determine the first exposure volume ratio.
[0227] In this embodiment, the proportion of the first exposure amount is determined by obtaining the sum of the first exposure amounts of the sample contents with click-through rates not greater than the first click-through rate and the sum of the first exposure amounts of all the sample contents. For example, the values of the click-through rates and exposure amounts of each sample content in the basic feature sample set for pushing are shown in Chart 2.
[0228] Click volume First exposure volume Sample content a 500 3600 Sample content b 120 1500 Sample content c 88 1300 Sample content d 360 2800 Sample content e 32 510 Sample content f 10 200 Sample content g 210 1005
[0229] Table 2
[0230] Based on Table 2, if the click-through rate of sample content b is used as the first click-through rate, the sample contents with click-through rates not greater than the first click-through rate include sample content b, sample content c, sample content e, and sample content f. The first sum of the first exposure amounts corresponding to these four sample objects is 1500 + 1300 + 510 + 200 = 3510, and the second sum of the first exposure amounts of all the sample contents is 3600 + 1500 + 1300 + 2800 + 510 + 200 + 1005 = 10915. Therefore, the proportion of the first exposure amount corresponding to the first click-through rate is 3510 / 10915 = 0.32.
[0231] The method for determining the proportion of the first exposure amount in this embodiment directly calculates using the click-through rate and the first exposure amount of each sample content, improving the accuracy of determining the proportion of the first exposure amount.
[0232] In another embodiment, determining the proportion of the first exposure amount of the sample contents in the basic feature sample set for pushing with click-through rates not greater than the first click-through rate includes: substituting the first click-through rate into the first function to obtain the proportion of the first exposure amount. The first function indicates the mapping from the first click-through rate to the proportion of the first exposure amount.
[0233] The first function needs to be pre-fitted. As Figure 11 shown, the parameters in the first function are fitted by the following method:
[0234] Step 1110: Set multiple critical proportions;
[0235] Step 1120: Obtain the second click-through rate of each fitting content sample in the fitting content sample set;
[0236] Step 1130: Sort the fitting content samples in ascending order of the second click-through rate, and accumulate the second exposure amount proportions of the fitting content samples in the order from front to back. When the second exposure amount proportion reaches each critical proportion, record the second click-through rate as the critical click-through rate corresponding to the critical proportion;
[0237] Step 1140: Fit the parameters in the first function based on the critical proportion and the critical click-through rate corresponding to the critical proportion.
[0238] In this embodiment, first, a plurality of critical ratios are set, and the plurality of critical ratios are arranged in an arithmetic progression. For example, a plurality of critical ratios with a tolerance of 1% are set between 0 and 1, then the plurality of critical ratios are: 0, 1%, 2%, 3%,..., 99%, 1.
[0239] Obtain the second click volume of each fitting content sample in the fitting content sample set. The fitting content samples in the fitting content sample set are all the content obtained within a predetermined time period in the content push application. For example, all the content obtained by the content push application within one month. The second click volume and the second exposure volume of the fitting content sample are the number of clicks and the number of pushes of the sample within the predetermined time period. The sample content in the fitting content sample set and the push basic feature sample set may contain the same content or different content.
[0240] Sort the fitting content samples in ascending order of the second click volume, and accumulate the second exposure volume of each fitting content sample one by one according to the sorting, and calculate the second exposure volume ratio. The second exposure volume ratio is the ratio of the sum of the second exposure volumes of a plurality of fitting content samples to the sum of the second exposure volumes of all fitting content samples in the fitting content sample set. When the second exposure volume ratio reaches each critical ratio, record the corresponding second click volume, and use the second click volume as the critical click volume corresponding to the critical ratio. As Figure 12 shown, if the plurality of critical ratios are: 0, 1%, 2%, 3%,..., 99%, 1. After sorting the plurality of fitting content samples in ascending order of the first click volume of each fitting content sample, start calculating the second exposure volume ratio one by one from the fitting content sample A. The second exposure volume ratio of the fitting content sample A and the fitting content sample B is 0.2%, which does not reach the critical ratio; the second exposure volume ratio from the fitting content sample A to the fitting content sample C is 0.4%, which does not reach the critical ratio; the second exposure volume ratio from the fitting content sample A to the fitting content sample D is 1%, which reaches the critical ratio. Therefore, the second click volume (21) of the fitting content sample D is the critical click volume of 1%; continue to calculate backward until the second exposure volume ratio from the fitting content sample A to the fitting content sample F is 2%, which reaches the critical ratio. Therefore, the second click volume (42) of the fitting content sample F is the critical click volume of 2%. Calculate backward one by one in this way to find the critical click volume corresponding to each critical ratio in turn until the second exposure volume ratio from the fitting content sample A to the last fitting content sample is 1, which reaches the critical ratio. Therefore, the second click volume (2000658) of the fitting content sample N is the critical click volume of 1. Determine the critical click volume corresponding to each critical ratio through the above process, and based on the corresponding relationship between the critical click volume and the critical ratio, a relationship curve as Figure 13 shown can be obtained. Through Figure 13It can be seen that the horizontal axis is the critical click volume, and the vertical axis is the critical proportion. The critical proportion shows a positive correlation with the critical click volume within the range of 0 to 1.
[0241] After determining the corresponding relationship between the critical click volume and the critical proportion, the critical click volume and the critical proportion are used to fit the first function.
[0242] In one embodiment, as Figure 14 shown, based on the critical proportion and the critical click volume corresponding to the critical proportion, the parameters in the first function are fitted, including:
[0243] Step 1410: Substitute the critical click volume into the first function to obtain the actual proportion corresponding to the critical click volume;
[0244] Step 1420: Determine the parameters in the first function based on the first difference between the actual proportion corresponding to the critical click volume and the critical proportion.
[0245] To fit the parameters in the first function, the critical click volume is substituted into the first function. For example, if the first function is F(x), substitute the critical click volume s i corresponding to the i-th critical proportion into the first function to obtain the actual proportion as F(s i ).
[0246] In one embodiment, the parameters of the first function include a first parameter, a second parameter, and a third parameter. As Figure 15 shown, substituting the critical click volume into the first function to obtain the actual proportion corresponding to the critical click volume includes:
[0247] Step 1510: Calculate the first power of the critical click volume as the first intermediate value;
[0248] Step 1520: Calculate the second intermediate value based on the first intermediate value and the second parameter;
[0249] Step 1530: Calculate the second difference between the first constant and the second power of the predetermined base with the second intermediate value;
[0250] Step 1540: Take the third power of the second difference as the actual proportion.
[0251] The first function in this embodiment can be represented by Formula 3:
[0252]
[0253] In Formula 3, a is the first parameter, b is the second parameter, and c is the third parameter. Substitute the critical click volume s i as x into Formula 3, and first calculate the first intermediate value as s i a; Based on the first intermediate value and the second parameter, the second intermediate value is obtained as -bs i a ; The first constant can be set to 1, and the predetermined base can be set to e. Therefore, the second difference is Calculate the third power of the second difference to obtain the actual proportion F(s i ) as
[0254] In this embodiment, by adjusting the values of the three parameters, the first function that best fits the corresponding relationship between the critical click volume and the critical proportion is found, so that the first function can adapt to the data scenarios in different content push applications, improving the accuracy and flexibility of determining the first function.
[0255] In step 1420, use the i-th critical proportion t i Subtract the actual proportion F(s i ) to obtain the first difference. Since the critical proportion represents the accurate value of the corresponding second exposure proportion when the click volume is the critical click volume. Therefore, the purpose of fitting the first function is to make the actual proportion obtained by the first function based on s i as close as possible to t i , that is, by adjusting the parameters in the first function, making the first difference as small as possible.
[0256] In one embodiment, as Figure 16 shown, based on the first difference between the actual proportion corresponding to the critical click volume and the critical proportion, determine the parameters in the first function, including:
[0257] Step 1610, determine the sum of squares of the first differences corresponding to each critical click volume;
[0258] Step 1620, determine the parameter that minimizes the sum of squares as the parameter in the first function.
[0259] In this embodiment, the process of fitting the parameters in the first function can be represented by formula 4:
[0260]
[0261] In formula 4, ω is the parameter in the first function F(x), and M is the number of critical proportions. t i -F(s i ) is the first difference corresponding to the i-th critical click volume. Determine the sum of squares of the first differences corresponding to M critical click volumes, and determine the parameter ω that can minimize the sum of squares.
[0262] Based on the first function in the embodiment of steps 1510 - 1540, the process of fitting the parameters in the first function can be represented by formula 5:
[0263]
[0264] Through Equation 5, the parameters a, b, and c that minimize the sum of squares can be determined.
[0265] The advantage of this embodiment is that it can ensure the uniqueness of the fitted parameters and improve the accuracy of fitting the first function.
[0266] The curve obtained by fitting the parameters in the first function using the embodiment of steps 1410 - 1420 is as Figure 17 shown. By comparing Figure 13 with Figure 17 it can be seen that the curve of the first function corresponding to the critical click volume is basically coincident with the curve of the critical proportion. Therefore, using the embodiment of steps 1410 - 1420 to fit the first function can improve the fitting accuracy.
[0267] The advantage of using the first function to determine the first exposure proportion is that it does not require calculating a large amount of data, improving the efficiency of determining the first exposure proportion.
[0268] The second threshold in step 930 refers to the proportion of long-tail content among all sample contents in the push basic feature sample set. For example, if the second threshold is 0.2, it means that 20% of the sample contents are long-tail content. Since the number of long-tail contents in different content push applications is different, the corresponding second threshold can be determined for different content push application scenarios. Since long-tail content is the content with extremely low exposure volume in content push applications, the proportion of long-tail content among all contents can be determined using the cumulative exposure volume of the content.
[0269] In one embodiment, as Figure 18 shown, the second threshold is determined in the following manner:
[0270] Step 1810: Obtain the third click volume of each push content sample in the push content sample set;
[0271] Step 1820: Arrange the push content samples in descending order of the third click volume, and calculate the third exposure proportion of the push content samples in the order from front to back in the sorting, generating a mapping relationship between each ranking in the sorting and the corresponding third exposure proportion;
[0272] Step 1830: Determine the second threshold through the mapping relationship between the ranking and the third exposure proportion.
[0273] In this embodiment, the push content sample set is a set of contents obtained by the content push application within a predetermined time period before the current time point. The third click volume is the number of times the push content sample is clicked within the predetermined time period.
[0274] Arrange the push content samples in descending order according to the third click volume, and cumulatively calculate the proportion of the third exposure volume of the push content samples in the order from front to back. The proportion of the third exposure volume represents the sum of the third exposure volumes of the push content samples whose rankings are not lower than the ranking corresponding to the third click volume, accounting for the proportion of the sum of the third exposure volumes of all push content samples in the push content sample set. The process of calculating the proportion of the third exposure volume corresponding to each ranking is similar to that of obtaining the proportion of the second exposure volume corresponding to each second click volume in the foregoing embodiment. The difference is that the proportion of the second exposure volume is cumulatively calculated one by one in ascending order according to the second click volume, while the proportion of the third exposure volume is cumulatively calculated one by one in descending order according to the third click volume. Therefore, the specific process of obtaining the third exposure volume will not be elaborated here.
[0275] For example Figure 19 It represents the mapping relationship curve between the ranking and the proportion of the third exposure volume in a content push application. The coordinates represent the rankings arranged in descending order according to the third click volume of each push content sample. Figure 19 In the mapping relationship represented, it ranks from the 1st to the 1020943rd, that is, there are 1020943 push content samples in total. The ordinate represents the proportion of the third exposure volume. As the third click volume decreases, the growth trend of the corresponding proportion of the third exposure volume changes from steep to gentle. This represents that for each decrease in the ranking by one, the third exposure volume brought by the corresponding push content sample becomes less, so the growth rate of the proportion of the third exposure volume becomes slower and slower. For example, the proportion of the third exposure corresponding to the 1st is 0.01, and the proportion of the third exposure volume corresponding to the 2nd is 0.019. That is to say, the third exposure volume of the 1st alone accounts for 0.01 of the overall third exposure volume, while the third exposure volume of the 2nd alone accounts for 0.009 (0.019 - 0.01) of the overall third exposure volume. The third exposure volume of the 10000th is 0.15, and the proportion of the third exposure volume of the 10001st is 0.1525. That is to say, the third exposure volume of the push content sample of the 10001st accounts for 0.0025 of the overall, which is much smaller than the third exposure volumes of the 1st and 2nd.
[0276] In step 1830, the second threshold can be determined by the change in the proportion of the third exposure amount in the mapping relationship curve between the ranking and the proportion of the third exposure amount. Based on the change amount of the proportion of the third exposure amount between two adjacent rankings, the push content samples are divided into popular content and long-tail content. If the change amount between the proportion of the third exposure amount of ranking A and the next ranking B is less than a predetermined ratio, and the change amount between the proportion of the third exposure amount of ranking A and the previous ranking C is greater than the predetermined ratio, this indicates that starting from ranking A, the proportion of the third exposure amount of the subsequent push content samples suddenly drops. The coordinate point of ranking A in the mapping relationship curve can be an inflection point for dividing popular content and long-tail content. For example, the predetermined ratio is 0.01, the proportion of the third exposure amount corresponding to the 9,999th ranking is 0.84, the proportion of the third exposure amount corresponding to the 10,000th ranking is 0.852, and the proportion of the third exposure amount corresponding to the 10,001st ranking is 0.8524. The change amount of the proportion of the third exposure amount between the 9,999th ranking and the 10,000th ranking is 0.012, which is greater than 0.01; the change amount of the proportion of the third exposure amount between the 10,000th ranking and the 10,001st ranking is 0.004, which is less than 0.01. Therefore, the content before the 10,000th ranking can be used as popular content, and the content after the 10,000th ranking can be used as long-tail content.
[0277] Take the proportion of the third exposure amount corresponding to the long-tail content part as the second threshold. As can be seen from Figure 19 it, there are two obvious inflection points in the curve, namely point a and point b. These two inflection points divide the curve into three regions, X, Y, and Z, according to the change in the proportion of the third exposure amount. There are relatively few push content samples in region X, approximately the top 36,000 push content samples, but the proportion of the third exposure amount of these push content samples is the highest, accounting for 80%. It can be determined therefrom that the push content samples in region X are all popular content. The ranking of the push content samples in region Y is approximately from 36,000 to 100,000, a total of 64,000 push content samples, and the proportion of the third exposure amount of these push content samples is approximately 0.92 - 0.8 = 0.12. The ranking of the push content samples in region Z is approximately after the 100,000th ranking, a total of 920,943 push content samples, but the proportion of the third exposure amount of these push content samples is only 0.08. Since the proportion of the third exposure amount in regions Y and Z is extremely different from that in region X, regions Y and Z can be combined to jointly form long-tail content. That is to say, the proportion of long-tail content in this content push application is 0.2. Therefore, the second threshold is 0.2.
[0278] In this embodiment, the advantage of determining the second threshold based on the third click-through rate and the proportion of the third exposure amount of the push content samples is that different second thresholds can be determined for the content click-through rate and exposure amount in different content push applications, improving the flexibility of determining the second threshold.
[0279] If the proportion of the first exposure amount of the sample content is less than the second threshold, it can be indicated that the sample content belongs to long-tail content, and it is determined that the sample content meets the long-tail sample condition.
[0280] The advantage of the embodiments of steps 910-step 930 in determining that the sample content meets the long-tail sample condition is that it determines whether the sample content is long-tail content based on the click volume and exposure amount of the sample content, improving the accuracy of determining long-tail content.
[0281] In step 630, if it is determined that the sample content corresponding to the pushed basic feature sample meets the object interest condition and the sample content meets the long-tail sample condition, the long-tail task weight of the pushed basic feature sample is determined as the first value. In one embodiment, as Figure 20 shown, determining the long-tail task weight of the pushed basic feature sample as the first value includes:
[0282] Step 2010, determining a first candidate value based on the proportion of the first exposure amount;
[0283] Step 2020, determining the second value as the second candidate value;
[0284] Step 2030, determining the larger of the first candidate value and the second candidate value as the first value, and determining the long-tail task weight of the pushed basic feature sample as the first value.
[0285] This embodiment can be represented by formula 6:
[0286] w1 i = max{N1, N2} (Formula 6).
[0287] In formula 6, w1 i represents the first value corresponding to the long-tail task weight of the i-th pushed basic feature sample; N1 represents the first candidate value obtained based on the proportion of the first exposure amount of the pushed basic feature sample; N2 represents the second value, which is used as the second candidate value. The larger of the first candidate value and the second candidate value is used as the first value. Since the second value is the long-tail sample weight of the pushed basic feature sample when the sample content does not meet the object interest condition or the long-tail sample condition. In order to enable the probability prediction model to improve the learning of long-tail content without sacrificing the prediction efficiency of non-long-tail content, the second value can be set to 1 so that the prediction probability model can also obtain complete and valid information from the pushed basic feature samples where the non-long-tail content is located.
[0288] In one embodiment, as Figure 21 shown, determining the first candidate value based on the proportion of the first exposure amount includes:
[0289] Step 2110, obtaining the second threshold and the third threshold;
[0290] Step 2120: Determine a first coefficient based on a second threshold and a third threshold.
[0291] Step 2130: Add the result of multiplying the first coefficient by a first exposure ratio to the third threshold to obtain a first candidate value.
[0292] In this embodiment, the second threshold is the same as the second threshold in the foregoing embodiment, and is used to represent the proportion of long-tail content in all sample contents in the push base feature set. Therefore, it will not be elaborated here.
[0293] The third threshold is the maximum value of the preset long-tail task weight. If the long-tail task weight of a push base feature sample in the push base feature set is extremely large, it will cause the probability prediction model to overfit to this push base feature sample, resulting in a decrease in training accuracy. Therefore, it is necessary to preset the maximum value of the long-tail task weight to control the long-tail task weight of each push base feature sample within an interval.
[0294] In one embodiment, as Figure 22 shown, determining the first coefficient based on the second threshold and the third threshold includes:
[0295] Step 2210: Calculate a third difference between a second constant and the third threshold.
[0296] Step 2220: Divide the third difference by the second threshold to obtain the first coefficient.
[0297] In this embodiment, the second constant must be less than the third threshold. The process of determining the first coefficient can be expressed by Formula 7:
[0298]
[0299] In Formula 7, k is the first coefficient, 1 is set as the second constant, m is the third threshold, and g is the second threshold. For example, the second threshold is 0.2 and the third threshold is 3. Then the first coefficient is (1 - 3) / 0.2 = -10.
[0300] In this embodiment, since the second constant is less than the third threshold, the first coefficient must be negative. Therefore, the first coefficient obtained by Formula 7 can be inversely proportional to the first exposure ratio, that is, the smaller the first exposure ratio, the larger the first candidate value, and the longer the sample content in the push base feature sample is the long-tail. The accuracy of determining different long-tail task weights for different push base feature samples is improved.
[0301] After determining the first coefficient, the result of multiplying the first coefficient by the first exposure ratio is added to the third threshold to obtain the first candidate value. For example, if the first exposure ratio is 0.1, the first coefficient is -10, and the third threshold is 3, then the first candidate value is -10 * 0.1 + 3 = 2.
[0302] The first exposure ratio is calculated using the second threshold and the third threshold to obtain the first candidate value, which can not only consider the long-tail degree of the push-based feature samples but also ensure that the first candidate value does not exceed the maximum value of the preset long-tail task weight, improving the accuracy of determining the first candidate value.
[0303] Based on the above embodiments, the process of determining the first value can be expressed as Formula 8:
[0304]
[0305] In Formula 8, F(x) represents the first exposure ratio obtained based on the first click-through rate of the sample content in the push-based feature samples. The remaining parameters are the same as those in the foregoing embodiments and will not be elaborated here. As Figure 23 shown, it is a curve graph of the first value calculated using Formula 8 based on the first click-through rate of different push-based feature samples. Among them, the abscissa is the first click-through rate, and the ordinate is the first value. It can be seen from the graph that when the first click-through rate is larger, the first value is smaller, and at this time, the first value is determined as the first candidate value; when the first click-through rate reaches a certain value, the first value remains at the second candidate value.
[0306] Therefore, if the first exposure ratio is less than the second threshold, that is to say, the sample content belongs to long-tail content, then the first candidate value must be greater than the second candidate value. Moreover, the smaller the first exposure ratio, the longer the tail of the sample content, and the larger the corresponding first candidate value. If the sample content is not long-tail content, then the first candidate value must be less than or equal to the second candidate value. Thus, it can be seen that Formula 8 can indirectly determine whether the sample content meets the long-tail sample conditions, determine the long-tail task weight of the push-based feature samples that meet the long-tail sample conditions as the first candidate value, and vice versa, determine it as the second candidate value, that is, the second value.
[0307] After using Formula 8 to determine whether the sample content meets the long-tail sample conditions, the long-tail task weight of the push-based feature samples that do not meet the object interest conditions in the long-tail content can also be adjusted to the second value. According to one of the foregoing embodiments, the object interest condition is that the sample object has clicked on the sample content and the playback duration of the sample content is greater than the first threshold. Based on this, the long-tail task weight of the push-based feature samples that meet the object interest conditions is determined as the first value, and vice versa as the second value, which can be specifically expressed as Formula 9:
[0308]
[0309] In Formula 9, w i represents the long-tail task weight of the i-th push basic feature sample; y i = 1 indicates that the sample object clicks on the sample content; t i is the playback duration; α is the first threshold; the second value is preset to 1. Through Formula 9, the long-tail task weight of the push basic feature sample that meets the object interest condition can be determined as the first value, and the long-tail task weights of other push basic feature samples can be determined as the second value. For example, using Formula 8, the first value corresponding to the push basic feature sample A is 2, the first value corresponding to the push basic feature B is 1, and the first value corresponding to the push basic feature sample C is 2.5. Among them, the push basic feature sample A meets the object interest condition, so its long-tail task weight is 2; the push basic feature sample B also meets the object interest condition, so its long-tail task weight is 1; although the first value corresponding to the push basic feature sample C is higher, it does not meet the object interest condition, so its long-tail task weight is adjusted to 1.
[0310] Thus, using Formula 8 and Formula 9, the long-tail task weights of the push basic feature samples that meet the long-tail sample condition and the object interest condition can be directly determined as the first value, and vice versa, as the second value. It is not necessary to pre-determine whether the sample content in the push basic feature sample meets these two conditions, which improves the efficiency of determining the long-tail task weight.
[0311] Therefore, the advantage of using the embodiments of Steps 2010 - 2030 to determine the first value is that different first values can be directly determined for the push basic feature samples corresponding to the long-tail content and the non-long-tail content, which improves the efficiency of determining the long-tail task weight. Moreover, the larger the first exposure ratio, the larger the corresponding first value, which improves the accuracy of determining different long-tail task weights for different push basic feature samples.
[0312] In Step 640, based on the weighted sum of the long-tail task loss functions of the push basic feature samples using the long-tail task weights, a total loss function is generated. For example, the long-tail task loss function of the push basic feature sample A is 0.8, and the long-tail task weight is 2.5; the long-tail task loss function of the push basic feature sample B is 0.6, and the long-tail task weight is 1.5; the long-tail task loss function of the push basic feature sample C is 0.5, and the long-tail task weight is 1; the long-tail task loss function of the push basic feature sample D is 0.4, and the long-tail task weight is 2.2. The weighted sum result is: 0.8 * 2.5 + 0.6 * 1.5 + 0.5 * 1 + 0.4 * 2.2 = 4.28. The weighted sum result can be directly used as the total loss function.
[0313] Since the weighted sum result is the sum of the loss functions of long-tail tasks and has a relatively large value, directly using it to train the probability prediction model may lead to inaccurate training results. Therefore, the total loss function can be the weighted average obtained by dividing the weighted sum result by the number of push-based feature samples. Based on the above example, the total loss function is 4.28 / 4 = 1.07.
[0314] In the foregoing embodiment, the long-tail task loss function can be obtained by calculating the cross-entropy between the sample label and the second probability. Based on this, the formula for generating the total loss function can be as shown in Formula 10:
[0315]
[0316] In Formula 10, L(θ) represents the total loss function, N represents the number of push-based feature samples, w i represents the long-tail task weight of the i-th push-based feature sample, y i represents the sample label, represents the second probability.
[0317] In the foregoing embodiment, the push-based feature samples are input into the probability prediction model to obtain the long-tail task loss function and other task loss functions for pushing sample content to the sample object. Based on this, in one embodiment, as Figure 24 shown, based on the weighted sum of the long-tail task loss functions of the push-based feature samples using the long-tail task weights of the push-based feature samples, a total loss function is generated, including:
[0318] Step 2410: Perform a weighted sum on the long-tail task loss functions of the push-based feature samples using the long-tail task weights of the push-based feature samples to obtain a first result;
[0319] Step 2420: Perform a weighted sum on the other task loss functions of the push-based feature samples using the other task weights of the push-based feature samples to obtain a second result;
[0320] Step 2430: Generate a total loss function based on the first result and the second result.
[0321] In this embodiment, the process of obtaining the first result in Step 2410 is the same as the process of determining the total loss function only based on the long-tail task loss function in the foregoing embodiment, and will not be elaborated here.
[0322] In step 2420, the other task weights of each push basic feature sample and the other task loss functions can be used to perform a weighted sum to obtain a second result. For example, for the like task, the loss function of the like task for push basic feature sample A is 0.3, and the weight of the like task is 1.2; the loss function of the like task for push basic feature sample B is 0.4, and the weight of the like task is 1.5; the loss of the like task for push basic feature sample C is 0.8, and the weight of the like task is 2.1; the loss of the like task for push basic feature sample D is 0.1, and the weight of the like task is 1. The weighted sum result is 0.3 * 1.2 + 0.4 * 1.5 + 0.8 * 2.1 + 0.1 * 1 = 2.74. Obtaining the second result based on the weighted sum result is the same as the process of obtaining the first result in the foregoing embodiments, and will not be elaborated here. In the above example, the second result corresponding to the like task can be 2.74 / 4 = 0.685.
[0323] In step 2420, a total loss function can be generated based on the first result and the second result. The total loss function can be to calculate the average value of the first result and the second result. For example, if the first result is 1.07, the second result corresponding to the like task is 0.685, and the second result corresponding to the comment task is 0.521, then the total loss function is (1.07 + 0.685 + 0.521) / 3 = 0.759. The advantage of using the average value as the total loss function is that the loss functions of different tasks have the same influence on training the probability prediction model, improving the training fairness.
[0324] The total loss function can also be to calculate the weighted average value of the first result and the second result. Each task can correspond to a weight. For example, the task weight corresponding to the long-tail task is 0.5, the task weight corresponding to the like task is 0.3, and the task weight corresponding to the comment task is 0.2. Then the total loss function is 1.07 * 0.5 + 0.685 * 0.3 + 0.521 * 0.2 = 0.8447. The advantage of using the weighted average value as the total loss function is that different weights can be assigned to different tasks according to the actual training needs, improving the training flexibility and the training accuracy in practical applications.
[0325] In step 650, based on the total loss function, the probability prediction model is trained. The purpose of training the probability prediction model is to reduce the total loss so that the prediction result is as close as possible to the push basic feature sample. The training process can be to adjust the parameters of the probability prediction model in the direction of the gradient descent of the total loss function.
[0326] The advantages of the embodiments of steps 610 - 650 are that the trained probability prediction model can improve the prediction accuracy of the long-tail content that the prediction object is interested in without sacrificing the push efficiency of popular content, thereby improving the quantity and quality of the long-tail content pushed by using the probability prediction model.
[0327] Detailed description of step 330
[0328] In step 330, push the content to be pushed to the target object based on the first probability.
[0329] In one embodiment, pushing the content to be pushed to the target object based on the first probability includes:
[0330] Sort the content to be pushed in descending order according to the first probability;
[0331] Push the first predetermined number of the content to be pushed in the sorting to the target object.
[0332] For example, after sorting the content to be pushed according to the first probability, the first probability of the content to be pushed A is 0.88, the first probability of the content to be pushed B is 0.62, the first probability of the content to be pushed C is 0.42, the first probability of the content to be pushed D is 0.31, and the first probability of the content to be pushed E is 0.21. If the predetermined number of the content to be pushed to be pushed to the target object is 3, then push the content to be pushed A, the content to be pushed B, and the content to be pushed C to the target object.
[0333] The advantage of pushing the first predetermined number of the content to be pushed in the sorting to the target object is that it ensures that a sufficient number of the content to be pushed can be pushed to the target object, enables the probability prediction model to obtain more effective information for pushing, and further improves the personalized recommendation ability of the probability prediction model.
[0334] In one embodiment, pushing the content to be pushed to the target object based on the first probability includes:
[0335] Obtain a probability threshold;
[0336] Push the content to be pushed with the first probability not lower than the probability threshold to the target object.
[0337] Based on the above example, if the probability threshold is 0.5, then push the content to be pushed A and the content to be pushed B to the target object.
[0338] The advantage of pushing the content to be pushed to the target object based on the probability threshold is that it ensures that the content to be pushed to the target object is of interest to the target object, which is beneficial to improving the pushing accuracy.
[0339] Since in the foregoing embodiment, multiple target push basic features are input into the probability prediction model to obtain the first probability that the content to be pushed belongs to the long-tail content of interest to the target object and the third probability that the target object completes multiple other tasks for the content to be pushed. Based on this, pushing the content to be pushed to the target object based on the first probability includes: pushing the content to be pushed to the target object based on the first probability and the third probability of each other task.
[0340] In one embodiment, as Figure 25 shown, pushing the content to be pushed to the target object based on the first probability and the third probability of each other task includes:
[0341] Step 2510, determining the first task weight of the long-tail task corresponding to the first probability and the second task weights of multiple other tasks;
[0342] Step 2520, obtaining a multi-task score according to the first probability and the first task weight, the third probability of each other task and the second task weight;
[0343] Step 2530, pushing the content to be pushed to the target object according to the multi-task score.
[0344] In this embodiment, the long-tail task is the task for determining the first probability that the content to be pushed belongs to the long-tail content of interest to the target object in the probability prediction model. The first task weight indicates the influence degree of the long-tail task on the target object's preference for the content to be pushed. The second task weight indicates the influence degree of other tasks on the target object's preference for the content to be pushed.
[0345] In one embodiment, obtaining a multi-task score according to the first probability and the first task weight, the third probability of each other task and the second task weight includes:
[0346] Multiplying the first probability by the first task weight to obtain a first long-tail task result;
[0347] Performing a weighted sum of the third probability using the second task weight of each other task to obtain a first other task result;
[0348] Adding the first long-tail content result and the first other task result to obtain a multi-task score.
[0349] For example, the first task weight corresponding to the long-tail task is 0.3, the first probability is 0.8, and the first long-tail task result is 0.3 * 0.8 = 0.24. Among other tasks, the second task weight corresponding to the like task is 0.5, and the corresponding third probability is 0.6; the second task weight corresponding to the comment task is 0.2, and the corresponding third probability is 0.4; the first other task result is 0.5 * 0.6 + 0.2 * 0.4 = 0.38. Therefore, the multi-task score is 0.24 + 0.38 = 0.62.
[0350] The advantage of using the result of the weighted sum to determine the multi-task score is that
[0351] In another embodiment, obtaining a multi-task score according to the first probability and the first task weight, the third probability of each other task and the second task weight includes:
[0352] Perform a power operation with the first probability as the base and the first task weight as the exponent to obtain a second long-tail task result;
[0353] Perform a power operation with the third probability of each other task as the base and the second task weight as the exponent corresponding to the third probability, and multiply the power operation results corresponding to each other task to obtain a second other task result;
[0354] Multiply the second long-tail task result by the second other task result to obtain a multi-task score.
[0355] For example, the first task weight corresponding to the long-tail task is 0.3, the first probability is 0.8, and the second long-tail task result is 0.8^0.3 = 0.935. Among other tasks, the second task weight corresponding to the like task is 0.5, and the corresponding third probability is 0.6; the second task weight corresponding to the comment task is 0.2, and the corresponding third probability is 0.4; the second other task result is 0.6^0.5 * 0.4^0.2 = 0.645. Therefore, the multi-task score is 0.935 * 0.645 = 0.603.
[0356] Pushing the content to be pushed to the target object according to the multi-task score is similar to the process of pushing the content to be pushed to the target object based on the first probability in the foregoing embodiments, and will not be elaborated here.
[0357] Pushing the content to be pushed to the target object according to the first probability of the long-tail task and the first task weight, and the third probability and the second task weight of each other task can adjust the weights corresponding to different tasks according to the needs of actual applications to determine the influence of different tasks on content pushing, and improve the flexibility of pushing the content to be pushed to the target object.
[0358] The advantage of pushing the content to be pushed to the target object using the first probability and the third probability corresponding to other tasks is that it can be pushed in combination with the task completion situation of the target object for other tasks, improving the pushing accuracy.
[0359] Implementation details of the pushing processing method of the present disclosure embodiment
[0360] The following refers to Figure 26 , and details of the implementation of the pushing processing method of the present disclosure embodiment will be described in detail and exemplarily.
[0361] In step 2610, obtain multiple target push basic features, and the multiple target push basic features include target object features and content-to-be-pushed features.
[0362] In step 2620, multiple target push basic features are input into the probability prediction model to obtain a first probability that the content to be pushed belongs to the long-tail content of interest to the target object, and a third probability that the target object completes multiple other tasks for the content to be pushed.
[0363] Among them, the training details of the probability prediction model include:
[0364] Step 2621: Obtain a push basic feature sample set. Each push basic feature sample in the push basic feature sample set includes a sample object feature and a sample content feature;
[0365] Step 2622: Input the push basic feature sample into the probability prediction model to obtain a long-tail task loss function and other task loss functions for pushing the sample content to the sample object;
[0366] Step 2623: Determine that the sample object has clicked on the sample content; obtain the playback duration of the sample content by the sample object; determine that the playback duration is greater than the first threshold. If so, execute step 2624; if not, execute step 2627;
[0367] Step 2624: Determine the first click volume of the sample content;
[0368] Step 2625: Substitute the first click volume into the first function to obtain the first exposure volume ratio; if the first exposure volume ratio is less than the second threshold, determine that the sample content meets the long-tail sample condition. If so, execute step 2626; if not, execute step 2627; Among them, the fitting process of the first function includes:
[0369] Step 262501: Set multiple critical ratios, where the multiple critical ratios are arranged in an arithmetic progression; obtain the second click volume of each fitting content sample in the fitting content sample set; sort the fitting content samples in ascending order of the second click volume, and accumulate the second exposure volume ratios of the fitting content samples in the order from front to back, and record the second click volume when the second exposure volume ratio reaches each critical ratio as the critical click volume corresponding to the critical ratio;
[0370] Step 262502: The parameters include a first parameter, a second parameter, and a third parameter. Calculate the first power of the critical click volume as the first intermediate value; calculate the second intermediate value based on the first intermediate value and the second parameter; calculate the second difference between the first constant and the second intermediate value power of the predetermined base; calculate the third power of the second difference as the actual ratio;
[0371] Step 262503: Determine the sum of the squares of the first differences corresponding to each critical click volume; determine the parameters that minimize the sum of squares as the parameters in the first function;
[0372] Step 2626: Determine the long-tail task weight of the pushed basic feature sample as the first value;
[0373] Step 2627: Determine the long-tail task weight of the pushed basic feature sample as the second value;
[0374] Step 2628: Perform a weighted sum on the long-tail task loss function of the pushed basic feature sample using the long-tail task weight of the pushed basic feature sample to obtain a first result; perform a weighted sum on the other task loss functions of the pushed basic feature sample using the other task weights of the pushed basic feature sample to obtain a second result; generate a total loss function based on the first result and the second result;
[0375] Step 2629: Train a probability prediction model based on the total loss function.
[0376] In step 2630, determine the first task weight of the long-tail task corresponding to the first probability and the second task weights of multiple other tasks;
[0377] In step 2640, perform a power operation with the first probability as the base and the first task weight as the exponent to obtain a long-tail task result; perform a power operation with the third probability of each other task as the base and the second task weight as the exponent corresponding to the third probability, and multiply the power operation results corresponding to each other task to obtain an other task result; multiply the long-tail task result by the other task result to obtain a multi-task score.
[0378] In step 2650, push the content to be pushed to the target object according to the multi-task score.
[0379] After using the push processing method of the embodiments of the present disclosure, the push effect of a content push platform is as Figure 27A as Figure 27B shown. Figure 27A Shows the change in the content click volume before and after using the push processing method of the embodiments of the present disclosure. The content click volume is the number of times an object clicks on the content, and its increase can reflect an increase in the long-tail content that the object is interested in. After using the push processing method of the embodiments of the present disclosure (after September 1), the content click volume increased by an average of 3.87% compared to before September 1. Figure 27B Shows the change in the total content playback duration before and after using the push processing method of the embodiments of the present disclosure. The increase in the total content playback duration reflects an increase in the object's interest in the pushed content. After September 1, the total content playback duration increased by an average of 1.79% compared to before. Based on Figure 27A and Figure 27B the data, it can be seen that after applying the push processing method of the embodiments of the present disclosure, the number of long-tail content pushed can be increased, and at the same time, the interest correlation between the pushed long-tail content and the object can also be increased.
[0380] Description of the devices and equipment according to the embodiments of the present disclosure
[0381] It can be understood that although the steps in the above various flowcharts are sequentially shown according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this embodiment, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the above flowcharts may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0382] It should be noted that in each specific embodiment of the present application, when it comes to performing relevant processing based on data related to the target content characteristics such as target content attribute information or attribute information set, the permission or consent of the target content will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiments of the present application need to obtain target content attribute information, the separate permission or separate consent of the target content will be obtained by means of pop-up windows or jumping to a confirmation page. After clearly obtaining the separate permission or separate consent of the target content, the necessary target content-related data for the normal operation of the embodiments of the present application will be obtained.
[0383] Figure 28 FIG. 12 is a schematic structural diagram of a push processing device 2800 provided by an embodiment of the present disclosure. The push processing device 2800 includes:
[0384] An obtaining unit 2810, configured to obtain a plurality of target push basic features, where the plurality of target push basic features include target object features and to-be-pushed content features;
[0385] An input unit 2820, configured to input the plurality of target push basic features into a probability prediction model to obtain a first probability that the to-be-pushed content belongs to the long-tail content of interest to the target object. The probability prediction model is trained based on a total loss function, and the total loss function is a weighted sum of the long-tail task loss functions of the push basic feature samples in the push basic feature sample set using the long-tail task weights of the respective push basic feature samples. If the sample content corresponding to the push basic feature sample meets the object interest condition and the sample content meets the long-tail sample condition, the long-tail task weight of the push basic feature sample is determined to be a first value; otherwise, the long-tail task weight of the push basic feature sample is determined to be a second value, where the first value is greater than the second value;
[0386] A pushing unit 2830, configured to push content to be pushed for a target object based on a first probability.
[0387] Optionally, the probability prediction model is trained in the following manner:
[0388] Obtain a set of pushing basic feature samples, where each pushing basic feature sample in the set of pushing basic feature samples includes a sample object feature and a sample content feature;
[0389] Input the pushing basic feature samples into the probability prediction model to obtain a long-tail task loss function for pushing the sample content for the sample object;
[0390] If it is determined that the sample content corresponding to the pushing basic feature sample meets the object interest condition and it is determined that the sample content meets the long-tail sample condition, determine the long-tail task weight of the pushing basic feature sample as a first value; otherwise, determine the long-tail task weight of the pushing basic feature sample as a second value, where the first value is greater than the second value;
[0391] Generate a total loss function based on the weighted sum of the long-tail task loss functions of the pushing basic feature samples using the long-tail task weights of the pushing basic feature samples;
[0392] Train the probability prediction model based on the total loss function.
[0393] Optionally, determining that the sample content corresponding to the pushing basic feature sample meets the object interest condition includes:
[0394] Determine that the sample object has clicked on the sample content;
[0395] Obtain the playing duration of the sample content by the sample object;
[0396] Determine that the playing duration is greater than a first threshold.
[0397] Optionally, determining that the sample content meets the long-tail sample condition includes:
[0398] Determine the first click volume of the sample content;
[0399] Determine the first exposure volume ratio of the sample content in the set of pushing basic feature samples whose click volume is not greater than the first click volume;
[0400] If the first exposure volume ratio is less than a second threshold, determine that the sample content meets the long-tail sample condition.
[0401] Optionally, determining the first exposure volume ratio of the sample content in the set of pushing basic feature samples whose click volume is not greater than the first click volume includes:
[0402] Determine the first sum of the first exposure volumes of the sample content in the set of pushing basic feature samples whose click volume is not greater than the first click volume;
[0403] Determine the second sum of the first exposure amounts of each sample content in the push base feature sample set;
[0404] Based on the first sum and the second sum, determine the first exposure amount ratio.
[0405] Optionally, determining the first exposure amount ratio of the sample content in the push base feature sample set whose click amount is not greater than the first click amount includes:
[0406] Substitute the first click amount into the first function to obtain the first exposure amount ratio, where the parameters in the first function are fitted by the following method:
[0407] Set multiple critical ratios, where the multiple critical ratios are arranged in an arithmetic progression;
[0408] Obtain the second click amount of each fitting content sample in the fitting content sample set;
[0409] Sort the fitting content samples in ascending order of the second click amount, and accumulate the second exposure amount ratios of the fitting content samples in the order from front to back, and record the second click amount when the second exposure amount ratio reaches each critical ratio as the critical click amount corresponding to the critical ratio;
[0410] Based on the critical ratio and the critical click amount corresponding to the critical ratio, fit the parameters in the first function.
[0411] Optionally, based on the critical ratio and the critical click amount corresponding to the critical ratio, fitting the parameters in the first function includes:
[0412] Substitute the critical click amount into the first function to obtain the actual ratio corresponding to the critical click amount;
[0413] Based on the first difference between the actual ratio corresponding to the critical click amount and the critical ratio, determine the parameters in the first function.
[0414] Optionally, based on the first difference between the actual ratio corresponding to the critical click amount and the critical ratio, determining the parameters in the first function includes:
[0415] Determine the sum of the squares of the first differences corresponding to each critical click amount;
[0416] Determine the parameter that minimizes the sum of squares as the parameter in the first function.
[0417] Optionally, the parameters include a first parameter, a second parameter, and a third parameter;
[0418] Substitute the critical click amount into the first function to obtain the actual ratio corresponding to the critical click amount, including:
[0419] Calculate the first-parameter power of the critical click volume as the first intermediate value;
[0420] Calculate the second intermediate value based on the first intermediate value and the second parameter;
[0421] Calculate the second difference between the first constant and the second-intermediate-value power of the predetermined base number;
[0422] Take the third-parameter power of the second difference as the actual proportion.
[0423] Optionally, input the push base feature sample into the probability prediction model to obtain the long-tail task loss function for pushing sample content to the sample object, including:
[0424] Input the push base feature sample into the probability prediction model to obtain the long-tail task loss function and other task loss functions for pushing sample content to the sample object;
[0425] Generate the total loss function based on the weighted sum of the long-tail task loss function of the push base feature sample weighted by the long-tail task weight of the push base feature sample, including:
[0426] Perform a weighted sum of the long-tail task loss function of the push base feature sample using the long-tail task weight of the push base feature sample to obtain the first result;
[0427] Perform a weighted sum of the other task loss functions of the push base feature sample using the other task weights of the push base feature sample to obtain the second result;
[0428] Generate the total loss function based on the first result and the second result.
[0429] Optionally, determine the long-tail task weight of the push base feature sample as the first value, including:
[0430] Determine the first candidate value based on the first exposure proportion;
[0431] Determine the second value as the second candidate value;
[0432] Determine the larger of the first candidate value and the second candidate value as the first value, and determine the long-tail task weight of the push base feature sample as the first value.
[0433] Optionally, determine the first candidate value based on the first exposure proportion, including:
[0434] Obtain the second threshold and the third threshold;
[0435] Determine the first coefficient based on the second threshold and the third threshold;
[0436] Add the result of multiplying the first coefficient by the first exposure proportion to the third threshold to obtain the first candidate value.
[0437] Optionally, determining a first coefficient based on a second threshold and a third threshold includes:
[0438] Calculating a third difference between a second constant and the third threshold, where the second constant is less than the third threshold;
[0439] Dividing the third difference by the second threshold to obtain the first coefficient.
[0440] Optionally, inputting a push base feature sample into a probability prediction model to obtain a long-tail task loss function for pushing sample content to a sample object, including:
[0441] Inputting the push base feature sample into the probability prediction model to obtain a second probability that the sample content belongs to the long-tail content of interest to the sample object;
[0442] Obtaining a sample label of the push base feature sample, where the sample label indicates whether the sample object clicks on the sample content;
[0443] Based on the sample label and the second probability, obtaining a long-tail task loss function for pushing the sample content to the sample object.
[0444] Optionally, the input unit 2820 is specifically configured to:
[0445] Inputting multiple target push base features into the probability prediction model to obtain a first probability that the content to be pushed belongs to the long-tail content of interest to the target object, and a third probability that the target object completes multiple other tasks for the content to be pushed;
[0446] Pushing the content to be pushed to the target object based on the first probability, including:
[0447] Pushing the content to be pushed to the target object based on the first probability and the third probability of each other task.
[0448] Optionally, the input unit 2820 is specifically configured to:
[0449] Determining a first task weight of the long-tail task corresponding to the first probability, and second task weights of multiple other tasks;
[0450] Obtaining a multi-task score according to the first probability and the first task weight, the third probability of each other task and the second task weight;
[0451] Pushing the content to be pushed to the target object according to the multi-task score.
[0452] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the function of that module or unit.
[0453] Referring to Figure 29 , Figure 29 , FIG. 140 is a block diagram of a part of a terminal for implementing the push processing method according to an embodiment of the present disclosure. The terminal 140 includes components such as a radio frequency (RF) circuit 2910, a memory 2915, an input unit 2930, a display unit 2940, a sensor 2950, an audio circuit 2960, a wireless fidelity (WiFi) module 2970, a processor 2980, and a power supply 2990. Those skilled in the art can understand that Figure 29 The structure of the terminal 140 shown does not limit a mobile phone or a computer, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0454] The RF circuit 2910 can be used for receiving and transmitting signals during information reception or call processes. Specifically, after receiving the downlink information from the base station, it is given to the processor 2980 for processing; in addition, the uplink data designed is sent to the base station.
[0455] The memory 2915 can be used to store software programs and modules. The processor 2980 executes various functional applications and data processing of the content terminal 140 by running the software programs and modules stored in the memory 2915.
[0456] The input unit 2930 can be used to receive input digital or character information, and generate key signal inputs related to the settings and function controls of the content terminal 140. Specifically, the input unit 2930 may include a touch panel 2931 and other input devices 2932.
[0457] The display unit 2940 can be used to display input information or provided information, as well as various menus of the content terminal 140. The display unit 2940 may include a display panel 2941.
[0458] The audio circuit 2960, the speaker 2961, and the microphone 2962 can provide an audio interface.
[0459] In this embodiment, the processor 2980 included in the terminal 140 may execute the push processing method of the previous embodiment.
[0460] The terminal 140 of the embodiments of the present disclosure includes, but is not limited to, mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, aircraft, etc. The embodiments of the present invention can be applied to various scenarios, including but not limited to big data, artificial intelligence, etc.
[0461] Figure 30 The structure block diagram of part of the server 110 for implementing the push processing method of the embodiments of the present disclosure. The server 110 may vary greatly due to configuration or performance, and may include one or more central processing units (CPUs) 3022 (for example, one or more processors) and a memory 3032, and one or more storage media 3030 (for example, one or more mass storage devices) for storing application programs 3042 or data 3044. Among them, the memory 3032 and the storage media 3030 may be transient storage or persistent storage. The program stored in the storage media 3030 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server 110. Further, the central processing unit 3022 may be configured to communicate with the storage media 3030 and execute a series of instruction operations in the storage media 3030 on the server 110.
[0462] The server 110 may further include one or more power supplies 3026, one or more wired or wireless network interfaces 3050, one or more input / output interfaces 3058, and / or one or more operating systems 3041, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.
[0463] The central processing unit 3022 in the server 110 may be used to execute the push processing method of the embodiments of the present disclosure.
[0464] The embodiments of the present disclosure further provide a computer-readable storage medium, and the computer-readable storage medium is used to store program codes, and the program codes are used to execute the push processing methods of the foregoing various embodiments.
[0465] The embodiments of the present disclosure further provide a computer program product, and the computer program product includes a computer program. The processor of the computer device reads and executes the computer program, so that the computer device executes to implement the above-mentioned push processing method.
[0466] In the description of the present disclosure and the above-mentioned accompanying drawings, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar content and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprising" and "including" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0467] It should be understood that in the present disclosure, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated content and indicates that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the content before and after is an "or" relationship. "At least one (one) of the following" or its similar expression means any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0468] It should be understood that in the description of the embodiments of the present disclosure, the meaning of "a plurality (or multiple items)" is more than two. Understandings such as "greater than", "less than", "exceeding", etc. do not include the present number, and understandings such as "above", "below", "within", etc. include the present number.
[0469] In several embodiments provided by the present disclosure, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.
[0470] The unit described as a separate component may or may not be physically separated, and the component displayed as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0471] In addition, each functional unit in various embodiments of the present disclosure may be integrated in a processing unit, may exist separately as individual physical units, or two or more units may be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0472] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present disclosure. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0473] It should also be understood that the various embodiments provided in the present disclosure can be combined arbitrarily to achieve different technical effects.
[0474] The above is a specific description of the embodiments of the present disclosure, but the present disclosure is not limited to the above embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present disclosure, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present disclosure.
Claims
1. A push processing method, characterized in that, it includes: Obtain multiple target push basic features, where the multiple target push basic features include target object features and content features to be pushed; Input the multiple target push basic features into a probability prediction model to obtain a first probability that the content to be pushed belongs to the long-tail content of interest to the target object, where the probability prediction model is trained based on a total loss function, and the total loss function is the weighted sum of the long-tail task loss functions of each push basic feature sample in the push basic feature sample set using the long-tail task weights of the push basic feature samples. If the sample content corresponding to the push basic feature sample meets the object interest condition and the sample content meets the long-tail sample condition, the long-tail task weight of the push basic feature sample is determined to be a first value; otherwise, the long-tail task weight of the push basic feature sample is determined to be a second value, where the first value is greater than the second value; Push the content to be pushed to the target object based on the first probability.
2. The push processing method according to claim 1, characterized in that, the probability prediction model is trained in the following manner: Obtain a push basic feature sample set, where each push basic feature sample in the push basic feature sample set includes a sample object feature and a sample content feature; Input the push basic feature sample into the probability prediction model to obtain the long-tail task loss function for pushing the sample content to the sample object; If it is determined that the sample content corresponding to the push basic feature sample meets the object interest condition and it is determined that the sample content meets the long-tail sample condition, the long-tail task weight of the push basic feature sample is determined to be the first value; otherwise, the long-tail task weight of the push basic feature sample is determined to be the second value, where the first value is greater than the second value; Generate a total loss function based on the weighted sum of the long-tail task loss functions of the push basic feature samples using the long-tail task weights of the push basic feature samples; Train the probability prediction model based on the total loss function.
3. The push processing method according to claim 2, characterized in that, the determination that the sample content corresponding to the push basic feature sample meets the object interest condition includes: Determine that the sample object has clicked on the sample content; Obtain the playback duration of the sample content by the sample object; Determine that the playback duration is greater than a first threshold.
4. The push processing method according to claim 2, characterized in that, the determination that the sample content meets the long-tail sample condition includes: Determine the first click volume of the sample content; Determine the first exposure volume ratio of the sample content in the push basic feature sample set whose click volume is not greater than the first click volume; If the first exposure volume ratio is less than a second threshold, determine that the sample content meets the long-tail sample condition.
5. The push processing method according to claim 4, characterized in that, Determining the first exposure ratio of the sample content with a click volume not greater than the first click volume in the push basic feature sample set includes: Determining the first sum of the first exposure volumes of the sample content with a click volume not greater than the first click volume in the push basic feature sample set; Determining the second sum of the first exposure volumes of each sample content in the push basic feature sample set; Based on the first sum and the second sum, determining the first exposure ratio.
6. The push processing method according to claim 4, wherein, Determining the first exposure ratio of the sample content with a click volume not greater than the first click volume in the push basic feature sample set includes: Substituting the first click volume into a first function to obtain the first exposure ratio, wherein the parameters in the first function are fitted by the following method: Setting a plurality of critical ratios, wherein the plurality of critical ratios are arranged in an arithmetic progression; Obtaining the second click volume of each fitting content sample in the fitting content sample set; Sorting the fitting content samples in ascending order of the second click volume, and accumulating the second exposure ratios of the fitting content samples in the order from front to back according to the sorting, and recording the second click volume when the second exposure ratio reaches each critical ratio as the critical click volume corresponding to the critical ratio; Based on the critical ratio and the critical click volume corresponding to the critical ratio, fitting the parameters in the first function.
7. The push processing method according to claim 6, wherein, The fitting the parameters in the first function based on the critical ratio and the critical click volume corresponding to the critical ratio includes: Substituting the critical click volume into the first function to obtain the actual ratio corresponding to the critical click volume; Based on the first difference between the actual ratio corresponding to the critical click volume and the critical ratio, determining the parameters in the first function.
8. The push processing method according to claim 7, wherein, The determining the parameters in the first function based on the first difference between the actual ratio corresponding to the critical click volume and the critical ratio includes: Determining the sum of squares of the first differences corresponding to each critical click volume; Determining the parameter that minimizes the sum of squares as the parameter in the first function.
9. The push processing method according to claim 7, wherein, The parameters include a first parameter, a second parameter, and a third parameter; The substituting the critical click volume into the first function to obtain the actual ratio corresponding to the critical click volume includes: Calculating the first power of the critical click volume as a first intermediate value; Based on the first intermediate value and the second parameter, calculating a second intermediate value; Calculating the second difference between a first constant and the second intermediate power of a predetermined base number; Taking the third power of the second difference as the actual ratio.
10. The push processing method according to claim 2, wherein, Inputting the push basic feature sample into the probability prediction model to obtain the long-tail task loss function for pushing the sample content to the sample object includes: Inputting the push basic feature sample into the probability prediction model to obtain the long-tail task loss function for pushing the sample content to the sample object and other task loss functions; The generating of the total loss function based on the weighted sum of the long-tail task loss function of the push basic feature sample by the long-tail task weight of the push basic feature sample includes: Performing a weighted sum on the long-tail task loss function of the push basic feature sample with the long-tail task weight of the push basic feature sample to obtain a first result; Performing a weighted sum on the other task loss function of the push basic feature sample with the other task weight of the push basic feature sample to obtain a second result; Generating the total loss function based on the first result and the second result.
11. The push processing method according to claim 4, wherein, The determining of the long-tail task weight of the push basic feature sample as the first value includes: Determining a first candidate value based on the first exposure ratio; Determining the second value as a second candidate value; Determining the larger one of the first candidate value and the second candidate value as the first value, and determining the long-tail task weight of the push basic feature sample as the first value.
12. The push processing method according to claim 11, wherein, The determining of the first candidate value based on the first exposure ratio includes: Obtaining the second threshold and the third threshold; Determining a first coefficient based on the second threshold and the third threshold; Adding the result of multiplying the first coefficient by the first exposure ratio to the third threshold to obtain the first candidate value.
13. The push processing method according to claim 12, wherein, The determining of the first coefficient based on the second threshold and the third threshold includes: Calculating the third difference between the second constant and the third threshold, the second constant being less than the third threshold; Using the third difference divided by the second threshold to obtain the first coefficient.
14. The push processing method according to claim 2, wherein, The inputting of the push basic feature sample into the probability prediction model to obtain the long-tail task loss function for pushing the sample content to the sample object includes: Inputting the push basic feature sample into the probability prediction model to obtain a second probability that the sample content belongs to the long-tail content of interest to the sample object; Obtaining the sample label of the push basic feature sample, the sample label indicating whether the sample object clicks the sample content; Based on the sample label and the second probability, obtaining the long-tail task loss function for pushing the sample content to the sample object.
15. The push processing method according to claim 1, wherein, Inputting the multiple target push basic features into a probability prediction model to obtain a first probability that the content to be pushed belongs to the long-tail content of interest to the target object includes: Inputting the multiple target push basic features into a probability prediction model to obtain a first probability that the content to be pushed belongs to the long-tail content of interest to the target object, and a third probability that the target object completes multiple other tasks for the content to be pushed; Pushing the content to be pushed to the target object based on the first probability includes: Pushing the content to be pushed to the target object based on the first probability and the third probability of each of the other tasks.
16. The push processing method according to claim 15, wherein, Pushing the content to be pushed to the target object based on the first probability and the third probability of each of the other tasks includes: Determining a first task weight of the long-tail task corresponding to the first probability, and second task weights of the multiple other tasks; Obtaining a multi-task score according to the first probability and the first task weight, the third probability of each of the other tasks and the second task weight; Pushing the content to be pushed to the target object according to the multi-task score.
17. A push processing device, wherein, it includes: An acquisition unit, configured to acquire multiple target push basic features, where the multiple target push basic features include target object features and content-to-be-pushed features; An input unit, configured to input the multiple target push basic features into a probability prediction model to obtain a first probability that the content to be pushed belongs to the long-tail content of interest to the target object, where the probability prediction model is trained based on a total loss function, and the total loss function is a weighted sum of the long-tail task loss functions of the push basic feature samples in the push basic feature sample set weighted by the long-tail task weights of the push basic feature samples, where if the sample content corresponding to the push basic feature sample meets the object interest condition and the sample content meets the long-tail sample condition, the long-tail task weight of the push basic feature sample is determined as a first value, otherwise, the long-tail task weight of the push basic feature sample is determined as a second value, where the first value is greater than the second value; A push unit, configured to push the content to be pushed to the target object based on the first probability.
18. An electronic device, including a memory and a processor, where the memory stores a computer program, wherein, When the processor executes the computer program, it implements the push processing method according to any one of claims 1 to 16.
19. A computer-readable storage medium, where the storage medium stores a computer program, wherein, When the computer program is executed by a processor, it implements the push processing method according to any one of claims 1 to 16.
20. A computer program product comprising a computer program which is read and executed by a processor of a computer device, such that the computer device performs the push processing method according to any one of claims 1 to 16.