Push processing method, related device and medium

By generating and fusing multiple feature cross vectors, the problem of insufficient measurement of feature cross results in the prior art is solved, and higher push accuracy is achieved.

CN120075286APending Publication Date: 2025-05-30TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311627054.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-29
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art cannot effectively measure the feature impact between low-order feature crossover results and high-order feature crossover results, resulting in low push accuracy.

Method used

By obtaining the basic features of the target push, a cross feature vector of multiple feature combinations is generated, and these vectors and all cross feature vectors are input into the fusion model to obtain the probability of pushing content to the target object.

Benefits of technology

Effectively measure the feature impact between low-order and high-order feature crossover results, improving push accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120075286A_ABST
    Figure CN120075286A_ABST
Patent Text Reader

Abstract

The invention provides a push processing method, a related device and a medium. The method comprises the following steps: acquiring a first number of target pushing basic features; based on the first number of target pushing basic features, generating a second number of orders of cross feature vectors corresponding to the target pushing basic feature combination; inputting the first number of target pushing basic features into a feature full-cross model to obtain a full-cross feature vector; and inputting the first number of target pushing basic feature vectors, a plurality of second number of order cross feature vectors corresponding to the plurality of target pushing basic feature combinations, and the full cross feature vector into a fusion model to obtain a first probability of pushing the to-be-pushed content to the target object, and pushing the to-be-pushed content to the target object based on the first probability. According to the embodiment of the invention, the feature influence between the low-order cross feature result and the high-order cross feature result can be effectively measured, so that the pushing accuracy is improved. The embodiment of the invention is applied to big data, data processing and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of big data, and particularly to a push processing method, related device, and medium. Background Art

[0002] In current Internet applications, push processing generally inputs target object features and features of content to be pushed (such as short videos, articles, etc.) into a push model to determine the content to be pushed to the target object. When delivering features to the push model, sometimes not only the features themselves need to be input into the push model, but also the influence brought by the interaction between features needs to be considered. For example, the cross feature of feature A and feature B reflects the influence brought by the interaction between feature A and feature B. High-order feature crossing refers to crossing all basic features input into the push model, and low-order feature crossing refers to crossing a part of the basic features input into the push model.

[0003] In the prior art, the processing of low-order feature crossing and high-order feature crossing is usually carried out separately by different sub-model structures, and after obtaining the low-order cross feature result and the high-order cross feature result, the low-order cross feature result and the high-order cross feature result are directly added or concatenated. However, directly adding or concatenating will lose the semantic meaning of the context of high and low-order features, resulting in the inability to effectively measure the feature influence between the low-order cross feature result and the high-order cross feature result.

[0004] It can be seen that the prior art cannot effectively measure the feature influence between the low-order cross feature result and the high-order cross feature result, resulting in low push accuracy. Summary of the Invention

[0005] Embodiments of the present disclosure provide a push processing method, related device, and medium, which can effectively measure the feature influence between the low-order feature crossing result and the high-order feature crossing result, thereby improving push accuracy.

[0006] According to one aspect of the present disclosure, there is provided a push processing method, including:

[0007] Obtain a first number of target push basic features, where the first number of the target push basic features includes target object features and features of content to be pushed;

[0008] Based on the first number of the target push basic features, obtain a plurality of target push basic feature combinations, each of the target push basic feature combinations including a second number of the target push basic features selected from the first number of the target push basic features, and the second number is less than the first number;

[0009] Generate a second-order cross feature vector corresponding to the target push basic feature combination;

[0010] Input the first number of the target push basic features into the feature full-cross model to obtain the full-cross feature vector;

[0011] Input the first number of target push basic feature vectors corresponding to the first number of the target push basic features, the second number of cross-order feature vectors corresponding to multiple combinations of the target push basic features, and the full-cross feature vector into the fusion model to obtain the first probability of pushing the content to be pushed for the target object, and push the content to be pushed for the target object based on the first probability.

[0012] According to one aspect of the present disclosure, there is provided a push processing device, including:

[0013] A first acquisition unit, configured to acquire the first number of target push basic features, where the first number of the target push basic features includes target object features and content-to-be-pushed features;

[0014] A second acquisition unit, configured to acquire multiple combinations of target push basic features based on the first number of the target push basic features, where each combination of the target push basic features includes the second number of the target push basic features selected from the first number of the target push basic features, and the second number is less than the first number;

[0015] A first generation unit, configured to generate the second number of cross-order feature vectors corresponding to the combination of the target push basic features;

[0016] A second generation unit, configured to input the first number of the target push basic features into the feature full-cross model to obtain the full-cross feature vector;

[0017] A prediction unit, configured to input the first number of target push basic feature vectors corresponding to the first number of the target push basic features, the second number of cross-order feature vectors corresponding to multiple combinations of the target push basic features, and the full-cross feature vector into the fusion model to obtain the first probability of pushing the content to be pushed for the target object, and push the content to be pushed for the target object based on the first probability.

[0018] Optionally, the prediction unit is specifically configured to:

[0019] Take each of the first number of the target push basic feature vectors, the second number of cross-order feature vectors, and the full-cross feature vector as an input vector, and determine an input vector weight vector, where the input vector weight vector indicates the input vector weight of each input vector;

[0020] Using the weights of the input vectors, perform a weighted sum on the first output vector of the fusion model for the input vectors to obtain a second output vector for pushing the content to be pushed to the target object;

[0021] Based on the second output vector, determine the first probability.

[0022] Optionally, the prediction unit is further specifically configured to:

[0023] Obtain a first weight matrix and a first offset vector, where the number of rows of the first weight matrix is the number of input vectors, the number of columns of the first weight matrix is the number of input vectors multiplied by the dimension of the input vectors, and the dimension of the first offset vector is the number of input vectors;

[0024] Multiply the first weight matrix by the transpose of the concatenated vector of the multiple input vectors and add the first offset vector to obtain a first sum vector;

[0025] Perform a first exponential normalization on the elements of the first sum vector to obtain the input vector weight vector.

[0026] Optionally, the prediction unit is further specifically configured to:

[0027] Obtain a first weight vector and a first offset, where the dimension of the first weight vector is equal to the dimension of the input vectors;

[0028] Multiply the first weight vector by the transpose of the second output vector and add the first offset to obtain a first score;

[0029] Perform a second exponential normalization on the first score to obtain the first probability.

[0030] Optionally, the fusion model includes multiple fusion layers. The bottommost fusion layer among the multiple fusion layers inputs the multiple input vectors and generates layer output vectors corresponding to each input vector. The other fusion layers among the multiple fusion layers receive the layer output vectors of the next lower fusion layer corresponding to the input vectors and generate layer output vectors of the fusion layer corresponding to the input vectors;

[0031] The first output vector is generated by the prediction unit in the following manner:

[0032] Based on the input vectors and the layer output vectors corresponding to the input vectors output by each fusion layer, determine the first output vector.

[0033] Optionally, the prediction unit is further specifically configured to:

[0034] Add a fusion layer below the fusion model, and the layer output vector corresponding to the input vector output by the added fusion layer is equal to the input vector;

[0035] Determine the layer weights of each fusion layer after adding the fusion layer;

[0036] Use the layer weights to perform a weighted sum on the layer output vectors corresponding to the input vector output by each fusion layer after adding the fusion layer to obtain the first output vector.

[0037] Optionally, the prediction unit is further specifically configured to:

[0038] Obtain a second weight matrix and a second offset vector, where the number of rows of the second weight matrix is the number of fusion layers after adding the fusion layer, the number of columns of the second weight matrix is the dimension of the input vector, and the second offset vector is the number of fusion layers after adding the fusion layer;

[0039] Multiply the second weight matrix by the transpose of the input vector and add the second offset vector to obtain a second sum vector;

[0040] Perform a first exponential normalization on the elements of the second sum vector to obtain the layer weight vector, and the layer weight vector indicates the layer weights of each fusion layer.

[0041] Optionally, the fusion layer includes a multi-head attention model and a feed-forward network, and the layer output vector is generated by the prediction unit through the following process:

[0042] Input the layer input vector corresponding to the input vector into the attention model to obtain the attention output vector corresponding to the input vector;

[0043] Input the attention output vector corresponding to the input vector into the feed-forward network to obtain the layer output vector corresponding to the input vector.

[0044] Optionally, the prediction unit is further specifically configured to:

[0045] Add the layer input vector corresponding to the input vector and the attention output vector corresponding to the input vector to obtain a third sum vector;

[0046] Perform layer normalization on the third sum vector to obtain a first normalized vector;

[0047] Input the first normalized vector into the feed-forward network to obtain the layer output vector corresponding to the input vector.

[0048] Optionally, the prediction unit is further specifically configured to:

[0049] Input the first normalized vector into the feed-forward network to obtain a feed-forward vector corresponding to the input vector;

[0050] Add the first normalized vector and the feed-forward vector corresponding to the input vector to obtain a fourth sum vector;

[0051] Perform layer normalization on the fourth sum vector to obtain the layer output vector corresponding to the input vector.

[0052] Optionally, the prediction unit is further specifically configured to:

[0053] Obtain a third weight matrix and a third offset vector, where the number of rows of the third weight matrix is the second dimension, the number of columns of the third weight matrix is the dimension of the input vector, and the dimension of the third offset vector is the second dimension;

[0054] Multiply the third weight matrix by the transpose of the first normalized vector and add the third offset vector to obtain a fifth sum vector;

[0055] Apply an activation function to the fifth sum vector and determine a feed-forward vector corresponding to the input vector based on the result of the activation function.

[0056] Optionally, the prediction unit is further specifically configured to:

[0057] Obtain a fourth weight matrix and a fourth offset vector, where the number of rows of the fourth weight matrix is the dimension of the input vector, the number of columns of the third weight matrix is the second dimension, and the dimension of the fourth offset vector is equal to the dimension of the input vector;

[0058] Multiply the fourth weight matrix by the transpose of the result of the activation function and add the fourth offset vector to obtain the feed-forward vector corresponding to the input vector.

[0059] Optionally, the prediction unit is further specifically configured to:

[0060] Perform layer normalization on the layer input vector corresponding to the input vector to obtain a layer input vector after layer normalization;

[0061] Input the layer input vector after layer normalization into the attention model to obtain the attention output vector corresponding to the input vector.

[0062] Optionally, the prediction unit is further specifically configured to:

[0063] Determine the mean and variance of each element in the layer input vector corresponding to the input vector;

[0064] Determine the first sum of the variance and the smoothing amount;

[0065] Subtract the mean from each element of the layer input vector corresponding to the input vector, and divide by the first sum to obtain the layer-normalized layer input vector.

[0066] Optionally, the prediction unit is further specifically configured to:

[0067] Subtract the mean from each element of the layer input vector corresponding to the input vector, and divide by the first sum to obtain a vector to be linearly processed;

[0068] Multiply each element of the vector to be linearly processed by a first parameter and add a second parameter to obtain the layer-normalized layer input vector.

[0069] Optionally, the attention model includes multiple attention heads; the prediction unit is further specifically configured to:

[0070] Input the layer-normalized layer input vector into the attention head to obtain the attention scores output by the attention head, where the multiple attention scores output by the multiple attention heads form an attention score vector;

[0071] Multiply the attention score vector by a fifth weight matrix to obtain the attention output vector corresponding to the input vector, where the number of rows of the fifth weight matrix is the number of attention heads, and the number of columns of the fifth weight matrix is the dimension of the input vector.

[0072] Optionally, the prediction unit is further specifically configured to:

[0073] Multiply a sixth weight matrix by the transpose of the layer-normalized layer input vector to obtain a first intermediate vector, where the number of rows of the sixth weight matrix is a third dimension, and the number of columns of the sixth weight matrix is the dimension of the input vector;

[0074] Multiply a seventh weight matrix by the transpose of each layer output vector of the previous fusion layer to obtain a second intermediate vector corresponding to each layer output vector of the previous fusion layer, where the number of rows of the seventh weight matrix is the third dimension, and the number of columns of the seventh weight matrix is the dimension of the input vector;

[0075] Multiply an eighth weight matrix by the transpose of the layer-normalized layer input vector to obtain a third intermediate vector, where the number of rows of the eighth weight matrix is a fourth dimension, the fourth dimension is the dimension of the input vector divided by the number of attention heads, and the number of columns of the eighth weight matrix is the dimension of the input vector;

[0076] Determine the attention score based on the first intermediate vector, the second intermediate vector corresponding to each layer output vector of the previous fusion layer, and the third intermediate vector.

[0077] Optionally, the prediction unit is further specifically configured to:

[0078] Multiply the first intermediate vector by the transpose of the second intermediate vector corresponding to each layer output vector of the previous fusion layer respectively to obtain the influence score of each layer output vector of the previous fusion layer on the layer input vector;

[0079] Divide the influence score of each layer output vector of the previous fusion layer on the layer input vector by the third dimension respectively to obtain the attenuation influence score of each layer output vector of the previous fusion layer on the layer input vector, where the attenuation influence scores of multiple layer output vectors of the previous fusion layer on the layer input vector form an attenuation influence score vector;

[0080] Multiply the attenuation influence score vector by the elements with the same serial number in the third intermediate vector, and accumulate the multiplication results to obtain the attention score.

[0081] Optionally, the first generation unit is specifically configured to:

[0082] Vectorize the target push basic features in the target push basic feature combination into the target push basic feature vector;

[0083] Multiply each target push basic feature vector corresponding to each target push basic feature in the target push basic feature combination element by element to obtain the second-order cross feature vector.

[0084] Optionally, after vectorizing the target push basic features in the target push basic feature combination into the target push basic feature vector, the push processing device further includes a pooling unit for:

[0085] Determine that the dimension of the target push basic feature vector is greater than the first dimension;

[0086] Perform pooling processing on the target push basic feature vector so that the dimension of the target push basic feature vector after pooling processing is equal to the first dimension.

[0087] Optionally, the feature full-cross model includes multiple feature full-cross models corresponding to multiple tasks;

[0088] The second generation unit is specifically configured to: for each of the multiple tasks, input the first number of the target push basic features into the feature full-crossing model corresponding to the task, and obtain the full-crossing feature vector corresponding to the task.

[0089] Optionally, the fusion model includes multiple fusion models corresponding to multiple tasks;

[0090] The prediction unit is further specifically configured to:

[0091] For each of the multiple tasks, input the first number of the target push basic feature vectors, the multiple second-order crossing feature vectors corresponding to the multiple target push basic feature combinations, and the full-crossing feature vector corresponding to the task into the fusion model corresponding to the task, and obtain the first probability for the task of pushing the to-be-pushed content to the target object;

[0092] Based on the first probabilities for each of the tasks, push the to-be-pushed content to the target object.

[0093] Optionally, the prediction unit is further specifically configured to:

[0094] Perform a weighted sum of the first probabilities for each of the tasks to obtain a second probability;

[0095] Based on the second probability, push the to-be-pushed content to the target object.

[0096] According to one aspect of the present disclosure, there is provided an electronic device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the above-mentioned push processing method is implemented.

[0097] According to one aspect of the present disclosure, there is provided a computer-readable storage medium, where the storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned push processing method is implemented.

[0098] According to one aspect of the present disclosure, there is provided a computer program product, which includes a computer program, and the computer program is read and executed by a processor of a computer device, so that the computer device executes the above-mentioned push processing method.

[0099] In the embodiments of the present disclosure, the target push basic features include target object features and content-to-be-pushed features. The first number of target push basic feature vectors corresponding to the first number of target push basic features can reflect the features of the target object itself or the features of the content-to-be-pushed itself. Therefore, the target push basic feature vector can also be called a first-order cross feature vector. The multiple second-number-order cross feature vectors (such as second-order cross feature vectors, third-order cross feature vectors, etc.) corresponding to the combination of multiple target push basic features can reflect the interaction effects between some of the target push basic features. For example, the interaction effects between target object features, the interaction effects between target object features and content-to-be-pushed features, and the interaction effects between content-to-be-pushed features and content-to-be-pushed features. The full cross feature vector output by the feature full cross model can reflect the interaction effects between the first number of target push basic features, that is, the interaction effects between all target object features and content-to-be-pushed features. Therefore, the full cross feature vector can also be called a full-order cross feature vector. When fusing the feature cross vectors of different orders above, instead of adding or splicing these vectors, these vectors are input into a fusion model to obtain the first probability of pushing the content-to-be-pushed to the target object. The fusion model can fully explore the context semantics between the feature cross vectors of different orders, so as to obtain a relatively accurate first probability. Finally, based on the relatively accurate first probability, the content-to-be-pushed is pushed to the target object, thereby improving the push accuracy.

[0100] Other features and advantages of the present disclosure will be described in the following specification, and, in part, will be obvious from the specification, or will be understood by implementing the present disclosure. The objectives and other advantages of the present disclosure can be realized and obtained through the structures specifically pointed out in the specification, the claims, and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0101] The drawings are used to provide a further understanding of the technical solutions of the present disclosure, and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the technical solutions of the present disclosure, and do not constitute a limitation to the technical solutions of the present disclosure.

[0102] Figure 1 is a framework diagram of the system to which the push processing method according to the embodiment of the present disclosure is applied;

[0103] Figures 2A - 2C is a schematic interface diagram of the scenario where the embodiment of the present disclosure is applied to view the video content push message in the instant short message application;

[0104] Figure 3 is a flowchart of the push processing method according to an embodiment of the present disclosure;

[0105] Figure 4It is a schematic diagram of the characteristics of the target object and the content to be pushed;

[0106] Figure 5 It is a schematic diagram of the specific implementation of obtaining the basic feature combination of the target push;

[0107] Figure 6 It is a schematic diagram of the specific implementation of determining the first probability based on the basic features of the target push;

[0108] Figure 7 It is Figure 3 The flowchart of step 350 in determining the first probability;

[0109] Figure 8 It is Figure 3 The schematic diagram of the specific implementation of step 350 in determining the first probability;

[0110] Figure 9 It is Figure 7 The flowchart of step 710 in determining the weight vector of the input vector;

[0111] Figure 10 It is Figure 7 The flowchart of step 730 in determining the first probability;

[0112] Figure 11 It is a schematic diagram that the fusion model includes multiple fusion layers;

[0113] Figure 12 It is a flowchart of using multiple fusion layers to determine the first output vector;

[0114] Figure 13 It is a flowchart of using the layer weights and layer output vectors of multiple fusion layers to determine the first output vector;

[0115] Figure 14 It is Figure 12 The flowchart of step 1220 in determining the layer weights;

[0116] Figure 15A It is a schematic diagram of the structure of the fusion layer;

[0117] Figure 15B It is a schematic diagram of the structure of the multi-head attention model;

[0118] Figure 16 It is a flowchart of determining the layer output vector based on the attention model and the feed-forward network;

[0119] Figure 17 It is Figure 16 The flowchart of step 1610 in determining the attention output vector;

[0120] Figure 18 It is Figure 17Flowchart of step 1710 in determining the input vector after layer normalization;

[0121] Figure 19 Yes Figure 17 Flowchart of step 1720 in determining the attention output vector;

[0122] Figure 20 Yes Figure 19 Flowchart of step 1910 in determining the attention score;

[0123] Figure 21 Yes Figure 16 Flowchart of determining the layer output vector based on the attention model and the feed - forward network on the basis of

[0124] Figure 22 Yes Figure 21 Another flowchart of determining the layer output vector based on the attention model and the feed - forward network on the basis of

[0125] Figure 23 Yes Figure 22 Flowchart of step 2210 in determining the feed - forward vector corresponding to the input vector;

[0126] Figure 24 Yes

[0127] Figure 25 Yes

[0128] Figure 26 Yes Figure 3 Terminal structure diagram for executing the push - processing method shown in

[0129] Figure 27 Yes Figure 3 Server structure diagram for executing the push - processing method shown in Detailed implementation

[0130] In order to make the objectives, technical solutions and advantages of the present disclosure more clear and understandable, the present disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present disclosure and are not used to limit the present disclosure.

[0131] Before further elaborating on the embodiments of the present disclosure, the nouns and terms involved in the embodiments of the present disclosure are described. The nouns and terms involved in the embodiments of the present disclosure are applicable to the following explanations:

[0132] Deep Neural Networks (DNN): It is a multi-layer unsupervised neural network. The deep neural network uses the output features of the previous layer as the input of the next layer for feature learning. After layer-by-layer feature mapping, the features of the existing space samples are mapped to another feature space, so as to learn better feature expressions for the existing inputs. The deep neural network has multiple non-linear mapping feature transformations and can fit highly complex functions.

[0133] Push: It refers to the method by which a system or platform actively sends information, content, or services to users without the users' explicit requests or searches. Push processing methods usually perform targeted delivery based on personal information such as users' interests, preferences, and historical behaviors to provide personalized content pushes.

[0134] Non-linear activation function (sigmoid function): It is also called the logistic function or S-shaped function. The sigmoid function maps the input value to an output value between 0 and 1, that is, the range of the output value is between 0 and 1. When the input value approaches negative infinity, the output value approaches 0; when the input value approaches positive infinity, the output value approaches 1.

[0135] Exponential normalization (softmax function): It compresses a K-dimensional vector containing arbitrary real numbers into another K-dimensional vector, so that the range of each element is between (0, 1), and the sum of all elements is 1.

[0136] System architecture and scenario description applied in the embodiments of the present disclosure

[0137] Figure 1 It is a system architecture diagram applied to the push processing method according to the embodiments of the present disclosure. It includes: object terminal 110, Internet 120, gateway 130, and push processing server 140.

[0138] The object terminal 110 is a device for the object to view the reminder messages corresponding to the pushed content. It includes various forms such as desktop computers, laptops, PDAs (Personal Digital Assistants), mobile phones, in-vehicle terminals, home theater terminals, and dedicated terminals. Additionally, it can be a single device or a collection of multiple devices. For example, multiple devices are connected through a local area network and share a display device for collaborative work, jointly constituting a terminal. The object terminal 110 can also communicate with the Internet 120 in a wired or wireless manner to exchange data.

[0139] The gateway 130 is also known as an internetwork connector or protocol converter. The gateway 130 realizes network interconnection at the transport layer and is a computer system or device that acts as a converter. Between two systems using different communication protocols, data formats, or languages, and even with completely different architectures, the gateway 130 is a translator. At the same time, the gateway 130 can also provide filtering and security functions. Messages sent by the object terminal 110 to the push processing server 140 need to be sent to the corresponding push processing server 140 through the gateway 130. Messages sent by the push processing server 140 to the object terminal 110 also need to be sent to the corresponding object terminal 110 through the gateway 130.

[0140] The push processing server 140 refers to a computer system that can provide content message push services to the object terminal 110. Compared with the object terminal 110, the push processing server 140 has higher requirements in terms of stability, security, performance, etc. The push processing server 140 can be a high-performance computer in a network platform, a cluster of multiple high-performance computers, a part (such as a virtual machine) allocated from a high-performance computer, a combination of parts (such as virtual machines) allocated from multiple high-performance computers, etc. The push processing server 140 can also communicate with the Internet 120 through wired or wireless means to exchange data.

[0141] The embodiments of the present disclosure can be applied in various scenarios, such as Figures 2A - 2C the scenario of viewing video content push messages in an instant short message application as shown, etc.

[0142] As Figure 2A shown, in the interface of the instant short message application in the object terminal 110, the instant short message application can be used to conduct instant messaging with multiple contacts and can also be used to watch videos. After triggering the "Message" option in the lower function bar, a short message list is displayed. The short message list contains message bars for multiple contacts. At 11:30 am, in the "Short Video" option of the lower function bar, there is a content prompt ( Figure 2A the small black dot in the upper right corner of "Short Video" in ), and this content prompt is used to remind the object that there is new content pushed. The object can trigger the "Short Video" option in the lower function bar to watch the recommended content. The object can also trigger the "Short Video Highlights" push bar to watch the pushed content.

[0143] As Figure 2B shown, at 12:00 pm, after the object views the "small black dot" in the interface of the instant short message application, the object selects to trigger the "Short Video" in the lower function bar.

[0144] As Figure 2CAs shown, after triggering the "Short Video" option, the instant messaging application opens the content push interface and plays the pushed content, which is a video V. The content push interface also shows the number of likes (67,000), comments (2,450), and favorites (1,422) that the video received before 12:00 am. The user can also like, favorite, and comment on the video in the interface. The video content pushed to the user is determined by the push processing server 140.

[0145] In current push processing, the push processing model often predicts the score or probability of pushing the content to be pushed to the target object based on the feature cross - result between the target object features and the content - to - be - pushed features, and then determines the order of pushing the content to be pushed to the target object (such as the order of pushing short videos) according to these scores and probabilities. However, often after obtaining the low - order feature cross - result and high - order feature cross - result based on the target object features and the content - to - be - pushed features, the low - order feature cross - result and high - order feature cross - result are directly added or concatenated, and then the score or probability is determined based on the added or concatenated feature cross - result. However, directly adding or concatenating will lose the context semantics of the high - and low - order features, resulting in the inability to effectively measure the feature influence between the low - order feature cross - result and the high - order feature cross - result, leading to low push accuracy. In the embodiments of the present disclosure, when performing push, multiple vectors reflecting different - order feature cross - results are uniformly input into the fusion model, and the fusion model can fully mine the context semantics between these vectors, so as to obtain a higher - accuracy score or probability. Finally, based on the higher - accuracy score or probability, the content to be pushed is pushed to the target object, improving the push accuracy.

[0146] General description of the embodiments of the present disclosure

[0147] According to an embodiment of the present disclosure, a push processing method is provided.

[0148] The push processing method is a process of determining the target content for the target object and pushing the reminder message of the target content to the object terminal 110 of the target object. The target object refers to the object that wants to view the target content. The target content can be of various types such as pictures, videos, articles, applications, etc.

[0149] In current push processing, the push processing model often predicts the score or probability of pushing the content to be pushed to the target object based on the feature cross - result between the target object features and the content - to - be - pushed features, and then determines the order of pushing the content to be pushed to the target object (such as the order of pushing short videos) according to these scores and probabilities.

[0150] When predicting the score or probability of pushing the content to be pushed to the target object based on the feature cross - result between the target object feature and the content - to - be - pushed feature, the prior art proposes directly adding or concatenating the low - order feature cross - result and the high - order feature cross - result, and then determining the score or probability based on the added or concatenated feature cross - result. For the addition method, specifically, the low - order feature cross - result and the high - order feature cross - result are summed at the final output layer of the model. The disadvantage is that it ignores the perception of the context of other feature cross - terms by each independent feature cross - term. In different samples, the same cross - term with the same value may have completely different impacts on the prediction result, and this simple summation method cannot model this effect. At the same time, there are generally significant differences between high - order and low - order terms in terms of quantity, magnitude, model differences, and the emphasis on representing vector extraction. This simple addition will also cause unexpected adverse effects on the stability of model training. For the concatenation method, specifically, the low - order feature cross - result and the original features are concatenated together and then unified implicit non - linear fusion is performed using DNN. The disadvantage is that the low - order cross - result focuses on the overall cross - modeling of the representation vectors of feature dimensions, while DNN performs implicit non - linear high - order cross - at the single - element granularity for the concatenated input vectors, which will break through the boundaries between input features and is not conducive to the extraction of low - order cross - information between features. Moreover, after being processed by DNN, some low - order signals that are effective for the final prediction will also be lost. It can be seen that direct addition or concatenation will lose the context semantics of high - and low - order features, resulting in the inability to effectively measure the feature influence between the low - order feature cross - result and the high - order feature cross - result.

[0151] The push - processing method of the embodiments of the present disclosure is executed on the server 140. After determining the target content to be pushed, the reminder message corresponding to the target content is pushed to the object terminal 110 through the network and the Internet 120.

[0152] As Figure 3 shown, according to an embodiment of the present disclosure, the push - processing method includes:

[0153] Step 310, obtain a first number of target push basic features, where the first number of target push basic features include target object features and content - to - be - pushed features;

[0154] Step 320, based on the first number of target push basic features, obtain a plurality of target push basic feature combinations, and each target push basic feature combination includes a second number of target push basic features selected from the first number of target push basic features, and the second number is less than the first number;

[0155] Step 330, generate a second - order cross - feature vector corresponding to the target push basic feature combination;

[0156] Step 340: Input the first number of target push basic features into the feature full-crossing model to obtain a full-crossing feature vector;

[0157] Step 350: Input the first number of target push basic feature vectors corresponding to the first number of target push basic features, the second number of order-crossing feature vectors corresponding to multiple combinations of target push basic features, and the full-crossing feature vector into the fusion model to obtain the first probability of pushing the content to be pushed to the target object, and push the content to be pushed to the target object based on the first probability.

[0158] The above steps 310 - 350 are described generally below.

[0159] In step 310, the target object features and the content to be pushed features are features for predicting different aspects of the content that the target object will watch.

[0160] The target object features refer to the features that the target object itself has. Usually, target objects with a certain type of common target object features may show certain commonalities in the content they watch. Therefore, using the target object features for push processing prediction helps to improve the effect of push processing. For example, Figure 4 as shown, the target object features include educational level features, work years features, work features, hobby features, historical watched content features, etc.

[0161] Regarding the educational level features, objects with the same educational level tend to choose the same type of content related to the educational level with a higher probability. In an example, an object with a university educational level may tend to watch video content about postgraduate school introductions.

[0162] Regarding the work years features, objects with a certain work years tend to choose the same type of content related to that work years with a higher probability. In an example, an object with 10 years of work experience may tend to watch video content about career promotions.

[0163] Regarding the work features, objects with the same work tend to choose the same type of content that is often needed to watch in that work with a higher probability. In an example, an object whose work is a lawyer tends to watch video content about popularizing the law.

[0164] Regarding the hobby features, objects with the same hobby may tend to watch the same content. In an example, an object whose hobby is swimming likes to watch swimming teaching video content.

[0165] Regarding the historical watched content features, an object may like to watch content with the same type of historical watched content features. In an example, for an object who has watched a 60 - second swimming teaching video, it is very likely that they want to watch video content about swimming venue introductions.

[0166] In one embodiment, to obtain the target object features of a target object, it can be done by obtaining the target object features pre - autonomously selected by the target object, or by obtaining the target object features statistically derived from the viewing content log of the target object.

[0167] Note that when obtaining the target object features of the target object, the consent of the target object should be sought in advance. Moreover, the collection, use, and processing of these object features will comply with relevant laws, regulations, and standards. When seeking the consent of the target object, individual permission or individual consent of the target object can be obtained through methods such as pop - up windows or redirecting to a confirmation page.

[0168] The features of the content to be pushed are the features inherent in the content to be pushed. Usually, content to be pushed with the same type of features of the content to be pushed is more likely to be viewed by the same target object. Therefore, using the features of the content to be pushed for push - processing prediction helps to improve the effect of push - processing. For example Figure 4 As shown, the features of the content to be pushed include content type features, content length features, like count features, comment count features, and favorite count features. The features of the content to be pushed can also include content title features, content cover features, etc.

[0169] In step 310 above, the target object features are used to represent the target object. However, in actual situations, the target scenario where the target object is located can also represent the target object. Therefore, in one embodiment, in addition to obtaining the target object features of the target object, the target scenario features are also obtained. The target scenario features refer to the features of the target scenario where the target object is located. Usually, the target object may view similar content in the same target scenario. For example, when the target scenario of the target object is an office, the target object may be more likely to view push content related to work; when at home, the target object may be more likely to view push content related to their own preferences. Therefore, using the target scenario features for push - processing prediction helps to improve the effect of push - processing. The target scenario features include geographical region features, current time features, current push day type features (weekday, weekend, or holiday), weather features (sunny, cloudy, or rainy), etc. In one example, the geographical region feature is City A, the current time feature is 1:10 pm, the current push day type is weekend, and the weather feature is rainy.

[0170] When obtaining the target scenario features where the target object is located, the consent of the target object should be obtained in advance. Similar to the above, it will not be elaborated here.

[0171] The characteristics of the current geographical area refer to the characteristics of the geographical area where the target object is currently located. This geographical area can be an administrative geographical area, such as City A, City B, etc., or a geographical area divided by entities on a map, such as the XX Shopping Mall area, the XX Hospital area, etc. For administrative geographical areas, it has an impact on the push content that the target object is interested in. For example, if the target object is in City A, the push content can be tourist attractions or food cultures related to City A. For geographical areas divided by entities, it also has an impact on the push content that the target object is interested in. For example, if the target object is in the XX Shopping Mall area, the push content can be activity information or special projects in the mall; if the target object is in the XX Hospital area, the push content can be content related to health preservation.

[0172] The characteristics of the current push day type refer to the characteristics of the type of day when the target object views the push content. It is divided into weekdays, weekends, and festivals, etc. The target object may want to view different types of push content on different types of days. For example, on weekdays (Monday to Friday), the target object is often at work and may want to view push content related to work; on weekends (Saturday and Sunday), when the target object does not need to work, the push content it may want to view may be related to its own hobbies; on festivals, the target object may want to view push content related to festival cultures.

[0173] The advantage of the above embodiments is that by using the target object characteristics, the characteristics of the content to be pushed, and the target scenario characteristics for predicting the target content to be pushed to the target object, it helps to improve the accuracy of the push processing.

[0174] In one embodiment, the target push basic characteristics are vectorized to obtain a target push basic feature vector. A vector is an array composed of numerical values in different dimensions, and it is a point in a multi-dimensional space. The line segment between this point and the origin in the multi-dimensional coordinate system has a magnitude and a direction, and this magnitude and direction are the magnitude and direction of the vector. Each numerical value in the vector is the point value projected by this point on the corresponding coordinate axis in the multi-dimensional coordinate system, that is, a vector element. The vector elements can be numerical values or symbols, etc. The vector elements of the target push basic feature vector in the embodiments of the present disclosure are numerical values.

[0175] In one embodiment, the target push basic feature vector corresponding to the target push basic characteristics can be determined by looking up a table.

[0176] In another embodiment, if the target push basic characteristic is a character type characteristic, the character type characteristic is input into an embedding layer to obtain a target push basic feature vector corresponding to the character type characteristic; if the target push basic characteristic is a numerical type characteristic, the numerical type characteristic is numerically mapped to obtain a target push basic feature vector corresponding to the numerical type characteristic.

[0177] In step 320, the target push basic feature combination refers to a combination composed of a second number of target push basic features selected from the first number of target push basic features. The second number is less than the first number. For example, referring to Figure 5 , if the first number is 10, the second number can be an integer less than 10, such as 2, 3, 4, etc.

[0178] It can be understood that if the target push basic feature combination includes the second number of target push basic features, then the target push basic feature combination is a second-order push basic feature combination. For example, if the second number is 2, then the target push basic feature combination is a second-order push basic feature combination. Another example is that if the second number is 3, then the target push basic feature combination is a third-order push basic feature combination. Another example is that if the second number is 4, then the target push basic feature combination is a fourth-order push basic feature combination.

[0179] In one example, referring to Figure 5 , a specific implementation schematic diagram of obtaining 9 target push basic feature combinations based on 10 target push basic features is shown. The 10 target push basic features include education level feature, work years feature, job feature, hobby feature, historical viewing content feature, content type feature, content length feature, like count feature, comment count feature, and favorite count feature. Selecting 2 target push basic features from the 10 target push basic features can obtain a second-order push basic feature combination. The second-order push basic feature combination can include (education level feature, work years feature), (job feature, hobby feature), (historical viewing content feature, content type feature), (content length feature, like count feature), (comment count feature, favorite count feature), etc. Selecting 3 target push basic features from the 10 target push basic features can obtain a third-order push basic feature combination. The third-order push basic feature combination can include (education level feature, work years feature, job feature), (job feature, hobby feature, historical viewing content feature), (historical viewing content feature, content type feature, content length feature), (like count feature, comment count feature, favorite count feature), etc.

[0180] It should be noted that in the same embodiment, the second numbers of different target push basic feature combinations can be the same or different. For example, in the above embodiment, the second numbers include 2 and 3, so the target push basic feature combination includes a second-order push basic feature combination and a third-order push basic feature combination. However, in other embodiments, the second number can only include 2, then the target push basic feature combination is a second-order push basic feature combination. The second number can only include 3, then the target push basic feature combination is a third-order push basic feature combination. The embodiments of the present disclosure do not make specific limitations on the value of the second number, and it can be set according to actual needs.

[0181] In step 330, the second-order cross feature vectors corresponding to the target push basic feature combinations can be generated by using a feature interaction method. The feature interaction method refers to multiplying two or more features to achieve a non-linear transformation of the sample space and increase the non-linear ability of the model. For example, if the second number is 2, the vectors corresponding to 2 target push basic features are multiplied to obtain second-order cross feature vectors. Another example is that if the second number is 3, the vectors corresponding to 3 target push basic features are multiplied to obtain third-order cross feature vectors.

[0182] In one embodiment, step 330 includes:

[0183] Vectorize the target push basic features in the target push basic feature combination into target push basic feature vectors;

[0184] Multiply each target push basic feature vector corresponding to each target push basic feature in the target push basic feature combination element by element to obtain second-order cross feature vectors.

[0185] In this embodiment, the target push basic feature combination includes the second number of target push basic features. One target push basic feature corresponds to one target push basic feature vector. Therefore, one target push basic feature combination corresponds to the second number of target push basic feature vectors.

[0186] The advantage of this embodiment is that the advantage of using different vectorization methods for different types of target push basic features is that the features can be vectorized according to the characteristics of different types of target push basic features, ensuring the accuracy of the feature vectorization results.

[0187] In one embodiment, after vectorizing the target push basic features in the target push basic feature combination into target push basic feature vectors, the push processing method of this embodiment further includes:

[0188] Determine that the dimension of the target push basic feature vector is greater than the first dimension;

[0189] Perform pooling processing on the target push basic feature vector so that the dimension of the target push basic feature vector after pooling processing is equal to the first dimension.

[0190] In this embodiment, the first dimension refers to a value for determining whether pooling processing needs to be performed on the target push base feature vector. If the dimension of the target push base feature vector is greater than the first dimension, then pooling processing needs to be performed on the target push base feature vector; if the dimension of the target push base feature vector is less than or equal to the first dimension, then pooling processing does not need to be performed on the target push base feature vector. Generally speaking, if the dimension of the target push base feature vector is less than the first dimension, it needs to be enlarged / extended to the first dimension.

[0191] Pooling, also known as pooling, its essence is actually sampling. For the input target push base feature, pooling can select a certain way to perform dimensionality reduction and compression on it to speed up the operation speed. Pooling includes average pooling and maximum pooling, etc.

[0192] The advantage of the above embodiment is that the dimensionality of the target push base feature vector is adjusted by using the pooling operation, so that each target push base feature vector has the same dimension, and the efficiency of generating the second-order cross feature vector based on the target push base feature vector is improved.

[0193] The element-wise multiplication between two target push base feature vectors can also be called the Hadamard product. The Hadamard product refers to the operation of element-wise multiplication of two vectors with the same dimension. If there are two vectors A and B of the same size, their element-wise multiplication is denoted as A⊙B. A⊙B results in vector M1, and each element of vector M1 is obtained by multiplying the corresponding elements of vector A and vector B. For example, assume there are two 1×2 vectors A and B: vector A = [a 11 , a 12 , vector B = [b 11 , b 12 . The element-wise multiplication result of vector A and vector B is vector M1, vector M1 = A⊙B = [a 11 * b 11 , a 12 * b 12 . Another example, assume there are three 1×2 vectors A, B, and C: vector A = [a 11 , a 12 , vector B = [b 11 , b 12 , vector C = [c 11 , c 12 . The element-wise multiplication result of vector A, vector B, and vector C is vector M2, vector M2 = A⊙B⊙C = [a 11 * b 11 * c 11 , a 12 * b 12 * c12 . The vector M1 belongs to the second-order cross feature vector. The vector M2 belongs to the third-order cross feature vector.

[0194] It can be understood that the second-order cross feature vectors may include second-order cross feature vectors, or third-order cross feature vectors, or cross feature vectors of higher orders than the third order (for example, fourth-order cross feature vectors, fifth-order cross feature vectors, etc.). The second-order cross feature vectors may also include second-order cross feature vectors, third-order cross feature vectors, and cross feature vectors of higher orders than the third order at the same time.

[0195] The advantages of the above embodiments of obtaining the second-order cross feature vectors based on element-wise multiplication are that element-wise multiplication can introduce non-linear relationships, thereby better capturing the complex interaction features among the basic features of each target push; through element-wise multiplication, the interaction information between the original features can be combined to obtain a richer and more comprehensive feature representation, which helps to improve the expression ability of the model.

[0196] In step 340, the first number of basic target push features are input into the feature full-cross model to obtain a full-cross feature vector. The full-cross feature vector refers to a complex and rich feature representation vector obtained from the first number of basic target push features. The feature full-cross model is a machine learning model that cross-combines features by fully combining the input features to obtain a more complex and rich feature representation. In the feature full-cross model, complex interaction relationships between features can be captured, thereby improving the expression ability and prediction accuracy of the model.

[0197] The feature full-cross model includes models such as DNN (Deep-Learning Neural Network), FM (Factorization Machines), and FFM (Field-aware Factorization Machines). These models perform well in processing high-dimensional sparse data and can effectively mine the interaction relationships between features, so they have important application values in the field of recommendation processing, etc.

[0198] In one embodiment, the feature full-cross model includes multiple feature full-cross models corresponding to multiple tasks. In this embodiment, step 340 includes: for each task among the multiple tasks, inputting the first number of basic target push features into the feature full-cross model corresponding to the task to obtain a full-cross feature vector corresponding to the task.

[0199] A task refers to a matter related to the content to be pushed, which is used to measure the degree of interest of the target object in the content to be pushed. For example, the task is whether the target object likes the content to be pushed, or whether the target object collects the content to be pushed, etc. The impact of the target push basic features on each task is not the same, so a corresponding feature full-cross model is set for each task. Therefore, for the same first number of target push basic features, the full-cross feature vectors output by the feature full-cross models of different tasks are generally not the same. In this way, when performing push processing for multiple tasks, the push accuracy is further improved.

[0200] In step 350, the first number of target push basic feature vectors corresponding to the first number of target push basic features, the multiple second number order cross feature vectors corresponding to the multiple target push basic feature combinations, and the full-cross feature vectors are input into the fusion model to obtain the first probability of pushing the content to be pushed to the target object, and the content to be pushed is pushed to the target object based on the first probability.

[0201] A fusion model refers to a model that integrates and fuses multiple features to improve the model's expression ability and prediction performance. More specifically, a fusion model usually refers to a model that fuses features from different levels or different branches. Fusion models usually include residual connection network models, multi-branch models, and deep fusion models. In the Residual Network (ResNet) model, features from different levels are added or connected to achieve the purpose of feature fusion, which helps to alleviate gradient disappearance and train deep networks. In the Multi-Branch Model, a neural network with multiple branches can be designed, and each branch learns different levels or different types of features. Finally, the features learned by these branches are fused to obtain a more comprehensive and rich feature representation. In the DeepFusion Model, features from different levels or different branches are deeply fused to obtain a more complex and rich feature representation.

[0202] In one embodiment, the fusion model includes multiple fusion models corresponding to multiple tasks. In this embodiment, step 350 includes:

[0203] For each task among the multiple tasks, the first number of target push basic feature vectors, the multiple second number order cross feature vectors corresponding to the multiple target push basic feature combinations, and the full-cross feature vectors corresponding to the task are input into the fusion model corresponding to the task to obtain the first probability for the task of pushing the content to be pushed to the target object;

[0204] Based on the first probabilities for each task, the content to be pushed is pushed to the target object.

[0205] In this embodiment, each piece of content to be pushed has a corresponding first probability under each task. For example, for a video content, the first probability that the target object likes the video content and the first probability that the target object collects the video content are not necessarily the same. Pushing the content to be pushed to the target object based on the first probabilities of multiple tasks is equivalent to comprehensively pushing content based on various preferences of the target object for the same content to be pushed, improving the accuracy of pushing.

[0206] In one embodiment, pushing the content to be pushed to the target object based on the first probabilities for each task includes:

[0207] Performing a weighted sum of the first probabilities for each task to obtain a second probability;

[0208] Pushing the content to be pushed to the target object based on the second probability.

[0209] The influence degrees of different tasks on the content to be pushed are not the same. For example, the target object generally likes and / or comments on video content, but rarely collects and / or forwards it. Therefore, relatively high weights can be set for collection and forwarding, and relatively low weights can be set for like and comment. In this way, by analyzing the preferences of the target object for the content to be pushed through multiple tasks and the weights of each task, the obtained second probability is more comprehensive and accurate, improving the accuracy of the push processing.

[0210] The advantage of the embodiment of step 310-step 350 is that, in the embodiment of the present disclosure, the target push basic features include target object features and to-be-pushed content features. The first number of target push basic feature vectors corresponding to the first number of target push basic features can reflect the features of the target object itself, or the features of the to-be-pushed content itself. Therefore, the target push basic feature vector can also be called a first-order cross feature vector. Multiple second-order cross feature vectors (for example, second-order cross feature vectors, third-order cross feature vectors, etc.) corresponding to multiple target push basic feature combinations can reflect the interaction between a part of the target push basic features. For example, the interaction between target object features and target object features, the interaction between target object features and to-be-pushed content features, and the interaction between to-be-pushed content features and to-be-pushed content features. The full cross feature vector can reflect the interaction between the first number of target push basic features, that is, the interaction between multiple target object features and multiple to-be-pushed content features. Therefore, the full cross feature vector can also be called a full-order cross feature vector. When the above feature cross vectors of different orders are fused, these vectors are not added or concatenated, but these vectors are input into the fusion model to obtain the first probability of pushing the content to be pushed for the target object. The fusion model can fully mine the contextual semantics between feature cross vectors of different orders, thereby obtaining a first probability with higher accuracy. Finally, the content to be pushed is pushed to the target object based on the first probability with higher accuracy, thereby improving the push accuracy.

[0211] The above is a general description of steps 310 to 350. Since steps 310, 320, 330 and 340 have been described in detail in the above general description, the specific implementation of step 350 will be described in detail below.

[0212] Detailed description of step 350

[0213] In step 350, a first number of target push basic feature vectors corresponding to a first number of target push basic features, a plurality of second-order cross feature vectors corresponding to a plurality of target push basic feature combinations, and a full cross feature vector are input into a fusion model to obtain a first probability of pushing the content to be pushed to the target object, and the content to be pushed is pushed to the target object based on the first probability.

[0214] In one example, assuming that there are N (a first number) target push basic feature vectors, the dimensions of which are all set to d, the nth target push basic feature vector is recorded as x n Assuming the second number is 2, for the i-th feature and the j-th feature, the second-order cross feature vector is f ij =x i ⊙x j There are a total of N features Second-order cross feature vectors. For the part of the feature full-cross model, its input is the result of concatenating these N target push basic feature vectors, that is, X = [..., x n ,...]. Wherein, [·] represents vector concatenation, and the input dimension of the concatenated sample is denoted as d X = Nd, and the full-cross feature vector output by the feature full-cross model is denoted as f Q = h(X). h(·) represents the model processing process of the feature full-cross model, and its last layer also outputs a d-dimensional vector without applying an activation function. The N target push basic feature vectors, second-order cross feature vectors, and 1 full-cross feature vector form an input vector sequence, that is, The number of vectors in the input vector sequence is Input the input vector sequence into the fusion model to obtain the first probability of pushing the content to be pushed to the target object, and push the content to be pushed to the target object based on the first probability.

[0215] In an example, referring to Figure 6 , Figure 6 shows a specific implementation diagram of determining the first probability based on the target push basic features. The first number of target push basic features includes target object features, target scene features, and content to be pushed features. In this example, the first number is 3 and the second number is 2. Input the 3 target push basic features into the vectorization layer to obtain 3 target push basic feature vectors ( Figure 6 shows the target push basic feature vector x 1 , the target push basic feature vector x 2 , and the target push basic feature vector x 3 ). If the dimension of a certain target push basic feature vector is greater than the first dimension, then input the target push basic feature vector into the pooling layer so that the dimension of the pooled target push basic feature vector is equal to the first dimension.

[0216] Continuing to refer to Figure 6 , based on the 3 target push basic features, 3 target push basic feature combinations can be obtained, and each target push basic feature combination includes 2 target push basic features. Based on the 3 target push basic feature combinations, the second number order cross feature vectors obtained include 3 second-order cross feature vectors. Specifically, based on the target push basic feature vector x 1 , and the target push basic feature vector x 2 , the second-order cross feature vector f 12 is obtained. Based on the target push basic feature vector x 1 , and the target push basic feature vector x 3 , the second-order cross feature vector f13 Based on the target push basic feature vector x 2 and the target push basic feature vector x 3 a second-order cross feature vector f is obtained 23 .

[0217] Continue to refer to Figure 6 and input the target push basic feature vector x 1 the target push basic feature vector x 2 and the target push basic feature vector x 3 into the splicing layer, and then input the vector output by the splicing layer into the feature full-cross model to obtain the full-cross feature vector f Q .

[0218] Continue to refer to Figure 6 and input the target push basic feature vector x 1 the target push basic feature vector x 2 the target push basic feature vector x 3 the second-order cross feature vector f 12 the second-order cross feature vector f 13 the second-order cross feature vector f 23 and the full-cross feature vector f Q into the fusion model together, and then input the vector output by the fusion model into the linear layer to obtain the first probability. Finally, push the content to be pushed to the target object based on the first probability.

[0219] In one embodiment, referring to Figure 7 step 350 includes:

[0220] Step 710: Take each of the first number of target push basic feature vectors, multiple second number order cross feature vectors, and the full-cross feature vector as an input vector, determine the input vector weight vector, and the input vector weight vector indicates the input vector weight of each input vector;

[0221] Step 720: Use the input vector weights to perform a weighted sum on the first output vector of the fusion model for the input vector to obtain a second output vector for pushing the content to be pushed to the target object;

[0222] Step 730: Determine the first probability based on the second output vector.

[0223] In step 710, the input vector weight vector includes multiple input vector weights corresponding to the input vectors. Referring to the above example, based on the target push basic feature vector x 1 the target push basic feature vector x 2 the target push basic feature vector x 3 the second-order cross feature vector f 12, the second-order cross feature vector f 13 , the second-order cross feature vector f 23 , and the full cross feature vector f Q , to obtain 7 input vectors. Based on 2 input vectors, the determined input vector weight vector is a vector with a dimension of 7. For example, the input vector weight vector is [0.1, 0.1, 0.1, 0.15, 0.15, 0.15, 0.25]. For different input vectors, their input vector weights may be different. This is for content push for the target object, and the contributions of different input vectors may be different.

[0224] In step 720, the input vectors are input into the fusion model, and the output is the first output vector. Refer to Figure 8 , if the input vector is obtained based on the target push basic feature vector, the first output vector can represent the self-feature information of the target push basic feature. If the input vector is obtained based on the second-order cross feature vector, the first output vector can represent the interaction feature information between the second number of target push basic features. If the input vector is obtained based on the full cross feature vector, the first output vector can represent the complex interaction feature information between the first number of target push basic features. It can be seen that each first output vector obtained in this embodiment can already represent rich and complex feature information. The first output vectors can be added together, so as to obtain a second vector based on the added vector. However, based on this, this embodiment fully considers that the input vector weight obtained based on the input vector can measure the influence degree of the input vector on pushing the content to be pushed for the target object. Therefore, the first output vector is weighted and summed with the input vector weight to obtain the second output vector. Obviously, in addition to being able to represent rich and complex feature information, the second output vector can also be flexibly adjusted based on business requirements or the characteristics of feature dynamic change over time, and has high universality.

[0225] In step 730, based on the second output vector, the first probability is determined. The first probability corresponding to the second output vector can be determined by looking up a table or using a calculation formula. For example, refer to Figure 8 , the second output vector is input into the linear layer (the linear layer includes a calculation formula) to obtain the first probability. The embodiments of the present application do not make specific limitations on the determination method.

[0226] The advantage of the embodiments of steps 710-730 is that the first output vector is weighted and summed based on the input vector weight vector, which reflects the different influences of different input vectors on pushing the content to be pushed for the target object, and improves the push accuracy.

[0227] In one embodiment, refer to Figure 9 , step 710 includes:

[0228] Step 910: Obtain a first weight matrix and a first offset vector;

[0229] Step 920: Multiply the first weight matrix by the transpose of the concatenated vector of multiple input vectors, and add the first offset vector to obtain a first sum vector;

[0230] Step 930: Perform a first exponential normalization on the elements of the first sum vector to obtain an input vector weight vector.

[0231] In this embodiment, the first weight matrix refers to the matrix used for weight multiplication. The number of rows of the first weight matrix is the number of input vectors. The number of columns of the first weight matrix is the number of input vectors multiplied by the dimension of the input vectors. The first offset vector refers to the vector that adds an offset to the result of weight multiplication to obtain the first sum vector. The dimension of the first offset vector is the number of input vectors. Performing a first exponential normalization on the elements of the first sum vector to obtain an input vector weight vector. Functions capable of performing the first exponential normalization include functions such as Softmax and Sigmoid.

[0232] In an example, the first weight matrix W G is a trainable parameter matrix with a dimension of T×Td. The first offset vector is B G which is a trainable bias parameter vector with a dimension of T×1. The dimension of the input vectors is d. The concatenated vector obtained by concatenating T input vectors Concatenated vector has a dimension of 1×Td. Multiply the first weight matrix W G by the transpose of the concatenated vector and add the first offset vector to obtain the first sum vector expressed as: Then perform a first exponential normalization on the elements of the first sum vector to obtain an input vector weight vector Input vector weight vector has a dimension of T×1.

[0233] The advantages of the embodiment of steps 910-930 are that the first weight matrix can assign different weights to different input vectors, thereby making the model more flexible and capable of adapting to the complex relationships between different input vectors; the first exponential normalization helps to expand the differences between values, making the data easier to distinguish and process, and can also reduce the impact of outliers on the data distribution, making the data more stable. In summary, the accuracy of determining the input vector weight vector is improved.

[0234] In one embodiment, referring to Figure 10 , step 730 includes:

[0235] Step 1010: Obtain a first weight vector and a first offset;

[0236] Step 1020: Multiply the first weight vector by the transpose of the second output vector, and add the first offset to obtain the first score.

[0237] Step 1030: Perform second exponential normalization on the first score to obtain the first probability.

[0238] In this embodiment, the first weight vector refers to the vector used for weight multiplication. The dimension of the first weight vector is equal to the dimension of the input vector. The first offset refers to the parameter that adds an offset to the result of weight multiplication to obtain the first score. Performing second exponential normalization on the first score to obtain the first probability. Functions that can perform second exponential normalization include functions such as Softmax and Sigmoid.

[0239] In one example, the second output vector t refers to the t-th input vector. refers to the weight corresponding to the t-th input vector in the input vector weights. refers to the first output vector corresponding to the t-th input vector. The first output vector has a dimension of 1×d. The second output vector has a dimension of 1×d. T refers to the total number of input vectors. The first probability The first weight vector is a trainable parameter vector with a dimension of 1×d, and the first offset is a 1-dimensional trainable bias parameter.

[0240] The advantages of the embodiment of steps 1010 - 1030 are that the first weight vector can assign different weights to different vector elements of the second input vector, so as to adapt to the complex relationships between different vector elements; the second exponential normalization helps to expand the differences between values, making the data easier to distinguish and process, and can also reduce the influence of outliers on the data distribution, making the data more stable. In summary, the accuracy of determining the first probability is improved.

[0241] In one embodiment, the fusion model includes multiple fusion layers. The bottommost fusion layer in the multiple fusion layers inputs multiple input vectors and generates layer output vectors corresponding to each input vector. The other fusion layers in the multiple fusion layers receive the layer output vectors of the next lower fusion layer corresponding to the input vectors and generate layer output vectors of the fusion layer corresponding to the input vectors. At this time, the first output vector is generated in the following manner: Based on the input vectors and the layer output vectors of each fusion layer corresponding to the input vectors, determine the first output vector.

[0242] It should be noted that in practical applications, the number of fusion layers can be set according to requirements. For example, it can be set to 3, 4, or 5, etc.

[0243] In one example, referring to Figure 11 , the fusion model includes three fusion layers, namely, fusion layer R1, fusion layer R2, and fusion layer R3. Multiple input vectors (including a first number of target push base feature vectors, multiple second-order cross feature vectors, and full cross feature vectors) are input into fusion layer R1 to obtain multiple layer output vectors of fusion layer R1. The multiple layer output vectors of fusion layer R1 are input into fusion layer R2 to obtain multiple layer output vectors of fusion layer R2. The multiple layer output vectors of fusion layer R2 are input into fusion layer R3 to obtain multiple layer output vectors of fusion layer R3. Then, based on the input vectors, and the multiple layer output vectors of fusion layer R1, the multiple layer output vectors of fusion layer R2, and the multiple layer output vectors of fusion layer R3, a first output vector is determined.

[0244] The advantage of this embodiment of determining the first output vector based on the fusion model including multiple fusion layers is that it can fully model / preserve the upper and lower semantic relationships between each input vector, improving the push accuracy.

[0245] In one embodiment, referring to Figure 12 , determining the first output vector based on the input vectors and the layer output vectors corresponding to the input vectors output by each fusion layer includes:

[0246] Step 1210: Add a fusion layer below the fusion model. The layer output vector corresponding to the input vector output by the added fusion layer is equal to the input vector;

[0247] Step 1220: Determine the layer weights of each fusion layer after adding the fusion layer;

[0248] Step 1230: Perform a weighted sum on the layer output vectors corresponding to the input vectors output by each fusion layer after adding the fusion layer using the layer weights to obtain the first output vector.

[0249] In one example, referring to Figure 13, the fusion model includes 3 fusion layers, namely, fusion layer R1, fusion layer R2, and fusion layer R3. Add a fusion layer R0 below the fusion model. The layer output vector corresponding to the input vector output by the fusion layer R0 is equal to the input vector. Each fusion layer has a layer weight, and the layer weights of each fusion layer are generally different. Multiply the layer output vector of the fusion layer R0 by the layer weight of the fusion layer R0 to obtain the weighted layer output vector of the fusion layer R0. Multiply the layer output vector of the fusion layer R1 by the layer weight of the fusion layer R1 to obtain the weighted layer output vector of the fusion layer R1. Multiply the layer output vector of the fusion layer R2 by the layer weight of the fusion layer R2 to obtain the weighted layer output vector of the fusion layer R2. Multiply the layer output vector of the fusion layer R3 by the layer weight of the fusion layer R3 to obtain the weighted layer output vector of the fusion layer R3. Add the weighted layer output vector of the fusion layer R0, the weighted layer output vector of the fusion layer R1, the weighted layer output vector of the fusion layer R2, and the weighted layer output vector of the fusion layer R3 to obtain the first output vector.

[0250] In one embodiment, referring to Figure 14 , step 1220 includes:

[0251] Step 1410, obtain a second weight matrix and a second offset vector;

[0252] Step 1420, multiply the second weight matrix by the transpose of the input vector and add the second offset vector to obtain a second sum vector;

[0253] Step 1430, perform a first exponential normalization on the elements of the second sum vector to obtain a layer weight vector, where the layer weight vector indicates the layer weights of each fusion layer.

[0254] In this embodiment, the second weight matrix refers to the matrix used for weight multiplication. The number of rows of the second weight matrix is the number of fusion layers after adding the fusion layer, and the number of columns of the second weight matrix is the dimension of the input vector. The first offset vector refers to the vector that adds an offset to the result of weight multiplication to obtain the second sum vector. The dimension of the second offset vector is the number of fusion layers after adding the fusion layer. Perform a first exponential normalization on the elements of the second sum vector to obtain a layer weight vector. The functions capable of performing the first exponential normalization include functions such as Softmax and Sigmoid.

[0255] In one example, the fusion model includes L fusion layers. The second weight matrix W t is a trainable parameter matrix with a dimension of (L + 1) × d. The second offset vector B t is a trainable bias parameter vector with a dimension of (L + 1) × 1. The input vector sequence The t-th input vector in the input vector sequence is denoted as The dimension of each input vector is 1×d. Multiply it with the transpose of the second weight matrix W t and add the second offset vector to obtain the second sum vector denoted as Then, perform the first exponential normalization on the elements of the second sum vector to obtain the layer weight vector The layer weight vector has a dimension of (L + 1)×1.

[0256] The advantage of the embodiment of this step 1410 - 1430 is that the second weight matrix can assign different weights to different input vectors, making the model more flexible and capable of adapting to the complex relationships between different input vectors; the first exponential normalization helps to expand the differences between numerical values, making the data easier to distinguish and process, and can also reduce the impact of outliers on the data distribution, making the data more stable. In summary, the accuracy of determining the layer weight vector is improved.

[0257] In one embodiment, referring to Figure 15A , the fusion layer includes a multi - head attention model and a feed - forward network. Figure 15A The shown fusion layer also includes two residual connections and two normalization layers, which will be described in detail later and will not be introduced here. Referring to Figure 16 , the layer output vector in step 1230 is generated through the following process:

[0258] Step 1610, input the layer input vector corresponding to the input vector into the attention model to obtain the attention output vector corresponding to the input vector;

[0259] Step 1620, input the attention output vector corresponding to the input vector into the feed - forward network to obtain the layer output vector corresponding to the input vector.

[0260] Specifically, the attention model refers to the ability to focus and concentrate on specific information when processing information. In this embodiment, the attention model is used to extract features from the layer input vector of the input vector, aiming to select / screen out the information that is beneficial to the context semantics between features. Common attention models include self-attention models, convolutional attention models, and attention mechanisms, etc. After obtaining the attention output vector corresponding to the input vector, the attention output vector corresponding to the input vector is input into the feedforward network. The feedforward network (Feedforward Neural Network) is a basic artificial neural network model, also known as the Multilayer Perceptron (MLP). This network model consists of a multi-layer structure composed of multiple neurons. Each layer of neurons is connected to the next layer, and the information is transmitted from the input layer through the hidden layer to the output layer without forming a loop. The characteristic of the feedforward network is that the information flow only propagates forward without generating cycles or feedback, so it is also called a forward propagation network. In this embodiment, the feedforward network is used to add non-linear information to the attention output vector. The obtained layer output vector can represent the complex non-linear relationship between features while retaining the rich context semantics between features.

[0261] The advantage of the embodiment of steps 1610 - 1620 is that by using the attention model and the feedforward network, the feature representation ability of the layer output vector is improved, so that the layer output vector contains rich context semantic information and complex non-linear relationships, thereby improving the push accuracy.

[0262] In one embodiment, referring to Figure 17 , step 1610 includes:

[0263] Step 1710, performing layer normalization on the layer input vector corresponding to the input vector to obtain the layer input vector after layer normalization:

[0264] Step 1720, inputting the layer input vector after layer normalization into the attention model to obtain the attention output vector corresponding to the input vector.

[0265] In this embodiment, normalization is a data preprocessing technique aimed at scaling data into a similar range so that the model can better process the data. Normalization is commonly used in machine learning and deep learning to ensure consistent numerical ranges between different features, thereby improving the performance and convergence speed of the model. In this embodiment, layer normalization is used to scale the vector elements of the layer input vector into the range of 0 to 1. Then, the layer input vector after layer normalization is input into the attention model. This can reduce the sensitivity of the model to different feature scales, thus improving the performance and generalization ability of the model. It should be noted that each input vector is obtained by interacting different numbers of target push basic features. Therefore, there are significant differences in scale and distribution between input vectors and between individual vector elements of input vectors. Therefore, in this embodiment, layer normalization is performed on the layer input vectors entering each fusion layer of the attention model, which can help the attention model better extract the rich feature information contained in the layer input vectors and obtain better feature representations.

[0266] The advantage of the embodiment of steps 1710-1720 is that by using layer normalization, the accuracy of the attention model in processing layer input vectors is improved, which is particularly useful when processing layer input vectors with different scales and distributions in this embodiment.

[0267] In one embodiment, referring to Figure 18 , step 1710 includes:

[0268] Step 1810, determining the mean and variance of each element in the layer input vector corresponding to the input vector;

[0269] Step 1820, determining the first sum of the variance and the smoothing amount;

[0270] Step 1830, subtracting the mean from each element of the layer input vector corresponding to the input vector and dividing by the first sum to obtain the layer input vector after layer normalization.

[0271] In one example, the t-th vector in the input vector sequence is denoted as The layer input vector of the input vector is first subjected to layer normalization (Layer Normalization, LN). The t-th layer input vector after layer normalization where ε is the smoothing amount. μ t is the mean calculated from all vector elements, that is, σ t is the standard deviation calculated from all vector elements, that is, Therefore, for the input vector sequence After performing LN processing, the layer-normalized layer input vector sequence S = (... , S t , ...).

[0272] The advantage of the embodiment of step 1810-1830 is that the mean and variance are important statistics of the data, and layer normalization through the mean and variance can better retain the distribution characteristics of each element in the layer input vector.

[0273] In one embodiment, step 1830 includes:

[0274] Subtract the mean from each element of the layer input vector corresponding to the input vector, and divide by the first sum to obtain the vector to be linearly processed;

[0275] Multiply each element of the vector to be linearly processed by the first parameter, and add the second parameter to obtain the layer-normalized layer input vector.

[0276] In this embodiment, wherein, the dimension of S t is d S = d. α is the trainable first parameter. β is the trainable second parameter.

[0277] The advantage of this embodiment is that on the basis that layer normalization through the mean and variance can better retain the distribution characteristics of each element in the layer input vector, the first parameter and the second parameter are used to increase the non-linear relationship, so that the layer-normalized input vector contains richer context semantic information between features.

[0278] In one embodiment, the attention model includes multiple attention heads. As Figure 19 shown, step 1720 includes:

[0279] Step 1910, input the layer-normalized layer input vector into the attention head to obtain the attention scores output by the attention head, where the multiple attention scores output by the multiple attention heads form an attention score vector;

[0280] Step 1920, multiply the attention score vector by the fifth weight matrix to obtain the attention output vector corresponding to the input vector, where the number of rows of the fifth weight matrix is the number of attention heads, and the number of columns of the fifth weight matrix is the dimension of the input vector.

[0281] In step 1910, each attention head separately focuses on the features of the layer-normalized layer input vector to obtain the attention scores. The attention scores can represent the feature information of each layer-normalized layer input vector itself. The attention score vector composed of multiple attention scores can represent the context semantic information between each vector on the basis of retaining the feature information of each layer-normalized layer input vector itself.

[0282] In one embodiment, referring to Figure 15B , the multi-head attention model includes H attention heads. Figure 15B The multi-head attention model shown also includes a first linear layer, a second linear layer, a third linear layer, a concatenation layer, and a fourth linear layer, which will not be described in detail here. Referring to Figure 20 , step 1910 includes:

[0283] Step 2010, multiply the transpose of the layer-normalized layer input vector by the sixth weight matrix (through H attention heads) to obtain a first intermediate vector, where the number of rows of the sixth weight matrix is the third dimension and the number of columns of the sixth weight matrix is the dimension of the input vector;

[0284] Step 2020, multiply the transpose of each layer output vector of the previous fusion layer by the seventh weight matrix (through H attention heads) to obtain a second intermediate vector corresponding to each layer output vector of the previous fusion layer, where the number of rows of the seventh weight matrix is the third dimension and the number of columns of the seventh weight matrix is the dimension of the input vector;

[0285] Step 2030, multiply the transpose of the layer-normalized layer input vector by the eighth weight matrix (through H attention heads) to obtain a third intermediate vector, where the number of rows of the eighth weight matrix is the fourth dimension, the fourth dimension is the dimension of the input vector divided by the number of attention heads, and the number of columns of the eighth weight matrix is the dimension of the input vector;

[0286] Step 2040, determine the attention scores based on the first intermediate vector, the second intermediate vectors corresponding to each layer output vector of the previous fusion layer, and the third intermediate vector.

[0287] In this embodiment, the sixth weight matrix, the seventh weight matrix, and the eighth weight matrix are all pre-set matrices for matrix multiplication. Generally, the feature information focused on during feature extraction by different weight matrices is not the same. Based on this principle, this embodiment uses three weight matrices (i.e., the sixth weight matrix, the seventh weight matrix, and the eighth weight matrix) to perform matrix operations on the normalized input vector at different levels, thereby obtaining attention scores with richer context semantics and further obtaining an attention score vector with higher accuracy.

[0288] In one embodiment, referring to Figure 15B, in addition to including H attention heads, the multi-head attention model further includes a first linear layer, a second linear layer, and a third linear layer. The first linear layer performs a linear transformation on the layer-normalized layer input vector before entering the sixth weight matrix. The second linear layer performs a linear transformation on the layer output vector before entering the seventh weight matrix. The third linear layer performs a linear transformation on the layer-normalized layer input vector before entering the eighth weight matrix. This can further improve the performance of the multi-head attention model. The multi-head attention model also includes a concatenation layer and a fourth linear layer. The concatenation layer is used to form an attention score vector by concatenating the H attention scores output by the H attention heads. The fourth linear layer applies a linear transformation to the attention score vector, further improving the performance of the multi-head attention model.

[0289] In one example, the layer-normalized layer input vector sequence S = (... , S t ,...). The sixth weight matrix of the h-th attention head has a dimension of d QK ×d S . d QK refers to the third dimension, and d S refers to the dimension of the input vector. The first intermediate vector The first intermediate vector has a dimension of d QK ×1. The seventh weight matrix of the h-th attention head has a dimension of d QK ×d S . d QK refers to the third dimension, and d S refers to the dimension of the input vector. The second intermediate vector corresponding to each layer output vector of the previous fusion layer has a dimension of d QK ×1. The eighth weight matrix of the h-th attention head has a dimension of d V ×d S . d V refers to the fourth dimension, and d V = d S / H. H refers to the number of attention heads. d QK can generally be taken to be the same as d V . d S refers to the dimension of the input vector. The third intermediate vector The third intermediate vector has a dimension of d V ×1.

[0290] In one embodiment, step 2040 includes:

[0291] Multiply the first intermediate vector by the transpose of the second intermediate vector corresponding to each layer output vector of the previous fusion layer respectively to obtain the influence score of each layer output vector of the previous fusion layer on the layer input vector;

[0292] Divide the influence score of each layer output vector of the previous fusion layer on the layer input vector by the third dimension respectively to obtain the decay influence score of each layer output vector of the previous fusion layer on the layer input vector, where the decay influence scores of multiple layer output vectors of the previous fusion layer on the layer input vector form a decay influence score vector;

[0293] Multiply the decay influence score vector by the elements with the same serial number in the third intermediate vector and accumulate the multiplication results to obtain the attention score.

[0294] In this embodiment, the calculation process of the decay influence score vector can be shown as follows: refers to the decay influence score vector output by the h-th attention head for the layer input vector S after layer normalization t output. refers to the first intermediate vector. refers to the second intermediate vector corresponding to the layer output vector of the i-th fusion layer. refers to the influence score of the layer output vector of the i-th fusion layer on the t-th layer input vector. refers to the decay influence score of the layer output vector of the i-th fusion layer on the t-th layer input vector. For the t-th layer input vector, the calculation process of the attention score of the h-th attention head can be shown as follows: refers to the attention score output by the h-th attention head for the t-th layer input vector. T refers to the total number of layer input vectors. refers to the decay influence score vector composed of the decay influence scores of the i-th fusion layer on the layer input vector. V i h refers to the third intermediate vector under the h-th attention head of the i-th fusion layer.

[0295] It should be noted that the above softmax is an exponential normalization function used to perform exponential normalization on the decay influence score vector. In some other embodiments, softmax can be removed, that is, there is no need to perform exponential normalization on the decay influence score vector.

[0296] The advantage of this embodiment is that by using the correlation weight distribution of the input vector of the t-th layer among the input vectors of T layers under the h-th attention head, the attenuation influence score vector is first determined, and then the attention score is determined. In this way, the attention score can contain the characteristic information of the layer input vectors themselves after normalization for each layer, and can also represent the context semantic information between the layer input vectors after normalization for each layer.

[0297] To further model the context semantics between features, a fifth weight matrix is also introduced to perform multi-level conversions such as elements and dimensions on the attention score vector, so that the obtained attention output vector contains more abundant context semantic information between features.

[0298] In one example, the attention model includes H attention heads. Each attention head calculates the attention score for the layer input vector after layer normalization to obtain the attention score. The H attention scores output by the H attention heads form an attention score vector. Referring to the above, the sequence of layer input vectors after layer normalization is S=(..., S t ,...). The layer input vector S t After being processed by the h-th attention head, the obtained attention score is The layer input vector S t after normalization is concatenated with the H attention scores of the H attention heads, and the obtained attention score vector is denoted as The number of rows of the attention score vector is 1, and the number of columns is H. The fifth weight matrix W A The number of rows of the attention score vector is H, and the number of columns is d S =d. The attention output vector

[0299] In one embodiment, referring to Figure 15A , the fusion layer includes, in addition to the multi-head attention model and the feed-forward network, a first residual connection layer and a first normalization layer. Referring to Figure 21 , the layer output vector in step 1230 is generated through the following process:

[0300] Step 1610, input the layer input vector corresponding to the input vector into the attention model to obtain the attention output vector corresponding to the input vector;

[0301] Step 2110, (through the first residual connection layer) add the layer input vector corresponding to the input vector and the attention output vector corresponding to the input vector to obtain a third sum vector;

[0302] Step 2120, (through the first normalization layer) perform layer normalization on the third sum vector to obtain a first normalized vector;

[0303] Step 2130: Input the first normalized vector into the feed-forward network to obtain a layer output vector corresponding to the input vector.

[0304] The advantages of the embodiments of this Step 1610 and Steps 2110 - 2130 are that the first residual connection layer and the first normalization layer can help the gradient propagate more easily in the model. This helps alleviate the problem of vanishing gradients commonly found in deep networks, making the model more stable and effective.

[0305] In one embodiment, referring to Figure 15A , in addition to the multi-head attention model, the first residual connection layer, the first normalization layer, and the feed-forward network, the fusion layer further includes a second residual connection layer and a second normalization layer. Referring to Figure 22 , the layer output vector in Step 1230 is generated through the following process:

[0306] Step 1610: Input the layer input vector corresponding to the input vector into the attention model to obtain an attention output vector corresponding to the input vector;

[0307] Step 2110: Add (through the first residual connection layer) the layer input vector corresponding to the input vector and the attention output vector corresponding to the input vector to obtain a third sum vector;

[0308] Step 2120: Perform layer normalization (through the first normalization layer) on the third sum vector to obtain a first normalized vector;

[0309] Step 2210: Input the first normalized vector into the feed-forward network to obtain a feed-forward vector corresponding to the input vector;

[0310] Step 2220: Add (through the second residual connection layer) the first normalized vector and the feed-forward vector corresponding to the input vector to obtain a fourth sum vector;

[0311] Step 2230: Perform layer normalization (through the second normalization layer) on the fourth sum vector to obtain a layer output vector corresponding to the input vector.

[0312] The advantages of the embodiments of this Step 1610, Steps 2110 - 2120, and Steps 2210 - 2230 are that the fusion layer composed of the multi-head attention model, the first residual connection layer, the first normalization layer, the feed-forward network, the second residual connection layer, and the second normalization layer improves the feature representation ability of the layer output vector, enabling the layer output vector to contain rich context semantic information and complex non-linear relationships, thereby improving the push accuracy.

[0313] In one embodiment, referring to Figure 23 , Step 2210 includes:

[0314] Step 2310: Obtain a third weight matrix and a third offset vector, where the number of rows of the third weight matrix is the second dimension, the number of columns of the third weight matrix is the dimension of the input vector, and the dimension of the third offset vector is the second dimension;

[0315] Step 2320: Multiply the third weight matrix by the transpose of the first normalized vector and add the third offset vector to obtain a fifth sum vector;

[0316] Step 2330: Apply an activation function to the fifth sum vector and determine a feed-forward vector corresponding to the input vector based on the result of the activation function.

[0317] The advantages of the embodiments of steps 2310 - 2330 are that the third weight matrix can assign different weights to different first normalized vectors, making the model more flexible and capable of adapting to the complex relationships between different first normalized vectors; applying the activation function helps to improve the non-linear relationship between vectors and can fully reflect the complex relationships between different first normalized vectors, improving the accuracy of determining the feed-forward vector.

[0318] In one embodiment, step 2330 includes:

[0319] Obtain a fourth weight matrix and a fourth offset vector, where the number of rows of the fourth weight matrix is the dimension of the input vector, the number of columns of the third weight matrix is the second dimension, and the dimension of the fourth offset vector is equal to the dimension of the input vector;

[0320] Multiply the fourth weight matrix by the transpose of the result of the activation function and add the fourth offset vector to obtain a feed-forward vector corresponding to the input vector.

[0321] The advantages of this embodiment are that the fourth weight matrix can assign different weights to the result of the activation function, making the model more flexible and capable of further reflecting the complex relationships between different vectors, improving the accuracy of determining the feed-forward vector.

[0322] In one example, referring to the above example, for the layer input vector S after layer normalization t , the attention output vector output by the attention model Add the layer input vector corresponding to the input vector to the attention output vector corresponding to the input vector to obtain a third sum vector Perform layer normalization on the third sum vector S' t to obtain a first normalized vector LN(S' t ). Input the first normalized vector LN(S' t ) into the feed-forward network to obtain a feed-forward vector FFN(LN(S't )) = W 2 ·[ReLU(W 1 ·LN(S′ t ) + b 1 )] + b 2 。S′ t refers to the third sum vector. LN(S′ t ) refers to the first normalized vector. W 1 refers to the third weight matrix. W 1 has a dimension of d FF ×d S 。d FF refers to the second dimension. d S refers to the dimension of the input vector. b 1 refers to the third offset vector. b 1 has a dimension of d FF ×1. ReLU refers to the activation function. Other activation functions can also be selected according to requirements. W 2 refers to the fourth weight matrix. W 2 has a dimension of d S ×d FF 。d FF refers to the second dimension. d S refers to the dimension of the input vector. b 2 refers to the fourth offset vector. b 2 has a dimension of d S ×1. Add the first normalized vector LN(S′ t ) to the feed-forward vector FFN(LN(S′ t )) corresponding to the input vector to obtain the fourth sum vector as LN(S′ t ) + FFN(LN(S′ t ))。Perform layer normalization on the fourth sum vector to obtain the layer output vector corresponding to the input vector

[0323] After obtaining the layer weights and the layer output vectors corresponding to the input vector for each output of the fusion layer through the above embodiments, in step 1230, use the layer weights to perform a weighted sum on the layer output vectors corresponding to the input vector for each output of the fusion layer after adding the fusion layer to obtain the first output vector.

[0324] In one example, based on the foregoing, the fusion layer processes the input sequence to obtain the output sequence and the layer weight vector The input sequence The t-th vector in is denoted as The dimension of each input vector is 1×d. The second weight matrix W tis a trainable parameter matrix with dimensions of (L + 1) × d. The second offset vector B t is a trainable bias parameter vector with dimensions of (L + 1) × 1. Assume that the fusion model includes L fusion layers, and an additional fusion layer is added below the fusion model. Therefore, there are a total of L + 1 fusion layers. The additional fusion layer mentioned above is denoted as the 0th fusion layer. And the layer output vector corresponding to the input vector output by the 0th fusion layer is equal to the input vector That is After being processed by L + 1 fusion layers, the input vector There are a total of L + 1 output vectors represents the output vector output by the l-th fusion layer represents the layer weight of the l-th fusion layer. The first output vector

[0325] The advantage of the embodiment of step 1210 - 1230 is that, on the basis of ensuring that the fusion layer fully models / preserves the hierarchical semantics between each input vector, the layer weights of each fusion layer can be flexibly adjusted, so that the first output vector has better feature ability and generalization ability, thereby improving the push accuracy

[0326] After obtaining the first output vector and the input vector weights as described above, in one example, the second output vector t refers to the t-th input vector refers to the weight corresponding to the t-th input vector in the input vector weights refers to the first output vector corresponding to the t-th input vector. The first output vector has dimensions of 1 × d. The second output vector has dimensions of 1 × d. T refers to the total number of input vectors. The first probability The first weight vector is a trainable parameter vector with dimensions of 1 × d, and the first offset is a 1-dimensional trainable bias parameter. Finally, the content to be pushed is pushed to the target object based on the first probability p

[0327] The training process of the fusion model in the embodiments of the present disclosure

[0328] In one embodiment, as Figure 11 shown, the fusion model includes multiple fusion layers ( Figure 11 shows 3 fusion layers, but the actual number can be set according to requirements). As Figure 15A shown, the fusion layer includes a multi-head attention model, a first residual connection layer, a first normalization layer, a feed-forward network, a second residual connection layer, and a second normalization layer. As Figure 15BAs shown in the figure, the multi-head attention model includes a first linear layer, a second linear layer, a third linear layer, H attention heads, a splicing layer, and a fourth linear layer. The introduction of the specific structure in the fusion model has been detailed above and will not be elaborated here.

[0329] In one embodiment, the fusion model is trained in the following manner:

[0330] Obtain a sample set, which includes multiple samples. Each sample includes a first number of sample push basic features. Among them, the first number of sample push basic features include sample object features and sample content features. The sample has a label, which indicates whether to push the sample content to the sample object;

[0331] For each sample, based on the first number of sample push basic features, obtain multiple combinations of sample push basic features. Each combination of sample push basic features includes a second number of sample push basic features selected from the first number of sample push basic features, and the second number is less than the first number;

[0332] Generate a second-order cross feature vector corresponding to the combination of sample push basic features;

[0333] Input the first number of sample push basic features into the feature full-cross model to obtain a full-cross feature vector;

[0334] Input the first number of sample push basic feature vectors corresponding to the first number of sample push basic features, the multiple second-order cross feature vectors corresponding to the multiple combinations of sample push basic features, and the full-cross feature vector into the fusion model to obtain a second probability of pushing the sample content to the sample object;

[0335] Calculate the loss function based on the second probability and the label;

[0336] Train the fusion model based on the loss function.

[0337] The above steps will be described in detail below.

[0338] The sample set refers to a set used to store multiple samples. The sample push basic features are the same as the above-mentioned target push basic features, except that here they all exist as training samples. For the sake of saving space, they will not be elaborated here. The sample push basic features also include sample scenario features. The sample scenario features here are similar to the above-mentioned target scenario features, except that here they all exist as training samples. For the sake of saving space, they will not be elaborated here. Therefore, obtaining the sample scenario features where the sample object is located can improve the accuracy of feature fusion of the trained fusion model, thereby improving the push accuracy.

[0339] The process of generating the second-order cross feature vectors and the full cross feature vectors based on the sample push basic features is basically the same as the above steps 320 - 340. Only here they all exist as training samples. For the sake of saving space, it will not be elaborated here. The second probability predicted by the fusion model is the same as the first probability in the above step 350. Only here they all exist as training samples. For the sake of saving space, it will not be elaborated here.

[0340] If the sample push basic features include sample object features and sample content features, the label indicates whether to push the sample content to the sample object. If the sample push basic features include sample object features, sample scenario features, and sample content features, the label indicates whether to push the sample content to the sample object in the sample scenario. For example, the label of a sample is {Content A: 1, Content B: 0}, where the label of 1 represents that the push is completed, and the label of 0 represents that the push is not completed. That is to say, in this example, when the sample object is in the sample scenario, Content A has been pushed to the sample object, but Content B has not been pushed to the sample object.

[0341] Specifically, in one embodiment, the loss function calculated based on the second probability and the label is as follows: where is the loss function, y is the label, and p is the second probability. If the label indicates that the push is completed, then the training label is 1, and vice versa, the training label is 0.

[0342] It should be noted that the above loss function is not unique. For example, logarithmic operations or linear operations can be added to the above loss function. This embodiment does not make specific limitations on this.

[0343] After determining the loss function, the fusion model can be trained based on the loss function, that is, the parameters of the fusion model are adjusted. Specifically, a first threshold can be set in advance. When the loss function is less than the first threshold, the training process ends. When the loss function is greater than or equal to the first threshold, the parameters of the fusion model are adjusted until the loss function is less than the first threshold.

[0344] The advantage of this embodiment is that by constructing the loss function to train the fusion model, the accuracy of feature fusion of the trained model is improved.

[0345] Implementation detail diagram of the push processing method in the embodiments of the present disclosure

[0346] Next, refer to Figure 24 , and detail the implementation details of the push processing method in the embodiments of the present disclosure by way of example.

[0347] In step 2410, obtain the first number of target push basic features.

[0348] In step 2420, based on the first number of target push basic features, multiple combinations of target push basic features are obtained. Each combination of target push basic features includes the second number of target push basic features selected from the first number of target push basic features, where the second number is less than the first number; a second-order cross feature vector corresponding to the combination of target push basic features is generated.

[0349] In step 2430, the first number of target push basic features are input into the feature full-cross model to obtain a full-cross feature vector.

[0350] In step 2440, each of the first number of target push basic feature vectors, multiple second-order cross feature vectors, and the full-cross feature vector is used as an input vector to obtain a first weight matrix and a first offset vector. Here, the number of rows of the first weight matrix is the number of input vectors, the number of columns of the first weight matrix is the number of input vectors multiplied by the dimension of the input vector, and the dimension of the first offset vector is the number of input vectors; multiply the first weight matrix by the transpose of the concatenated vector of multiple input vectors and add the first offset vector to obtain a first sum vector; perform a first exponential normalization on the elements of the first sum vector to obtain an input vector weight vector.

[0351] In step 2450, a fusion layer is added below the fusion model. The layer output vector corresponding to the input vector output by the added fusion layer is equal to the input vector; the layer weights of each fusion layer after adding the fusion layer are determined; the layer weights are used to perform a weighted sum on the layer output vectors corresponding to the input vector output by each fusion layer after adding the fusion layer to obtain a first output vector.

[0352] In step 2460, a weighted sum is performed on the first output vector of the fusion model for the input vector using the input vector weight to obtain a second output vector for pushing the content to be pushed to the target object;

[0353] In step 2470, a first weight vector and a first offset are obtained. Here, the dimension of the first weight vector is equal to the dimension of the input vector; multiply the first weight vector by the transpose of the second output vector and add the first offset to obtain a first score; perform a second exponential normalization on the first score to obtain a first probability.

[0354] Description of the device and equipment in the embodiments of the present disclosure

[0355] It can be understood that although the steps in each of the above flowcharts are shown sequentially according to the indication of the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless there is a clear description in this embodiment, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the above flowchart may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of the steps or stages in other steps or other steps.

[0356] It should be noted that in each specific embodiment of the present application, when it comes to performing relevant processing based on data related to the target content characteristics such as target content attribute information or attribute information set, the permission or consent of the target content will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiment of the present application needs to obtain target content attribute information, it will obtain the separate permission or separate consent of the target content through methods such as pop-up windows or jumping to a confirmation page. After clearly obtaining the separate permission or separate consent of the target content, the necessary target content-related data for the normal operation of the embodiment of the present application will be obtained.

[0357] Figure 25 FIG. 2500 is a schematic structural diagram of a push processing device 2500 provided by an embodiment of the present disclosure. The push processing device 2500 includes:

[0358] A first acquisition unit 2510, configured to acquire a first number of target push basic features, where the first number of target push basic features includes target object features and content features to be pushed;

[0359] A second acquisition unit 2520, configured to acquire a plurality of target push basic feature combinations based on the first number of target push basic features, where each target push basic feature combination includes a second number of target push basic features selected from the first number of target push basic features, and the second number is less than the first number;

[0360] A first generation unit 2530, configured to generate a second-order cross feature vector corresponding to the target push basic feature combination;

[0361] A second generation unit 2540, configured to input the first number of target push basic features into a feature full-cross model to obtain a full-cross feature vector;

[0362] A prediction unit 2550, configured to input the first number of target push basic feature vectors corresponding to the first number of target push basic features, the second number of order cross feature vectors corresponding to the multiple target push basic feature combinations, and the full cross feature vector into a fusion model, to obtain a first probability of pushing content to be pushed to a target object, and based on the first probability, push the content to be pushed to the target object.

[0363] Optionally, the prediction unit 2550 is specifically configured to:

[0364] Take each of the first number of target push basic feature vectors, the multiple second number of order cross feature vectors, and the full cross feature vector as an input vector, determine an input vector weight vector, where the input vector weight vector indicates the input vector weight of each input vector;

[0365] Use the input vector weights to perform a weighted sum on the first output vector of the fusion model for the input vector, to obtain a second output vector of pushing content to be pushed to the target object;

[0366] Based on the second output vector, determine the first probability.

[0367] Optionally, the prediction unit 2550 is further specifically configured to:

[0368] Obtain a first weight matrix and a first offset vector, where the number of rows of the first weight matrix is the number of input vectors, the number of columns of the first weight matrix is the number of input vectors multiplied by the dimension of the input vector, and the dimension of the first offset vector is the number of input vectors;

[0369] Multiply the first weight matrix by the transpose of the concatenated vector of the multiple input vectors, and add the first offset vector to obtain a first sum vector;

[0370] Perform a first exponential normalization on the elements of the first sum vector to obtain the input vector weight vector.

[0371] Optionally, the prediction unit 2550 is further specifically configured to:

[0372] Obtain a first weight vector and a first offset, where the dimension of the first weight vector is equal to the dimension of the input vector;

[0373] Multiply the first weight vector by the transpose of the second output vector, and add the first offset to obtain a first score;

[0374] Perform a second exponential normalization on the first score to obtain the first probability.

[0375] Optionally, the fusion model includes multiple fusion layers. The bottommost fusion layer among the multiple fusion layers takes multiple input vectors as input and generates layer output vectors corresponding to each input vector. The other fusion layers among the multiple fusion layers receive the layer output vectors of the next lower fusion layer corresponding to the input vectors and generate layer output vectors of the fusion layer corresponding to the input vectors;

[0376] The first output vector is generated by the prediction unit 2550 in the following manner:

[0377] Based on the input vectors and the layer output vectors corresponding to the input vectors output by each fusion layer, determine the first output vector.

[0378] Optionally, the prediction unit 2550 is further specifically configured to:

[0379] Add a fusion layer below the fusion model, and the layer output vector corresponding to the input vector output by the added fusion layer is equal to the input vector;

[0380] Determine the layer weights of each fusion layer after adding the fusion layer;

[0381] Use the layer weights to perform a weighted sum on the layer output vectors corresponding to the input vector output by each fusion layer after adding the fusion layer to obtain the first output vector.

[0382] Optionally, the prediction unit 2550 is further specifically configured to:

[0383] Obtain a second weight matrix and a second offset vector, where the number of rows of the second weight matrix is the number of fusion layers after adding the fusion layer, the number of columns of the second weight matrix is the dimension of the input vector, and the second offset vector is the number of fusion layers after adding the fusion layer;

[0384] Multiply the second weight matrix by the transpose of the input vector and add the second offset vector to obtain a second sum vector;

[0385] Perform a first exponential normalization on the elements of the second sum vector to obtain a layer weight vector, and the layer weight vector indicates the layer weights of each fusion layer.

[0386] Optionally, the fusion layer includes a multi-head attention model and a feed-forward network, and the layer output vector is generated by the prediction unit 2550 through the following process:

[0387] Input the layer input vector corresponding to the input vector into the attention model to obtain an attention output vector corresponding to the input vector;

[0388] Input the attention output vector corresponding to the input vector into the feed-forward network to obtain a layer output vector corresponding to the input vector.

[0389] Optionally, the prediction unit 2550 is further specifically configured to:

[0390] Add the layer input vector corresponding to the input vector and the attention output vector corresponding to the input vector to obtain a third sum vector;

[0391] Perform layer normalization on the third sum vector to obtain a first normalized vector;

[0392] Input the first normalized vector into a feed-forward network to obtain a layer output vector corresponding to the input vector.

[0393] Optionally, the prediction unit 2550 is further specifically configured to:

[0394] Input the first normalized vector into a feed-forward network to obtain a feed-forward vector corresponding to the input vector;

[0395] Add the first normalized vector and the feed-forward vector corresponding to the input vector to obtain a fourth sum vector;

[0396] Perform layer normalization on the fourth sum vector to obtain a layer output vector corresponding to the input vector.

[0397] Optionally, the prediction unit 2550 is further specifically configured to:

[0398] Obtain a third weight matrix and a third offset vector, where the number of rows of the third weight matrix is the second dimension, the number of columns of the third weight matrix is the dimension of the input vector, and the dimension of the third offset vector is the second dimension;

[0399] Multiply the third weight matrix by the transpose of the first normalized vector and add the third offset vector to obtain a fifth sum vector;

[0400] Apply an activation function to the fifth sum vector and determine a feed-forward vector corresponding to the input vector based on the result of the activation function.

[0401] Optionally, the prediction unit 2550 is further specifically configured to:

[0402] Obtain a fourth weight matrix and a fourth offset vector, where the number of rows of the fourth weight matrix is the dimension of the input vector, the number of columns of the third weight matrix is the second dimension, and the dimension of the fourth offset vector is equal to the dimension of the input vector;

[0403] Multiply the fourth weight matrix by the transpose of the result of the activation function and add the fourth offset vector to obtain a feed-forward vector corresponding to the input vector.

[0404] Optionally, the prediction unit 2550 is further specifically configured to:

[0405] Perform layer normalization on the layer input vector corresponding to the input vector to obtain a layer input vector after layer normalization;

[0406] Input the layer input vector after layer normalization into the attention model to obtain the attention output vector corresponding to the input vector.

[0407] Optionally, the prediction unit 2550 is further specifically configured to:

[0408] Determine the mean and variance of each element in the layer input vector corresponding to the input vector;

[0409] Determine the first sum of the variance and the smoothing amount;

[0410] Subtract the mean from each element of the layer input vector corresponding to the input vector and divide by the first sum to obtain the layer input vector after layer normalization.

[0411] Optionally, the prediction unit 2550 is further specifically configured to:

[0412] Subtract the mean from each element of the layer input vector corresponding to the input vector and divide by the first sum to obtain the vector to be linearly processed;

[0413] Multiply each element of the vector to be linearly processed by the first parameter and add the second parameter to obtain the layer input vector after layer normalization.

[0414] Optionally, the attention model includes multiple attention heads; the prediction unit 2550 is further specifically configured to:

[0415] Input the layer input vector after layer normalization into the attention head to obtain the attention scores output by the attention head, where the multiple attention scores output by the multiple attention heads form an attention score vector;

[0416] Multiply the attention score vector by the fifth weight matrix to obtain the attention output vector corresponding to the input vector, where the number of rows of the fifth weight matrix is the number of attention heads, and the number of columns of the fifth weight matrix is the dimension of the input vector.

[0417] Optionally, the prediction unit 2550 is further specifically configured to:

[0418] Multiply the sixth weight matrix by the transpose of the layer input vector after layer normalization to obtain the first intermediate vector, where the number of rows of the sixth weight matrix is the third dimension, and the number of columns of the sixth weight matrix is the dimension of the input vector;

[0419] Multiply the seventh weight matrix by the transpose of each layer output vector of the previous fusion layer respectively to obtain the second intermediate vector corresponding to each layer output vector of the previous fusion layer, where the number of rows of the seventh weight matrix is the third dimension, and the number of columns of the seventh weight matrix is the dimension of the input vector;

[0420] Multiply the eighth weight matrix by the transpose of the layer input vector after layer normalization to obtain a third intermediate vector, where the number of rows of the eighth weight matrix is the fourth dimension, the fourth dimension is the dimension of the input vector divided by the number of attention heads, and the number of columns of the eighth weight matrix is the dimension of the input vector;

[0421] Determine attention scores based on the first intermediate vector, the second intermediate vectors corresponding to each layer output vector of the previous fusion layer, and the third intermediate vector.

[0422] Optionally, the prediction unit 2550 is further specifically configured to:

[0423] Multiply the first intermediate vector by the transposes of the second intermediate vectors corresponding to each layer output vector of the previous fusion layer respectively to obtain the influence scores of each layer output vector of the previous fusion layer on the layer input vector;

[0424] Divide the influence scores of each layer output vector of the previous fusion layer on the layer input vector by the third dimension respectively to obtain the decay influence scores of each layer output vector of the previous fusion layer on the layer input vector, where the decay influence scores of multiple layer output vectors of the previous fusion layer on the layer input vector form a decay influence score vector;

[0425] Multiply the decay influence score vector by the elements with the same serial numbers in the third intermediate vector and accumulate the multiplication results to obtain attention scores.

[0426] Optionally, the first generation unit 2530 is specifically configured to:

[0427] Vectorize the target push base features in the target push base feature combination into target push base feature vectors;

[0428] Multiply the target push base feature vectors corresponding to the respective target push base features in the target push base feature combination element by element to obtain a second-order cross feature vector.

[0429] Optionally, after vectorizing the target push base features in the target push base feature combination into target push base feature vectors, the push processing device further includes a pooling unit (not shown) for:

[0430] Determine that the dimension of the target push base feature vector is greater than the first dimension;

[0431] Perform pooling processing on the target push base feature vector so that the dimension of the pooled target push base feature vector is equal to the first dimension.

[0432] Optionally, the feature full-cross model includes multiple feature full-cross models corresponding to multiple tasks;

[0433] The second generation unit 2540 is specifically configured to: for each of the multiple tasks, input the first number of target push basic feature inputs into the feature full-crossing model corresponding to the task, and obtain the full-crossing feature vector corresponding to the task.

[0434] Optionally, the fusion model includes multiple fusion models corresponding to multiple tasks;

[0435] The prediction unit 2550 is further specifically configured to:

[0436] For each of the multiple tasks, input the first number of target push basic feature vectors, the multiple second-number-order crossing feature vectors corresponding to the multiple target push basic feature combinations, and the full-crossing feature vector corresponding to the task into the fusion model corresponding to the task, and obtain the first probability for the task of pushing the content to be pushed for the target object;

[0437] Based on the first probabilities for each task, push the content to be pushed for the target object.

[0438] Optionally, the prediction unit 2550 is further specifically configured to:

[0439] Perform a weighted sum of the first probabilities for each task to obtain a second probability;

[0440] Based on the second probability, push the content to be pushed for the target object.

[0441] Refer to Figure 26 , Figure 26 FIG. is a block diagram of a part of the object terminal 110 for implementing the push processing method of the embodiments of the present disclosure. The terminal includes: a Radio Frequency (RF) circuit 2610, a memory 2615, an input unit 2630, a display unit 2640, a sensor 2650, an audio circuit 2660, a wireless fidelity (WiFi) module 2670, a processor 2680, and a power supply 2690, etc. Those skilled in the art can understand that Figure 26 The structure of the object terminal 110 shown does not constitute a limitation on a mobile phone or a computer, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0442] The RF circuit 2610 can be used for receiving and sending signals during information reception or call processes. Specifically, after receiving the downlink information of the base station, it is given to the processor 2680 for processing; in addition, the uplink data designed is sent to the base station.

[0443] The memory 2615 can be used to store software programs and modules. The processor 2680 executes various functional applications and data processing of the content terminal by running the software programs and modules stored in the memory 2615.

[0444] The input unit 2630 can be used to receive input digital or character information and generate key signal inputs related to the settings and function controls of the content terminal. Specifically, the input unit 2630 can include a touch panel 2631 and other input devices 2632.

[0445] The display unit 2640 can be used to display the input information or provided information and various menus of the content terminal. The display unit 2640 can include a display panel 2941.

[0446] The audio circuit 2660, speaker 2661, and microphone 2662 can provide an audio interface.

[0447] In this embodiment, the processor 2680 included in the object terminal 110 can execute the push processing method of the previous embodiment.

[0448] The object terminal 110 of the embodiments of the present disclosure includes, but is not limited to, mobile phones, computers, intelligent voice interaction devices, intelligent home appliances, vehicle-mounted terminals, aircraft, etc. The embodiments of the present invention can be applied to various scenarios, including but not limited to content recommendation, data screening, etc.

[0449] Figure 27 It is a structural block diagram of a part of the push processing server 140 for implementing the push processing method of the embodiments of the present disclosure. The push processing server 140 may vary greatly due to configuration or performance differences and may include one or more central processing units (CPUs) 2722 (for example, one or more processors) and a memory 2732, and one or more storage media 2730 (for example, one or more mass storage devices) storing application programs 2742 or data 2744. Among them, the memory 2732 and the storage media 2730 can be transient storage or persistent storage. The program stored in the storage media 2730 can include one or more modules (not shown in the figure), and each module can include a series of instruction operations on the server. Further, the central processor 2722 can be configured to communicate with the storage media 2730 and execute a series of instruction operations in the storage media 2730 on the server.

[0450] The push processing server 140 may also include one or more power supplies 2726, one or more wired or wireless network interfaces 2750, one or more input / output interfaces 2758, and / or one or more operating systems 2741, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, and so on.

[0451] The central processing unit 2722 in the push processing server 140 may be used to execute the push processing method of the embodiments of the present disclosure.

[0452] The embodiments of the present disclosure also provide a computer-readable storage medium for storing program codes for executing the push processing methods of the foregoing various embodiments.

[0453] The embodiments of the present disclosure also provide a computer program product including a computer program. The processor of the computer device reads and executes the computer program, so that the computer device executes to implement the above-mentioned push processing method.

[0454] It should be understood that in the present disclosure, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar contents and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "including" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0455] It should be understood that in the present disclosure, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated contents, indicating that there can be three relationships. For example, "A and / or B" may mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated contents before and after are an "or" relationship. "At least one (piece) of the following" or its similar expression means any combination of these items, including any combination of single item (piece) or plural items (pieces). For example, at least one (piece) of a, b or c may mean: a, b, c, "a and b", "a and c", "b and c", or "a, b and c", where a, b, c can be single or multiple.

[0456] It should be understood that in the description of the embodiments of the present disclosure, the meaning of "a plurality of (or multiple)" is more than two. Greater than, less than, exceeding, etc. are understood not to include the recited number, and above, below, within, etc. are understood to include the recited number.

[0457] In several embodiments provided by the present disclosure, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in an electrical, mechanical, or other form.

[0458] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0459] In addition, in the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.

[0460] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in the various embodiments of the present disclosure. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes.

[0461] It should also be understood that the various embodiments provided by the present disclosure can be combined arbitrarily to achieve different technical effects.

[0462] The above is a specific description of the embodiments of the present disclosure. However, the present disclosure is not limited to the above embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present disclosure, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present disclosure.

Claims

1. A push processing method, characterized in that, it includes: Obtain a first number of target push basic features, where the first number of the target push basic features include target object features and content features to be pushed; Based on the first number of the target push basic features, obtain a plurality of target push basic feature combinations, and each of the target push basic feature combinations includes a second number of the target push basic features selected from the first number of the target push basic features, and the second number is less than the first number; Generate a second-order cross feature vector corresponding to the target push basic feature combination; Input the first number of the target push basic features into a feature full-cross model to obtain a full-cross feature vector; Input the first number of target push basic feature vectors corresponding to the first number of the target push basic features, the plurality of second-order cross feature vectors corresponding to the plurality of the target push basic feature combinations, and the full-cross feature vector into a fusion model to obtain a first probability of pushing the content to be pushed for the target object, and push the content to be pushed for the target object based on the first probability.

2. The push processing method according to claim 1, characterized in that, The step of inputting the first number of target push basic feature vectors corresponding to the first number of the target push basic features, the plurality of second-order cross feature vectors corresponding to the plurality of the target push basic feature combinations, and the full-cross feature vector into a fusion model to obtain a first probability of pushing the content to be pushed for the target object includes: Take each of the first number of the target push basic feature vectors, the plurality of second-order cross feature vectors, and the full-cross feature vector as an input vector, and determine an input vector weight vector, where the input vector weight vector indicates the input vector weight of each input vector; Use the input vector weights to perform a weighted sum on the first output vector of the fusion model for the input vector to obtain a second output vector for pushing the content to be pushed for the target object; Determine the first probability based on the second output vector.

3. The push processing method according to claim 2, characterized in that, The step of determining the input vector weight vector includes: Obtain a first weight matrix and a first offset vector, where the number of rows of the first weight matrix is the number of the input vectors, the number of columns of the first weight matrix is the number of the input vectors multiplied by the dimension of the input vector, and the dimension of the first offset vector is the number of the input vectors; Multiply the first weight matrix by the transpose of the concatenated vector of the plurality of the input vectors and add the first offset vector to obtain a first sum vector; Perform a first exponential normalization on the elements of the first sum vector to obtain the input vector weight vector.

4. The push processing method according to claim 2, characterized in that, The step of determining the first probability based on the second output vector includes: Obtain a first weight vector and a first offset, where the dimension of the first weight vector is equal to the dimension of the input vector; Multiply the first weight vector by the transpose of the second output vector and add the first offset to obtain a first score; Perform a second exponential normalization on the first score to obtain the first probability.

5. The push processing method according to claim 2, wherein, The fusion model includes a plurality of fusion layers. The bottommost one of the plurality of fusion layers inputs a plurality of the input vectors and generates a layer output vector corresponding to each of the input vectors. The other fusion layers in the plurality of fusion layers receive the layer output vectors of the fusion layer below corresponding to the input vectors and generate the layer output vectors of the fusion layer corresponding to the input vectors; The first output vector is generated by the following method: Determine the first output vector based on the input vector and the layer output vectors corresponding to the input vector output by each fusion layer.

6. The push processing method according to claim 5, wherein, The determining the first output vector based on the input vector and the layer output vectors corresponding to the input vector output by each fusion layer includes: Add a fusion layer below the fusion model, and the layer output vector corresponding to the input vector output by the added fusion layer is equal to the input vector; Determine the layer weights of each fusion layer after adding the fusion layer; Perform a weighted sum on the layer output vectors corresponding to the input vector output by each fusion layer after adding the fusion layer using the layer weights to obtain the first output vector.

7. The push processing method according to claim 6, wherein, The determining the layer weights of each fusion layer after adding the fusion layer includes: Obtain a second weight matrix and a second offset vector, where the number of rows of the second weight matrix is the number of fusion layers after adding the fusion layer, the number of columns of the second weight matrix is the dimension of the input vector, and the second offset vector is the number of fusion layers after adding the fusion layer; Multiply the second weight matrix by the transpose of the input vector and add the second offset vector to obtain a second sum vector; Perform a first exponential normalization on the elements of the second sum vector to obtain the layer weight vector, and the layer weight vector indicates the layer weights of each fusion layer.

8. The push processing method according to claim 6, wherein, The fusion layer includes a multi-head attention model and a feed-forward network, and the layer output vector is generated through the following process: Input the layer input vector corresponding to the input vector into the attention model to obtain an attention output vector corresponding to the input vector; Input the attention output vector corresponding to the input vector into the feed-forward network to obtain the layer output vector corresponding to the input vector.

9. The push processing method according to claim 8, wherein, Inputting the attention output vector corresponding to the input vector into the feed-forward network to obtain the layer output vector corresponding to the input vector includes: Adding the layer input vector corresponding to the input vector and the attention output vector corresponding to the input vector to obtain a third sum vector; Performing layer normalization on the third sum vector to obtain a first normalized vector; Inputting the first normalized vector into the feed-forward network to obtain the layer output vector corresponding to the input vector.

10. The push processing method according to claim 9, wherein, Inputting the first normalized vector into the feed-forward network to obtain the layer output vector corresponding to the input vector includes: Inputting the first normalized vector into the feed-forward network to obtain a feed-forward vector corresponding to the input vector; Adding the first normalized vector and the feed-forward vector corresponding to the input vector to obtain a fourth sum vector; Performing layer normalization on the fourth sum vector to obtain the layer output vector corresponding to the input vector.

11. The push processing method according to claim 10, wherein, Inputting the first normalized vector into the feed-forward network to obtain a feed-forward vector corresponding to the input vector includes: Obtaining a third weight matrix and a third offset vector, wherein the number of rows of the third weight matrix is the second dimension, the number of columns of the third weight matrix is the dimension of the input vector, and the dimension of the third offset vector is the second dimension; Multiplying the third weight matrix by the transpose of the first normalized vector and adding the third offset vector to obtain a fifth sum vector; Applying an activation function to the fifth sum vector and determining a feed-forward vector corresponding to the input vector based on the result of the activation function.

12. The push processing method according to claim 11, wherein, Determining a feed-forward vector corresponding to the input vector based on the result of the activation function includes: Obtaining a fourth weight matrix and a fourth offset vector, wherein the number of rows of the fourth weight matrix is the dimension of the input vector, the number of columns of the third weight matrix is the second dimension, and the dimension of the fourth offset vector is equal to the dimension of the input vector; Multiplying the fourth weight matrix by the transpose of the result of the activation function and adding the fourth offset vector to obtain the feed-forward vector corresponding to the input vector.

13. The push processing method according to claim 8, wherein, Inputting the layer input vector corresponding to the input vector into the attention model to obtain the attention output vector corresponding to the input vector includes: Performing layer normalization on the layer input vector corresponding to the input vector to obtain a layer-normalized layer input vector; Inputting the layer-normalized layer input vector into the attention model to obtain the attention output vector corresponding to the input vector.

14. The push processing method according to claim 13, wherein, Performing layer normalization on the layer input vector corresponding to the input vector to obtain a layer-normalized layer input vector includes: Determining the mean and variance of each element in the layer input vector corresponding to the input vector; Determining the first sum of the variance and the smoothing amount; Subtracting the mean from each element of the layer input vector corresponding to the input vector and dividing by the first sum to obtain the layer-normalized layer input vector.

15. The push processing method according to claim 14, wherein, subtracting the mean from each element of the layer input vector corresponding to the input vector and dividing by the first sum to obtain the layer-normalized layer input vector includes: Subtracting the mean from each element of the layer input vector corresponding to the input vector and dividing by the first sum to obtain a vector to be linearly processed; Multiplying each element of the vector to be linearly processed by a first parameter and adding a second parameter to obtain the layer-normalized layer input vector.

16. The push processing method according to claim 13, wherein, the attention model includes a plurality of attention heads; inputting the layer-normalized layer input vector into the attention model to obtain the attention output vector corresponding to the input vector includes: Inputting the layer-normalized layer input vector into the attention head to obtain the attention scores output by the attention head, wherein a plurality of attention scores output by the plurality of attention heads form an attention score vector; Multiplying the attention score vector by a fifth weight matrix to obtain the attention output vector corresponding to the input vector, wherein the number of rows of the fifth weight matrix is the number of attention heads, and the number of columns of the fifth weight matrix is the dimension of the input vector.

17. A push processing device, wherein, it includes: A first acquisition unit for acquiring a first number of target push basic features, wherein the first number of target push basic features include target object features and content features to be pushed; A second acquisition unit for acquiring a plurality of target push basic feature combinations based on the first number of target push basic features, each target push basic feature combination including a second number of target push basic features selected from the first number of target push basic features, and the second number is less than the first number; A first generation unit for generating a second-order cross feature vector corresponding to the target push basic feature combination; A second generation unit for inputting the first number of target push basic features into a feature full-cross model to obtain a full-cross feature vector; A prediction unit for inputting the first number of target push basic feature vectors corresponding to the first number of target push basic features, the plurality of second-order cross feature vectors corresponding to the plurality of target push basic feature combinations, and the full-cross feature vector into a fusion model to obtain a first probability of pushing the content to be pushed for the target object, and pushing the content to be pushed for the target object based on the first probability.

18. An electronic device, comprising a memory and a processor, the memory storing a computer program, characterized in that, when the processor executes the computer program, the push processing method according to any one of claims 1 to 16 is implemented.

19. A computer-readable storage medium, the storage medium storing a computer program, characterized in that, when the computer program is executed by a processor, the push processing method according to any one of claims 1 to 16 is implemented.

20. A computer program product, the computer program product comprising a computer program, the computer program being read and executed by a processor of a computer device, so that the computer device executes the push processing method according to any one of claims 1 to 16.