Push processing method, related device and medium
By obtaining the feature vectors of the target object and the content to be pushed, and using the combined cross network and probability prediction model, the resource consumption and accuracy problems caused by invalid cross features in push processing are solved, achieving more efficient content push.
Patent Information
- Application Number
- CN202410010124.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-03
- Publication Date
- 2025-07-04
AI Technical Summary
The prior art has problems such as excessive invalid crossover features in push processing, resulting in increased resource consumption and reduced push accuracy.
By obtaining the feature vectors of the target object and the content to be pushed, using the combined cross network and probability prediction model, the combination of effective cross features and push probability calculation are performed, reducing the processing of invalid cross features and improving push accuracy.
Improve the accuracy of push processing, reduce resource consumption, and achieve more efficient content push.
Smart Images

Figure CN120256848A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of big data, and particularly to a push processing method, related apparatus, and medium. Background Art
[0002] In current push processing, in order to improve push accuracy, not only are single push basic features (such as object features or content features) used for pushing content to an object, but also multi-layer cross networks are used to perform push processing on cross features obtained from push basic features (such as features obtained by crossing every two push basic features). The cross features of object features and content features are more effective for push processing. In related technologies, a large number of invalid cross features that cross between object features or cross between content features are generated during multi-order feature crossing. The invalid cross features are not very helpful for push accuracy, but increase the resource consumption of the push processing model, thus affecting the push processing efficiency. In order to remove the invalid cross features, related technologies propose to rely on manual mining of effective cross features. However, this will very likely result in the omission of effective cross features, leading to insufficient learning of the push processing model and low push accuracy. Summary of the Invention
[0003] Embodiments of the present disclosure provide a push processing method, related apparatus, and medium, which can improve the accuracy of push processing and reduce processing resource consumption.
[0004] According to one aspect of the present disclosure, there is provided a push processing method, including:
[0005] Obtain a plurality of push basic features, where the plurality of push basic features include at least one target object feature and at least one content feature to be pushed;
[0006] Based on at least one of the target object features, determine a first feature vector corresponding to each target object feature, and based on at least one of the content features to be pushed, determine a second feature vector corresponding to each content feature to be pushed;
[0007] For each of the first feature vectors, combine the first feature vector with each of the second feature vectors respectively, and input each combination of the first feature vector and the second feature vector into a combined cross network to obtain a combined cross feature vector;
[0008] Input the combined cross feature vector into a probability prediction model to obtain a first probability of pushing the content to be pushed to the target object, and push the content to be pushed to the target object based on the first probability.
[0009] According to one aspect of the present disclosure, there is provided a push processing apparatus, including:
[0010] An acquisition unit, configured to acquire a plurality of push basic features, where the plurality of push basic features include at least one target object feature and at least one content-to-be-pushed feature;
[0011] A determination unit, configured to determine a first feature vector corresponding to each target object feature based on at least one of the target object features, and determine a second feature vector corresponding to each content-to-be-pushed feature based on at least one of the content-to-be-pushed features;
[0012] A combination unit, configured to, for each of the first feature vectors, combine the first feature vector with each of the second feature vectors respectively, and input each combination of the first feature vector and the second feature vectors into a combined cross network to obtain a combined cross feature vector;
[0013] A push unit, configured to input the combined cross feature vector into a probability prediction model to obtain a first probability of pushing the content-to-be-pushed for the target object, and push the content-to-be-pushed for the target object based on the first probability.
[0014] Optionally, the determination unit is specifically configured to:
[0015] Vectorize at least one of the target object features into at least one target object feature vector;
[0016] Cascade at least one of the target object feature vectors and input them into a first cross network to obtain a first intermediate vector;
[0017] Obtain the first feature vector corresponding to each target object feature from the first intermediate vector.
[0018] Optionally, the first cross network includes a plurality of first layers;
[0019] The determination unit is specifically configured to:
[0020] Cascade at least one of the target object feature vectors to obtain a first cascaded vector;
[0021] Input the first cascaded vector into the first of the first layers to obtain an output vector of the first of the first layers;
[0022] Through the other first layers except the first of the first layers among the plurality of first layers, obtain the output vectors of the other first layers based on the output vector of the previous first layer and the first cascaded vector, where the output vector of the last of the first layers is used as the first intermediate vector.
[0023] Optionally, the determination unit is specifically configured to:
[0024] Obtain the first weight matrix and the first bias vector of other first layers among the multiple first layers except the first first layer, where the number of rows and columns of the first weight matrix is the dimension of the first concatenated vector, and the dimension of the first bias vector is the dimension of the first concatenated vector;
[0025] Based on the output vector of the previous first layer, the first concatenated vector, the first weight matrix, and the first bias vector, obtain the output vectors of the other first layers.
[0026] Optionally, the determining unit is specifically configured to:
[0027] Multiply the output vector of the previous first layer by the first weight matrix and add the first bias vector to obtain a second intermediate vector;
[0028] Multiply the first concatenated vector and the second intermediate vector bit by bit to obtain a third intermediate vector;
[0029] Add the third intermediate vector to the output vector of the previous first layer to obtain the output vectors of the other first layers.
[0030] Optionally, the first concatenated vector is obtained by concatenating a first number of target object feature vectors with the same dimension;
[0031] The determining unit is specifically configured to:
[0032] Input the first concatenated vector into the first first layer, and generate an output feature vector for each target object feature vector;
[0033] Concatenate the first number of output feature vectors according to the sorting of the target object feature vectors in the first concatenated vector to obtain the output vector of the first first layer;
[0034] For each target object feature vector in the first concatenated vector, based on the target object feature vector, the output vector of the previous first layer, the first weight matrix, and the first bias vector, obtain the output feature vector of the target object feature vector in the other first layers;
[0035] Concatenate the first number of output feature vectors according to the sorting of the target object features in the first vector to obtain the output vectors of the other first layers.
[0036] Optionally, the first weight matrix is composed of a second number of first sub - weight matrices, where the second number is equal to the square of the first number, and the number of rows and columns of each first sub - weight matrix is the dimension of the target object feature vector; the first bias vector is composed of the first number of first sub - bias vectors, and the dimension of the first sub - bias vector is equal to the dimension of the target object feature vector;
[0037] The determining unit is specifically configured to:
[0038] For each target object feature vector in the first concatenated vector, obtain the first number of target first sub - weight matrices from the second number of first sub - weight matrices, and obtain the target first sub - bias vector from the first number of first sub - bias vectors;
[0039] Based on the target object feature vector, the first number of target first sub - weight matrices, the target first sub - bias vector, and the output vector of the previous first layer, obtain the output feature vector of the target object feature vector in other first layers.
[0040] Optionally, the determining unit is specifically configured to:
[0041] Multiply each of the first number of output feature vectors in the output vector of the previous first layer by the first number of target sub - weight matrices one by one to obtain the first number of fourth intermediate vectors;
[0042] Multiply each of the first number of fourth intermediate vectors by the target object feature vector bit by bit and then sum to obtain a first sum vector;
[0043] Multiply the target object feature vector by the target first sub - bias vector bit by bit to obtain a fifth intermediate vector;
[0044] Add the first sum vector, the fifth intermediate vector, and the output feature vector corresponding to the target object feature vector in the output vector of the previous first layer to obtain the output feature vector of the target object feature vector in other first layers.
[0045] Optionally, the determining unit is specifically configured to:
[0046] Obtain the first low - rank matrix and the second low - rank matrix of other first layers except the first first layer among multiple first layers, where the number of rows of the first low - rank matrix and the second low - rank matrix is equal to the dimension of the first concatenated vector, and the number of columns is less than the first dimension of the first concatenated vector;
[0047] For each of the other said first layers, multiply the first low-rank matrix by the transpose of the second low-rank matrix to obtain the first weight matrix.
[0048] Optionally, the push processing device further includes:
[0049] An input unit, configured to input a plurality of the push base features into a full-cross network to obtain a full-cross feature vector;
[0050] The push unit is specifically configured to:
[0051] Input the combined cross feature vector and the full-cross feature vector into a probability prediction model to obtain a first probability of pushing the content to be pushed for the target object, and push the content to be pushed for the target object based on the first probability.
[0052] Optionally, the push unit is specifically configured to:
[0053] Add the combined cross feature vector and the full-cross feature vector to obtain a final cross feature vector;
[0054] Based on the final cross feature vector, obtain a first probability of pushing the content to be pushed for the target object.
[0055] Optionally, the push unit is specifically configured to:
[0056] Obtain a first weight vector and a first bias amount, where the dimension of the first weight vector is equal to the dimension of the final cross feature vector;
[0057] Multiply the first weight vector by the transpose of the final cross feature vector and add the first bias amount to obtain a first score;
[0058] Perform exponential normalization on the first score to obtain a first probability of pushing the content to be pushed for the target object.
[0059] Optionally, the determining unit is specifically configured to:
[0060] If the target object feature is a fixed-length sequence feature, obtain a feature mapping reference table corresponding to the target object feature, and based on the feature mapping reference table, map the target object feature to the corresponding target object feature vector;
[0061] If the target object feature is a variable-length sequence feature, perform character mapping on each character in the target object feature to obtain a variable-length feature vector, and perform pooling processing on the variable-length feature vector to obtain a target object feature vector corresponding to the target object feature and having a second dimension.
[0062] Optionally, the determining unit is specifically configured to:
[0063] Vectorize at least one of the content features to be pushed into at least one content feature vector to be pushed;
[0064] Concatenate at least one of the content feature vectors to be pushed and input them into a second cross network to obtain a sixth intermediate vector;
[0065] Obtain the second feature vector corresponding to each of the content features to be pushed from the sixth intermediate vector.
[0066] Optionally, the first feature vector and the second feature vector have the same dimension;
[0067] The combining unit is specifically configured to:
[0068] Multiply the first feature vector and the second feature vector bit by bit to obtain a combined cross feature vector.
[0069] Optionally, the probability prediction model includes multiple first task prediction models corresponding to multiple tasks, and the first task prediction model is used to predict the probability that the target object completes the task for the content to be pushed;
[0070] The pushing unit is specifically configured to:
[0071] Input the combined cross feature vector into multiple first task prediction models to obtain multiple first task completion probabilities corresponding to the multiple tasks;
[0072] Based on the multiple first task completion probabilities, obtain a first probability of pushing the content to be pushed for the target object.
[0073] Optionally, the probability prediction model includes gating nodes corresponding to multiple tasks and a second task prediction model;
[0074] The input unit is specifically configured to:
[0075] Input multiple pushing basic features into multiple full cross networks to obtain multiple full cross feature vectors;
[0076] The pushing unit is specifically configured to:
[0077] Input the combined cross feature vector and multiple full cross feature vectors into the gating nodes corresponding to the tasks to obtain seventh intermediate vectors corresponding to the tasks;
[0078] Input the seventh intermediate vector into the second task prediction model corresponding to the task to obtain a second task completion probability corresponding to the task;
[0079] Based on the second task completion probabilities corresponding to the multiple tasks, a first probability of pushing the content to be pushed to the target object is obtained.
[0080] According to one aspect of the present disclosure, an electronic device is provided, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above-mentioned push processing method is implemented.
[0081] According to one aspect of the present disclosure, a computer-readable storage medium is provided. The storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned push processing method is implemented.
[0082] According to one aspect of the present disclosure, a computer program product is provided. The computer program product includes a computer program, and the computer program is read and executed by a processor of a computer device, so that the computer device executes the above-mentioned push processing method.
[0083] In the embodiments of the present disclosure, based on at least one target object feature in the push basic features, a first feature vector corresponding to each target object feature is determined. In this way, the first feature vector corresponding to each determined target object feature is not only related to the target object feature, but also reflects the influence of other target object features. Similarly, the second feature vector corresponding to each content feature to be pushed determined in the embodiments of the present disclosure is not only related to the content feature to be pushed, but also reflects the influence of other content features to be pushed. Then, each first feature vector and each second feature vector are combined pairwise, and all the obtained combinations are input into a combined cross network to obtain combined cross feature vectors. Since the first feature vector corresponds to the target object feature and the second feature vector corresponds to the content feature to be pushed, all the pairwise combinations obtained reflect all the pairwise combinations of each target object feature and each content feature to be pushed. The cross feature of the object feature and the content feature is more effective for push processing than the cross feature between object features or the cross feature between content features, and the mutual influence between target object features and the mutual influence between content features to be pushed are also reflected in the first feature vector and the second feature vector. Therefore, using the obtained combined cross feature vectors in this way can improve the accuracy of push processing more. And the processing of the combined cross network does not involve the processing between target object features and the processing between the content to be pushed, greatly reducing the consumption of processing resources.
[0084] Other features and advantages of the present disclosure will be described in the following specification, and part of them will become obvious from the specification, or be understood by implementing the present disclosure. The objectives and other advantages of the present disclosure can be realized and obtained through the structures specifically pointed out in the specification, the claims and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0085] The accompanying drawings are used to provide a further understanding of the technical solutions of the present disclosure, and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the technical solutions of the present disclosure, and do not constitute a limitation to the technical solutions of the present disclosure.
[0086] Figure 1 It is a schematic diagram of the architecture of the system to which the push processing method according to the embodiment of the present disclosure is applied;
[0087] Figures 2A - 2B It is a schematic diagram of the interface when the embodiment of the present disclosure is applied to the video content push scenario;
[0088] Figure 3 It is a flowchart of the push processing method according to an embodiment of the present disclosure;
[0089] Figure 4 It is a schematic diagram of obtaining a first feature vector corresponding to each target object feature and a second feature vector corresponding to each content to be pushed according to an embodiment of the present disclosure;
[0090] Figure 5 It is Figure 3 A flowchart of step 320 in
[0091] Figure 6 It is a schematic diagram of obtaining a first feature vector corresponding to each target object feature based on at least one target object feature according to an embodiment of the present disclosure;
[0092] Figure 7 It is a flowchart of vectorizing the target object feature according to an embodiment of the present disclosure;
[0093] Figure 8 It is an example diagram of the feature mapping reference table according to an embodiment of the present disclosure;
[0094] Figure 9 It is a schematic diagram of obtaining a variable-length feature vector after character mapping for each character in the variable-length feature vector according to an embodiment of the present disclosure;
[0095] Figure 10A It is a process of pooling the variable-length feature vector using the average pooling method according to an embodiment of the present disclosure;
[0096] Figure 10B It is a process of pooling the variable-length feature vector using the maximum pooling method according to an embodiment of the present disclosure;
[0097] Figure 11 It is a schematic diagram of mapping the target object feature to the target object feature vector according to an embodiment of the present disclosure;
[0098] Figure 12 It is a schematic diagram of cascading target object feature vectors into a first cascaded vector according to an embodiment of the present disclosure;
[0099] Figure 13 It is Figure 5 a flowchart of step 520 in
[0100] Figure 14 It is a schematic diagram of obtaining a first intermediate vector from the first cascaded vector through multiple first layers in the first cross network according to an embodiment of the present disclosure;
[0101] Figure 15 It is Figure 13 a flowchart of step 1330 in
[0102] Figure 16 It is Figure 15 a flowchart of step 1510 in
[0103] Figure 17 It is Figure 15 a flowchart of step 1520 in
[0104] Figure 18 It is Figure 13 a flowchart of step 1320 in
[0105] Figure 19 It is Figure 15 a flowchart of step 1520 in
[0106] Figure 20 It is a schematic diagram of each first layer obtaining an output feature vector for each target object feature vector and cascading the output feature vectors to obtain an output vector of the first layer according to an embodiment of the present disclosure;
[0107] Figure 21 It is a schematic diagram of dividing a first weight matrix into multiple first sub - weight matrices according to an embodiment of the present disclosure;
[0108] Figure 22 It is a schematic diagram of dividing a first bias vector into multiple first sub - bias vectors according to an embodiment of the present disclosure;
[0109] Figure 23 It is Figure 19 a flowchart of step 1910 in
[0110] Figure 24 It is a schematic diagram of obtaining a target first sub - weight matrix from the first weight matrix and obtaining a target first sub - bias vector from the first bias vector according to an embodiment of the present disclosure;
[0111] Figure 25 is Figure 23 a flowchart of step 2320 in
[0112] Figure 26 a schematic diagram of multiplying the output vector of the upper layer by the target first sub - weight matrix one by one according to an embodiment of the present disclosure;
[0113] Figure 27 is Figure 3 another flowchart of step 320 in
[0114] Figure 28 a schematic diagram of pairwise combining each first feature vector with each second feature vector according to an embodiment of the present disclosure;
[0115] Figure 29 a schematic diagram of the overall process of obtaining the first probability based on the target object feature and the feature of the content to be pushed according to an embodiment of the present disclosure;
[0116] Figure 30 is Figure 3 a flowchart of adding a full cross - network to the push processing process shown in
[0117] Figure 31 is Figure 30 a flowchart of step 345 in
[0118] Figure 32 another schematic diagram of the overall process of obtaining the first probability based on the target object feature and the feature of the content to be pushed according to an embodiment of the present disclosure;
[0119] Figure 33 is Figure 31 a flowchart of step 3120 in
[0120] Figure 34 is Figure 3 a flowchart of step 340 in
[0121] Figure 35 a schematic diagram of inputting the combined cross - feature vector into the first task prediction models corresponding to multiple tasks to obtain the first probability according to an embodiment of the present disclosure;
[0122] Figure 36 is Figure 30 another flowchart of step 345 in
[0123] Figure 37 a schematic diagram of determining the second task completion probabilities corresponding to multiple tasks based on the combined cross - feature vector and the full cross - feature vector and then obtaining the first probability according to an embodiment of the present disclosure;
[0124] Figure 38 It is a flowchart for training a probability prediction model;
[0125] Figure 39 It is an implementation detail diagram of the push processing method of the embodiment of the present disclosure.
[0126] Figure 40 is a module diagram of a push processing device according to an embodiment of the present disclosure;
[0127] Figure 41 According to one embodiment of the present disclosure, Figure 3 The terminal structure diagram of the push processing method shown;
[0128] Figure 42 According to one embodiment of the present disclosure, Figure 3 The server structure diagram of the push processing method shown. DETAILED DESCRIPTION
[0129] In order to make the purpose, technical solution and advantages of the present disclosure more clear, the present disclosure is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present disclosure and are not used to limit the present disclosure.
[0130] Before further describing the embodiments of the present disclosure in detail, the nouns and terms involved in the embodiments of the present disclosure are described. The nouns and terms involved in the embodiments of the present disclosure are subject to the following interpretations:
[0131] Deep Neural Networks (DNN): is a multi-layer unsupervised neural network. The deep neural network uses the output features of the previous layer as the input of the next layer for feature learning. After layer-by-layer feature mapping, the features of the existing spatial samples are mapped to another feature space, so as to learn to have better feature expression for the existing input. The deep neural network has multiple nonlinear mapping feature transformations and can fit highly complex functions.
[0132] Push: refers to the method by which a system or platform proactively sends information, content or services to users without explicit requests or searches by users. Push processing methods are usually targeted based on personal information such as user interests, preferences, historical behavior, etc. to provide personalized content push.
[0133] Non-linear activation function (sigmoid function): Also known as the logistic function or S-shaped function. The sigmoid function maps the input value to an output value between 0 and 1, that is, the range of the output value is between 0 and 1. When the input value approaches negative infinity, the output value approaches 0; when the input value approaches positive infinity, the output value approaches 1.
[0134] System architecture and scenario description applied in the embodiments of the present disclosure
[0135] Figure 1 It is a system architecture diagram applied to the push processing method according to the embodiments of the present disclosure. It includes: object terminal 140, Internet 130, gateway 120, and server 110.
[0136] Server 110 refers to a computer system that can provide content push services to object terminal 140. Compared with object terminal 140, server 110 has higher requirements in terms of stability, security, performance, etc. Server 110 can be a high-performance computer in a network platform, a cluster of multiple high-performance computers, a part (such as a virtual machine) allocated from a high-performance computer, a combination of parts (such as virtual machines) allocated from multiple high-performance computers, etc. Server 110 can also communicate with Internet 130 by wired or wireless means to exchange data.
[0137] Gateway 120 is also called an inter-network connector and protocol converter. Gateway 120 realizes network interconnection at the transport layer and is a computer system or device that acts as a converter. Between two systems using different communication protocols, data formats, or languages, and even with completely different architectures, gateway 120 is a translator. At the same time, gateway 120 can also provide filtering and security functions. The message sent by object terminal 140 to server 110 needs to be sent to the corresponding server 110 through gateway 120. The message sent by server 110 to object terminal 140 also needs to be sent to the corresponding object terminal 140 through gateway 120.
[0138] Object terminal 140 is a device used by the object to view the pushed content. It includes various forms such as desktop computers, laptops, PDAs (Personal Digital Assistants), mobile phones, in-vehicle terminals, home theater terminals, dedicated terminals, etc. In addition, it can be a single device or a set composed of multiple devices. For example, multiple devices are connected through a local area network and share a display device for collaborative work, jointly constituting a terminal. Object terminal 140 can also communicate with Internet 130 by wired or wireless means to exchange data.
[0139] The embodiments of the present disclosure can be applied in various scenarios, such as Figures 2A - 2B The scenario of viewing the pushed video content in an instant short message application as shown, etc.
[0140] As shown Figure 2A in the interface of the instant messaging application in the object terminal 140, the instant messaging application can be used to conduct instant messaging with multiple contacts and can also be used to watch videos. After triggering the "Message" option in the function bar below, a short message list with multiple contacts is displayed. The short message list contains message bars for multiple contacts. After triggering the "Video" option in the function bar below, the interface of the instant messaging application is as Figure 2B shown
[0141] In Figure 2B , the instant messaging application opens the video interface and plays the pushed video content. The interface also shows the number of likes (53,600), comments (8,512), and favorites (3,621) obtained by the currently pushed video content before 13:00. The object can also perform operations such as liking, commenting, or favoriting on the video content it is interested in on the interface.
[0142] General description of embodiments of the present disclosure
[0143] According to an embodiment of the present disclosure, a push processing method is provided.
[0144] The push processing method is a process of pushing content of interest to the target object to the terminal 140 of the target object. The target object is the object that wants to view the content to be pushed. Embodiments of the present disclosure can be applied to the push of different types of content, and the content to be pushed can be various types such as pictures, videos, articles, applications, etc.
[0145] During the push processing, the interest correlation degree between the target object and the content to be pushed can be predicted by using the object features of the target object and the content features of the content to be pushed. The higher the interest correlation degree between the content pushed to the target object and the target object, the higher the accuracy of the push processing. In order to improve the push accuracy, in the current push processing, not only single object features or single content features are used for prediction, but also multi-order cross features between different features are used for prediction. The multi-order cross features can reflect the mutual influence between different features. Therefore, in order to predict the interest correlation degree between the target object and the content to be pushed, the multi-order cross features between the object features and the content features can bring more valuable information to the prediction accuracy. Such multi-order cross features between the object features and the content features are called effective cross features. However, currently, when performing multi-order feature crossing, a large number of cross features between object features or cross features between content features will be generated, and these cross features are called invalid cross features. This is because these cross features can only be used to reflect the mutual influence between object features or the mutual influence between content features, and are of little help for predicting the interest correlation degree between the target object and the content to be pushed. A large number of invalid cross features will also increase the resource consumption during the push processing. In order to remove the influence of invalid cross features on the push processing, the related art currently also proposes to rely on manual mining of effective cross features. However, since the frequency of occurrence of some feature combinations between object features and content features in the push event samples is relatively sparse, these feature combinations are easily omitted during the manual mining process, resulting in insufficient learning of the push processing model and low push accuracy.
[0146] In the embodiment of the present disclosure, when performing push processing, the accuracy of the push processing can be improved by comprehensively constructing effective cross features. At the same time, the invalid cross features are reduced during the push processing, greatly reducing the resource consumption during the push processing.
[0147] The push processing method of the embodiment of the present disclosure can be executed by the terminal 140 or by the server 110. When executed by the server 110, after the execution is completed, the content to be pushed determined to be pushed to the target object will be transmitted to the object terminal 140 through the Internet 130, and the terminal 140 will display the content to be pushed to the target object.
[0148] As Figure 3 shown, according to an embodiment of the present disclosure, the push processing method includes:
[0149] Step 310, obtain a plurality of push basic features;
[0150] Step 320: Based on at least one target object feature, determine the first feature vector corresponding to each target object feature, and based on at least one content feature to be pushed, determine the second feature vector corresponding to each content feature to be pushed;
[0151] Step 330: For each first feature vector, combine the first feature vector with each of the second feature vectors respectively, and input each combination of the first feature vector and the second feature vector into the combined cross network to obtain a combined cross feature vector;
[0152] Step 340: Input the combined cross feature vector into a probability prediction model to obtain a first probability of pushing the content to be pushed to the target object, and push the content to be pushed to the target object based on the first probability.
[0153] The above steps 310 - 340 are described in detail below.
[0154] Detailed description of step 310
[0155] In step 310, obtain multiple basic push features, where the multiple basic push features include at least one target object feature and at least one content feature to be pushed.
[0156] The target object feature is a feature related to the target object. The target object feature can be a feature that the target object itself has. For example, the educational level feature, industry feature, hobby feature, etc. of the target object.
[0157] Generally, target objects with a certain type of common target object feature may show certain commonalities in their preferences for content. For example, for educational level, objects with a higher educational level tend to be more interested in content with a higher professional depth and stronger professionalism. Therefore, the pushed content may be more inclined to be in-depth and professional content within the object's professional field. For industry features, objects within the same industry tend to watch content related to that industry. For example, objects in the education industry may be more concerned about relevant educational popular science content; objects in the medical industry may be more concerned about health care-related content. For hobby features, objects with the same hobby may tend to like the same content. For example, objects who like music tend to watch music-related content; for objects who like traveling, it is more likely to recommend travel guide content for them.
[0158] Thus, using the features that the target object itself has for predicting the content to be pushed helps to predict the content that the target object is interested in by analyzing the characteristics of the target object itself, improving the accuracy of the push processing.
[0159] Note that when obtaining the features inherent in the target object as the target object features, the consent of the target object should be obtained in advance. Moreover, the collection, use, and processing of these object features will comply with relevant laws, regulations, and standards. When obtaining the consent of the target object, separate permission or separate consent of the target object can be obtained through methods such as pop-up windows or redirecting to a confirmation page.
[0160] In one embodiment, obtaining the features inherent in the target object as the target object features can be achieved through obtaining registration information. Here, the registration information can be the registration information of the target object on the object terminal 140, or the registration information of the target object in the content push application, etc. When the target object first uses the object terminal 140, the object terminal 140 may require the target object to register, and relevant object information needs to be filled in as the registration information during registration. In this way, the target object features can be obtained from the object terminal 140 used by the target object. When the target object first uses the content push application, registration is also required. The content push application may require the target object to fill in relevant object information as the registration information. Therefore, the target object features can also be obtained by using the registration information of the target object in the content push application. The advantage of obtaining the target object features through registration information is: convenient and fast, with high acquisition efficiency. When obtaining the registration information of the target object, as mentioned above, the consent of the target object should also be obtained in advance, which will not be elaborated here.
[0161] In another embodiment, obtaining the features inherent in the target object as the target object features can be based on the object actions of the target object in the content push application. Object actions refer to dynamic events in the content push application that are dominated by the target object. For example, the target object interacts with other objects, or the target object searches for content in the content push application, etc. By counting the interaction events between the target object and other objects, the interaction frequency can be used as a target object feature; by querying the search records of the target object in the content push application, the object search records can also be used as a target object feature. Object actions can directly or indirectly reflect the needs of the target object when using the content push application. For example, a target object with a high interaction frequency has a high social need when using the content push application and may be more interested in content that can provide interaction and communication. Another example is that if the search records of the target object contain a large amount of content related to medical and health, then such content can be pushed to the target object. Thus, obtaining the target object features based on the object actions of the target object in the content push application can more comprehensively understand the object's needs and improve the accuracy of pushing the content to be pushed to the target object. When obtaining the object actions of the target object, as mentioned above, the consent of the target object should also be obtained in advance, which will not be elaborated here.
[0162] The target object characteristics may also include the scene characteristics of the scene where the target object is viewing the content to be pushed. For the same object in different scenarios, the content it wants to view may also be different. For example, when the object is in a work scenario, it may want to view content related to work; when the object is at home, it may want to view content related to its own preferences. The scene characteristics include: when determining the content to be pushed to the target object, the geographical area characteristics, time period characteristics, input day type characteristics (working day, weekend, festival), etc. where the target object is located.
[0163] When obtaining the scene characteristics of the target object, the consent of the target object must be obtained in advance. Similar to the above, it will not be elaborated here.
[0164] The geographical area characteristics refer to the characteristics of the geographical area where the target object is located. This geographical area can be an administrative geographical area, such as City A, City B, etc.; it can also be a geographical area divided by entities on the map, such as the XX School area, the XX Shopping Mall area, etc. The administrative geographical area has an impact on the content that the object wants to view. For example, when the object is in City A, it may want to view content related to tourist attractions or food culture in City A. The geographical area divided by entities also has an impact on the content that the object wants to view. For example, when the object is in the XX Shopping Mall area, it may want to view the activity information or special projects in the mall.
[0165] The geographical area characteristics where the target object is located can be obtained from the positioning information of the target object's object terminal 140. The object terminal 140 of the target object has positioning devices such as GPS and Beidou, and this positioning device can continuously obtain the positioning information where the object terminal 140 is located. The geographical area where the target object is located when viewing the pushed content can be determined according to this positioning information as the geographical area characteristics. For example, according to the GPS positioning information obtained by the positioning device and mapping it to an electronic map, it is found that it belongs to City A, then the geographical area characteristics when viewing the pushed content are obtained as "City A".
[0166] The time period characteristics refer to the time period within a day when the target object is viewing the content to be pushed. For example, 9:00 - 11:00, 11:00 - 13:00, 13:00 - 18:00, 18:00 - 22:00, etc. Since the activities of the object are different in different time periods of a day and the purpose of viewing content may also be different, the pushed content will be affected by the time period characteristics. For example, from 9:00 to 11:00, the object may be at work, and during this time period, the object is very likely to want to view content related to work; from 18:00 to 22:00, the object is usually resting, and during this time period, the object is very likely to want to view content related to its own interests.
[0167] The input date type feature refers to the feature of the type of day when the target object views the content to be pushed. It is divided into weekdays, weekends, festivals, etc. The types of content that the object wants to view may be different on different types of days. For example, on weekdays (Monday to Friday), the object is often at work and may want to view work-related content; on weekends (Saturday and Sunday), when the object does not need to work, it may want to view content related to going out for fun; on festivals, the object may want to view content related to festival culture.
[0168] The time period feature and the input date type feature can be obtained from the system time of the object terminal 140 of the target object. The system time generally includes information such as year, month, day, day of the week, hour, minute, and second. The input time period feature can be determined from the information of hour, minute, and second. For example, if the current time is 10:22:54, it is determined that it belongs to 9:00 - 11:00, that is, the input time period feature is 9:00 - 11:00. If the date is Sunday, December 3, 2023, it is determined that the current input date type feature is the weekend.
[0169] Obtaining the scene feature of the target object's location as the target object feature takes into account the influence of the scene factor on the target object's preferences, which is beneficial to improving the accuracy of pushing the content to be pushed to the target object.
[0170] The content feature to be pushed is used to represent the characteristics of multiple aspects of the content to be pushed. For example, content type, content length, content release time, and content keywords, etc. For the same content feature to be pushed, if the feature values are different, the preferences of the target object for the content to be pushed may also be different. For example, different objects may have different preferences for the content length. Taking video content as an example, some objects may like to watch short video content of less than 3 minutes, while some objects may like to watch long video content of more than 10 minutes. Usually, the content to be pushed with the same type of content feature to be pushed is likely to be watched by the same target object. Therefore, using the content feature to be pushed for content pushing helps to improve the effect of pushing processing.
[0171] Note that when obtaining the content feature to be pushed, the consent of the content publisher of the content to be pushed should be obtained in advance. The specific method is the same as the method of obtaining the consent of the target object, which will not be elaborated here. Moreover, the collection, use, and processing of these content features to be pushed will comply with relevant laws, regulations, and standards.
[0172] Among the multiple push basic features, it contains at least one target object feature and at least one content feature to be pushed, ensuring that the push processing can be carried out based on the correlation degree between the target object and the content to be pushed.
[0173] Detailed description of step 320
[0174] In step 320, based on at least one target object feature, a first feature vector corresponding to each target object feature is determined, and based on at least one content feature to be pushed, a second feature vector corresponding to each content feature to be pushed is determined.
[0175] The first feature vector corresponding to each target object feature can reflect the mutual influence between this target object feature and other target object features. The second feature vector corresponding to each content feature to be pushed can reflect the mutual influence between this content feature to be pushed and other content features to be pushed. When determining the first feature vector and the second feature vector, the target object features and the content features to be pushed do not affect each other. As Figure 4 shown, the target object features include: "education", "loving music", and "City A". Based on these three target object features, the first feature vector A corresponding to the "education" feature, the first feature vector B corresponding to the "loving music" feature, and the first feature vector C corresponding to the "City A" feature can be obtained. The content features to be pushed include: "health preservation" and "8 minutes". Based on these two content features to be pushed, the second feature vector A corresponding to the "health preservation" feature and the second feature vector B corresponding to the "8 minutes" feature can be obtained.
[0176] In one embodiment, as Figure 5 shown, based on at least one target object feature, determining a first feature vector corresponding to each target object feature includes:
[0177] Step 510: Vectorize at least one target object feature into at least one target object feature vector;
[0178] Step 520: Concatenate at least one target object feature vector, input it into the first cross network, and obtain a first intermediate vector;
[0179] Step 530: Obtain the first feature vector corresponding to each target object feature from the first intermediate vector.
[0180] In this embodiment, the overall process of determining the first feature vector corresponding to each target object feature can be as Figure 6 shown. Vectorize each target object feature to obtain the target object feature vector corresponding to each target object feature; concatenate all target object feature vectors and input them into the first cross network to obtain a first intermediate vector. The first intermediate vector is composed of the first feature vectors corresponding to the target object features. Therefore, the first feature vector corresponding to each target object feature can be obtained from the first intermediate vector.
[0181] A vector is an array composed of values in different dimensions and is a point in a multi-dimensional space. The line segment between this point and the origin in the multi-dimensional coordinate system has a magnitude and a direction, which are the magnitude and direction of the vector. Each value in the vector is the point value projected by this point onto the corresponding coordinate axis in the multi-dimensional coordinate system, that is, a vector element. Vector elements can be numerical values or symbols, etc. A target object feature corresponds to a target object feature vector.
[0182] In one embodiment, the target object feature vector corresponding to the target object feature can be determined by looking up a table. In the target object feature vector mapping table, the mapping relationship between each target object feature and the corresponding target object feature vector is stored.
[0183] In one embodiment, as Figure 7 shown, vectorizing at least one target object feature into at least one target object feature vector includes:
[0184] Step 710: If the target object feature is a fixed-length sequence feature, obtain the feature mapping reference table corresponding to the target object feature, and based on the feature mapping reference table, map the target object feature to the corresponding target object feature vector;
[0185] Step 720: If the target object feature is a variable-length sequence feature, perform character mapping on each character in the target object feature to obtain a variable-length feature vector, and perform pooling processing on the variable-length feature vector to obtain the target object feature vector corresponding to the target object feature and having a second dimension.
[0186] In this embodiment, the target object features can be divided into fixed-length sequence features and variable-length sequence features.
[0187] A fixed-length sequence feature means that the feature value string content and length of the target object feature are determined, and a target object feature vector with a fixed dimension can also be obtained based on the target object feature. For example, the feature values of the education level feature can include: "primary school", "middle school", "junior college", "undergraduate", "master", and "doctor", etc. The feature values are determined, and the corresponding target object feature vectors can be directly mapped according to these feature values. The mapping relationship between the target object feature and the target object feature vector can be presented through a feature mapping reference table.
[0188] As Figure 8The feature mapping reference table of education level features is shown. Each education level feature value corresponds to a vector. If the target object feature is "primary school", then the corresponding target object feature vector is [2,3,4,5,3,0,0,9]; if the target object feature is "undergraduate", then the corresponding target object feature vector is [3,13,5,9,7,0,0,6]. The advantage of determining the target object feature vector corresponding to the fixed-length sequence feature based on the feature mapping reference table is that the search speed is fast and the vectorization efficiency is high.
[0189] An indefinite-length sequence feature refers to a target object feature whose feature value content is uncertain and whose string length is also uncertain. It is impossible to directly find the corresponding target object feature vector through feature value mapping. For example, the object profile feature in the target object feature. The object profile feature can be customized content of the object, not a preset fixed value. For example, an object profile feature is "love to travel". For indefinite-length sequence features, in this embodiment, each character can be mapped, and the mapping results of each character can be concatenated to obtain an indefinite-length feature vector. Figure 9 As shown, the four characters in “爱旅行” are mapped to vector representations respectively, “热” is mapped to vector [1,1,3], “爱” is mapped to vector [2,5,2], “旅” is mapped to vector [5,4,3], and “行” is mapped to vector [8,5,3]; these four vectors are concatenated to obtain the indefinite-length feature vector [1,1,3,2,5,2,5,4,3,8,5,3] corresponding to “爱旅行”.
[0190] Since the dimensions of the variable-length feature vectors corresponding to different variable-length sequence features may be different, the different dimensions of different target object feature vectors may cause the push processing process to shift to the target object features with longer dimensions. Therefore, the variable-length feature vectors are pooled into target object feature vectors with a second dimension. The second dimension is a preset dimension for variable-length sequence features. Pooling all variable-length sequence features into a unified second-dimensional vector is conducive to improving push accuracy.
[0191] Pooling refers to the dimensionality reduction and compression of input variable-length feature vectors. Pooling methods include average pooling and maximum pooling.
[0192] In one embodiment, Figure 10AAs shown, the variable-length feature vector is [1, 1, 3, 2, 5, 2, 5, 4, 3, 8, 5, 3], with a dimension of 12, and the preset second dimension is 8. To pool the 12-dimensional variable-length feature vector into an 8-dimensional target object feature vector, the size of the pooling kernel required is 5*1, and the stride is 1. Starting from the first position of the variable-length feature vector, using the pooling kernel, we get [1, 1, 3, 2, 5]. Using average pooling, we need to calculate the average value of this vector, resulting in (1 + 1 + 3 + 2 + 5) / 5 = 2.4. Moving the pooling kernel one position backward according to the stride, we get [1, 3, 2, 5, 2], and using average pooling gives (1 + 3 + 2 + 5 + 2) / 5 = 2.6. Similarly, based on the vector [3, 2, 5, 2, 5], the average value 3.4 can be obtained; based on the vector [2, 5, 2, 5, 4], the average value 3.4 can be obtained; based on the vector [5, 2, 5, 4, 3], the average value 3.8 can be obtained; based on the vector [2, 5, 4, 3, 8], the average value 4.4 can be obtained; based on the vector [5, 4, 3, 8, 5], the average value 5 can be obtained; based on the vector [4, 3, 8, 5, 3], the average value 4.6 can be obtained. Therefore, after pooling the variable-length feature vector [1, 1, 3, 2, 5, 2, 5, 4, 3, 8, 5, 3] using average pooling, the resulting target object feature vector is [2.4, 2.6, 3.4, 3.6, 3.8, 4.4, 5, 4.6].
[0193] In another embodiment, as Figure 10B shown, different from Figure 10A the average pooling used, Figure 10B max pooling is adopted, that is, finding the maximum value from the vectors obtained by the pooling kernel. Based on the vector [1, 1, 3, 2, 5], the maximum value 5 can be obtained; based on [1, 3, 2, 5, 2], the maximum value 5 can be obtained; based on the vector [3, 2, 5, 2, 5], the maximum value 5 can be obtained; based on the vector [2, 5, 2, 5, 4], the maximum value 5 can be obtained; based on the vector [5, 2, 5, 4, 3], the maximum value 5 can be obtained; based on the vector [2, 5, 4, 3, 8], the maximum value 8 can be obtained; based on the vector [5, 4, 3, 8, 5], the maximum value 8 can be obtained; based on the vector [4, 3, 8, 5, 3], the maximum value 8 can be obtained. Therefore, after pooling the variable-length feature vector [1, 1, 3, 2, 5, 2, 5, 4, 3, 8, 5, 3] using max pooling, the resulting target object feature vector is [5, 5, 5, 5, 5, 8, 8, 8].
[0194] In steps 710 - 720 of the above embodiments, the fixed-length sequence features are mapped to fixed-length target object feature vectors, and the variable-length sequence features are mapped to target object feature vectors in the second dimension through pooling processing. Different methods are used for vectorization processing of different types of target object features to control the dimensions of the target object feature vectors, improving the uniformity of vectorization processing of the target object features.
[0195] Vectorize at least one target object feature into at least one target object feature vector, and combine Figure 11 As shown, there are a total of three target object features, including: industry feature, hobby feature, and geographical area feature. After vectorizing these three target object features, the target object feature vector with the industry feature of "education" is [3, 1, 2, 5, 2, 0, 1, 0], the target object feature vector with the hobby feature of "music" is [8, 6, 2, 0, 1, 4, 9, 9], and the target object feature vector with the geographical area feature of "City A" is [5, 3, 1, 4, 0, 0, 2, 6].
[0196] In step 520, cascade at least one target object feature vector and input it into the first cross network to obtain a first intermediate vector.
[0197] Cascading at least one target object feature vector means splicing at least one target object feature vector to form a vector. Specifically, assume that N target object feature vectors include x1, x2,..., x n , then the cascaded vector X = [x1, x2,..., x n . For example Figure 12 As shown, the target object feature vectors include: [3, 1, 2, 5, 2, 0, 1, 0], [8, 6, 2, 0, 1, 4, 9, 9], and [5, 3, 1, 4, 0, 0, 2, 6]. After cascading the three target object feature vectors end to end, a vector is formed as [3, 1, 2, 5, 2, 0, 1, 0, 8, 6, 2, 0, 1, 4, 9, 9, 5, 3, 1, 4, 0, 0, 2, 6].
[0198] The dimension of the cascaded vector is the sum of the dimensions of at least one target object feature vector. If the dimensions of each target object feature vector are the same, then the dimension calculation of the cascaded vector X is shown in formula 1:
[0199] d x = N × d (formula 1).
[0200] In formula 1, d xis the dimension of the cascaded vector X, N is the number of target object feature vectors, and d is the dimension of the target object feature vectors. For example, if the number of target object feature vectors is 3 and the dimension is 8, then the dimension of the cascaded vector is 8 * 3 = 24.
[0201] Since the cascaded vector is composed of at least one target object feature vector, the first cross network can obtain the mutual influence between the target object feature vectors based on the cascaded vector to obtain a first intermediate vector.
[0202] In one embodiment, the first cross network includes multiple first layers. As Figure 13 shown, cascading at least one target object feature vector and inputting it into the first cross network to obtain a first intermediate vector includes:
[0203] Step 1310: Cascading at least one target object feature vector to obtain a first cascaded vector;
[0204] Step 1320: Inputting the first cascaded vector into the first first layer to obtain an output vector of the first first layer;
[0205] Step 1330: Through the other first layers in the multiple first layers except the first first layer, based on the output vector of the previous first layer and the first cascaded vector, obtain the output vectors of the other first layers, where the output vector of the last first layer is used as the first intermediate vector.
[0206] In this embodiment, the structure of the first cross network is as Figure 14 shown. The first cross network includes multiple first layers, and each first layer is used to perform a feature cross on the target object feature vectors. After the first cross network inputs the first cascaded vector and undergoes multi-order feature crosses of multiple first layers, a first intermediate vector is obtained.
[0207] In step 1310, at least one target object feature vector is cascaded to obtain a first cascaded vector. This process has been described in detail in the foregoing content. The first cascaded vector is the cascaded vector, and for the sake of brevity, it will not be repeated here.
[0208] In step 1320, the first cascaded vector is input into the first first layer to obtain an output vector of the first first layer.
[0209] In step 1330, for the other first layers except the first first layer, based on the output vector of the previous first layer and the first cascaded vector, obtain the output vectors of the other first layers. The output vector of the last first layer is used as the first intermediate vector output by the first cross network.
[0210] The calculation process when the first first layer performs feature crossing with other first layers can be the same. Other first layers obtain an output vector based on the output vector of the previous first layer and the first concatenated vector, and the first first layer obtains an output vector based on the first concatenated vector. It can also be considered that the first first layer obtains an output vector based on its previous first layer (the first concatenated vector) and the first concatenated vector. Thus, the feature crossing process of the first first layer and other first layers can adopt the same implementation manner.
[0211] Since replacing the output vector of the previous first layer in step 1330 with the first concatenated vector can represent the implementation manner of step 1320. Therefore, the specific process of step 1320 is similar to that of step 1330. To save space, the details of step 1320 will not be described in detail. The following is a detailed description of the implementation manner of step 1330.
[0212] In one embodiment, as Figure 15 shown, through other first layers among multiple first layers except the first first layer, based on the output vector of the previous first layer and the first concatenated vector, obtain the output vector of other first layers, including:
[0213] Step 1510: Obtain the first weight matrix and the first bias vector of other first layers among multiple first layers except the first first layer;
[0214] Step 1520: Based on the output vector of the previous first layer, the first concatenated vector, the first weight matrix, and the first bias vector, obtain the output vector of other first layers.
[0215] In this embodiment, there are a first weight matrix and a first bias vector for calculating the output vector in the first layer. The first weight matrix is a trainable parameter matrix, and the number of rows and columns of the first weight matrix is the dimension of the first concatenated vector. The first bias vector is a trainable parameter vector, and the dimension of the first bias vector is the dimension of the first concatenated vector. For example, if the dimension of the first concatenated vector is 4, then the first weight matrix is a 4×4 matrix, and the first bias vector is a vector with a dimension of 4.
[0216] Push event data often has the characteristic of low rank. Low rank is divided into two cases. The first case is that different push events are similar. For example, if the objects in push event A and push event B have the same or similar object characteristics, and the content pushed to the objects is also the same, then the benefit of using these two push events for training the push model simultaneously is relatively small. The second case is that the influence of different features on push preferences is the same or similar. For example, the object characteristics include "industry characteristics" and "occupation characteristics", and the influence of these two characteristics on object preferences is not very different in most cases. Therefore, considering the influence of these two characteristics on push preferences simultaneously has a small benefit. When training the push model, if the training sample set containing multiple push event samples is full rank, it means that each push event sample in the training sample is different, and the difference of each feature in the push event sample is large, and each can bring effective information for the training of the push model. However, the training sample set in practical applications is often low rank.
[0217] Based on the low rank of the push event data, in one embodiment, as Figure 16 shown, obtaining the first weight matrix of other first layers except the first first layer among multiple first layers includes:
[0218] Step 1610: Obtain the first low rank matrix and the second low rank matrix of other first layers except the first first layer among multiple first layers;
[0219] Step 1620: For each other first layer, multiply the first low rank matrix by the transpose of the second low rank matrix to obtain the first weight matrix.
[0220] The number of rows of the first low rank matrix and the second low rank matrix is equal to the dimension of the first concatenated vector, and the number of columns is less than the first dimension of the first concatenated vector. For example, if the dimension of the first concatenated vector is 4, then the number of rows of the first low rank matrix and the second low rank matrix is 4; the first dimension is less than the first concatenated vector, specifically, it can be determined according to the rank of the training sample set. The first low rank matrix and the second low rank matrix are both trainable parameter matrices.
[0221] The process of obtaining the first weight matrix based on the first low rank matrix and the second low rank matrix can be expressed by formula 2:
[0222] W l =Q l ·[R l T (Formula 2)
[0223] In formula 2, W l represents the first weight matrix, Q l represents the first low rank matrix, R l Represents the second low-rank matrix. Multiply the first low-rank matrix by the transpose of the second low-rank matrix to obtain the first weight matrix. For example, the number of rows of the first low-rank matrix and the second low-rank matrix is 4, and the number of columns is 2. The first low-rank matrix is
[0224] The second low-rank matrix is Then the first weight matrix is
[0225] In this embodiment, since the push event data has the characteristic of low rank, a high-dimensional first weight matrix can be obtained based on two low-rank matrices with a relatively small number of parameters, which is beneficial to improving the feature crossing efficiency of the first layer.
[0226] For each other first layer, based on the output vector of the previous first layer, the first concatenated vector, the first weight matrix, and the first bias vector, obtain the output vector of this first layer.
[0227] In one embodiment, as Figure 17 shown, based on the output vector of the previous first layer, the first concatenated vector, the first weight matrix, and the first bias vector, obtain the output vector of the other first layer, including:
[0228] Step 1710: Multiply the output vector of the previous first layer by the first weight matrix and add the first bias vector to obtain a second intermediate vector;
[0229] Step 1720: Multiply the first concatenated vector and the second intermediate vector bit by bit to obtain a third intermediate vector;
[0230] Step 1730: Add the third intermediate vector to the output vector of the previous first layer to obtain the output vector of the other first layer.
[0231] In this embodiment, the process of obtaining the output vector of the other first layer can be expressed as Formula 3:
[0232]
[0233] In Formula 3, l represents the serial number of the first layer. The first layer in the first cross network is numbered starting from 0. That is to say, for the first first layer, l is 1, for the second first layer, l is 2, and so on. X l represents the input vector of the first layer l, that is, the output vector of the previous first layer l - 1, X l+1 represents the output vector of the first layer l; X 0 represents the input vector of the first first layer, that is, the first concatenated vector; W l represents the first weight matrix of the first layer l; B lRepresents the first bias vector of the first layer l. Is the bitwise multiplication operator. For example,
[0234] Based on Formula 3, to calculate the output vectors of other first layers, first, the output vector of the previous first layer needs to be multiplied by the first weight matrix, and the result of the multiplication is added to the first bias vector to obtain the second intermediate vector. For example, to calculate the output vector of the second first layer (l = 1), the output vector of the previous first layer is [2, 6, 0, 4], and the first weight matrix is The first bias vector is [0, 3, 2, 4]. The second intermediate quantity is: The first concatenated vector is multiplied bitwise by the second intermediate vector to obtain the third intermediate vector. For example, if the first concatenated vector is [3, 1, 1, 4], then the third intermediate vector is: [3, 1, 1, 4] ⊙ [6, 23, 12, 48] = [36, 23, 12, 192]. Finally, the third intermediate vector is added to the output vector of the previous first layer to obtain the output vector of this first layer. Therefore, the output vector of the second first layer (l = 1) is: [36, 23, 12, 192] + [2, 6, 0, 4] = [38, 29, 12, 196].
[0235] The embodiments of steps 1710 - 1730 perform feature crossing based on the mutual influence between all target object features, so that the finally obtained first intermediate vector can represent the multi - order feature crossing result of the overall target object features, with high calculation efficiency and can fully obtain the associations between target object features.
[0236] In another embodiment, the first concatenated vector is obtained by concatenating the first number of target object feature vectors with the same dimension. For example, a first concatenated vector with a dimension of 4 is obtained by concatenating 2 target object feature vectors with a dimension of 2. Based on this, as Figure 18 shown, inputting the first concatenated vector into the first first layer to obtain the output vector of the first first layer, including:
[0237] Step 1810: Input the first concatenated vector into the first first layer, and generate output feature vectors for each target object feature vector;
[0238] Step 1820: Concatenate the first number of output feature vectors according to the sorting of the target object feature vectors in the first concatenated vector to obtain the output vector of the first first layer;
[0239] Based on steps 1810 - 1820, as Figure 19 shown, based on the output vector of the previous first layer, the first concatenated vector, the first weight matrix, and the first bias vector, obtain the output vectors of other first layers, including:
[0240] Step 1910: For each target object feature vector in the first-level concatenated vector, based on the target object feature vector, the output vector of the previous first layer, the first weight matrix, and the first bias vector, obtain the output feature vector of the target object feature vector in other first layers;
[0241] Step 1920: Concatenate the first number of output feature vectors according to the sorting of the target object features in the first vector to obtain the output vector of other first layers.
[0242] In this embodiment, the first layer performs feature crossing based on each target object feature vector to obtain output feature vectors, and concatenates the output feature vectors corresponding to the first number of target object feature vectors again to obtain the output feature vector of this first layer.
[0243] As Figure 20 shown, the first-level concatenated vector is formed by concatenating the target object feature vector A, the target object feature vector B, and the target object feature vector C. After the first-level concatenated vector is input into the first first layer L1 in the first cross network, L1 will split the first-level concatenated vector, and then obtain the L1 output feature vector A corresponding to the target object feature vector A, the L1 output feature vector B corresponding to the target object feature vector B, and the L1 output feature vector C corresponding to the target object feature vector C. Concatenate the L1 output feature vector A, the L1 output feature vector B, and the L1 output feature vector C to obtain the L1 output vector of the first first layer. Input the L1 output vector into the second first layer L2, and L2 will obtain the L2 output feature vector A corresponding to the target object feature vector A, the L2 output feature vector B corresponding to the target object feature vector B, and the L2 output feature vector C corresponding to the target object feature vector C based on the L1 output vector. Concatenate the L2 output feature vector A, the L2 output feature vector B, and the L2 output feature vector C to obtain the L2 output vector of the second first layer. Input the L2 output vector into the next first layer, and so on, until the first intermediate vector is obtained. The first intermediate vector is obtained by concatenating the output feature vectors corresponding to each target object feature in the last first layer.
[0244] Similar to the foregoing embodiment, the implementation manners of steps 1810-step 1820 and the implementation manners of steps 1910-step 1920 are similar. The following is a detailed description of the embodiment of steps 1910-step 1920.
[0245] In step 1910, for each target object feature vector, in the first layer, based on the target object feature vector, the output vector of the previous first layer, the first weight matrix, and the first bias vector, obtain the output feature vector corresponding to the target object feature vector.
[0246] In one embodiment, the first weight matrix is composed of a second number of first sub - weight matrices, where the second number is equal to the square of the first number. The number of rows and columns of each first sub - weight matrix is the dimension of the target object feature vector. For example, if the first number is 2, that is, there are 2 target object feature vectors with the same dimension. Then the second number is 2^2 = 4, and the first weight matrix is composed of 4 first sub - weight matrices. The dimension of the 2 target object feature vectors is 3, so the number of rows and columns of the first sub - weight matrix is 3. The dimension of the first concatenated vector is 3*2 = 6, so the number of rows and columns of the first weight matrix is 6. Therefore, the 4 first sub - weight matrices in the first weight matrix are arranged in a 2*2 manner. As Figure 21 shown, four first sub - weight matrices a, b, c, and d with the number of rows and columns both being 3 are arranged in a 2*2 manner to form a first weight matrix with the number of rows and columns both being 6.
[0247] The first bias vector is composed of a first number of first sub - bias vectors, and the dimension of the first sub - bias vector is equal to the dimension of the target object feature vector. For example Figure 22 shown, the first number is 2, and the dimension of the target object feature vector is 3, then the dimension of the first bias vector is 6, which is composed of 2 first sub - bias vectors with the dimension of 3.
[0248] As Figure 23 shown, for each target object feature vector in the first concatenated vector, based on the target object feature vector, the output vector of the previous first layer, the first weight matrix, and the first bias vector, the output feature vector of the target object feature vector in other first layers is obtained, including:
[0249] Step 2310: For each target object feature vector in the first concatenated vector, obtain a first number of target first sub - weight matrices from the second number of first sub - weight matrices, and obtain the target first sub - bias vector from the first number of first sub - bias vectors;
[0250] Step 2320: Based on the target object feature vector, the first number of target first sub - weight matrices, the target first sub - bias vector, and the output vector of the previous first layer, obtain the output feature vector of the target object feature vector in other first layers.
[0251] In step 2310, for each target object feature vector in the first concatenated vector, obtain a first number of target first sub - weight matrices from the first sub - weight matrices, and obtain one target first sub - bias vector from the first sub - bias vectors. The target first sub - weight matrix and the target first sub - bias vector are specifically used for the feature cross - of the target object feature vector.
[0252] In one implementation, obtaining the first number of target first sub-weight matrices from the second number of first sub-weight matrices and obtaining the target first sub-bias vector from the first number of first sub-bias vectors can be based on the position of the target object feature vector in the first concatenated vector.
[0253] As Figure 24 shown, the first sub-weight matrices in the first weight matrix W are identified as w 11 , w 12 , w 21 and w 22 according to their arrangement positions; the first sub-bias vectors in the first bias vector B are identified as b1 and b2 according to their arrangement positions. The first concatenated vector is obtained by concatenating the target object feature vectors x1 and x2. For each target object feature vector, based on its position in the first concatenated vector, the first sub-weight matrix in the first weight matrix with the same column number as this position is obtained as the target first sub-weight matrix, and the first sub-bias vector with the same position as this position in the first bias vector is obtained as the target first sub-bias vector. That is to say, for the target object feature vector x1, its target first sub-weight matrices are w 11 and w 21 , and the target first sub-bias vector is b1; for the target object feature vector x2, its target first sub-weight matrices are w 12 and w 22 , and the target first sub-bias vector is b2.
[0254] The advantage of obtaining the target first sub-weight matrix and the target first sub-bias vector based on the position of the target object feature vector in the first concatenated vector is high acquisition efficiency and convenience for training.
[0255] In step 2320, in the first layer, based on the target object feature vector, the target first sub-weights according to the town, the target first sub-bias vector, and the output vector of the previous first layer, the output feature vector of the target object feature in the first layer is obtained.
[0256] In one embodiment, as Figure 25 shown, based on the target object feature vector, the first number of target first sub-weight matrices, the target first sub-bias vector, and the output vector of the previous first layer, the output feature vector of the target object feature vector in other first layers is obtained, including:
[0257] Step 2510: Multiply each of the first number of output feature vectors in the output vector of the previous first layer by the first number of target first sub-weight matrices one by one to obtain the first number of fourth intermediate vectors;
[0258] Step 2520: Multiply each of the first number of fourth intermediate vectors by the target object feature vector bit by bit and then sum them to obtain the first sum vector;
[0259] Step 2530: Multiply the target object feature vector and the target first sub-bias vector bit by bit to obtain a fifth intermediate vector;
[0260] Step 2540: Add the first sum vector, the fifth intermediate vector, and the output feature vector corresponding to the target object feature vector in the output vector of the previous first layer to obtain the output feature vector of the target object feature vector in other first layers.
[0261] In step 2510, the output vector of the previous first layer is obtained by concatenating the first number of output feature vectors obtained from the previous first layer. Multiply the first number of output feature vectors in the output vector of the previous first layer by the first number of target first sub-weight matrices one by one to obtain the first number of fourth intermediate vectors. As Figure 26 shown, to calculate the output feature vector corresponding to the target object feature vector x1 in the first layer l, the target object feature vector x1 is [2, 1, 1]. The previous first layer is the l - 1 layer, and the output feature vector corresponding to the target object feature vector x1 in the output vector (the input vector of the l layer) of the l - 1 layer is [1, 0, 6], and the output feature vector corresponding to the target object feature vector x2 is [2, 2, 1]. The target first sub-weight matrix w 11 is The target first sub-weight matrix w 21 is Multiply by w 11 , and the obtained fourth intermediate vector is [0, 9, 10]; multiply by w 21 , and the obtained fourth intermediate vector is [9, 12, 9].
[0262] In step 2520, sum the first number of fourth intermediate vectors after multiplying them bit by bit with the target object feature vector one by one. That is, multiply the target object feature vector [2, 1, 1] by the above fourth intermediate vector [0, 9, 10] bit by bit to get [0, 9, 10]; multiply it by the above fourth intermediate vector [9, 12, 9] bit by bit to get [18, 12, 9]. Sum [0, 9, 10] and [18, 12, 9] to obtain the first sum vector as [18, 21, 19].
[0263] In step 2530, multiply the target object feature vector x1 and the target first sub-bias vector bit by bit. If the target first sub-bias vector is [1, 3, 2], the fifth intermediate vector obtained by multiplying the first sub-bias vector and the above target object feature vector [2, 1, 1] bit by bit is [2, 3, 2].
[0264] In step 2540, add the first sum vector, the fifth intermediate vector, and the output feature vector corresponding to the target object feature vector in the output vector of the previous first layer to obtain the output feature vector of the target object feature vector in this first layer. The above first sum vector is [18, 21, 19], the fifth intermediate vector is [2, 3, 2], and the output feature vector of the target object feature vector x1 in the previous first layer (layer l - 1) is [1, 0, 6]. Therefore, the output feature vector of the target object feature vector x1 in the current first layer (layer l) is [18, 21, 19] + [2, 3, 2] + [1, 0, 6] = [21, 24, 27].
[0265] In summary, the embodiments of steps 2510 - 2540 can be represented by formula 4:
[0266]
[0267] Formula 4 calculates the output feature vector of the i-th target object feature in the first layer l in the first concatenated vector. Where x i represents the i-th target object feature vector in the first concatenated vector; represents the output feature vector of the i-th target object feature vector in the first layer l; represents the output feature vector of the target object feature vector i in the previous first layer (the previous first layer of the first layer l), i.e., layer l - 1; represents the target first bias vector of the i-th target object feature in the first layer l; N represents the number of target object features in the first concatenated vector, and the value range of j is between [1, N]. Therefore, represents the output feature vector of the j-th target object feature in the first layer l - 1; represents the j-th target first sub-weight matrix corresponding to the i-th target object feature vector in the first layer l. It can be obtained from formula 3 that The result obtained is the fourth intermediate vector obtained by multiplying the output feature vector corresponding to the j-th target object feature vector in the first layer l - 1 by the j-th target first sub-weight matrix corresponding to the i-th target object feature vector in the first layer l. The result obtained is the first sum vector obtained by summing the products of the i-th target object feature vector and N fourth intermediate vectors. The result obtained is the fifth intermediate vector obtained by multiplying the i-th target object feature vector bit by bit with the target first bias vector. Therefore, adding the first sum vector, the output feature vector of the target object feature vector i in the first layer l - 1, and the fifth intermediate vector, the obtained is the output feature vector of the target object feature i in the first layer l.
[0268] The output feature vector of the target object feature vector xi at the first layer l is obtained through Equation 4 It is necessary to utilize the output feature vector of the first layer l-1 And the output feature vector of the first layer l-1 It is necessary to utilize the output feature vector of the first layer l-2 Obtained. Therefore, further expanding Equation 4 gives Equation 5:
[0269]
[0270] Equation 5 is obtained by substituting the calculation processes of and in Equation 4. Based on Equation 4,
[0271] wherein, the meanings of k and h are the same as that of j, and the value ranges of k and h are also within [1, N]
[0272] between. The difference between Equation 5 and Equation 4 is only that Equation 4 is the feature cross process of layer l, obtained based on the output feature vector of layer l-1, while Equation 5 fuses the feature cross processes of layer l and layer l-1, and replaces the output feature vector of the first layer l-1 with the feature cross process of the first layer l-1. Therefore, the meanings of the various parameters in Equation 5 are similar to those in Equation 4, and will not be elaborated here
[0273] After organizing Equation 5 and combining like terms, Equation 6 can be obtained:
[0274]
[0275] Since every time passing through a first layer is to perform a feature cross on the target object feature vector, the more the number of first layers passed through, the higher the feature cross order. Specifically, the target object feature vector itself is a first-order feature vector; the first layer (L) in the first cross network starts counting from 0. After passing through the first first layer (l0), performing a feature cross, the obtained output feature vector is a second-order feature vector for the target object feature vector; inputting the second-order feature cross result into the second first layer (l1) for another cross, the obtained output feature vector is a third-order feature vector for the target object feature vector; and so on. The last first layer is the (M + 1)-th first layer (l M )), then the output feature vector obtained by l M is the (M + 2)-th order feature vector for the target object feature vector
[0276] It can be seen from Equation 6 that In the case where the target object feature vector xi does not cross with any feature vectors, therefore, this can be regarded as the first-order feature vector based on the target object feature vector xi itself. In this case, the target object feature vector xi will perform feature crossing with the target object feature vector xj, and since xj is also a first-order feature vector, therefore, it can be regarded as the second-order feature vector based on the target object feature vector xi. is the target object feature vector x i which is the l-order feature vector output by the first layer l - 2. In this case, the target object feature vector x i crosses with and since is the l-order feature vector of the target object feature vector x j therefore, it can be regarded as the l + 1-order feature vector based on the target object feature vector x i Finally, In this case, first, the target object feature vector x j crosses with the target object feature vector x k which is the l-order feature vector output by the first layer l - 2 to obtain the l + 1-order feature vector of the target object feature vector x j ; then, the target object feature vector x i crosses with the l + 1-order feature vector of the target object feature vector x j to obtain the l + 2-order feature vector of the target object feature vector x i For example, there are 4 first layers (L = [0, 3]), and the first feature vector corresponding to the target object feature vector output by the last first layer l3 is obtained by adding the target object feature vector itself (first-order feature vector), the second-order feature vector of the target object feature vector (obtained through l0), the third-order feature vector (obtained through l1), the fourth-order feature vector (obtained through l2), and the fifth-order feature vector (obtained through l3). That is to say, when there are L first layers, the sequence of the last first layer is L - 1, and the output feature vector of the target object feature vector x i obtained in the first layer L - 1 is composed of the combination of the first-order feature vector to the L + 1-order feature vector of the target object feature vector x i
[0277] It can be seen that the target object feature vector obtained through the calculation process of the embodiments from step 2510 to step 2540 in the other first layers of the output feature vectors can indicate that the output feature vector is obtained based on the multi-order feature vectors from the previous first layer to the current first layer, so as to ensure that the first feature vector corresponding to the target object feature obtained by the last first layer fully integrates the results of multi-order feature crossing and can accurately reflect the mutual influence between the target object features.
[0278] In the embodiments of step 2310 to step 2320, the first weight matrix and the first bias vector are partitioned according to the target object feature vector, and the output feature vector of the target object feature vector in the first layer is obtained based on the target first sub-weight matrix and the target first sub-bias vector. In this way, the parameters in the first weight matrix and the first bias vector can be learned for the target object feature vector, and then a more accurate feature crossing result can be obtained.
[0279] In step 1920, since each target object feature vector will obtain a corresponding output feature vector in the other first layers. Therefore, the first number of output feature vectors are cascaded according to the sorting of the target object feature vector in the first concatenated vector, and the output vector of the other first layers can be obtained. The process of cascading the first number of target object feature vectors is similar to the process of cascading at least one target object feature vector in the foregoing embodiments, and will not be elaborated here.
[0280] The advantages of the embodiments of step 1910 to step 1920 are that feature crossing is performed on each target object feature vector, so that the crossing result can fully reflect the mutual influence between the target object feature and other target feature vectors, and then a more accurate first feature vector can be obtained.
[0281] Step 1810 to step 1820 are the previous steps of step 1910 to step 1920. Therefore, the advantages are the same as those of the embodiments of step 1910 to step 1920, and will not be elaborated here.
[0282] In the embodiments of step 1510 to step 1520, the first weight matrix and the first bias vector are obtained after training. Therefore, using the first weight matrix, the first bias vector, the output vector of the previous first layer, and the first concatenated vector to obtain the output vector of the other first layers is based on the push information in the actual application scenario, and the feature crossing accuracy is relatively high, and then the push accuracy can be improved.
[0283] In the embodiment of steps 1310 - 1330, the first cross network includes multiple cascaded first layers for performing high-order feature crossing on the target object feature vectors. Performing high-order feature crossing can deeply mine the associations between the target object feature vectors, which is beneficial to improving the accuracy of pushing.
[0284] In step 530, from the first intermediate vector, obtain the first feature vector corresponding to each target object feature vector. Based on the cascaded vector [x1, x2, ……, x n to obtain the first intermediate vector X L which is [[x1] L , [x2] L’ , ……, [x n L’ . Among them, the first feature vector corresponding to x1 is [x1] L , the first feature vector corresponding to x2 is [x2] L , and so on. Thus, it can be seen that the first intermediate vector is obtained by cascading the first feature vectors corresponding to at least one target object feature vector. Therefore, based on the cascading position of the target object feature vector in the first cascaded vector, the vector at the corresponding cascading position in the first intermediate vector can be obtained as the first feature vector of the target object.
[0285] In the embodiment of steps 510 - 520, after vectorizing at least one target object feature and cascading them, inputting the cascaded vector into the first cross network can enable the first cross network to fully obtain the mutual influence between the target object features. The first feature vector obtained based on the first cross network can accurately reflect the association between the corresponding target object feature and other target object features, improving the accuracy of determining the first feature vector corresponding to each target object feature.
[0286] Similar to the embodiment of steps 510 - 520, as Figure 27 shown, based on at least one feature of the content to be pushed, determining the second feature vector corresponding to each feature of the content to be pushed includes:
[0287] Step 2710: Vectorize at least one feature of the content to be pushed into at least one feature vector of the content to be pushed;
[0288] Step 2720: Cascade at least one feature vector of the content to be pushed and input it into the second cross network to obtain a sixth intermediate vector;
[0289] Step 2730: From the sixth intermediate vector, obtain the second feature vector corresponding to each feature of the content to be pushed.
[0290] The embodiment of step 2710 - step 2720 is a process of determining the second feature vector corresponding to each feature of the content to be pushed. It is similar to the process of determining the first feature vector corresponding to each feature of the target object in step 510 - step 520, and will not be elaborated here. Similarly, this embodiment can improve the accuracy of determining the second feature vector corresponding to each feature of the content to be pushed.
[0291] Detailed description of step 330
[0292] In step 330, for each first feature vector, the first feature vector is combined with each second feature vector respectively, and each combination of the first feature vector and the second feature vector is input into the combined cross network to obtain a combined cross feature vector.
[0293] When forming combinations, each first feature vector is combined with each second feature vector respectively. If the number of first feature vectors is M and the number of second feature vectors is N, then the number of formed combinations is M * N. For example Figure 28 As shown, the first feature vector U1, the first feature vector U2, and the first feature vector U3 are successively combined with the second feature vector I1 and the second feature vector I2, and the formed combinations include: (the first feature vector U1, the second feature vector I1), (the first feature vector U1, the second feature vector I2), (the first feature vector U2, the second feature vector I1), (the first feature vector U2, the second feature vector I2), (the first feature vector U3, the second feature vector I1), and (the first feature vector U3, the second feature vector I2).
[0294] Each combination is input into the combined cross network to obtain a combined cross feature vector. The combined cross feature vector is the cross result of the first feature vector and the second feature vector. Since the first feature vector represents the interaction between one target object feature and other target object features, and the second feature vector represents the interaction between one feature of the content to be pushed and other features of the content to be pushed. Therefore, the combined cross feature vector is the high - order fusion between the target object feature and the feature of the content to be pushed.
[0295] In one embodiment, the dimensions of the first feature vector and the second feature vector are the same. Based on this, inputting each combination of the first feature vector and the second feature vector into the combined cross network to obtain a combined cross feature vector includes: multiplying the first feature vector and the second feature vector bit by bit to obtain a combined cross feature vector.
[0296] The process by which the combined cross network obtains the combined cross feature vector based on the first feature vector and the second feature vector can be expressed by formula 7:
[0297]
[0298] In Formula 7, represents the first eigenvector p, represents the second eigenvector q, f pq represents the combined cross eigenvector of the first eigenvector p and the second eigenvector q. For example, if the first eigenvector p is [2, 3, 5, 1, 7, 1, 5, 4] and the second eigenvector is [3, 2, 1, 2, 6, 3, 2, 1], then their combined cross eigenvector is [2 * 3, 3 * 2, 5 * 1, 1 * 2, 7 * 6, 1 * 3, 5 * 2, 4 * 1] = [6, 6, 5, 2, 42, 3, 10, 4].
[0299] In the foregoing embodiment, the first eigenvector is obtained based on L first layers in the first cross network, and the value of the first eigenvector is obtained by adding the first-order eigenvector to the (L + 1)-order eigenvector of the target object eigenvector, and can be specifically expressed as: Wherein, represents the target object eigenvector x p of the first-order eigenvector, represents the second-order eigenvector, and so on, represents the (L + 1)-order eigenvector.
[0300] Similarly, the second eigenvector is obtained based on L second layers in the second cross network, and the value of the second eigenvector is obtained by adding the first-order eigenvector to the (L + 1)-order eigenvector of the content feature vector to be pushed, and can be specifically expressed as: Wherein, represents the content feature vector x to be pushed q of the first-order eigenvector, represents the second-order eigenvector, and so on, represents the (L + 1)-order eigenvector.
[0301] Also, because Therefore, Formula 7 can be further expressed as Formula 8:
[0302]
[0303] It can be seen from Formula 8 that f pq reflects the result obtained by combining and crossing the multi-order eigenvectors in the first eigenvector and the multi-order eigenvectors in the second eigenvector and then accumulating.
[0304] Therefore, in the above embodiments, the first feature vector and the second feature vector are multiplied bit by bit to obtain a combined cross feature vector, which is the result of accumulating the combined cross of each order feature vector in the first feature vector and the second feature vector. In this way, the feature cross between the first feature vector and the second feature vector is more comprehensive, and the mutual influence between the target object features and the features of the content to be pushed can be deeply obtained, which helps to improve the push accuracy.
[0305] Detailed description of step 340
[0306] In step 340, the combined cross feature vector is input into the probability prediction model, and the first probability of pushing the content to be pushed to the target object can be obtained, and the content to be pushed is pushed to the target object based on the first probability.
[0307] In step 330, a combined cross feature vector will be obtained for each combination of the first feature vector and the second feature vector. Therefore, the combined cross feature vectors received in the probability prediction model are a vector sequence composed of multiple combined cross feature vectors. The vector sequence can be expressed as F=(f1, f2, …, f n ), where f1, f2, …, f n represent N combined feature vectors in sequence.
[0308] In one embodiment, the dimensions of each combined feature vector are the same. Therefore, inputting the combined cross feature vector into the probability prediction model to obtain the first probability of pushing the content to be pushed to the target object includes:
[0309] Adding multiple combined cross feature vectors to obtain a final cross feature vector;
[0310] Based on the final cross feature vector, obtain the first probability of pushing the content to be pushed to the target object.
[0311] The process of obtaining the final cross feature vector can be expressed by formula 9:
[0312]
[0313] In formula 9, represents the final cross feature vector. For example, if the combined cross feature vectors in the vector sequence are: [1, 5, 0, 3, 2, 2], [0, 2, 5, 3, 1, 2], [3, 3, 5, 0, 2, 1] and [6, 1, 1, 8, 3, 2], then the final cross feature vector is equal to the direct addition of each combined cross feature vector, and the result is [10, 11, 11, 14, 8, 7].
[0314] After obtaining the final cross feature vector, the first probability of pushing content to the target object is obtained based on the final cross feature vector.
[0315] Based on the foregoing embodiments, the process of obtaining the first probability of pushing content for the target object based on the target object features and the features of the content to be pushed may be as follows Figure 29 As shown, after the target object features U1, target object features U2, the features I1 of the content to be pushed, and the features I2 of the content to be pushed are vectorized, the target object feature vectors U1, target object feature vectors U2, the feature vectors I1 of the content to be pushed, and the feature vectors I2 of the content to be pushed are obtained. After concatenating the target object feature vector U1 and the target object feature vector U2, they are input into the first cross network to obtain the first feature vector [U1] corresponding to the target object feature vector U1 L and the first feature vector [U2] corresponding to the target object feature vector U2 L . After concatenating the feature vector I1 of the content to be pushed and the feature vector I2 of the content to be pushed, they are input into the second cross network to obtain the first feature vector [I1] corresponding to the feature vector I1 of the content to be pushed L and the first feature vector [I2] corresponding to the feature vector I2 of the content to be pushed L . Each first feature vector and the second feature vector are combined in sequence to obtain ([U1] L , [I1] L ), ([U1] L , [I2] L ), ([U2] L , [I1] L ) and ([U2] L , [I2] L ) four combinations. The first feature vector and the second feature vector in each combination are combined and crossed to obtain the corresponding combined cross feature vector. The combined cross feature vector is input into the probability prediction model to obtain the first probability.
[0316] In another embodiment, as Figure 30 shown, for each first feature vector, after combining the first feature vector with each of the second feature vectors respectively and inputting each combination of the first feature vector and the second feature vector into the combined cross network to obtain the combined cross feature vector, it further includes:
[0317] Step 335: Input multiple push basic features into the full cross network to obtain the full cross feature vector.
[0318] Based on this, inputting the combined cross feature vector into the probability prediction model to obtain the first probability of pushing the content to be pushed for the target object includes:
[0319] Step 345: Input the combined cross feature vector and the full cross feature vector into the probability prediction model to obtain the first probability of pushing the content to be pushed for the target object, and push the content to be pushed for the target object based on the first probability.
[0320] In this embodiment, the full-cross network obtains the mutual influence between all the push basic features through implicit feature crossing.
[0321] In one embodiment, multiple push basic features are input into the full-cross network to obtain a full-cross feature vector, including:
[0322] Vectorize multiple push basic features to obtain multiple push basic feature vectors;
[0323] Concatenate multiple push basic feature vectors to obtain a second concatenated vector;
[0324] Input the second concatenated vector into the full-cross network to obtain a full-cross feature vector.
[0325] The process of vectorizing and then concatenating multiple push basic features is the same as the process of vectorizing and then concatenating the target object features in the foregoing embodiment, and will not be elaborated here.
[0326] The full-cross network can perform crossing on a single feature dimension in the second concatenated vector, break the boundaries between push basic features, and fully fuse the semantics between multiple push basic features. The full-cross network can be a deep neural network (DNN). The full-cross feature vector obtained by the full-cross network based on the second concatenated vector can be expressed as: f d = h(S). Where S represents the second concatenated vector, h(·) represents the feature crossing processing process of the full-cross network, and f d represents the full-cross feature vector.
[0327] The process of obtaining the full-cross feature vector does not interfere with the process of obtaining the combined cross feature vector. In step 345, the combined cross feature vector and the full-cross feature vector are jointly input into the probability prediction model to obtain a first probability. The vector sequence received by the probability prediction model can be expressed as: F = (f1, f2,..., f n , f d ). Where f1, f2,..., f n represent N combined feature vectors in sequence, and f d represents the full-cross feature vector.
[0328] In one embodiment, as Figure 31 shown, input the combined cross feature vector and the full-cross feature vector into the probability prediction model to obtain the first probability of pushing the content to be pushed to the target object, including:
[0329] Step 3110, add the combined cross feature vector and the full-cross feature vector to obtain a final cross feature vector;
[0330] Step 3120: Obtain a first probability of pushing the content to be pushed to the target object based on the final cross feature vector.
[0331] The final cross feature vector obtained in step 3110 can be expressed as formula 10:
[0332]
[0333] Referring to formula 10, the combined cross feature vectors in the vector sequence are: [1, 5, 0, 3, 2, 2], [0, 2, 5, 3, 1, 2], [3, 3, 5, 0, 2, 1] and [6, 1, 1, 8, 3, 2], the full cross feature vector is [2, 5, 1, 0, 3, 7], so the final cross feature vector is [12, 16, 12, 14, 11, 14].
[0334] After adding the full cross network in the push processing process of this embodiment of the present disclosure, the process of obtaining the first probability of pushing content to the target object based on the target object feature and the content feature to be pushed can be as Figure 32 shown. On the basis of Figure 29 , cascade the target object feature vector U1, the target object feature vector U2, the content feature vector I1 to be pushed and the content feature vector I2 to be pushed, and input them into the full cross network to obtain the full cross feature vector. Input the four combined cross feature vectors and the full cross feature vector into the probability prediction model to obtain the first probability.
[0335] In this embodiment, adding a full cross network in the push processing process and jointly determining the first probability of pushing the content to be pushed to the target object based on the combined cross feature vector and the full cross vector has the advantage that, in addition to considering the mutual influence between the target object feature and the content feature to be pushed in the push processing process, it can deeply mine the semantic fusion between all push basic features, so that the push processing process can more comprehensively cover the associations between all features and improve the accuracy of the push processing.
[0336] Regardless of whether the full cross network is included in the push processing process, the probability prediction model will obtain the first probability of pushing the content to be pushed to the target object according to the final cross feature vector.
[0337] In one embodiment, as Figure 33 shown, obtaining the first probability of pushing the content to be pushed to the target object based on the final cross feature vector includes:
[0338] Step 3310: Obtain a first weight vector and a first bias;
[0339] Step 3320: Multiply the first weight vector by the transpose of the final cross feature vector and add the first bias to obtain a first score;
[0340] Step 3330: Perform exponential normalization on the first score to obtain the first probability of pushing the content to be pushed to the target object.
[0341] In this embodiment, the dimension of the first weight vector is equal to the dimension of the final cross feature vector. Both the first weight vector and the first bias are trainable parameters. For example, if the final cross feature vector is [1, 3, 1, 0] with a dimension of 4, the first weight vector is [0.5, 0.1, 0.2, 0.1], and the first bias is 0.2.
[0342] In step 3320, multiply the first weight vector by the transpose of the final cross feature vector and add the first bias vector to obtain the first score. Therefore, based on the above example, the first score obtained is:
[0343]
[0344] In step 3330, perform exponential normalization on the first score to obtain the first probability. Exponential normalization can be achieved using a non-linear activation function (sigmoid function). The specific form of the sigmoid function can be expressed as Equation 11:
[0345]
[0346] where x is the input of the sigmoid function, which is the first score; f(x) is the first probability obtained based on the first score. Therefore, if the first score is 1.2, then the first probability is
[0347] In summary, the process of obtaining the first probability in the embodiments of steps 3310 - 3330 can be expressed as Equation 12:
[0348]
[0349] In Equation 12, p is the first probability; is the first weight vector; is the final cross feature vector; is the first bias; G(·) represents the exponential normalization process, which can be sigmoid(·).
[0350] The advantage of obtaining the first probability in steps 3310 - 3330 is that the calculation process is differentiable, and the model training efficiency based on the gradient is relatively high.
[0351] In practical applications, determining the content to be pushed to the target object often requires multiple tasks for evaluation. A task refers to an operation that the target object performs on the pushed content, such as: liking, favoriting, commenting, watching completely, etc.
[0352] Thus, in one embodiment, the probability prediction model includes multiple first task prediction models corresponding to multiple tasks, and the first task prediction model is used to predict the probability that the target object completes the task for the content to be pushed. For example, the first task prediction model corresponding to the like task is used to predict the probability that the target object likes the content to be pushed.
[0353] As Figure 34 shown, inputting the combined cross feature vector into the probability prediction model to obtain the first probability of pushing the content to be pushed to the target object includes:
[0354] Step 3410: Input the combined cross feature vector into multiple first task prediction models to obtain multiple first task completion probabilities corresponding to multiple tasks;
[0355] Step 3420: Based on the multiple first task completion probabilities, obtain the first probability of pushing the content to be pushed to the target object.
[0356] The probability prediction model includes multiple first task prediction models. Input multiple combined cross feature vectors into each first task prediction model to obtain the first task completion probability that the target object completes the corresponding task for the content to be pushed. As Figure 35 shown, input the combined cross feature vector A, the combined cross feature vector B, the combined cross feature vector C, and the combined cross feature vector D into the first task prediction model A to obtain the first task completion probability A; input into the first task prediction model B to obtain the first task completion probability B; input into the first task prediction model C to obtain the first task completion probability C.
[0357] The process of each first task prediction model obtaining the first task completion probability can be the same as the process of directly obtaining the first probability based on the combined cross feature vector in the foregoing embodiment. Details are not described herein again.
[0358] After determining the first task completion probability corresponding to each first task prediction model, the first probability of pushing the content to be pushed to the target object can be obtained by calculating the average value. For example, if the first task completion probability A is 0.56, the first task completion probability B is 0.44, and the first task completion probability C is 0.36, then the first probability is equal to (0.56 + 0.44 + 0.36) / 3 = 0.453. The advantage of using the average value calculation is that the influence of each task on pushing the content to be pushed to the target object is the same, which improves the fairness of the push processing based on multiple tasks.
[0359] To determine the first probability based on the completion probabilities of multiple first tasks, the method of calculating the weighted average can also be used. The task weights are determined based on the impact degree of each task on the target object, and the weighted average of the first task completion probabilities is calculated based on the task weights of multiple tasks. For example, the task weight of task A is 0.2, and the corresponding first task completion probability A is 0.56; the task weight of task B is 0.3, and the corresponding first task completion probability B is 0.44; the task weight of task C is 0.5, and the corresponding first task completion probability C is 0.36. Therefore, the first probability is equal to 0.2 * 0.56 + 0.3 * 0.44 + 0.5 * 0.36 = 0.424. The advantage of using the weighted average calculation is that the corresponding task weights can be set based on the impact degree of each task on the target object, which improves the flexibility and accuracy of the push processing based on multiple tasks.
[0360] In the embodiment of steps 3410 - 3420, the first probability can be obtained based on the task completion probabilities of multiple tasks that the target object completes for the content to be pushed, and then the first probability is used to push to the target object, which improves the push accuracy.
[0361] In the foregoing embodiment, a fully cross network is also applied in the push processing. The fully cross network can also be applied to the push processing based on multiple tasks. Based on this, in another embodiment, multiple push basic features are input into the fully cross network to obtain a fully cross feature vector, including: inputting multiple push basic features into multiple fully cross networks to obtain multiple fully cross feature vectors.
[0362] After each fully cross network receives multiple push basic features, the corresponding fully cross feature vector is obtained. Multiple fully cross networks can be used to analyze and calculate each dimension of the push basic features to balance the sharing and mutual exclusion between multiple tasks.
[0363] The probability prediction model includes gating nodes corresponding to multiple tasks and a second task prediction model. As Figure 36 shown, inputting the combined cross feature vector and the fully cross feature vector into the probability prediction model to obtain the first probability of pushing the content to be pushed to the target object, including:
[0364] Step 3610: Input the combined cross feature vector and multiple fully cross feature vectors into the gating nodes corresponding to the tasks to obtain the seventh intermediate vector corresponding to the tasks;
[0365] Step 3620: Input the seventh intermediate vector into the second task prediction model corresponding to the tasks to obtain the second task completion probability corresponding to the tasks;
[0366] Step 3630: Based on the second task completion probabilities corresponding to multiple tasks, obtain the first probability of pushing the content to be pushed to the target object.
[0367] In this embodiment, each task has a corresponding gating node and a second task prediction model. The gating node receives the combined cross feature vector and multiple full cross feature vectors to obtain the seventh intermediate vector corresponding to the task. As Figure 37 shown, the combined cross feature vector A, the combined cross feature vector B, the combined cross feature vector C, the full cross feature vector A generated by the full cross network a, and the full cross feature vector B generated by the full cross network b are sequentially input into the gating node A, the gating node B, and the gating node C. The seventh intermediate vector output by each gating node is input into the second task prediction model corresponding to the task to obtain the corresponding second task completion probability. Finally, the first probability is obtained based on the second task completion probabilities of multiple tasks.
[0368] Among the gating nodes corresponding to different tasks, the vector weights for the same combined cross feature vector or full cross feature vector may be different. This is because the same combined cross feature vector or full cross feature vector may contribute differently to different tasks. Therefore, when the combined cross feature vector and multiple full cross feature vectors are input into the gating node corresponding to the task, the gating node can calculate the weighted sum of the combined cross feature vector and the full cross feature vectors based on the vector weights of each combined cross feature vector and each full cross feature vector to obtain the seventh intermediate vector corresponding to the task.
[0369] After obtaining the seventh intermediate vector, the seventh intermediate vector is input into the second task prediction model corresponding to the task to obtain the second task completion probability corresponding to the task.
[0370] The process by which the second task prediction model obtains the second task completion probability based on the seventh intermediate vector may be the same as the process by which the probability prediction model obtains the first probability based on the final cross feature vector in the foregoing embodiment, and will not be elaborated here.
[0371] After obtaining the second task completion probability corresponding to each task, the first probability can be obtained based on the second task completion probabilities of multiple tasks. This process is the same as the process of obtaining the first probability based on the first task completion probabilities of multiple tasks in the foregoing embodiment, and will not be elaborated here.
[0372] In the embodiments of steps 3610 - 3630, feature crossing is performed from multiple dimensions of the push base features using multiple full cross feature vectors, and the gating nodes of multiple tasks are used to evaluate the influence of different vectors on the task, which is beneficial to improving the accuracy of determining the task completion probabilities of each task.
[0373] After determining the first probability, it is necessary to push the content to be pushed to the target object based on the first probability.
[0374] In one embodiment, pushing the content to be pushed to the target object based on the first probability includes:
[0375] Sort the content to be pushed in descending order of the first probability;
[0376] Push the top predetermined number of the content to be pushed in the sorting to the target object.
[0377] For example, after sorting the content to be pushed according to the first probability, the first probability of the content to be pushed A is 0.82, the first probability of the content to be pushed B is 0.68, the first probability of the content to be pushed C is 0.39, the first probability of the content to be pushed D is 0.31, and the first probability of the content to be pushed E is 0.19. If the predetermined number of the content to be pushed to be pushed to the target object is 3, then push the content to be pushed A, the content to be pushed B, and the content to be pushed C to the target object.
[0378] The advantage of pushing the top predetermined number of the content to be pushed in the sorting to the target object is that it ensures that a sufficient number of the content to be pushed can be pushed to the target object, enables the probability prediction model to obtain more effective information for pushing, and further improves the personalized recommendation ability of the probability prediction model.
[0379] In one embodiment, pushing the content to be pushed to the target object based on the first probability includes:
[0380] Obtain a probability threshold;
[0381] Push the content to be pushed whose first probability is not lower than the probability threshold to the target object.
[0382] Based on the above example, if the probability threshold is 0.6, then push the content to be pushed A and the content to be pushed B to the target object.
[0383] The advantage of pushing the content to be pushed to the target object based on the probability threshold is that it ensures that the content to be pushed to the target object is of interest to the target object, which is beneficial to improving the pushing accuracy.
[0384] The training process of the probability prediction model in the embodiments of the present disclosure
[0385] In the embodiments of the present disclosure, in order to obtain an accurate first probability based on the combined cross feature vector, it is necessary to train the probability prediction model.
[0386] In one embodiment, as Figure 38 shown, the probability prediction model is specifically trained in the following manner:
[0387] Step 3810: Obtain a push event sample, where the push event sample includes at least one sample object feature, at least one sample content feature, and an event label;
[0388] Step 3820: Based on at least one sample object feature and at least one sample content feature, use a probability prediction model to determine a second probability;
[0389] Step 3830: Based on the second probability and the event label, determine an error function for the push event sample;
[0390] Step 3840: Train the probability prediction model based on the error function.
[0391] The push event sample includes at least one sample object feature, at least one sample content feature, and an event label. Among them, the sample object feature is the same as the target object feature, and the sample content feature is the same as the content feature to be pushed. Details are not elaborated here. The event label indicates whether the sample object performs a certain task on the sample content. For example, viewing the sample content, or liking the sample content, etc. If so, the event label is 1, otherwise it is 0.
[0392] The process of determining the second probability based on at least one sample object feature and at least one sample content feature using a probability prediction model is the same as the process of obtaining the first probability in steps 310 - 340 described above, and details are not elaborated here.
[0393] The error function refers to the gap between the second probability and the event label. The smaller the value of the error function, the more accurate the prediction result. The error function for the push event sample can be determined based on the second probability and the event label as shown in Formula 13:
[0394] L = -y log p - (1 - y) log(1 - p) (Formula 13)
[0395] In Formula 13, L represents the error function of the push event sample, y represents the event label (0 or 1), and p represents the second probability. For example, if the event label of a push event sample is 1 and the second probability is 0.5, then the error function of this push event sample is equal to -1·log(0.5) = 0.3.
[0396] After determining the error function, train the probability prediction model based on the error function. The purpose of training is to reduce the error function, and a reduction in the error function represents a more accurate prediction result. Therefore, the training process is to update the parameters in the probability prediction model so that the updated model will reduce the finally generated error function.
[0397] The probability prediction model can be trained based on the error function using the gradient descent method so that the parameters in the probability prediction model can be updated in the direction of the gradient descent of the error function.
[0398] Implementation details diagram of the push processing method of the present disclosure embodiment
[0399] The following refers toFigure 39 , the implementation details of the push processing method according to the embodiments of the present disclosure are described in detail by way of example.
[0400] In step 3910, a plurality of push basic features are obtained, where the plurality of push basic features include at least one target object feature and at least one content feature to be pushed.
[0401] In step 3920, at least one target object feature is vectorized into at least one target object feature vector; the at least one target object feature vector is cascaded and input into the first cross network to obtain a first intermediate vector; from the first intermediate vector, a first feature vector corresponding to each target object feature is obtained.
[0402] At least one content feature to be pushed is vectorized into at least one content feature vector to be pushed; the at least one content feature vector to be pushed is cascaded and input into the second cross network to obtain a sixth intermediate vector; from the sixth intermediate vector, a second feature vector corresponding to each content feature to be pushed is obtained.
[0403] In step 3930, for each first feature vector, the first feature vector is combined with each of the second feature vectors respectively, and each combination of the first feature vector and the second feature vector is input into the combination cross network; the first feature vector and the second feature vector are multiplied bit by bit to obtain a combined cross feature vector.
[0404] In step 3940, the plurality of push basic features are input into the full cross network to obtain a full cross feature vector.
[0405] In step 3950, the combined cross feature vector and the full cross feature vector are input into the probability prediction model; the combined cross feature vector and the full cross feature vector are added to obtain a final cross feature vector; based on the final cross feature vector, a first probability of pushing the content to be pushed to the target object is obtained.
[0406] In step 3960, the content to be pushed is pushed to the target object based on the first probability.
[0407] Description of the device and equipment according to the embodiments of the present disclosure
[0408] It can be understood that although the steps in each of the above flowcharts are shown sequentially according to the indication of the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless there is a clear indication in this embodiment, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the above flowchart may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0409] It should be noted that in each specific implementation manner of this application, when it comes to performing relevant processing based on data related to the target content characteristics such as target content attribute information or attribute information set, permission or consent for the target content will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiment of this application needs to obtain target content attribute information, it will obtain the separate permission or separate consent for the target content through methods such as pop-up windows or jumping to a confirmation page. After clearly obtaining the separate permission or separate consent for the target content, the necessary target content-related data for the normal operation of the embodiment of this application will be obtained.
[0410] Figure 40 It is a schematic structural diagram of a push processing device 4000 provided by an embodiment of the present disclosure. The push processing device 4000 includes:
[0411] An acquisition unit 4010, configured to acquire a plurality of push basic features, where the plurality of push basic features include at least one target object feature and at least one content feature to be pushed;
[0412] A determination unit 4020, configured to determine a first feature vector corresponding to each target object feature based on at least one target object feature, and determine a second feature vector corresponding to each content feature to be pushed based on at least one content feature to be pushed;
[0413] A combination unit 4030, configured to, for each first feature vector, combine the first feature vector with each of the second feature vectors respectively, and input each combination of the first feature vector and the second feature vector into a combination cross network to obtain a combined cross feature vector;
[0414] A push unit 4040, configured to input the combined cross feature vector into a probability prediction model to obtain a first probability of pushing the content to be pushed to the target object, and push the content to be pushed to the target object based on the first probability.
[0415] Optionally, the determination unit 4020 is specifically configured to:
[0416] Vectorize at least one target object feature into at least one target object feature vector;
[0417] Cascade at least one target object feature vector, input it into the first cross network, and obtain a first intermediate vector;
[0418] Obtain a first feature vector corresponding to each target object feature from the first intermediate vector.
[0419] Optionally, the first cross network includes multiple first layers;
[0420] The determination unit 4020 is specifically configured to:
[0421] Cascade at least one target object feature vector to obtain a first cascaded vector;
[0422] Input the first cascaded vector into the first first layer to obtain an output vector of the first first layer;
[0423] Through the other first layers except the first first layer among the multiple first layers, based on the output vector of the previous first layer and the first cascaded vector, obtain the output vectors of the other first layers, where the output vector of the last first layer is used as the first intermediate vector.
[0424] Optionally, the determination unit 4020 is specifically configured to:
[0425] Obtain the first weight matrix and the first bias vector of the other first layers except the first first layer among the multiple first layers, where the number of rows and columns of the first weight matrix is the dimension of the first cascaded vector, and the dimension of the first bias vector is the dimension of the first cascaded vector;
[0426] Based on the output vector of the previous first layer, the first cascaded vector, the first weight matrix, and the first bias vector, obtain the output vectors of the other first layers.
[0427] Optionally, the determination unit 4020 is specifically configured to:
[0428] Multiply the output vector of the previous first layer by the first weight matrix and add the first bias vector to obtain a second intermediate vector;
[0429] Multiply the first cascaded vector and the second intermediate vector bit by bit to obtain a third intermediate vector;
[0430] Add the third intermediate vector and the output vector of the previous first layer to obtain the output vectors of the other first layers.
[0431] Optionally, the first cascaded vector is obtained by cascading the first number of target object feature vectors with the same dimension;
[0432] The determination unit 4020 is specifically configured to:
[0433] Input the first cascaded vector into the first first layer, and generate an output feature vector for each target object feature vector;
[0434] Cascade the first number of output feature vectors according to the sorting of the target object feature vectors in the first cascaded vector to obtain the output vector of the first first layer;
[0435] For each target object feature vector in the first cascaded vector, based on the target object feature vector, the output vector of the previous first layer, the first weight matrix, and the first bias vector, obtain the output feature vectors of the target object feature vector in other first layers;
[0436] Cascade the first number of output feature vectors according to the sorting of the target object features in the first vector to obtain the output vectors of other first layers.
[0437] Optionally, the first weight matrix is composed of a second number of first sub - weight matrices, where the second number is equal to the square of the first number, and the number of rows and columns of each first sub - weight matrix is the dimension of the target object feature vector; the first bias vector is composed of a first number of first sub - bias vectors, and the dimension of the first sub - bias vector is equal to the dimension of the target object feature vector;
[0438] The determination unit 4020 is specifically configured to:
[0439] For each target object feature vector in the first cascaded vector, obtain a first number of target first sub - weight matrices from the second number of first sub - weight matrices, and obtain the target first sub - bias vector from the first number of first sub - bias vectors;
[0440] Based on the target object feature vector, the first number of target first sub - weight matrices, the target first sub - bias vector, and the output vector of the previous first layer, obtain the output feature vectors of the target object feature vector in other first layers.
[0441] Optionally, the determination unit 4020 is specifically configured to:
[0442] Multiply each of the first number of output feature vectors in the output vector of the previous first layer by the first number of target sub - weight matrices one by one to obtain a first number of fourth intermediate vectors;
[0443] Sum the results of multiplying each of the first number of fourth intermediate vectors by the target object feature vector bit by bit to obtain a first sum vector;
[0444] Multiply the target object feature vector by the target first sub - bias vector bit by bit to obtain a fifth intermediate vector;
[0445] Add the first sum vector, the fifth intermediate vector, and the output feature vector corresponding to the target object feature vector in the output vectors of the previous first layer to obtain the output feature vector of the target object feature vector in other first layers.
[0446] Optionally, the determining unit 4020 is specifically configured to:
[0447] Obtain the first low-rank matrix and the second low-rank matrix of other first layers except the first first layer among multiple first layers, where the number of rows of the first low-rank matrix and the second low-rank matrix is equal to the dimension of the first concatenated vector, and the number of columns is less than the first dimension of the first concatenated vector;
[0448] For each other first layer, multiply the first low-rank matrix by the transpose of the second low-rank matrix to obtain the first weight matrix.
[0449] Optionally, the pushing processing device 4000 further includes:
[0450] An input unit (not shown), configured to input multiple pushing basic features into the full cross network to obtain a full cross feature vector;
[0451] The pushing unit 4040 is specifically configured to:
[0452] Input the combined cross feature vector and the full cross feature vector into the probability prediction model to obtain a first probability of pushing the content to be pushed for the target object, and push the content to be pushed for the target object based on the first probability.
[0453] Optionally, the pushing unit 4040 is specifically configured to:
[0454] Add the combined cross feature vector and the full cross feature vector to obtain a final cross feature vector;
[0455] Based on the final cross feature vector, obtain a first probability of pushing the content to be pushed for the target object.
[0456] Optionally, the pushing unit 4040 is specifically configured to:
[0457] Obtain a first weight vector and a first bias amount, where the dimension of the first weight vector is equal to the dimension of the final cross feature vector;
[0458] Multiply the first weight vector by the transpose of the final cross feature vector and add the first bias amount to obtain a first score;
[0459] Perform exponential normalization on the first score to obtain a first probability of pushing the content to be pushed for the target object.
[0460] Optionally, the determining unit 4020 is specifically configured to:
[0461] If the target object feature is a fixed-length sequence feature, obtain the feature mapping reference table corresponding to the target object feature, and based on the feature mapping reference table, map the target object feature to the corresponding target object feature vector;
[0462] If the target object feature is a variable-length sequence feature, perform character mapping on each character in the target object feature to obtain a variable-length feature vector, and perform pooling processing on the variable-length feature vector to obtain a target object feature vector corresponding to the target object feature and having a second dimension.
[0463] Optionally, the determining unit 4020 is specifically configured to:
[0464] Vectorize at least one content feature to be pushed into at least one content feature vector to be pushed;
[0465] Cascade at least one content feature vector to be pushed and input it into a second cross network to obtain a sixth intermediate vector;
[0466] Obtain a second feature vector corresponding to each content feature to be pushed from the sixth intermediate vector.
[0467] Optionally, the first feature vector and the second feature vector have the same dimension;
[0468] The combining unit 4030 is specifically configured to:
[0469] Multiply the first feature vector and the second feature vector bit by bit to obtain a combined cross feature vector.
[0470] Optionally, the probability prediction model includes multiple first task prediction models corresponding to multiple tasks, and the first task prediction model is used to predict the probability that the target object completes the task for the content to be pushed;
[0471] The pushing unit 4040 is specifically configured to:
[0472] Input the combined cross feature vector into multiple first task prediction models to obtain multiple first task completion probabilities corresponding to multiple tasks;
[0473] Based on the multiple first task completion probabilities, obtain a first probability of pushing the content to be pushed to the target object.
[0474] Optionally, the probability prediction model includes gating nodes corresponding to multiple tasks and a second task prediction model;
[0475] The input unit (not shown) is specifically configured to:
[0476] Input multiple push base features into multiple full cross networks to obtain multiple full cross feature vectors;
[0477] The pushing unit 4040 is specifically configured to:
[0478] Input the combined cross feature vector and multiple full cross feature vectors into the gated node corresponding to the task to obtain the seventh intermediate vector corresponding to the task;
[0479] Input the seventh intermediate vector into the second task prediction model corresponding to the task to obtain the second task completion probability corresponding to the task;
[0480] Based on the second task completion probabilities corresponding to multiple tasks, obtain the first probability of pushing the content to be pushed to the target object.
[0481] Refer to Figure 41 , Figure 41 FIG. 14 is a block diagram of a part of the object terminal 140 for implementing the push processing method according to an embodiment of the present disclosure. The terminal includes: a radio frequency (RF) circuit 4110, a memory 4115, an input unit 4130, a display unit 4140, a sensor 4150, an audio circuit 4160, a wireless fidelity (WiFi) module 4170, a processor 4180, and a power supply 4190, etc. Those skilled in the art can understand that Figure 41 the structure of the object terminal 140 shown does not limit the mobile phone or computer, and may include more or fewer components than shown, or combine some components, or have different component arrangements.
[0482] The RF circuit 4110 can be used for receiving and transmitting signals during information reception or call processes. Specifically, after receiving the downlink information from the base station, it is given to the processor 4180 for processing; in addition, the uplink data designed is sent to the base station.
[0483] The memory 4115 can be used to store software programs and modules. The processor 4180 executes various functional applications and data processing of the content terminal by running the software programs and modules stored in the memory 4115.
[0484] The input unit 4130 can be used to receive input digital or character information, and generate key signal inputs related to the settings and function controls of the content terminal. Specifically, the input unit 4130 can include a touch panel 4131 and other input devices 4132.
[0485] The display unit 4140 can be used to display the input information or the provided information and various menus of the content terminal. The display unit 4140 can include a display panel 4141.
[0486] The audio circuit 4160, the speaker 4161, and the microphone 4162 can provide an audio interface.
[0487] In this embodiment, the processor 4180 included in the object terminal 140 may execute the push processing method of the previous embodiment.
[0488] The object terminal 140 of the embodiments of the present disclosure includes, but is not limited to, mobile phones, computers, intelligent voice interaction devices, intelligent home appliances, vehicle-mounted terminals, aircraft, etc. The embodiments of the present invention can be applied to various scenarios, including but not limited to content recommendation, data screening, etc.
[0489] Figure 42 It is a structural block diagram of a part of the server 110 for implementing the push processing method of the embodiments of the present disclosure. The server 110 may vary greatly due to configuration or performance differences, and may include one or more central processing units (CPUs) 4222 (for example, one or more processors) and a memory 4232, and one or more storage media 4230 (for example, one or more mass storage devices) for storing application programs 4242 or data 4244. Among them, the memory 4232 and the storage media 4230 may be transient storage or persistent storage. The program stored in the storage media 4230 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Further, the central processing unit 4222 may be configured to communicate with the storage media 4230 and execute a series of instruction operations in the storage media 4230 on the server.
[0490] The server 110 may further include one or more power supplies 4226, one or more wired or wireless network interfaces 4250, one or more input / output interfaces 4258, and / or, one or more operating systems 4241, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.
[0491] The central processing unit 4222 in the server 110 may be used to execute the push processing method of the embodiments of the present disclosure.
[0492] The embodiments of the present disclosure further provide a computer-readable storage medium, which is used to store program codes, and the program codes are used to execute the push processing methods of the foregoing embodiments.
[0493] The embodiments of the present disclosure further provide a computer program product, which includes a computer program. The processor of the computer device reads and executes the computer program, so that the computer device executes the push processing method described above.
[0494] In the description of the present disclosure and the above-mentioned accompanying drawings, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar content and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "comprise" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0495] It should be understood that in the present disclosure, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated content and indicates that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the content before and after is an "or" relationship. "At least one (one)" or its similar expression below refers to any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0496] It should be understood that in the description of the embodiments of the present disclosure, the meaning of "a plurality (or multiple items)" is more than two. Understandings such as greater than, less than, exceeding, etc. do not include the number itself, and understandings such as above, below, within, etc. include the number itself.
[0497] In the embodiments of the present disclosure, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the function of that module or unit.
[0498] In several embodiments provided by the present disclosure, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.
[0499] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0500] In addition, in each embodiment of the present disclosure, the functional units can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0501] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present disclosure. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0502] It should also be understood that the various embodiments provided by the embodiments of the present disclosure can be combined arbitrarily to achieve different technical effects.
[0503] The above is a specific description of the embodiments of the present disclosure. However, the present disclosure is not limited to the above embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present disclosure, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present disclosure.
Claims
1. A push processing method, characterized in that, Including: Obtain a plurality of push basic features, where the plurality of push basic features include at least one target object feature and at least one content feature to be pushed; Based on at least one of the target object features, determine a first feature vector corresponding to each target object feature, and based on at least one of the content features to be pushed, determine a second feature vector corresponding to each content feature to be pushed; For each of the first feature vectors, combine the first feature vector with each of the second feature vectors respectively, and input each combination of the first feature vector and the second feature vector into a combined cross network to obtain a combined cross feature vector; Input the combined cross feature vector into a probability prediction model to obtain a first probability of pushing the content to be pushed for the target object, and push the content to be pushed for the target object based on the first probability.
2. The push processing method according to claim 1, wherein Including: The determining, based on at least one of the target object features, a first feature vector corresponding to each target object feature includes: Vectorize at least one of the target object features into at least one target object feature vector; Concatenate at least one of the target object feature vectors and input them into a first cross network to obtain a first intermediate vector; From the first intermediate vector, obtain the first feature vector corresponding to each target object feature.
3. The push processing method according to claim 2, wherein The vectorizing at least one of the target object features into at least one target object feature vector includes: If the target object feature is a fixed-length sequence feature, obtain a feature mapping reference table corresponding to the target object feature, and based on the feature mapping reference table, map the target object feature to the corresponding target object feature vector; If the target object feature is a variable-length sequence feature, perform character mapping on each character in the target object feature to obtain a variable-length feature vector, and perform pooling processing on the variable-length feature vector to obtain a target object feature vector corresponding to the target object feature and having a second dimension.
4. The push processing method according to claim 2, wherein Including: The first cross network includes a plurality of first layers; The concatenating at least one of the target object feature vectors and inputting them into a first cross network to obtain a first intermediate vector includes: Concatenate at least one of the target object feature vectors to obtain a first concatenated vector; Input the first concatenated vector into the first of the first layers to obtain an output vector of the first of the first layers; Through the other first layers among the plurality of first layers except the first of the first layers, based on the output vector of the previous first layer and the first concatenated vector, obtain the output vectors of the other first layers, where the output vector of the last of the first layers is used as the first intermediate vector.
5. The push processing method according to claim 4, wherein The obtaining, through the other first layers among the plurality of first layers except the first of the first layers, the output vectors of the other first layers based on the output vector of the previous first layer and the first concatenated vector includes: Obtain the first weight matrix and the first bias vector of other first layers among the multiple first layers except the first first layer, where the number of rows and columns of the first weight matrix is the dimension of the first concatenated vector, and the dimension of the first bias vector is the dimension of the first concatenated vector; Based on the output vector of the previous first layer, the first concatenated vector, the first weight matrix, and the first bias vector, obtain the output vectors of the other first layers.
6. The push processing method according to claim 5, wherein The obtaining the output vectors of the other first layers based on the output vector of the previous first layer, the first concatenated vector, the first weight matrix, and the first bias vector includes: Multiply the output vector of the previous first layer by the first weight matrix and add the first bias vector to obtain a second intermediate vector; Multiply the first concatenated vector and the second intermediate vector bit by bit to obtain a third intermediate vector; Add the third intermediate vector to the output vector of the previous first layer to obtain the output vectors of the other first layers.
7. The push processing method according to claim 5, wherein The first concatenated vector is obtained by concatenating the first number of target object feature vectors with the same dimension; Input the first concatenated vector into the first first layer to obtain the output vector of the first first layer, including: Input the first concatenated vector into the first first layer, and for each target object feature vector, generate an output feature vector; Concatenate the first number of output feature vectors according to the sorting of the target object feature vectors in the first concatenated vector to obtain the output vector of the first first layer; The obtaining the output vectors of the other first layers based on the output vector of the previous first layer, the first concatenated vector, the first weight matrix, and the first bias vector includes: For each target object feature vector in the first concatenated vector, based on the target object feature vector, the output vector of the previous first layer, the first weight matrix, and the first bias vector, obtain the output feature vector of the target object feature vector in the other first layers; Concatenate the first number of output feature vectors according to the sorting of the target object features in the first vector to obtain the output vectors of the other first layers.
8. The push processing method according to claim 7, wherein The first weight matrix is composed of the second number of first sub-weight matrices, where the second number is equal to the square of the first number, and the number of rows and columns of each first sub-weight matrix is the dimension of the target object feature vector; the first bias vector is composed of the first number of first sub-bias vectors, and the dimension of the first sub-bias vector is equal to the dimension of the target object feature vector; The obtaining the output feature vector of the target object feature vector in the other first layers based on the target object feature vector, the output vector of the previous first layer, the first weight matrix, and the first bias vector for each target object feature vector in the first concatenated vector includes: For each of the target object feature vectors in the first cascaded vector, obtain the first number of target first sub-weight matrices from the second number of the first sub-weight matrices, and obtain a target first sub-bias vector from the first number of the first sub-bias vectors; Based on the target object feature vector, the first number of the target first sub-weight matrices, the target first sub-bias vector, and the output vector of the previous first layer, obtain the output feature vector of the target object feature vector in other first layers.
9. The push processing method according to claim 8, wherein The obtaining the output feature vector of the target object feature vector in other first layers based on the target object feature vector, the first number of the target first sub-weight matrices, the target first sub-bias vector, and the output vector of the previous first layer includes: Multiply the first number of output feature vectors in the output vector of the previous first layer by the first number of the target sub-weight matrices one by one to obtain the first number of fourth intermediate vectors; Sum the results of multiplying the first number of the fourth intermediate vectors by the target object feature vector bit by bit one by one to obtain a first sum vector; Multiply the target object feature vector by the target first sub-bias vector bit by bit to obtain a fifth intermediate vector; Add the first sum vector, the fifth intermediate vector, and the output feature vector corresponding to the target object feature vector in the output vector of the previous first layer to obtain the output feature vector of the target object feature vector in other first layers.
10. The push processing method according to claim 5, characterized in that, The obtaining the first weight matrices of other first layers except the first first layer among the multiple first layers includes: Obtain a first low-rank matrix and a second low-rank matrix of other first layers except the first first layer among the multiple first layers, where the number of rows of the first low-rank matrix and the second low-rank matrix is equal to the dimension of the first cascaded vector, and the number of columns is less than the first dimension of the first cascaded vector; For each other first layer, multiply the first low-rank matrix by the transpose of the second low-rank matrix to obtain the first weight matrix.
11. The push processing method according to claim 1, wherein After, for each of the first feature vectors, combining the first feature vector with each of the second feature vectors respectively, and inputting each combination of the first feature vector and the second feature vector into a combined cross network to obtain combined cross feature vectors, the method further includes: Input the multiple push base features into a full cross network to obtain full cross feature vectors; The inputting the combined cross feature vectors into a probability prediction model to obtain a first probability of pushing the content to be pushed for the target object, and pushing the content to be pushed for the target object based on the first probability includes: Input the combined cross feature vectors and the full cross feature vectors into a probability prediction model to obtain a first probability of pushing the content to be pushed for the target object, and push the content to be pushed for the target object based on the first probability.
12. The push processing method according to claim 11, wherein Inputting the combined cross feature vector and the full cross feature vector into a probability prediction model to obtain a first probability of pushing the content to be pushed to the target object, includes: Adding the combined cross feature vector and the full cross feature vector to obtain a final cross feature vector; Based on the final cross feature vector, obtaining a first probability of pushing the content to be pushed to the target object.
13. The push processing method according to claim 12, wherein The obtaining, based on the final cross feature vector, of a first probability of pushing the content to be pushed to the target object, includes: Obtaining a first weight vector and a first bias, where the dimension of the first weight vector is equal to the dimension of the final cross feature vector; Multiplying the first weight vector by the transpose of the final cross feature vector and adding the first bias to obtain a first score; Performing exponential normalization on the first score to obtain a first probability of pushing the content to be pushed to the target object.
14. The push processing method according to claim 1, wherein The determining, based on at least one of the content features to be pushed, of a second feature vector corresponding to each of the content features to be pushed, includes: Vectorizing at least one of the content features to be pushed into at least one content feature vector to be pushed; Cascading at least one of the content feature vectors to be pushed and inputting them into a second cross network to obtain a sixth intermediate vector; From the sixth intermediate vector, obtaining the second feature vector corresponding to each of the content features to be pushed.
15. A push processing device, characterized in that, Includes: An obtaining unit, configured to obtain a plurality of push basic features, where the plurality of push basic features include at least one target object feature and at least one content feature to be pushed; A determining unit, configured to determine, based on at least one of the target object features, a first feature vector corresponding to each of the target object features, and determine, based on at least one of the content features to be pushed, a second feature vector corresponding to each of the content features to be pushed; A combining unit, configured to, for each of the first feature vectors, combine the first feature vector with each of the second feature vectors respectively, input each combination of the first feature vector and the second feature vector into a combined cross network to obtain a combined cross feature vector; A pushing unit, configured to input the combined cross feature vector into a probability prediction model to obtain a first probability of pushing the content to be pushed to the target object, and push the content to be pushed to the target object based on the first probability.
16. An electronic device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the push processing method according to any one of claims 1 to 14.
17. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the push processing method according to any one of claims 1 to 14.
18. A computer program product, the computer program product includes a computer program, the computer program is read and executed by a processor of a computer device, so that the computer device executes the push processing method according to any one of claims 1 to 14.