Method, device, equipment, medium and program product for recall processing of information to be pushed

By calculating and correcting the interest scores of information to be pushed, and combining the recall requirements, filtering out information that does not need to be pushed, solving the problem of insufficient personalization and precision in the existing technology, and achieving a more efficient user experience.

CN114547396BActive Publication Date: 2025-06-24ECARX (HUBEI) TECHCO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210193411.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-28
Publication Date
2025-06-24
Estimated Expiration
2042-02-28

AI Technical Summary

Technical Problem

When the prior art pushes information to users, it fails to effectively perform personalized and accurate secondary screening, resulting in content that is of interest to users being pushed and affecting the user experience.

Method used

By obtaining the information feature table and the user interest feature table, the interest scores of the information to be pushed are calculated, and the scores are corrected using the time attenuation model. Finally, information that does not need to be pushed is filtered out based on the revised score and recall requirements.

Benefits of technology

It realizes accurate recall of push information, reduces uninterested information pushed to users, improves the personalization and accuracy of information push, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114547396B_ABST
    Figure CN114547396B_ABST
Patent Text Reader

Abstract

The present application provides a method, apparatus, device, medium and program product for recall processing of information to be pushed. By obtaining an information feature table and a user interest feature table, then determining a first interest score table of multiple pieces of information to be pushed according to the information feature table and the interest feature table, and then using a preset time decay model to correct the first interest score table according to the time information in the multiple pieces of information to be pushed and the target push time to determine a second interest score table, and finally determining recall information that does not need to be pushed to the user from each piece of information to be pushed according to the second interest score table and a preset recall requirement. The technical problem of how to accurately recall information to be pushed is solved, a large amount of information that the user is not interested in is avoided from being pushed to the user, and the technical effect of personalized and accurate push for each user is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer data processing, and particularly to a method, apparatus, device, medium and program product for recall processing of information to be pushed. Background Art

[0002] With the continuous development of Internet technology, pushing information to users on mobile terminals or in-vehicle terminals has become an important means of information dissemination.

[0003] Currently, with the continuous enrichment of various goods and services, the amount of information to be pushed to users is increasing. However, due to the increasing demand for information personalization by users, when the existing technology pushes information to users, it generally directly pushes the information after screening out the information to be pushed, without performing a secondary screening for personalization and accuracy.

[0004] This results in a large number of pushed messages containing a lot of content that users are not interested in. That is, when the information coverage is too large, users will think that the information push is not accurate enough. Another situation is that if the amount of pushed information is huge, but its coverage is too small, users will also think that the information push is not personalized enough, only pushing common hot content. These will seriously affect the user experience.

[0005] Therefore, how to accurately recall a large number of information to be pushed before pushing it to users has become an urgent technical problem to be solved. Summary of the Invention

[0006] This application provides a method, apparatus, device, medium and program product for recall processing of information to be pushed, so as to solve the technical problem of how to accurately recall the information to be pushed.

[0007] In a first aspect, this application provides a method for recall processing of information to be pushed, including:

[0008] Obtain an information feature table and an interest feature table of a user, where the information feature table is used to characterize the feature attributes corresponding to multiple pieces of information to be pushed, and the interest feature table is used to characterize the user's interest preferences for each piece of information in the information feature table;

[0009] Determine a first interest score table for multiple pieces of information to be pushed according to the information feature table and the interest feature table;

[0010] Perform time decay correction on the first interest score table according to the time record in each piece of information to be pushed and the target push time, so as to determine a second interest score table;

[0011] Determine each recall information that does not need to be pushed to the user from each piece of information to be pushed according to the second interest score table and a preset recall requirement.

[0012] In a possible design, according to the time records in each piece of information to be pushed and the target push time, perform a time decay correction on the first interest score table to determine the second interest score table, including:

[0013] Respectively determine the time differences between the time records of each piece of information to be pushed and the target push time;

[0014] Input each time difference and the preset time sensitivity coefficient into the preset time decay model for calculation to determine multiple time decay parameters;

[0015] Multiply each time decay parameter by the corresponding interest score in the first interest score table to determine the second interest score table.

[0016] In a possible design, obtain the interest feature table of the user, including:

[0017] Obtain multiple user characteristics of the user;

[0018] Through neural network inference on multiple user characteristics, enable the neural network to determine the mapping relationship between each user characteristic and the preset information classification and grouping rules according to the requirements of the word tree classification structure, so as to determine the interest feature table, where each information classification in the preset information classification and grouping rules corresponds to at least one information group, and the information classifications and the information groups are linearly independent of each other.

[0019] In a possible design, through neural network inference on multiple user characteristics, enable the neural network to determine the mapping relationship between each user characteristic and the information classification and grouping rules, including:

[0020] Use the feature extraction layer in the neural network to convert multiple user characteristics of the user into user feature information, and the user feature information represents multiple user characteristics in the form of a vector or a matrix;

[0021] Input the user feature information into the information classification fully connected layer to obtain the first mapping result of the user feature information mapped to each information classification;

[0022] Input the first mapping result into the information grouping fully connected layer to determine the second mapping result of the first mapping result mapped to each information group, where the information classification fully connected layer and the information grouping fully connected layer respectively establish multiple tensor parameters during the processing according to the requirements of the word tree classification structure, and the tensor parameters are used to represent the corresponding relationship between the category features of the information classification fully connected layer and the group features of the information grouping fully connected layer;

[0023] Use the decision layer in the neural network to process the second mapping result to determine the interest feature table.

[0024] In a possible design, the method further includes: obtaining a neural network for inferring a mapping relationship, including:

[0025] Obtaining sample user features;

[0026] Using a feature extraction layer in the neural network to extract sample user feature information from the sample user features, where the sample user feature information represents multiple sample user features in the form of a vector or a matrix;

[0027] Inputting the sample user feature information into a full connection layer for information classification and a full connection layer for information grouping for processing to determine a first sample mapping result and a second sample mapping result. The first sample mapping result is used to represent the mapping relationship of the sample user feature information mapped to each information classification label space through a first tensor parameter, and the second sample mapping result is used to represent the mapping relationship of the first sample mapping result mapped to each information grouping label space through a second tensor parameter; wherein, the first tensor parameter includes multiple first sub-tensor parameters, each first sub-tensor parameter is used to represent the category features in each information classification, and the second tensor parameter includes multiple second sub-tensor parameters, each second sub-tensor parameter is used to represent the group features in each information grouping;

[0028] Using a decision function to process the processing results of the full connection layer for information classification and the full connection layer for information grouping to determine a sample classification result, and adjusting various parameters in the neural network according to the sample classification result.

[0029] In a possible design, using a feature extraction layer in the neural network to extract feature information from the sample user feature vector includes:

[0030] Using multiple preset convolution kernels and a preset activation function to perform convolution processing and activation processing on the sample user feature vector to determine multiple first output vectors, where the first output vector is the output vector of the convolution layer in the neural network model; wherein, each first output vector corresponds to a preset convolution kernel, the convolution processing is used to amplify and / or extract the target feature values in the sample user feature vector, and the activation processing is used to increase the non-linearity of the output result of the convolution processing;

[0031] Performing pooling processing on the multiple first output vectors to determine multiple second output vectors of different scales, where the second output vector is the output vector of the pooling layer in the neural network model;

[0032] Fusing the multiple second output vectors of different scales into a third output vector of a preset scale as the sample user feature information.

[0033] In a possible design, adjusting various parameters in the neural network according to the classification result includes:

[0034] Calculate the loss function according to the sample classification result, the first tensor parameter, and the second tensor parameter;

[0035] Perform cyclic training to adjust the parameters in the neural network so that the value of the loss function tends to zero, and obtain linearly independent first tensor parameters and second tensor parameters.

[0036] Optionally, after making the value of the loss function tend to zero, it further includes:

[0037] Adjust the first tensor parameter to make each first subtensor parameter linearly independent;

[0038] And / or,

[0039] Adjust the second tensor parameter to make each second subtensor parameter linearly independent;

[0040] Wherein, when the first inner product between each first subtensor parameter is zero, each first subtensor parameter is linearly independent, and when the second inner product between the second subtensor parameters is zero, each second subtensor parameter is linearly independent.

[0041] In a second aspect, the present application provides a device for recall processing of information to be pushed, including:

[0042] An acquisition module, configured to acquire an information feature table and an interest feature table of a user, where the information feature table is used to characterize the feature attributes corresponding to multiple pieces of information to be pushed, and the interest feature table is used to characterize the user's interest preferences for each piece of information in the information feature table;

[0043] A processing module, configured to:

[0044] Determine a first interest score table for multiple pieces of information to be pushed according to the information feature table and the interest feature table;

[0045] Perform time decay correction on the first interest score table according to the time records in each piece of information to be pushed and the target push time to determine a second interest score table;

[0046] Determine each piece of recall information that does not need to be pushed to the user from each piece of information to be pushed according to the second interest score table and the preset recall requirement.

[0047] In a possible design, the processing module is configured to:

[0048] Determine the time difference between each time record and the target push time respectively;

[0049] Input each time difference and the preset time sensitivity coefficient into a preset time decay model for calculation to determine multiple time decay parameters;

[0050] Multiply each time decay parameter by the corresponding interest score in the first interest score table to determine the second interest score table.

[0051] In a possible design, the acquisition module is further configured to acquire multiple user characteristics of the user;

[0052] The processing module is further configured to:

[0053] Through neural network inference on multiple user characteristics, enable the neural network to determine the mapping relationship between each user characteristic and the preset information classification and grouping rules according to the requirements of the word tree classification structure, so as to determine the interest feature table, where each information classification in the preset information classification and grouping rules corresponds to at least one information group, and the information classifications and the information groups are linearly independent of each other.

[0054] In a possible design, the processing module is further configured to:

[0055] Use the feature extraction layer in the neural network to convert multiple user characteristics of the user into user feature information, and the user feature information represents multiple user characteristics in the form of a vector or a matrix;

[0056] Input the user feature information into the information classification fully connected layer to obtain the first mapping result of the user feature information mapped to each information classification;

[0057] Input the first mapping result into the information grouping fully connected layer to determine the second mapping result of the first mapping result mapped to each information group, where the information classification fully connected layer and the information grouping fully connected layer respectively establish multiple tensor parameters during the processing according to the requirements of the word tree classification structure, and the tensor parameters are used to represent the corresponding relationship between the category features of the information classification fully connected layer and the group features of the information grouping fully connected layer;

[0058] Use the decision layer in the neural network to process the second mapping result to determine the interest feature table.

[0059] In a possible design, the acquisition module is further configured to: acquire sample user characteristics;

[0060] The processing module is further configured to:

[0061] Use the feature extraction layer in the neural network to extract the sample user feature information in the sample user characteristics, and the sample user feature information represents multiple sample user characteristics in the form of a vector or a matrix;

[0062] Input the sample user feature information into the information classification fully-connected layer and the information grouping fully-connected layer for processing to determine the first sample mapping result and the second sample mapping result. The first sample mapping result is used to represent the mapping relationship of the sample user feature information mapped to each information classification label space through the first tensor parameter, and the second sample mapping result is used to represent the mapping relationship of the first sample mapping result mapped to each information grouping label space through the second tensor parameter. Among them, the first tensor parameter includes multiple first sub-tensor parameters, and each first sub-tensor parameter is used to represent the category feature in each information classification. The second tensor parameter includes multiple second sub-tensor parameters, and each second sub-tensor parameter is used to represent the group feature in each information grouping.

[0063] Use the decision function to process the processing results of the information classification fully-connected layer and the information grouping fully-connected layer to determine the sample classification result, and adjust the various parameters in the neural network according to the sample classification result.

[0064] In a possible design, the processing module is further configured to:

[0065] Use multiple preset convolution kernels and a preset activation function to perform convolution processing and activation processing on the sample user feature vector to determine multiple first output vectors. The first output vector is the output vector of the convolution layer in the neural network model. Among them, each first output vector corresponds to a preset convolution kernel. The convolution processing is used to amplify and / or extract the target feature values in the sample user feature vector, and the activation processing is used to increase the non-linearity of the output result of the convolution processing.

[0066] Perform pooling processing on the multiple first output vectors to determine multiple second output vectors of different scales. The second output vector is the output vector of the pooling layer in the neural network model.

[0067] Fuse the multiple second output vectors of different scales into a third output vector of a preset scale as the sample user feature information.

[0068] In a possible design, the processing module is further configured to:

[0069] Calculate the loss function according to the sample classification result and the first tensor parameter and the second tensor parameter.

[0070] Perform cyclic training to adjust the various parameters in the neural network to make the value of the loss function tend to zero, and obtain linearly independent first tensor parameters and second tensor parameters.

[0071] Optionally, the processing module is further configured to:

[0072] Adjust the first tensor parameter to make each first sub-tensor parameter linearly independent.

[0073] And / or

[0074] Adjust the second tensor parameters to make each second sub - tensor parameter linearly independent;

[0075] Among them, when the first inner product between each first sub - tensor parameter is zero, each first sub - tensor parameter is linearly independent, and when the second inner product between the second sub - tensor parameters is zero, each second sub - tensor parameter is linearly independent.

[0076] In a third aspect, the present application provides an electronic device, including:

[0077] A memory for storing program instructions;

[0078] A processor for calling and executing the program instructions in the memory and performing any possible recall processing method for the information to be pushed provided in the first aspect.

[0079] In a fourth aspect, the present application provides a storage medium, in which a computer program is stored, and the computer program is used to execute any possible recall processing method for the information to be pushed provided in the first aspect.

[0080] In a fifth aspect, the present application further provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements any possible recall processing system method for the information to be pushed provided in the first aspect.

[0081] The present application provides a method, device, equipment, medium and program product for recall processing of information to be pushed. By obtaining an information feature table and a user's interest feature table, then determining a first interest score table for multiple pieces of information to be pushed according to the information feature table and the interest feature table, and then using a preset time decay model to correct the first interest score table according to the time information in the multiple pieces of information to be pushed and the target push time to determine a second interest score table, and finally determining the recall information that does not need to be pushed to the user from each piece of information to be pushed according to the second interest score table and the preset recall requirement. It solves the technical problem of how to accurately recall the information to be pushed, avoids pushing a large amount of information that the user is not interested in, and achieves the technical effect of personalized and accurate pushing for each user. BRIEF DESCRIPTION OF THE DRAWINGS

[0082] The drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0083] Figure 1 It is a schematic flowchart of a method for recall processing of information to be pushed provided by an embodiment of the present application;

[0084] Figure 2A schematic flowchart of another method for recalling and processing information to be pushed provided for the implementation of this application;

[0085] Figure 3 A schematic diagram of the function image of the relu activation function;

[0086] Figure 4 A schematic diagram of a word tree structure provided by this application;

[0087] Figure 5 A schematic structural diagram of a device for recalling and processing information to be pushed provided for the embodiments of this application;

[0088] Figure 6 A schematic structural diagram of an electronic device provided by this application.

[0089] Through the above-mentioned drawings, specific embodiments of this application have been shown, and there will be more detailed descriptions hereinafter. These drawings and textual descriptions are not intended to limit the scope of the concept of this application in any way, but to illustrate the concept of this application to those skilled in the art by referring to specific embodiments. Detailed implementation manners

[0090] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts, including but not limited to combinations of multiple embodiments, fall within the scope of protection of this application.

[0091] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of this application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0092] In an information recommendation system, a single information recommendation often provides a vast amount of information to be pushed, such as a large number of candidate items and / or candidate services. This requires recall processing of the information to be pushed, quickly screening out a small number of items and / or services that users may be interested in, for pre-screening in the subsequent fine-ranking stage and reducing the amount of data to be processed by subsequent algorithms.

[0093] It should be noted that the amount of data to be processed in the recall stage of candidate information is very large, and high processing speed is required.

[0094] Existing solutions are mainly divided into content-based recall and collaborative filtering-based recall:

[0095] 1) Content-based recall:

[0096] Recall is performed based on the similarity of item content. For example, if user A likes item a, and items b, c, d are highly similar to item a, then b, c, d are recommended to user A. This method does not require using data of other users and only uses the similarity relationship between items for matching. Since user characteristics cannot be utilized, the recall effect is limited.

[0097] 2) Collaborative filtering:

[0098] The correlation between users and items is used for recommendation to improve the scalability and robustness of model recommendation. Collaborative filtering recommendation can be divided into traditional methods and deep learning-based methods.

[0099] Traditional methods include: various calculation methods based on the user-item id co-occurrence matrix or rating matrix, such as the nearest neighbor method, matrix factorization, FM (Factorization Machine) algorithm, FFM (Field-aware Factorization Machines) algorithm, etc. This method generally only uses the user and item ids as input information and is relatively fast, but due to the loss of user and item information, the accuracy and reliability are relatively low.

[0100] Deep learning methods include: the twin tower model, youtube DNN (Deep Neural Networks), etc. The principle is to map user and item information into dense vector embeddings, measure the closeness of the connection between users and items by calculating the distance between embeddings (Euclidean distance, cosine angle, etc.), and recommend the top K items with higher closeness to users. The twin tower model needs to train each group of user-item combinations, save the embeddings of users and items, calculate the user's embedding U user embedding mapping vector during inference, and then calculate the distance between this vector and the embedding I item embedding mapping vector of all items in the library. This method requires a lot of space to store the embeddings of massive items, and the embeddings need to be calculated one by one for new users, so the cold start calculation is large. Youtube DNN transforms the recall problem into a classification problem of user input and item output. It uses the IDs of a large number of items to be selected as the output tensor for training. During inference, the parameters of the fully connected layer before the output layer are used as the user embedding mapping vector, and the connection parameter W between this layer and the output layer is used as the item embedding mapping vector of all items. Through vector dot product calculation, a series of cosine scores are obtained, and the items to be recommended are selected according to the size of the values. This method model needs to build a huge output layer. After adding new items, the model needs to be rebuilt for training. In addition, it is not interpretable enough, and it is difficult to establish the potential mapping relationship between users and item features.

[0101] In summary, the existing solutions have the following shortcomings and problems:

[0102] 1. Traditional methods lose user information and item attributes, and have low accuracy and reliability;

[0103] 2. The dual-tower model is not cold-start friendly, requires large amounts of computation, and requires a large amount of space to store the embeddings of a large number of items;

[0104] 3. YouTube DNN does not consider the characteristics of the items themselves, making it difficult to explore the potential mapping relationship between user and item characteristics and user preferences. In addition, the number of tensors in the output layer is too large, making the model too bloated.

[0105] In order to solve the above technical problems, the invention concept of this application is:

[0106] By using the powerful feature expression and integration capabilities of neural networks, we can establish connections between user and item features and mine key features that users care about. This approach is highly explanatory and can easily solve the cold start problem for new users.

[0107] 1. Establish a deep neural network model between user and item features, mine the item features that users care about and like, and recall a large number of items based on the users' interests;

[0108] 2. Adopt a word tree hierarchical structure, consider the internal linear independence between categories, make the modeling more accurate and more in line with the actual situation; construct a time decay coefficient, and incorporate the time information of items into the recall processing flow, that is, multiply the calculated item score matrix by the time decay coefficient.

[0109] 3. Facilitate the recommendation of new items to users.

[0110] Before specifically introducing the recall processing method for the information to be pushed provided by this application, it is necessary to explain the terms involved in this application:

[0111] The "information" recorded below includes: item information, service information, multimedia information, etc.

[0112] Both the Onehot one-hot encoding method and the Mutihot multi-hot encoding method are ways to vectorize features. Whether it is a label or the feature value of an attribute, it can be processed in this way. For example, for humans, the gender attribute includes: male and female, and the personality attribute includes: optimistic, pessimistic, kind, numb. If a person is male and has an optimistic and kind personality, then this person can be represented by vectors as: gender: [1, 0]; personality: [1, 0, 1, 0]. In this way, the vectorization of the gender attribute is the Onehot method, and the personality is the Mutihot method.

[0113] The following will introduce in detail how this application realizes the transmission of information through touch.

[0114] Figure 1 It is a schematic flowchart of a recall processing method for the information to be pushed provided by an embodiment of this application. As Figure 1 shown, the specific steps of this recall processing method for the information to be pushed include:

[0115] S101. Obtain an information feature table and a user interest feature table.

[0116] In this step, the information feature table is used to represent the feature attributes corresponding to multiple pieces of information to be pushed, and the interest feature table is used to represent the user's interest preferences for each piece of information in the information feature table.

[0117] In this embodiment, for the sake of easy understanding, the following takes the information to be pushed as various items as an example to illustrate the information feature table. All the information to be pushed is statistically processed in a certain statistical format to obtain the information feature table shown in Table 1:

[0118]

[0119] Table 1

[0120] Then, according to the table header, the information feature table is represented by a matrix in Multihot form, as shown in formula (*):

[0121]

[0122] Specifically, in Table 1, each row is the id of an item and the feature keywords, i.e., feature attributes. If it contains the keywords in the list header, the corresponding position is 1, otherwise it is 0. For example, A1236 male football star is represented as [0, 1, 0, ……, 1].

[0123] Similarly, the user's interest feature table can also be converted into a corresponding matrix using the same principle.

[0124] S102. Determine a first interest score table for multiple to-be-pushed messages according to the information feature table and the interest feature table.

[0125] In this embodiment, specifically, Table 1 forms the Multihot matrix of information features, i.e., the matrix shown in formula (*). The product of this matrix and the transpose of the matrix corresponding to the user's interest feature table gives the vector group of the user's interest scores for each item, such as [0.1, 0.3, 0.2, ……]. The combination of each interest score vector is the first interest score table expressed in matrix form.

[0126] In a possible design, in order to avoid excessive memory occupation when processing a large amount of information data, it can be calculated in blocks. For example, the feature matrix is divided into blocks by rows, and the product of each block and the transpose of the interest score vector is calculated. After the calculation is completed, they are spliced, such as: [0.22, 0.43, 0.31], [0.12, 0.32, 0.28], ……, and after the calculation, it is spliced into [0.22, 0.43, 0.31, 0.12, 0.32 ……].

[0127] S103. Perform time decay correction on the first interest score table according to the time records in each to-be-pushed message and the target push time to determine a second interest score table.

[0128] The purpose of this step is to model time to ensure that the latest items have more advantages in recommendation than the old items. The time record is also called a timestamp, which is the time identifier recorded when the to-be-pushed message is generated. According to this time record, the newness and oldness of the information or items can be distinguished.

[0129] In this step, it specifically includes:

[0130] Respectively determine the time difference between the time record of each to-be-pushed message and the target push time;

[0131] Input each time difference and the preset time sensitivity coefficient into a preset time decay model to determine multiple time decay parameters;

[0132] Multiply each time decay parameter by the corresponding interest score in the first interest score table to determine the second interest score table.

[0133] For the preset time decay model, one possible implementation is shown in formula (1):

[0134]

[0135] where Tap is the time decay parameter, alpha is the time sensitivity coefficient, and td is the time difference.

[0136] According to the time difference corresponding to each information to be pushed, calculate the corresponding time decay parameter, and multiply the time decay parameter corresponding to the item by each term in the matrix of the first interest score table, then the item score vector considering the time difference can be obtained, that is, the second interest score table expressed in matrix or vector form.

[0137] S104. Determine each recall information that does not need to be pushed to the user from each information to be pushed according to the second interest score table and the preset recall requirement.

[0138] In this step, determine the information to be pushed corresponding to the interest score that meets the preset recall requirement in the second interest score table as the recall information. In this way, the amount of information pushed to the user can be reduced, thus achieving the technical effect of accurate pushing.

[0139] Specifically, the interest scores in the second interest score table can be arranged in ascending order. Assuming that the number K of recall information is specified in the preset recall requirement, then determine the information to be pushed corresponding to the top K interest scores as the recall information.

[0140] In a possible design, the information to be pushed includes: item information, audio information, web page information, video information, picture information, etc.

[0141] After recalling the information to be pushed, what remains are the target information with a relatively high probability of being of interest to the user and a coverage that is neither too large nor too small. In this way, it pre-screens for the subsequent fine-rank stage of the target information on the user side, reduces the amount of data to be processed by the subsequent display rendering algorithm, reduces the lag of information pushing, and improves the user experience.

[0142] This embodiment provides a method for recall processing of information to be pushed. By obtaining an information feature table and a user interest feature table, then determining a first interest score table for multiple pieces of information to be pushed according to the information feature table and the interest feature table, and then using a preset time decay model to correct the first interest score table according to the time information in the multiple pieces of information to be pushed and the target push time to determine a second interest score table. Finally, according to the second interest score table and the preset recall requirements, recall information that does not need to be pushed to the user is determined from each piece of information to be pushed. This solves the technical problem of how to accurately recall information to be pushed, avoids pushing a large amount of information that the user is not interested in, and achieves the technical effect of personalized and accurate pushing for each user.

[0143] Figure 2 It is a schematic flowchart of another method for recall processing of information to be pushed provided by the implementation of this application. As Figure 2 shown, this embodiment of the application constructs a tree-shaped multi-label classification structure, so it is simply called DTN (Deep Tree Net). In this embodiment, the recall processing (matching) is divided into two stages:

[0144] Training stage: Taking user features as input, through the embedding layer and convolutional and pooling layers, they are transformed into a series of features. Then, according to the number of categories N involved in the information object, N category tensors are constructed. These category tensors are guaranteed to be linearly independent through the product loss function. Each category tensor is connected to the feature keywords F1 to Fn of the information object of this category. After fusing all the feature keywords together to form an output, through training, this output is the degree of interest of the user in each feature of the information object.

[0145] Inference stage: For the input user features, through model inference, the corresponding information object feature interest vector can be obtained. Dot multiply this vector with the matrix formed by the Multihot multi-hot vectors of all information objects in the library, and then multiply the result by the time decay coefficient. Sort the elements in the obtained score vector, and the information objects corresponding to the top K elements are the K information objects that are expected to be recalled.

[0146] The specific steps of this method for recall processing of information to be pushed include:

[0147] First, construct and train the neural network:

[0148] S201: Obtain sample user features.

[0149] In this step, the sample user features are the neural network training data confirmed through manual screening. The sample user features include: age, place of residence, gender, occupation, etc.

[0150] S202. Extract the sample user feature information in the sample user features using the feature extraction layer in the neural network.

[0151] In this step, the sample user feature information is embodied in the form of a vector or a matrix and is used to characterize multiple sample user features. The feature extraction layer includes: an input layer, an embedding layer, a convolutional layer, a pooling layer, and a first fusion layer.

[0152] Specifically, the input layer (also called the Onehot conversion layer) in the neural network converts the sample user features into digital feature vectors through the one-hot encoding method and the embedding mapping tool. That is, all elements contained in the input are all subjected to Onehot conversion to generate one-hot vectors for subsequent algorithm calculations. Onehot conversion example: Suppose the input is [1, 2, 3], after one-hot conversion, it is [1, 0, 0], [0, 1, 0], [0, 0, 1].

[0153] For ease of understanding, a specific example of Table 2 is given:

[0154] User Gender Age Occupation A Male 25 Teacher B Male 20 Lawyer C Female 10 Student D Female 56 Judge

[0155] Table 2

[0156] Then for user A, the one-hot vector obtained after one-hot conversion of gender can be expressed as [1, 0, 0, 0, 0], the one-hot vector obtained after one-hot conversion of age can be expressed as [0, 1, 0, 0, 0], and the one-hot vector obtained after one-hot conversion of occupation can be expressed as [0, 0, 1, 0, 0].

[0157] It should be noted that the total number of elements in the one-hot vector is determined by the types of all feature attributes, and for ease of merging, the scale size of each one-hot vector, that is, the total number of elements, is the same. In Table 2, the scale size of each one-hot vector is 5.

[0158] Optionally, it is also possible to obtain several recent behaviors of the user as user features. Generally speaking, it includes some static features of the user (such as occupation, age, gender, etc.) and dynamic features (the current behaviors have an impact on recommendations).

[0159] In this embodiment, after encoding multiple user features in the input layer of the preset neural network model, the obtained sparse user feature vectors, that is, one-hot vectors, can be subjected to embedding mapping processing through the embedding layer in the preset neural network model to obtain dense user feature vectors.

[0160] The embedding layer represents the meaning of each user ID with a multi-dimensional floating-point data, such as 32 dimensions. The embedding layer is obtained during the training phase and can be directly used during the inference phase. After the index array of the input layer, i.e., the one-hot vector, is processed by the word embedding layer, it becomes a multi-dimensional word vector, i.e., the user feature vector.

[0161] It should be noted that vectors are often used in machine learning, including the storage of features, optimized calculations, etc., all of which are inseparable from vectors. However, in specific implementations, two methods are often used to store vectors. One is to model vectors using the data structure of an array, which usually stores ordinary vectors and is also called a dense vector. The other is to model vectors using the data structure of a map, and most of the elements stored in this structure are equal to zero, and such a vector is called a sparse vector.

[0162] The reason for using two different storage structures is that features in machine learning are often elements in a high-dimensional space with thousands of components, and these components are obtained through discretization. Discretization means dividing a feature whose original value is a real number (such as a feature being price with a value of 475.2) into several intervals according to the value range (for example, the range is between 350 and 800), (such as dividing into intervals of every 10, i.e., 350 - 360, 360 - 370, 790 - 800), and the original one-dimensional feature is correspondingly discretized into several dimensions. If the price is in the interval of 470 - 480, the value of the corresponding dimension feature is 1, and the values of other dimension features are 0. Therefore, if sparse vectors are used for storage, not only space is saved, but also the efficiency will be improved in subsequent various vector operations and optimized calculations.

[0163] Next, enter the processing of the convolutional layer:

[0164] Using multiple preset convolutional kernels and a preset activation function, perform convolutional processing and activation processing on the sample user feature vectors to determine multiple first output vectors.

[0165] It should be noted that the first output vector is the output vector of the convolutional layer in the neural network model, and each first output vector corresponds to a preset convolutional kernel. The convolutional processing is used to amplify and / or extract the target feature values in the sample user feature vectors, and the activation processing is used to increase the non-linearity of the output result of the convolutional processing.

[0166] Specifically, for the convolutional processing, it is specifically implemented through the convolutional layer in the preset neural network model, and its role is to amplify and extract certain user features. For example, in this embodiment, the convolutional operation automatically extracts the mutual relationships between some user features in the user feature vector, and extracts the corresponding user features for subsequent serial number classification.

[0167] The convolutional kernels used in this embodiment include: [1, 128], [3, 128], [5, 128]. There can be multiple truncation positions (up to 128) in each convolutional kernel, and each convolutional calculation result outputs a result data. A convolutional pyramid can be constructed by connecting multiple multi-scale convolutional layers. Optionally, extracting local correlation features under different receptive fields can better represent the mutual relationship between items.

[0168] It should be noted that convolution is widely used in neural networks. The essence of convolution is a process of extracting features from input data using a kernel function, and the output is the extracted features (mapping). Convolution calculation is a process of multiplication and accumulation.

[0169] For activation processing, it is achieved through a preset activation function in the activation layer, such as the relu activation function. A neural network is a kind of connection, and the activation layer is used to introduce non-linearity. For example, if the outputs of the upper layer are X1, X2, X3, the essence of the fully connected layer processing can be expressed by the mathematical formula w1X1 + w2X2 + w3X3. Without the activation layer, the output of the fully connected layer is equivalent to a linear mapping of the upper layer, and its expressive ability is very limited.

[0170] Figure 3 It is a schematic diagram of the function image of the relu activation function. As Figure 3 shown, the horizontal axis of the relu activation function represents the input, and the vertical axis represents the output. The role of the activation function is to bring non-linear characteristics to the network. Layers such as convolution and fully connected layers themselves cannot bring non-linear characteristics. A neural network is like a function that transforms input data into the expected output data. A network without non-linear characteristics is not sufficient to complete this transformation.

[0171] It should be noted that the activation function in this step does not change the dimension of the data. The input and output are of the same size. For example, assume the output is still [68, 128], [67, 128], [66, 128].

[0172] Next, enter the processing of the pooling layer:

[0173] Perform pooling processing on multiple first output vectors to determine multiple second output vectors of different scales.

[0174] It should be noted that the second output vector is the output vector of the pooling layer in the neural network model.

[0175] The input of the pooling layer is the first output vector of the convolutional layer. The purpose of performing pooling processing is to select features with certain characteristics (such as maximum value, minimum value, average value) from the features extracted by the convolution operation, which can effectively reduce the number of feature parameters and has a regularization effect, and can improve the robustness of the model to a certain extent.

[0176] Specifically, in this embodiment, the maximum pooling method is adopted to extract the maximum value in the features, that is, the maximum positive information is extracted using max pooling MaxPooling.

[0177] Max pooling is to find the maximum value in a piece of data (matrix). For example, in a 4x4 piece of data, the entire matrix is replaced with the maximum value to extract the positive maximum feature, so as to select the most important user feature for the classification of the corresponding position number for subsequent number classification calculation.

[0178] Next, enter the processing of the first fusion layer:

[0179] Fuse multiple second output vectors of different scales into a third output vector of a preset scale as the sample user feature information.

[0180] It should be noted that the third output vector is used to determine the interest feature table after being processed by two fully connected layers.

[0181] Specifically, the first fusion layer in the preset neural network model combines the previous output results of different sizes, that is, the second output vectors. Assuming that the second output vectors output in the previous step are [1, 128], [1, 128], [1, 128], they are fused together to become a [1, K]-dimensional data, for example, the output is [1, 384].

[0182] S203. Input the sample user feature information into the information classification fully connected layer and the information grouping fully connected layer for processing.

[0183] In this step, after being processed by the information classification fully connected layer, a first sample mapping result is obtained, and this first sample mapping result is used to represent the mapping relationship between the sample user feature information and each information classification marker space through the first tensor parameter.

[0184] After being processed by the information grouping fully connected layer, a second sample mapping result is obtained, and this second sample mapping result is used to represent the mapping relationship between the first sample mapping result and each information grouping marker space through the second tensor parameter.

[0185] It should be noted that the first tensor parameter includes multiple first sub-tensor parameters, and each first sub-tensor parameter is used to represent the category feature in each information classification. The second tensor parameter includes multiple second sub-tensor parameters, and each second sub-tensor parameter is used to represent the group feature in each information grouping.

[0186] Specifically, the information classification fully connected layer fully connects the user features with the large category features to obtain the first marker space where the sample user features are mapped to the information classification.

[0187] The information grouping fully connected layer connects information classification with the keyword features of each information grouping. Thus, through the full connection between the features mapped to the first token space of information classification and the information grouping under the information classification, the user features are mapped to the second token space of the information grouping.

[0188] S204. Use the decision function to process the processing results of the information classification fully connected layer and the information grouping fully connected layer to determine the sample classification result, and adjust the various parameters in the neural network according to the sample classification result.

[0189] In this embodiment, the information keywords corresponding to each information grouping, that is, information features, are used as output tensor tensors, and are fully connected to each upper information classification respectively to construct a tree structure.

[0190] Figure 4 It is a schematic diagram of a word tree structure provided by this application. As Figure 4 shown, the information classification model 40 corresponds to multiple information classifications 41, and each information classification 41 corresponds to multiple information groupings 411. The relationship between the information classification 41 and the information grouping 42 is the word tree classification layer. Each subtree is a category (classify all information such as items and / or services), and each category connects the characteristic keywords of its corresponding information grouping. The characteristic keywords of information such as items and / or services are divided into several major categories, that is, information classifications, such as news, music; taking news as an example, news is divided into finance, sports, military, etc.

[0191] Essentially, these several information classifications or information groupings need to be linearly independent. For example, the keyword features of information such as items and / or services under the finance category are linearly independent of those of sports or military, etc.

[0192] To ensure linear independence, in this step, it is necessary to adjust the various parameters in the neural network according to the classification result, specifically including:

[0193] Calculate the loss function according to the sample classification result and the first tensor parameter and the second tensor parameter;

[0194] Perform cyclic training to adjust the various parameters in the neural network to make the value of the loss function tend to zero, and obtain the linearly independent first tensor parameter and the second tensor parameter.

[0195] Specifically, during the training process of the neural network, calculate the Product loss function, and make the value of the Product loss function tend to zero, which can ensure that both the major categories and the minor categories are linearly independent, that is, both the information classification and the information grouping are linearly independent.

[0196] The calculation method of the Product loss loss function is to perform cumulative multiplication on the tensor parameters representing various categories at this layer. The tensor is a multi-dimensional array used to represent features.

[0197] The Product loss loss function is shown in formula (2):

[0198]

[0199] Among them, the categorical_tensor_param is the tensor parameter representing each information classification and / or information grouping.

[0200] The Product loss loss function will be superimposed with the output binary categorical crossentropy loss to form a joint loss to participate in the gradient descent update, thereby realizing the linearly independent functions of the above-mentioned various information classifications and / or information groupings.

[0201] After that, the loss value is calculated through each tensor output by the second fusion layer and the sigmod layer for the vector obtained from the small class mapping space, that is, the second label space, so that the Product loss product loss tends to zero, which can ensure the linear independence of each tensor.

[0202] Among them, it is also necessary to take the inner product of the vector corresponding to a certain tensor of the large class, that is, the information classification, and the sum of the vectors corresponding to the remaining tensors. When this inner product is 0, the linear independence of each tensor can be ensured.

[0203] Connect the item keywords by constraining the non-linear relationship between the large classes, that is, the information classifications.

[0204] Then, the output results y1, y2,... yn of the upper layer are combined through the second fusion layer in the preset neural network model. Assuming that the output of the previous layer is

[128] ,

[128] ,

[128] , they are fused together to become a [K]-dimensional data. In this example, the output is

[384] .

[0205] Optionally, after making the value of the loss function tend to zero, it further includes:

[0206] Adjust the first tensor parameter to make each first sub-tensor parameter linearly independent;

[0207] and / or,

[0208] Adjust the second tensor parameter to make each second sub-tensor parameter linearly independent;

[0209] Among them, when the first inner product between each first sub-tensor parameter is zero, each first sub-tensor parameter is linearly independent, and when the second inner product between the second sub-tensor parameters is zero, each second sub-tensor parameter is linearly independent.

[0210] Next, through the decision layer in the neural network, the sigmoid decision function is added to the fully connected layer merged and fused by the second fusion layer. The sigmoid decision function is shown in formula (3):

[0211]

[0212] The output is a floating-point number between 0 and 1, which is generally used in a binary classification model.

[0213] This layer applies the sigmoid decision function to each output tensor. If the output is N tensors, it is equivalent to N binary classifiers, and the output results are like: [0.11, 0.32, 0.92, 0.72…0.21]. Each element can represent the score of the degree of interest in the features of the information. These binary classifiers are applicable to the multi-label prediction of this application, that is, to judge whether it is 0 or 1 for each label. 0 represents negation, and 1 represents affirmation of the classification of the feature keywords of the information to which the user features tend. Each output has a connection to this binary classifier, and the final loss value is the sum of the loss values corresponding to each binary classifier.

[0214] So far, the adjustment of various parameters in the neural network has been completed. Next, the trained neural network can be used to extract and recognize the features of the user to obtain a personalized interest feature table for each user.

[0215] S205. Obtain multiple user features of the user.

[0216] In this step, the input layer (also called the Onehot conversion layer) in the trained neural network takes multiple user features of the target user as input, such as age, place of residence, gender, occupation, etc.

[0217] S206. Use the feature extraction layer in the neural network to convert multiple user features of the user into user feature information.

[0218] S207. Through neural network inference on multiple user features, enable the neural network to determine the mapping relationship between each user feature and the preset information classification grouping rule according to the requirements of the word tree classification structure, so as to determine the interest feature table.

[0219] In this step, each information classification in the preset information classification grouping rule corresponds to at least one information group, and each information classification and each information group are linearly independent of each other.

[0220] In this embodiment, it specifically includes:

[0221] Using the feature extraction layer in the neural network to convert multiple user features of the user into user feature information, and the user feature information represents the multiple user features in the form of a vector or a matrix;

[0222] Inputting the user feature information into the information classification fully connected layer to obtain the first mapping result of the user feature information mapped to each information classification;

[0223] Inputting the first mapping result into the information grouping fully connected layer to determine the second mapping result of the first mapping result mapped to each information grouping. Among them, during the processing, the information classification fully connected layer and the information grouping fully connected layer respectively establish multiple tensor parameters according to the requirements of the word tree classification structure, and the tensor parameters are used to represent the corresponding relationship between the category features of the information classification fully connected layer and the group features of the information grouping fully connected layer;

[0224] Using the decision-making layer in the neural network to process the second mapping result to determine the interest feature table.

[0225] It should be noted that the specific implementation process of this step is basically the same as that of S203 - S204. The difference is that in S204, the neural network is trained, so various parameters in the neural network need to be modified, while in this step, various parameters in the neural network will not be modified, that is, the tensor parameters will not be corrected, but only the neural network is used for information classification and information grouping.

[0226] S208. Obtain the information feature table.

[0227] S209. Determine the first interest score table of multiple to-be-pushed messages according to the information feature table and the interest feature table.

[0228] S210. Perform time decay correction on the first interest score table according to the time records in each to-be-pushed message and the target push time to determine the second interest score table.

[0229] S211. Determine each recall message that does not need to be pushed to the user from each to-be-pushed message according to the second interest score table and the preset recall requirement.

[0230] In this embodiment, the noun explanations and implementation principles in S208 - S211 can refer to S101 - S104, which will not be elaborated here.

[0231] This embodiment provides a method for recall processing of information to be pushed. By obtaining an information feature table and a user interest feature table, then determining a first interest score table for multiple pieces of information to be pushed according to the information feature table and the interest feature table, and then using a preset time decay model to correct the first interest score table according to the time information in multiple pieces of information to be pushed and the target push time to determine a second interest score table. Finally, according to the second interest score table and a preset recall requirement, recall information that does not need to be pushed to the user is determined from each piece of information to be pushed. It solves the technical problem of how to accurately recall information to be pushed, avoids pushing a large amount of information that the user is not interested in, and achieves the technical effect of personalized and accurate push for each user.

[0232] Figure 5 This is a schematic structural diagram of a device for recall processing of information to be pushed provided by an embodiment of the present application. The device 500 for recall processing of information to be pushed can be implemented by software, hardware, or a combination of both.

[0233] As Figure 5 shown, the device 500 for recall processing of information to be pushed includes:

[0234] An acquisition module 501, configured to acquire an information feature table and a user interest feature table, where the information feature table is used to characterize the feature attributes corresponding to multiple pieces of information to be pushed, and the interest feature table is used to characterize the user's interest preferences for each piece of information in the information feature table;

[0235] A processing module 502, configured to:

[0236] Determine a first interest score table for multiple pieces of information to be pushed according to the information feature table and the interest feature table;

[0237] Perform time decay correction on the first interest score table according to the time records in each piece of information to be pushed and the target push time to determine a second interest score table;

[0238] Determine recall information that does not need to be pushed to the user from each piece of information to be pushed according to the second interest score table and a preset recall requirement.

[0239] In a possible design, the processing module 502 is configured to:

[0240] Respectively determine the time differences between each time record and the target push time;

[0241] Input each time difference and a preset time sensitivity coefficient into a preset time decay model for calculation to determine multiple time decay parameters;

[0242] Multiply each time decay parameter by the corresponding interest score in the first interest score table to determine the second interest score table.

[0243] In a possible design, the obtaining module 501 is further configured to obtain multiple user features of a user;

[0244] The processing module 502 is further configured to:

[0245] By performing neural network inference on multiple user features, the neural network determines the mapping relationship between each user feature and the preset information classification and grouping rules according to the requirements of the word tree classification structure, so as to determine an interest feature table, where each information classification in the preset information classification and grouping rules corresponds to at least one information group, and the information classifications and the information groups are linearly independent of each other.

[0246] In a possible design, the processing module 502 is further configured to:

[0247] Use the feature extraction layer in the neural network to convert multiple user features of the user into user feature information, and the user feature information represents the multiple user features in the form of a vector or a matrix;

[0248] Input the user feature information into the information classification fully connected layer to obtain a first mapping result of the user feature information mapped to each information classification;

[0249] Input the first mapping result into the information grouping fully connected layer to determine a second mapping result of the first mapping result mapped to each information group, where the information classification fully connected layer and the information grouping fully connected layer respectively establish multiple tensor parameters according to the requirements of the word tree classification structure during the processing, and the tensor parameters are used to represent the corresponding relationship between the category features of the information classification fully connected layer and the group features of the information grouping fully connected layer;

[0250] Use the decision-making layer in the neural network to process the second mapping result to determine an interest feature table.

[0251] In a possible design, the obtaining module 501 is further configured to: obtain sample user features;

[0252] The processing module 502 is further configured to:

[0253] Use the feature extraction layer in the neural network to extract sample user feature information from the sample user features, and the sample user feature information represents the multiple sample user features in the form of a vector or a matrix;

[0254] Input the sample user feature information into the information classification fully connected layer and the information grouping fully connected layer for processing to determine the first sample mapping result and the second sample mapping result. The first sample mapping result is used to represent the mapping relationship of the sample user feature information mapped to each information classification label space through the first tensor parameter, and the second sample mapping result is used to represent the mapping relationship of the first sample mapping result mapped to each information grouping label space through the second tensor parameter. Among them, the first tensor parameter includes multiple first sub-tensor parameters, and each first sub-tensor parameter is used to represent the category feature in each information classification. The second tensor parameter includes multiple second sub-tensor parameters, and each second sub-tensor parameter is used to represent the group feature in each information grouping.

[0255] Use the decision function to process the processing results of the information classification fully connected layer and the information grouping fully connected layer to determine the sample classification result, and adjust the various parameters in the neural network according to the sample classification result.

[0256] In a possible design, the processing module 502 is further configured to:

[0257] Use multiple preset convolution kernels and a preset activation function to perform convolution processing and activation processing on the sample user feature vector to determine multiple first output vectors. The first output vector is the output vector of the convolution layer in the neural network model. Among them, each first output vector corresponds to a preset convolution kernel. The convolution processing is used to amplify and / or extract the target feature values in the sample user feature vector, and the activation processing is used to increase the non-linearity of the output result of the convolution processing.

[0258] Perform pooling processing on the multiple first output vectors to determine multiple second output vectors of different scales. The second output vector is the output vector of the pooling layer in the neural network model.

[0259] Fuse the multiple second output vectors of different scales into a third output vector of a preset scale as the sample user feature information.

[0260] In a possible design, the processing module 502 is further configured to:

[0261] Calculate the loss function according to the sample classification result and the first tensor parameter and the second tensor parameter.

[0262] Perform cyclic training to adjust the various parameters in the neural network to make the value of the loss function tend to zero, and obtain linearly independent first tensor parameters and second tensor parameters.

[0263] Optionally, the processing module 502 is further configured to:

[0264] Adjust the first tensor parameter to make each first sub-tensor parameter linearly independent.

[0265] And / or,

[0266] Adjust the second tensor parameter to make each second sub-tensor parameter linearly independent;

[0267] Wherein, when the first inner product between each first sub-tensor parameter is zero, each first sub-tensor parameter is linearly independent, and when the second inner product between the second sub-tensor parameters is zero, each second sub-tensor parameter is linearly independent.

[0268] It should be noted that, Figure 5 The device provided by the illustrated embodiment can execute the method provided in any of the above method embodiments. The specific implementation principles, technical features, explanations of professional terms, and technical effects are similar, and will not be elaborated herein.

[0269] Figure 6 The following is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 6 shown, the electronic device 600 may include: at least one processor 601 and a memory 602. Figure 6 The electronic device shown is taken as an example with one processor.

[0270] The memory 602 is used to store a program. Specifically, the program may include program code, and the program code includes computer operation instructions.

[0271] The memory 602 may include a high-speed RAM memory, and may also include non-volatile memory, such as at least one disk memory.

[0272] The processor 601 is used to execute the computer execution instructions stored in the memory 602 to implement the methods described in the above method embodiments.

[0273] Wherein, the processor 601 may be a central processing unit (CPU for short), or an application specific integrated circuit (ASIC for short), or one or more integrated circuits configured to implement the embodiments of the present application.

[0274] Optionally, the memory 602 may be either independent or integrated with the processor 601. When the memory 602 is a device independent of the processor 601, the electronic device 600 may further include:

[0275] A bus 603 is used to connect the processor 601 and the memory 602. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc., but it does not mean that there is only one bus or one type of bus.

[0276] Optionally, in specific implementation, if the memory 602 and the processor 601 are integrated on a chip, the memory 602 and the processor 601 can complete communication through an internal interface.

[0277] The embodiment of the present application also provides a computer-readable storage medium, which may include: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc. Specifically, program instructions are stored in the computer-readable storage medium, and the program instructions are used for the methods in the above method embodiments.

[0278] The embodiment of the present application also provides a computer program product, including a computer program, which implements the methods in the above method embodiments when executed by a processor.

[0279] It should be understood that the present application is not limited to the exact structure already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.

[0280] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for recall processing of information to be pushed, characterized in that, Applied to a mobile terminal, the method includes: Obtain an information feature table and an interest feature table of a user, where the information feature table is used to characterize the feature attributes corresponding to multiple pieces of information to be pushed, and the interest feature table is used to characterize the user's interest preferences for each piece of information in the information feature table; Determine a first interest score table for multiple pieces of the information to be pushed according to the information feature table and the interest feature table; Perform time decay correction on the first interest score table according to the time records in each piece of the information to be pushed and the target push time to determine a second interest score table; Determine recall information that does not need to be pushed to the user from each piece of the information to be pushed according to the second interest score table and a preset recall requirement; The obtaining of the interest feature table of the user includes: Obtain multiple user features of the user; Through neural network inference on multiple user features, enable the neural network to determine the mapping relationship between each user feature and a preset information classification grouping rule according to the requirements of the word tree classification structure, so as to determine the interest feature table, where each information classification in the preset information classification grouping rule corresponds to at least one information group, and the information classifications and the information groups are linearly independent of each other.

2. The method for recall processing of information to be pushed according to claim 1, wherein The performing of time decay correction on the first interest score table according to the time records in each piece of the information to be pushed and the target push time to determine a second interest score table includes: Respectively determine the time difference between the time record of each piece of the information to be pushed and the target push time; Input each time difference and a preset time sensitivity coefficient into a preset time decay model for calculation to determine multiple time decay parameters; Multiply each time decay parameter by the corresponding interest score in the first interest score table to determine the second interest score table.

3. The method for recall processing of information to be pushed according to claim 1, wherein The enabling of the neural network to determine the mapping relationship between each user feature and the information classification grouping rule through neural network inference on the multiple user features includes: Use the feature extraction layer in the neural network to convert multiple user features of the user into user feature information, and the user feature information represents multiple user features in the form of a vector or a matrix; Input the user feature information into an information classification fully connected layer to obtain a first mapping result of the user feature information mapped to each information classification; Input the first mapping result into an information group fully connected layer to determine a second mapping result of the first mapping result mapped to each information group, where the information classification fully connected layer and the information group fully connected layer respectively establish multiple tensor parameters during the processing according to the requirements of the word tree classification structure, and the tensor parameters are used to characterize the corresponding relationship between the category features of the information classification fully connected layer and the group features of the information group fully connected layer; Use the decision-making layer in the neural network to process the second mapping result to determine the interest feature table.

4. The method for recall processing of information to be pushed according to claim 1, wherein The method further includes: obtaining the neural network for inferring the mapping relationship, including: Obtaining sample user features; Extracting sample user feature information in the sample user features by using a feature extraction layer in the neural network, where the sample user feature information represents multiple sample user features in the form of a vector or a matrix; Inputting the sample user feature information into a fully connected information classification layer and a fully connected information grouping layer for processing to determine a first sample mapping result and a second sample mapping result, where the first sample mapping result is used to represent the mapping relationship between the sample user feature information and each information classification label space through a first tensor parameter, and the second sample mapping result is used to represent the mapping relationship between the first sample mapping result and each information grouping label space through a second tensor parameter; wherein, the first tensor parameter includes multiple first sub-tensor parameters, and each first sub-tensor parameter is used to represent the category feature in each information classification, and the second tensor parameter includes multiple second sub-tensor parameters, and each second sub-tensor parameter is used to represent the group feature in each information grouping; Processing the processing results of the fully connected information classification layer and the fully connected information grouping layer by using a decision function to determine a sample classification result, and adjusting each parameter in the neural network according to the sample classification result.

5. The method for recall processing of information to be pushed according to claim 4, wherein The extracting the feature information in the sample user feature vector by using the feature extraction layer in the neural network includes: Performing convolution processing and activation processing on the sample user feature vector by using multiple preset convolution kernels and a preset activation function to determine multiple first output vectors, where the first output vector is the output vector of the convolution layer in the neural network model; wherein, each first output vector corresponds to the preset convolution kernel, the convolution processing is used to magnify and / or extract the target feature value in the sample user feature vector, and the activation processing is used to increase the non-linearity of the output result of the convolution processing; Performing pooling processing on the multiple first output vectors to determine multiple second output vectors with different scales, where the second output vector is the output vector of the pooling layer in the neural network model; Fusing the multiple second output vectors with different scales into a third output vector with a preset scale as the sample user feature information.

6. The method for recall processing of information to be pushed according to claim 4, wherein Adjusting each parameter in the neural network according to the classification result includes: Calculating a loss function according to the sample classification result, the first tensor parameter, and the second tensor parameter; Performing cyclic training to adjust each parameter in the neural network to make the value of the loss function tend to zero, and obtaining linearly independent first tensor parameter and second tensor parameter.

7. The method for recall processing of information to be pushed according to claim 6, wherein After making the value of the loss function tend to zero, it further includes: Adjusting the first tensor parameter to make each first sub-tensor parameter linearly independent; And / or Adjusting the second tensor parameter to make each second sub-tensor parameter linearly independent; Among them, when the first inner product between each of the first sub-tensor parameters is zero, each of the first sub-tensor parameters is linearly independent, and when the second inner product between the second sub-tensor parameters is zero, each of the second sub-tensor parameters is linearly independent.

8. An information recall processing device to be pushed, characterized in that, It includes: An acquisition module, configured to acquire an information feature table and an interest feature table of a user, where the information feature table is used to characterize the feature attributes corresponding to multiple pieces of information to be pushed, and the interest feature table is used to characterize the user's interest preferences for each piece of information in the information feature table; A processing module, configured to: Determine a first interest score table for multiple pieces of information to be pushed according to the information feature table and the interest feature table; Perform time decay correction on the first interest score table according to the time records in each piece of information to be pushed and the target push time, so as to determine a second interest score table; Determine each piece of recall information that does not need to be pushed to the user from each piece of information to be pushed according to the second interest score table and a preset recall requirement; The acquisition module is specifically configured to acquire multiple user features of the user; by performing neural network inference on the multiple user features, the neural network determines the mapping relationship between each user feature and a preset information classification and grouping rule according to the requirements of the word tree classification structure, so as to determine the interest feature table, where each information classification in the preset information classification and grouping rule corresponds to at least one information group, and each of the information classifications and each of the information groups are linearly independent.

9. An electronic device, characterized in that, It includes: A processor; And, A memory, configured to store a computer program of the processor; Wherein, the processor is configured to execute the method for recalling information to be pushed according to any one of claims 1 to 7 by executing the computer program.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method for recalling information to be pushed according to any one of claims 1 to 7.

11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the method for recalling information to be pushed according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Recommendation method and device, equipment and computer storage medium

    CN112395489A