Method and apparatus for pushing information
By using the capsule network to generate multiple user representation vectors, combined with the matching degree of information attribute feature vectors, the problem of convergence of information push content in the prior art is solved, and more accurate and rich information push is achieved.
Patent Information
- Application Number
- CN202110934824.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-16
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2041-08-16
AI Technical Summary
Existing information push methods based on big data can easily lead to convergence of push content, making it difficult to explore the potential preferences of users, and affecting the user experience.
By obtaining the user's behavior feature vector and attribute feature vector, a pre-trained capsule network is used to generate multiple user representation vectors, combining the attribute feature vectors of information, computing the matching degree to push information.
It improves the accuracy and content richness of information push, avoids the convergence of push information, and can generalize information without direct interaction among users, improving user experience.
Smart Images

Figure CN113609397B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of computer technologies, and more particularly, to methods and apparatuses for pushing information. Background Art
[0002] Existing big data-based information pushing methods usually first model a user's interests based on the user's historical data to accurately capture the user's interests and push content that the user is interested in. Therefore, modeling the user's interests is a very crucial step, which directly affects the subsequent information pushing effect.
[0003] Generally, information that the user has directly interacted with in the past (such as items and brands browsed or purchased) is pushed to the user based on the user's historical behavior data. However, this method is likely to lead to the convergence of the pushed content, and frequently pushing information that the user has interacted with will also affect the user experience to a certain extent. Based on this, how to explore the user's potential preferences and generalize the user's preferences is also a question worthy of consideration. Summary of the Invention
[0004] Embodiments of the present disclosure provide methods and apparatuses for pushing information.
[0005] In a first aspect, an embodiment of the present disclosure provides a method for pushing information. The method includes: obtaining a behavior feature vector and an attribute feature vector of a user, and respectively obtaining attribute feature vectors of each piece of information in an information set; inputting the behavior feature vector of the user into a pre-trained capsule network to generate at least two capsule vectors for characterizing the user's interests; respectively splicing the at least two capsule vectors with the attribute feature vector of the user, and generating at least two user representation vectors for characterizing the user according to the splicing result; determining a matching degree between the attribute feature vector of the information in the information set and the user representation vector, and selecting information from the information set for pushing according to the determined matching degree.
[0006] In a second aspect, an embodiment of the present disclosure provides an apparatus for pushing information. The apparatus includes: an obtaining unit configured to obtain a behavior feature vector and an attribute feature vector of a user, and respectively obtain attribute feature vectors of each piece of information in an information set; a generating unit configured to input the behavior feature vector of the user into a pre-trained capsule network to generate at least two capsule vectors for characterizing the user's interests; the generating unit is further configured to respectively splice the at least two capsule vectors with the attribute feature vector of the user, and generate at least two user representation vectors for characterizing the user according to the splicing result; a pushing unit configured to determine a matching degree between the attribute feature vector of the information in the information set and the user representation vector, and select information from the information set for pushing according to the determined matching degree.
[0007] In a third aspect, embodiments of the present disclosure provide an electronic device, which includes: one or more processors; a storage device for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the method described in any implementation manner of the first aspect.
[0008] In a fourth aspect, embodiments of the present disclosure provide a computer-readable medium, on which a computer program is stored, and when the computer program is executed by a processor, the method described in any implementation manner of the first aspect is implemented.
[0009] The method and apparatus for pushing information provided by the embodiments of the present disclosure perform multi-interest analysis on users based on the behavioral characteristics and attribute characteristics of users by using a capsule network to capture the multi-faceted interests of users, use multiple user representation vectors to more comprehensively represent users, and then based on the matching degree between the multiple user representation vectors and the attribute characteristics of the information, not only can the accuracy of the information pushed to the user be improved, but also some information that the user has no direct interaction with in history can be generalized, avoiding the situation where the pushed information received by the user converges. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objects, and advantages of the present disclosure will become more apparent:
[0011] Figure 1 is an exemplary system architecture diagram to which an embodiment of the present disclosure can be applied;
[0012] Figure 2 is a flowchart of an embodiment of the method for pushing information according to the present disclosure;
[0013] Figure 3 is a flowchart of another embodiment of the method for pushing information according to the present disclosure;
[0014] Figure 4 is a schematic diagram of an embodiment of the training network structure of the user representation model according to the present disclosure;
[0015] Figure 5 is a flowchart of still another embodiment of the method for pushing information according to the present disclosure;
[0016] Figure 6 is a schematic diagram of an application scenario of the method for pushing information according to the embodiments of the present disclosure;
[0017] Figure 7 is a schematic structural diagram of an embodiment of the apparatus for pushing information according to the present disclosure;
[0018] Figure 8 It is a schematic structural diagram of an electronic device suitable for implementing the embodiments of the present disclosure. Detailed implementation manners
[0019] The present disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only for explaining the relevant invention and not for limiting the invention. Additionally, it should be noted that for the sake of description, only the parts related to the relevant invention are shown in the drawings.
[0020] It should be noted that the data collection involved in the embodiments of the present disclosure (such as user attribute features, behavioral features, item information, etc.) is carried out on the basis of obtaining the authorization of the relevant subject and complies with the provisions of relevant laws and regulations.
[0021] It should be noted that, without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other. The present disclosure will be described in detail below with reference to the drawings and embodiments.
[0022] Figure 1 An exemplary architecture 100 of an embodiment of a method for pushing information or a device for pushing information to which the present disclosure can be applied is shown.
[0023] As Figure 1 shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0024] The terminal devices 101, 102, 103 interact with the server 105 through the network 104 to receive or send messages, etc. Various client applications may be installed on the terminal devices 101, 102, 103. For example, browser applications, search applications, instant messaging tools, social platforms, shopping applications, etc.
[0025] The terminal devices 101, 102, 103 may be hardware or software. When the terminal devices 101, 102, 103 are hardware, they may be various electronic devices, including but not limited to smart phones, tablet computers, e-book readers, laptop portable computers, and desktop computers, etc. When the terminal devices 101, 102, 103 are software, they may be installed in the above-listed electronic devices. It may be implemented as multiple software or software modules (such as multiple software or software modules for providing distributed services), or may be implemented as a single software or software module. No specific limitation is made here.
[0026] The server 105 may be a server that provides various services. For example, it is a server that provides backend support for client applications installed on the terminal devices 101, 102, and 103. The server 105 may generate at least two user representation vectors based on the behavior feature vectors and attribute feature vectors of the users corresponding to the terminal devices 101, 102, and 103, and select information from the information set according to the matching degree between the generated user representation vectors and the attribute feature vectors of the information in the information set, and push the information to the terminal devices 101, 102, and 103.
[0027] It should be noted that the method for pushing information provided by the embodiments of the present disclosure is generally executed by the server 105. Correspondingly, the device for pushing information is generally arranged in the server 105.
[0028] It should also be pointed out that information push applications may also be installed in the terminal devices 101, 102, and 103. The terminal devices 101, 102, and 103 may also process the behavior feature vectors and attribute feature vectors of the users and the attribute feature vectors of each piece of information in the information set based on the information push applications. At this time, the method for pushing information may also be executed by the terminal devices 101, 102, and 103. Correspondingly, the device for pushing information may also be arranged in the terminal devices 101, 102, and 103. At this time, the exemplary system architecture 100 may not include the server 105 and the network 104.
[0029] It should be noted that the server 105 may be hardware or software. When the server 105 is hardware, it may be implemented as a distributed server cluster composed of multiple servers, or as a single server. When the server 105 is software, it may be implemented as multiple software or software modules (such as multiple software or software modules for providing distributed services), or as a single software or software module. No specific limitation is made here.
[0030] It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in
[0031] Continue to refer to Figure 2 , which shows a flow 200 of an embodiment of the method for pushing information according to the present disclosure. The method for pushing information includes the following steps:
[0032] Step 201, obtain the behavior feature vector and attribute feature vector of the user, and respectively obtain the attribute feature vectors of each piece of information in the information set.
[0033] In this embodiment, the user can be any user. The user's behavioral feature vector can be used to characterize the user's behavioral features. Among them, the user's behavioral features can refer to the features of various behaviors of the user. In different application scenarios, the behaviors of the user can be the same or different. For example, for users of video applications, the user's behaviors include but are not limited to browsing, collecting, publishing, commenting, etc. Another example is that for users of shopping applications, the user's behaviors include but are not limited to browsing, clicking, adding to the shopping cart, collecting, purchasing, etc.
[0034] The user's attribute feature vector can be used to characterize the user's attribute features. Among them, the user's attribute features can refer to the features of various attributes of the user himself.
[0035] The information set can be composed of several pieces of information. The information in the information set can be various types of information. For example, each piece of information in the information set can be information used to indicate different items. The attribute feature vector of the information can be used to characterize the attribute features of the information. Among them, the attribute features of the information can refer to the features of various attributes of the information itself. The attributes of different types of information can be the same or different. Taking the information used to indicate an item as an example, the attributes of the information include but are not limited to the brand, category, price (such as the price of the item indicated by the information, the average price of all items under the brand or category indicated by the information of the item), popularity (such as the click-through rate, attention, etc. of the item indicated by the information), item richness, etc.
[0036] The execution subject of the method for pushing information to the user (such as Figure 1 the server 105 shown, etc.) can obtain the user's behavioral feature vector and attribute feature vector from local or other storage devices, etc., and obtain the attribute feature vectors of each piece of information in the information set. It should be noted that the user's behavioral feature vector, the user's attribute feature vector, and the attribute feature vectors of each piece of information in the information set can be obtained from the same data source or from different data sources.
[0037] The user's behavioral feature vector can be generated based on the user's behavioral feature data. Specifically, various existing vector encoding methods can be used to encode the user's behavioral data to obtain the user's behavioral feature vector. Among them, the user's behavioral data can be the feature values used to characterize the features of various behaviors of the user. For example, the user's behavioral data includes the brand and / or category of the item purchased by the user and the item name, purchase time, etc.
[0038] Similarly, the attribute feature vector of a user can be generated based on the user's attribute feature data. Specifically, various existing vector encoding methods can be used to encode the user's attribute feature data to obtain the user's attribute feature vector. Among them, the user's attribute feature data can be the attribute values used to represent various attributes of the user.
[0039] The attribute feature vector of information can be generated based on the attribute feature data of the information. Specifically, various existing vector encoding methods can be used to encode the attribute feature data of the information to obtain the attribute feature vector of the information. Among them, the attribute feature data of the information can be the attribute values used to represent various attributes of the information.
[0040] The behavioral feature vector of the user, the attribute feature vector of the user, and the attribute feature vector of the information can be generated by the above-mentioned execution entity using various methods, or can be generated by other electronic devices using various methods.
[0041] Step 202: Input the behavioral feature vector of the user into a pre-trained capsule network to generate at least two capsule vectors for representing the interests of the user.
[0042] In this embodiment, the capsule network is a new neural network different from the traditional neural network. The capsule network encodes both spatial information and the probability of the existence of an object, and the encoding is in the capsule vector. Generally, the modulus of the vector represents the probability of the existence of the feature, and the direction of the vector represents the pose information of the feature. A moving feature will change the capsule vector but does not affect the probability of the existence of the feature.
[0043] Specifically, the pre-trained capsule network can generate at least two capsule vectors according to the behavioral feature vector of the user. Each generated capsule vector can be used to represent different aspects of the user's interests respectively.
[0044] The training process of the capsule network can be completed by the above-mentioned execution entity, or can be completed by other electronic devices. Specifically, existing machine learning methods can be used to complete the training of the capsule network through pre-set training samples and loss functions.
[0045] As an example, the capsule vector can be obtained through the following formula:
[0046]
[0047]
[0048]
[0049]
[0050] Among them, represents the high-level capsule vector, represents the low-level capsule vector. Squash() represents the non-linear mapping function. S ij represents the mapping matrix to be learned between the high and low levels in the capsule network. w ij represents the weight connecting the high and low levels. b represents the routing mechanism between the high and low levels in the capsule network. i and k represent the serial numbers of the capsule vectors, and m represents the number of capsule vectors. Through several iterations of the above formula, a convergent capsule vector can be obtained.
[0051] Step 203: Concatenate at least two capsule vectors with the user's attribute feature vector respectively, and generate at least two user representation vectors for characterizing the user according to the concatenation result.
[0052] In this embodiment, for each of the at least two capsule vectors output by the capsule network, concatenate the capsule vector with the user's attribute feature vector to form the concatenated capsule vector corresponding to the capsule vector. Therefore, the concatenation result obtained by concatenating at least two capsule vectors with the user's attribute feature vector respectively may include the concatenated capsule vectors respectively corresponding to the at least two capsule vectors.
[0053] After obtaining the concatenation result, various methods can be further adopted according to the actual application requirements or application scenarios to generate user representation vectors according to the concatenation result. Among them, the user representation vector can be used to characterize the user. Each user representation vector among the at least two user representation vectors can be used to represent different aspects of the user's features respectively.
[0054] Generally, the concatenated capsule vectors included in the concatenation result may correspond one-to-one to the generated user representation vectors. For example, the at least two concatenated capsule vectors included in the concatenation result can be directly used as the user representation vectors.
[0055] Step 204: Determine the matching degree between the attribute feature vector of the information in the information set and the user representation vector, and select information from the information set for pushing according to the determined matching degree.
[0056] In this embodiment, for each piece of information in the information set, the matching degree between the attribute feature vector of the information and at least two user representation vectors can be determined respectively to obtain at least two matching degrees corresponding to the information. Among them, the matching degree between the attribute feature vector of the information and the user representation vector can be obtained by using various existing vector similarity calculation methods.
[0057] After obtaining at least two matching degrees corresponding to each piece of information, various methods can be flexibly adopted according to actual application requirements to select information from the information set for pushing. For example, the maximum value or average value of at least two matching degrees corresponding to each piece of information can be used as the target matching degree corresponding to the information first. Then, according to the magnitudes of the target matching degrees corresponding to each piece of information, the information whose corresponding target matching degree is greater than the preset matching degree threshold is selected from the information set for pushing.
[0058] Specifically, after selecting information from the information set, relevant information of the selected information can be pushed to the terminal device used by the user. For example, when the information in the information set indicates an item, relevant information of the item indicated by the selected information (such as an item introduction page, etc.) can be pushed to the user.
[0059] In some alternative implementation manners of this embodiment, the behavior data of the user can be obtained first, and then the behavior data of the user is input into a pre-trained user behavior feature extraction network to obtain a user behavior feature vector. Among them, the user behavior feature extraction network can be constructed based on the network structures of various existing feature extraction networks.
[0060] In some alternative implementation manners of this embodiment, the attribute data of the user can be obtained first, and then the attribute data of the user is input into a pre-trained user attribute feature extraction network to obtain a user attribute feature vector. Among them, the user attribute feature extraction network can be constructed based on the network structures of various existing feature extraction networks.
[0061] In some alternative implementation manners of this embodiment, for the information in the information set, the attribute data of the information can be obtained first, and then the attribute data of the information is input into a pre-trained information attribute feature extraction network to obtain an information attribute feature vector. Among them, the information attribute feature extraction network can be constructed based on the network structures of various existing feature extraction networks.
[0062] The above-mentioned user behavior feature extraction network, user attribute feature extraction network, and information attribute feature extraction network can adopt the same network structure or different network structures. The user behavior feature extraction network, user attribute feature extraction network, and information attribute feature extraction network can specifically be trained based on machine learning methods, respectively using preset training samples and loss functions.
[0063] In the prior art, generally based on the historical behavior data of users, a low-dimensional vector is used to represent user interests through simple statistical analysis or by using deep learning algorithms, etc. However, this method of representing users with a single vector is usually limited by the size of the vector dimension, and it is relatively difficult to comprehensively express the various interests or preferences of users. Moreover, this method of representing users can usually only determine the preferences of users for the information they have had direct interaction with in their history, and cannot generalize the information that users have no direct interaction with. Based on this, the method provided in the above embodiments of the present disclosure uses the behavioral feature data and attribute feature data of users, and uses a capsule network to generate multiple user representation vectors to more comprehensively express the various interests of users, and can generalize the information that users have not directly interacted with, so as to improve the content richness and accuracy of the push information received by users, and further help to improve the effective click-through rate of information and the browsing depth of effectively exposed users. In addition, generalizing more information that users have no direct interaction with helps to improve the interaction degree between some information and users, assist in the cold start of some information, etc., and improve the exposure of the generalized push information, so as to improve the average number of effective exposure categories of items as a whole, and at the same time help to improve the Session (session) order AUC (Area Under Curve) and Session click AUC in scenarios such as search fine ranking.
[0064] Further referring to Figure 3 , which shows the flow 300 of another embodiment of the method for pushing information. The flow 300 of the method for pushing information includes the following steps:
[0065] Step 301, obtain the behavioral feature vector and attribute feature vector of the user, and respectively obtain the attribute feature vectors of each piece of information in the information set.
[0066] Step 302, input the behavioral feature vector of the user into a pre-trained capsule network to generate at least two capsule vectors for representing the interests of the user.
[0067] Step 303, splice at least two capsule vectors with the attribute feature vector of the user respectively, and generate at least two user representation vectors for representing the user according to the splicing result by using an attention mechanism.
[0068] In this embodiment, after obtaining the splicing result including at least two spliced capsule vectors, an attention mechanism can be further used to generate the corresponding at least two user representation vectors. Specifically, various existing methods implemented based on the attention mechanism can be used to generate at least two user representation vectors.
[0069] The attention mechanism can be used to calculate the attention for each spliced capsule vector respectively to capture the relationship between the spliced capsule vectors, so as to calculate the corresponding user representation vector by weighting each spliced capsule vector.
[0070] Step 304: Determine the matching degree between the attribute feature vector of the information in the information set and the user representation vector, and select information from the information set for pushing according to the determined matching degree.
[0071] In this embodiment, the content not specifically described above can refer to Figure 2 the relevant descriptions in the corresponding embodiment, which will not be elaborated here.
[0072] In some optional implementation manners of this embodiment, after obtaining the splicing result including at least two spliced capsule vectors, the splicing result can be input into a pre-trained fully connected network to obtain at least two initial user representation vectors, and then the at least two initial user representation vectors can be input into an attention network implemented based on the attention mechanism to obtain at least two user representation vectors.
[0073] Among them, the fully connected network can include several fully connected layers. Specifically, the fully connected network is used to further perform feature mapping and dimension transformation on each spliced capsule vector to generate the initial user representation vectors corresponding to each spliced capsule vector respectively. Then, the attention network is used to perform weighted calculation on each initial user representation vector to obtain the user representation vectors corresponding to each initial user representation vector respectively.
[0074] Generally, the processing process of the attention network implemented based on the attention mechanism is usually to multiply each input initial user representation vector by three training matrices created during the training process to respectively generate three vectors: the Q (Query) vector, the K (Key) vector, and the V (Value) vector corresponding to the initial user representation vector. Then, calculate the inner product of the Q vector and the K vector corresponding to each initial user representation vector and perform normalization and other processing, and then multiply the V vector by the result after normalization as the weight. Finally, calculate the user representation vector output by each initial user representation vector after passing through the attention network through the weighted sum method.
[0075] Continue to refer to Figure 4 , Figure 4 which is a schematic diagram of an embodiment of the training network structure of the user representation model according to this embodiment. Among them, the user representation model can be composed of the above-mentioned user behavior feature extraction network, user attribute feature extraction network, capsule network, fully connected network, and attention network.
[0076] As Figure 4As shown, the user behavior feature extraction network can encode the input user behavior data to obtain a user behavior feature vector, and the capsule network can process the user behavior feature vector to generate at least two capsule vectors. The user attribute feature extraction network can encode the input user attribute data to obtain a user attribute feature vector. Then, at least two capsule vectors output by the capsule network can be respectively concatenated with the user attribute feature vector output by the user attribute feature extraction network to obtain at least two corresponding concatenated capsule vectors, which are input into a fully connected network MLP (Multi-Layer Perceptron) to obtain at least two initial user representation vectors, and then the Attention network is used to perform weighted processing on the at least two initial user representation vectors output by the MLP to obtain at least two corresponding user representation vectors. In addition, the information attribute feature extraction network can encode the input information attribute data to obtain an information attribute feature vector. Then, the information attribute feature vector can be used to perform matching calculations with the at least two user representation vectors finally output by the user representation model.
[0077] Take Figure 4 as an example, the user representation model can be trained through the following steps:
[0078] Step 1, obtain a sample set.
[0079] In this step, the sample set can include positive samples and negative samples. The positive samples can include the user's attribute data, the user's behavior data, and the attribute data of the target information. The negative samples can include the user's attribute data, the user's behavior data, and the attribute data of non-target information. Among them, the target information can include the subsequent behavior data corresponding to the user's behavior data. As an example, if the user's behavior data included in a positive sample is the user's behavior data at time point "T", then the target information can include the user's behavior data at time point "T + 1" to characterize the user's next behavior. Non-target information refers to any information other than the target information. The sample set can be obtained from any data source, or pre-collected or set by technicians.
[0080] Step 2, obtain the user representation model to be trained and the information attribute feature extraction network.
[0081] In this step, the user representation model to be trained can include the user behavior feature extraction network to be trained, the user attribute feature extraction network, the capsule network, the fully connected network (such as Figure 4 the MLP shown), and the attention network. The user representation model to be trained and the information attribute feature extraction network to be trained can be pre-set or built by technicians.
[0082] Step 3: Based on the machine learning method, use the sample set to train the initial user representation model and the initial information attribute feature extraction network.
[0083] In this step, specifically, the behavior data and attribute data of the user included in each sample in the sample set can be used as the input of the user behavior feature extraction network and the user attribute feature extraction network included in the user representation model to be trained respectively. At the same time, the attribute data of the target information or non-target information is used as the input of the information attribute feature extraction network to be trained, and the matching degree between the user representation vector actually output by the attention network included in the user representation model and the information attribute feature vector actually output by the information attribute feature extraction network is obtained. Then, according to the matching degree and the preset loss function, it is determined whether the training is completed. If it is determined that the training is not completed, the network parameters of the user representation model to be trained and the information attribute feature extraction network can be adjusted by using algorithms such as backpropagation and gradient descent according to the value of the loss function, and resampling can be continued for training until it is determined that the training is completed, and then the trained user representation model and the information attribute feature extraction network can be obtained. In the subsequent information push process, at least two user representation vectors can be obtained by using the trained user representation model, and the information attribute feature vectors corresponding to the information in the information set can be extracted by using the trained information attribute feature extraction network.
[0084] Among them, the loss function can be flexibly set by technicians according to actual application requirements. As an example, the objective function can be set based on the following formula:
[0085]
[0086]
[0087] Among them, L represents the loss function. u represents the user, and i and j represent information (including target information and non-target information). Pr(i|u) represents the matching degree (or interaction probability) between the user and the information. represents the user representation vector. represents the information attribute feature vector. τ represents the information data set. represents the interaction data set between the user and the information.
[0088] In some scenarios, the interaction data set between the user and the information and / or the information data set may be relatively large. Since the calculation process of the above loss function involves calculations such as cumulative summation, there will be a large computational overhead in these scenarios. Therefore, various techniques such as Sampled Softmax and nearest neighbor retrieval can be used for model training to reduce the computational overhead and improve the training efficiency.
[0089] To improve the model training effect, various methods can be flexibly adopted to complete model training during the specific training process. For example, according to the actual application scenario or application requirements, the sample set can be divided into a training set, an evaluation set, and a validation set in a specified proportion to train the model more effectively. Another example is that the training effect can be evaluated based on various indicators such as Hitrate (hit rate) and NDCG (Normalized Discounted Cumulative Gain). During the actual application process of the user representation model and the information attribute feature extraction model, the user representation model and the information attribute feature extraction model can also be continuously updated according to indicators such as CTR (Click-Through Rate) and CVR (Conversion Rate).
[0090] After generating multiple user representation vectors using the capsule network in the above embodiments of the present disclosure, the attention mechanism is used to perform attention calculation on the multiple user representation vectors, so as to avoid directly using the user representation vectors output by the capsule network for matching degree calculation, which may lose some information beneficial to information push, thereby ensuring that the information contained in the user representation vectors output by the capsule network can be more fully mined and utilized, and further helping to improve the subsequent information push effect.
[0091] Further referring to Figure 5 , which shows the flow 500 of another embodiment of the method for pushing information. The flow 500 of the method for pushing information includes the following steps:
[0092] Step 501, obtain a behavior feature vector for characterizing the real-time behavior features and historical behavior features of the user, as well as obtain the attribute feature vector of the user and the attribute feature vectors of each piece of information in the information set. For example, the behavior features of the user in the past two days can be obtained as the real-time behavior features.
[0093] In this embodiment, the behavior feature vector of the user can be used to characterize the real-time behavior features and historical behavior features of the user, so as to analyze the user's interests using more comprehensive behavior features, and further help improve the accuracy and timeliness of the subsequent obtained user representation vectors.
[0094] Step 502, input the behavior feature vector of the user into a pre-trained capsule network to generate at least two capsule vectors for characterizing the interests of the user.
[0095] Step 503, splice at least two capsule vectors with the attribute feature vector of the user respectively, and generate at least two user representation vectors for characterizing the user according to the splicing result.
[0096] Step 504: Determine the matching degree between the attribute feature vector of the information in the information set and the user representation vector, and select information from the information set for pushing according to the determined matching degree.
[0097] In this embodiment, the content not specifically described above may refer to Figure 2 and Figure 3 the relevant descriptions in the corresponding embodiments, which will not be elaborated here.
[0098] Continue to refer to Figure 6 , Figure 6 which is a schematic application scenario 600 of the method for pushing information according to this embodiment. In the Figure 6 application scenario, the user attribute data may include the first attribute value, the second attribute value, the third attribute value, and the fourth attribute value corresponding to the first attribute, the second attribute, the third attribute, and the fourth attribute respectively. The user behavior data includes the items and brands purchased. Among them, the purchased items include thermal protection items, down jackets, and mobile phones, and the purchased brands include "A", "B", "C", "D", "E", "F", and "G". Then, at least the user representation vector is obtained by using a pre-trained user representation model, such as the user interest clusters "User Embedding1", "User Embedding2",... "UserEmbeddingN" shown in the figure. At the same time, the information attribute features of the information in the information set are encoded by using a pre-trained information attribute feature extraction network to obtain the information attribute feature vector (such as the Brand Embedding shown in the figure). Among them, the information in the information set indicates items, and the information attribute data includes item brand information. Specifically, the item brand information includes the item category (such as first-level or second-level or third-level), the popularity of the item, and the richness of the item. Furthermore, the matching degree between the user representation vector and the information attribute feature vector can be calculated, and the items and brand information preferred by the user can be determined according to the calculation result. Furthermore, certain items of certain brands are pushed to the user based on the determined items and brand information preferred by the user.
[0099] Specifically, as shown in the figure, the determined items of user preference include thermal protection items of brands "A" and "B" that the user has purchased, down jackets of brands "C", "D", and "E", and mobile phones of brands "F" and "G". Generalized therefrom are thermal protection items of brand "D", heat patches of brand "H", masks of brand "I", down jackets of brands "J", "K", and "L", cotton-padded clothes of brands "D" and "M", jackets of brand "D", casual pants of brand "L", running shoes of brand "N", mobile phones of brands "O", "P", and "Q", creative accessories of brands "F" and "R", mobile power supplies of brand "F", chargers of brand "G", mobile phone cases or protective covers of brand "S", network boxes of brand "Q", etc., with which the user has had no direct interaction before. Furthermore, various thermal protection items, heat patches, masks, down jackets, cotton-padded clothes, jackets, sweatshirts, casual pants, running shoes, mobile phones, mobile power supplies, creative accessories, mobile phone cases, etc. of the above generalized brands can be pushed to the user.
[0100] The method provided by the above embodiments of the present disclosure combines the real-time behavior characteristics and historical behavior characteristics of the user to more comprehensively analyze the user's interest preferences, so as to generate a more accurate user representation vector, thereby improving the real-time push effect of information and avoiding problems such as poor timeliness of the pre-trained model.
[0101] Further referring to Figure 7 , as an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a device for pushing information. This device embodiment corresponds to Figure 2 the method embodiment shown, and this device can be specifically applied to various electronic devices.
[0102] As Figure 7 shown, the device 700 for pushing information provided in this embodiment includes an acquisition unit 701, a generation unit 702, and a push unit 703. Among them, the acquisition unit 701 is configured to acquire the behavior feature vector and attribute feature vector of the user, and respectively acquire the attribute feature vectors of each piece of information in the information set; the generation unit 702 is configured to input the behavior feature vector of the user into a pre-trained capsule network to generate at least two capsule vectors for characterizing the user's interests; splice at least two capsule vectors with the attribute feature vector of the user respectively, and generate at least two user representation vectors for characterizing the user according to the splicing result; the push unit 703 is configured to determine the matching degree between the attribute feature vector of the information in the information set and the user representation vector, and select information from the information set for pushing according to the determined matching degree.
[0103] In this embodiment, in the apparatus 700 for pushing information: the specific processing of the obtaining unit 701, the generating unit 702, and the pushing unit 703 and the technical effects brought by them can be respectively referred to Figure 2 the relevant descriptions of steps 201, 202, 203, and 204 in the corresponding embodiment, which will not be elaborated here.
[0104] In some alternative implementation manners of this embodiment, the above generating unit 702 is further configured to: according to the splicing result, generate at least two user representation vectors for characterizing the user by using an attention mechanism.
[0105] In some alternative implementation manners of this embodiment, the above generating unit 702 is further configured to: input the splicing result into a pre-trained fully connected network to obtain at least two initial user representation vectors; input the at least two initial user representation vectors into an attention network implemented based on an attention mechanism to obtain at least two user representation vectors.
[0106] In some alternative implementation manners of this embodiment, the above user behavior feature vector is used to characterize the real-time behavior feature and historical behavior feature of the user.
[0107] In some alternative implementation manners of this embodiment, the above obtaining unit 701 is further configured to: obtain the behavior data of the user, and input the behavior data of the user into a pre-trained user behavior feature extraction network to obtain the behavior feature vector of the user; obtain the attribute data of the user, and input the attribute data of the user into a pre-trained user attribute feature extraction network to obtain the attribute feature vector of the user; for the information in the information set, obtain the attribute data of the information, and input the attribute data of the information into a pre-trained information attribute feature extraction network to obtain the attribute feature vector of the information.
[0108] In some alternative implementation manners of this embodiment, the user representation model is trained through the following steps. The user representation model is composed of a user behavior feature extraction network, a user attribute feature extraction network, a capsule network, a fully connected network, and an attention network: obtain a sample set, where the sample set includes positive samples and negative samples. The positive samples include the attribute data of the user, the behavior data of the user, and the attribute data of the target information. The negative samples include the attribute data of the user, the behavior data of the user, and the attribute data of non-target information. The target information includes the subsequent behavior data corresponding to the behavior data of the user; obtain the user representation model to be trained and the information attribute feature extraction network; based on a machine learning method, use the sample set to train the user representation model to be trained and the information attribute feature extraction network.
[0109] The device provided by the above embodiments of the present disclosure obtains the behavior feature vector and attribute feature vector of the user, and respectively obtains the attribute feature vectors of each piece of information in the information set through an acquisition unit; a generation unit inputs the behavior feature vector of the user into a pre-trained capsule network to generate at least two capsule vectors for characterizing the user's interests; the generation unit is further configured to splice the at least two capsule vectors with the attribute feature vector of the user respectively, and generate at least two user representation vectors for characterizing the user according to the splicing result; a push unit determines the matching degree between the attribute feature vector of the information in the information set and the user representation vector, and selects information from the information set for pushing according to the determined matching degree, so as to use the capsule network to generate multiple user representation vectors to more comprehensively express the user's various interests, and can generalize information that the user has not directly interacted with, so as to improve the content richness and accuracy of the pushed information received by the user. In addition, generalizing more information that the user has not directly interacted with is also helpful to improve the interaction degree between some information and the user, and assist the cold start of some information, etc.
[0110] Reference is made below to Figure 8 , which shows a schematic structural diagram of an electronic device (such as Figure 1 the server in) 800 suitable for implementing the embodiments of the present disclosure. Figure 8 The server shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.
[0111] As Figure 8 shown, the electronic device 800 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 801, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage device 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the electronic device 800 are also stored. The processing device 801, the ROM 802, and the RAM 803 are connected to each other through a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0112] Generally, the following devices may be connected to the I / O interface 805: an input device 806 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 807 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 808 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 809. The communication device 809 may allow the electronic device 800 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 8An electronic device 800 with various devices is shown, but it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had. Figure 8 Each block shown in [it] may represent a device or, as required, multiple devices.
[0113] In particular, according to an embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a computer-readable medium, and the computer program contains program code for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device 809, or installed from a storage device 808, or installed from a ROM 802. When the computer program is executed by a processing device 801, the above functions defined in the methods of the embodiments of the present disclosure are performed.
[0114] It should be noted that the computer-readable medium described in the embodiments of the present disclosure may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the embodiments of the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the embodiments of the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, and the computer-readable signal medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0115] The above computer-readable medium may be included in the above electronic device; or may exist separately without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device is caused to: obtain the behavior feature vector and the attribute feature vector of the user, and respectively obtain the attribute feature vectors of the respective information in the information set; input the behavior feature vector of the user into a pre-trained capsule network to generate at least two capsule vectors for characterizing the interests of the user; splice the at least two capsule vectors with the attribute feature vector of the user respectively, and generate at least two user representation vectors for characterizing the user according to the splicing result; determine the matching degree between the attribute feature vector of the information in the information set and the user representation vector, and select information from the information set for pushing according to the determined matching degree.
[0116] Computer program code for performing the operations of the embodiments of the present disclosure may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, by connecting through the Internet using an Internet service provider).
[0117] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0118] The units involved in the embodiments described in this disclosure can be implemented in software or in hardware. The described units can also be provided in a processor. For example, it can be described as: a processor includes an acquisition unit, a generation unit, and a push unit. Among them, the names of these units do not constitute a limitation on the unit itself in some cases. For example, the acquisition unit can also be described as "a unit that acquires the behavior feature vector and attribute feature vector of a user, and respectively acquires the attribute feature vectors of each piece of information in the information set".
[0119] The above description is only a preferred embodiment of this disclosure and an explanation of the technical principles applied. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features having similar functions disclosed in the embodiments of this disclosure.
Claims
1. A method for pushing information, comprising: Obtain the user's behavioral feature vector and attribute feature vector, and respectively obtain the attribute feature vectors of each piece of information in the information set, where the attribute feature vector of the information is generated based on the attribute values representing various attributes of the information; Input the user's behavioral feature vector into a pre-trained capsule network to generate at least two capsule vectors for representing the user's interests; Concatenate the at least two capsule vectors with the user's attribute feature vector respectively, and generate at least two user representation vectors for representing the user according to the concatenation result; Determine the matching degree between the attribute feature vector of the information in the information set and the user representation vector, and select information from the information set for pushing according to the determined matching degree; Among them, the selecting information from the information set for pushing according to the determined matching degree includes: Taking the maximum value of at least two matching degrees corresponding to the information in the information set or the average value of at least two matching degrees as the target matching degree corresponding to the information; Select the information with the target matching degree greater than the preset matching degree threshold from the information set for pushing according to the magnitudes of the target matching degrees corresponding to each piece of information.
2. The method according to claim 1, wherein, The generating at least two user representation vectors for representing the user according to the concatenation result includes: Generating at least two user representation vectors for representing the user by using an attention mechanism according to the concatenation result.
3. The method according to claim 2, wherein, The generating at least two user representation vectors for representing the user by using an attention mechanism according to the concatenation result includes: Input the concatenation result into a pre-trained fully connected network to obtain at least two initial user representation vectors; Input the at least two initial user representation vectors into an attention network implemented based on the attention mechanism to obtain at least two user representation vectors.
4. The method according to claim 1, wherein, The user's behavioral feature vector is used to represent the user's real-time behavioral features and historical behavioral features.
5. The method according to claim 3, wherein, The obtaining the user's behavioral feature vector and attribute feature vector, and respectively obtaining the attribute feature vectors of each piece of information in the information set includes: Obtain the user's behavioral data, and input the user's behavioral data into a pre-trained user behavioral feature extraction network to obtain the user's behavioral feature vector; Obtain the user's attribute data, and input the user's attribute data into a pre-trained user attribute feature extraction network to obtain the user's attribute feature vector; For the information in the information set, obtain the attribute data of the information, and input the attribute data of the information into a pre-trained information attribute feature extraction network to obtain the attribute feature vector of the information.
6. The method according to claim 5, wherein, The user representation model is trained through the following steps, and the user representation model is composed of the user behavioral feature extraction network, user attribute feature extraction network, capsule network, fully connected network and attention network: Obtain a sample set, where the sample set includes positive samples and negative samples. The positive samples include the attribute data of the user, the behavior data of the user, and the attribute data of the target information. The negative samples include the attribute data of the user, the behavior data of the user, and the attribute data of non-target information. The target information includes the subsequent behavior data corresponding to the behavior data of the user; Obtain a user representation model to be trained and an information attribute feature extraction network; Based on a machine learning method, use the sample set to train the user representation model to be trained and the information attribute feature extraction network.
7. An apparatus for pushing information, comprising: An acquisition unit configured to obtain the behavior feature vector and attribute feature vector of the user, and respectively obtain the attribute feature vectors of each piece of information in the information set. The attribute feature vector of the information is generated based on the attribute values representing various attributes of the information; A generation unit configured to input the behavior feature vector of the user into a pre-trained capsule network to generate at least two capsule vectors for representing the interests of the user; The generation unit is further configured to splice the at least two capsule vectors with the attribute feature vector of the user respectively, and generate at least two user representation vectors for representing the user according to the splicing result; A push unit configured to determine the matching degree between the attribute feature vector of the information in the information set and the user representation vector, and select information from the information set for pushing according to the determined matching degree; Wherein, the push unit is further configured to: Take the maximum value of at least two matching degrees corresponding to the information in the information set or the average value of at least two matching degrees as the target matching degree corresponding to the information; Select the information in the information set with a target matching degree greater than a preset matching degree threshold for pushing according to the magnitudes of the target matching degrees corresponding to each piece of information.
8. The apparatus according to claim 7, wherein, The generation unit is further configured to: Generate at least two user representation vectors for representing the user by using an attention mechanism according to the splicing result.
9. The apparatus according to claim 7, wherein, The behavior feature vector of the user is used to represent the real-time behavior feature and historical behavior feature of the user.
10. An electronic device, comprising: One or more processors; A storage device storing one or more programs thereon; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-6.
11. A computer-readable medium having a computer program stored thereon, wherein, When the program is executed by the processor, it implements the method according to any one of claims 1-6.
Citation Information
Patent Citations
Method and device for generating information
CN112395490A
Session recommendation method based on multi-interest capsule network
CN112765461A