User feature extraction method and device, equipment and storage medium

CN112231572BActive Publication Date: 2026-09-18TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011162452.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-10-27
Publication Date
2026-09-18
Estimated Expiration
2040-10-27

AI Technical Summary

Technical Problem

[0004]因此,用户向量的提取至关重要,但目前方案提取得到的用户向量不够准确

Benefits of technology

[0018]By using a semantic extraction model to perform context-based feature extraction on the content browsing sequence of the target user, context-based content vectors are obtained for each content in the content preview sequence. Then, user vectors for the target user are constructed from these content vectors. This process fully considers the contextual information of the content in the content browsing sequence during the extraction of user vectors, so that the final user vectors can more accurately reflect the user's characteristics in dimensions such as relevance, order, or preference among the content viewed, thus improving the accuracy of user vectors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112231572B_ABST
    Figure CN112231572B_ABST
Patent Text Reader

Abstract

The application discloses a user feature extraction method and device, equipment and a storage medium, and relates to the technical field of machine learning of artificial intelligence. The method comprises the following steps: obtaining a content browsing sequence of a target user, wherein the content browsing sequence of the target user comprises n contents browsed by the target user, and n is a positive integer; performing context-based feature extraction processing on the content browsing sequence by using a semantic extraction model to obtain context-based content vectors corresponding to the n contents respectively; and generating a user vector of the target user according to the context-based content vectors corresponding to the n contents respectively. In the extraction process of the user vector, the context information of the contents in the content browsing sequence is fully considered, so that the finally obtained user vector can more accurately reflect the correlation, sequence or preference dimension features between the contents browsed by the user, and the accuracy of the user vector is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of machine learning technology in artificial intelligence, and in particular to a method, apparatus, device, and storage medium for extracting user features. Background Technology

[0002] Currently, in some content push scenarios (such as article push, news push, etc.), content that similar users have viewed is pushed to target users, thereby increasing the click-through rate of the pushed content.

[0003] In related technologies, machine learning techniques are used to determine user vectors that reflect the user characteristics of a target user based on the target user's historical browsing content. Then, based on the similarity between user vectors of different users, similar users of the target user are identified. After that, it is possible to push content viewed by similar users to the target user.

[0004] Therefore, user vector extraction is crucial, but the user vectors extracted by current solutions are not accurate enough. Summary of the Invention

[0005] This application provides a method, apparatus, device, and storage medium for extracting user features, which can improve the accuracy of the extracted user vectors and more accurately reflect user features. The technical solution is as follows:

[0006] According to one aspect of the embodiments of this application, a method for extracting user features is provided, the method comprising:

[0007] Obtain the content browsing sequence of the target user, wherein the content browsing sequence of the target user includes n pieces of content that the target user has browsed, where n is a positive integer;

[0008] The content viewing sequence is subjected to context-based feature extraction processing by a semantic extraction model to obtain context-based content vectors corresponding to the n contents; wherein, the context-based content vector refers to the feature vector representation of the contents considering the context information of the contents in the content viewing sequence.

[0009] Based on the context-based content vectors corresponding to the n content items, a user vector for the target user is generated, and the user vector for the target user is used to characterize the user features of the target user.

[0010] According to one aspect of the embodiments of this application, a user feature extraction apparatus is provided, the apparatus comprising:

[0011] The reading sequence acquisition module is used to acquire the content reading sequence of the target user, wherein the content reading sequence of the target user includes n pieces of content that the target user has read, and n is a positive integer;

[0012] The content vector extraction module is used to perform context-based feature extraction processing on the content viewing sequence through a semantic extraction model to obtain context-based content vectors corresponding to the n contents respectively; wherein, the context-based content vector refers to the feature vector representation that takes into account the context information of the content in the content viewing sequence;

[0013] The user vector generation module is used to generate the user vector of the target user based on the context-based content vectors corresponding to the n contents, and the user vector of the target user is used to characterize the user features of the target user.

[0014] According to one aspect of the embodiments of this application, a computer device is provided, the computer device including a processor and a memory, the memory storing at least one instruction, at least one program, code set or instruction set, the at least one instruction, the at least one program, the code set or instruction set being loaded and executed by the processor to implement the above-described method for extracting user features.

[0015] According to one aspect of the embodiments of this application, a computer-readable storage medium is provided, wherein at least one instruction, at least one program, code set, or instruction set is stored in the storage medium, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the above-described method for extracting user features.

[0016] According to one aspect of the embodiments of this application, a computer program product or computer program is provided, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the aforementioned user feature extraction method.

[0017] The technical solutions provided in this application embodiment may have the following beneficial effects:

[0018] By using a semantic extraction model to perform context-based feature extraction on the content browsing sequence of the target user, context-based content vectors are obtained for each content in the content preview sequence. Then, user vectors for the target user are constructed from these content vectors. This process fully considers the contextual information of the content in the content browsing sequence during the extraction of user vectors, so that the final user vectors can more accurately reflect the user's characteristics in dimensions such as relevance, order, or preference among the content viewed, thus improving the accuracy of user vectors. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 This is a schematic diagram of the implementation environment of a solution provided in one embodiment of this application;

[0021] Figure 2 This is a flowchart of a user feature extraction method provided in one embodiment of this application;

[0022] Figure 3 This is a flowchart of a user feature extraction method provided in another embodiment of this application;

[0023] Figure 4 This is a schematic diagram of a semantic extraction model provided in one embodiment of this application;

[0024] Figure 5 This is a schematic diagram of the training process of a semantic extraction model provided in one embodiment of this application;

[0025] Figure 6 This is a block diagram of a user feature extraction device provided in one embodiment of this application;

[0026] Figure 7 This is a block diagram of a user feature extraction device provided in another embodiment of this application;

[0027] Figure 8 This is a block diagram of a computer device provided in one embodiment of this application. Detailed Implementation

[0028] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0029] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.

[0030] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0031] Machine learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learn-by-doing.

[0032] With the research and advancement of artificial intelligence (AI) technology, AI is being studied and applied in various fields, such as smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, autonomous driving, drones, robots, smart healthcare, and smart customer service. It is believed that with the development of technology, AI will be applied in more fields and play an increasingly important role.

[0033] The solution provided in this application relates to machine learning technology in artificial intelligence. It uses machine learning technology to train a semantic extraction model, and then uses this semantic extraction model to extract the user vector of the target user based on the target user's content browsing sequence.

[0034] The method provided in this application can be executed by a computer device, which refers to an electronic device with data computing, processing, and storage capabilities. This computer device can be a terminal such as a PC (Personal Computer), tablet computer, smartphone, wearable device, or intelligent robot; or it can be a server. The server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.

[0035] The technical solution provided in this application embodiment performs context-based feature extraction processing on the content viewing sequence of the target user through a semantic extraction model, obtaining context-based content vectors corresponding to each content in the content preview sequence. Then, user vectors of the target user are constructed from these content vectors. This ensures that the contextual information of the content in the content viewing sequence is fully considered during the extraction of user vectors, so that the final user vectors can more accurately reflect the user's characteristics in dimensions such as relevance, order, or preference among the viewed content, thereby improving the accuracy of user vectors.

[0036] In one example, such as Figure 1 As shown, taking a content push system as an example, the system may include a terminal 10 and a server 20.

[0037] Terminal 10 can be an electronic device such as a mobile phone, tablet computer, PC, or wearable device. Users can access server 20 through terminal 10 and perform content viewing operations. For example, a client of the target application can be installed on terminal 10, allowing users to access server 20 and perform content viewing operations. The target application can be any application that provides content viewing functionality, such as reading applications, video applications, news and information applications, social applications, instant messaging applications, and lifestyle service applications; this embodiment does not limit the scope of the application.

[0038] The content offered to users by different applications may vary. For example, the content may include different categories such as books, articles, news, and videos, and this application embodiment does not limit this.

[0039] Server 20 can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server providing cloud computing services. Server 20 is used to provide background services for the client of the target application in terminal 10. For example, server 20 can be the background server of the aforementioned target application.

[0040] Terminal 10 and server 20 can communicate via a network.

[0041] In an exemplary embodiment, the server 20 may use the method described in the following embodiments to generate a user vector of the target user, and then determine similar users of the target user based on the user vector of the target user and the user vectors of other users. After that, it can push content viewed by similar users to the target user's terminal 10.

[0042] Please refer to Figure 2 This document illustrates a flowchart of a user feature extraction method according to an embodiment of this application. The execution entity for each step of this method can be a computer device, such as a server or terminal. The method may include the following steps (201-203):

[0043] Step 201: Obtain the content browsing sequence of the target user, which includes n pieces of content that the target user has viewed, where n is a positive integer.

[0044] Content viewed by the target user refers to content that the target user has viewed or read. The above n pieces of content are arranged in a certain order to form a content viewing sequence. For example, the content viewing sequence can be represented as (t1, t2, t3, t4, ..., tt...). n ), where t k This indicates the k-th content in the content viewing sequence (e.g., the content at the k-th position from left to right), where k is a positive integer less than or equal to n.

[0045] Step 202: The content viewing sequence is processed by a semantic extraction model to extract features based on context, resulting in context-based content vectors for each of the n contents.

[0046] In this embodiment, the feature extraction method used is a context-based feature extraction method, that is, when extracting the content vector, the context information of the content in the content viewing sequence is considered. A context-based content vector refers to a feature vector representation that takes into account the context information of the content in the content viewing sequence. Here, the context information of the content refers to other content before and / or after the content in the content viewing sequence.

[0047] Additionally, a semantic extraction model is a machine learning model used to extract content vectors. For example, this semantic extraction model could be the ALBERT model.

[0048] Step 203: Generate the target user's user vector based on the context-based content vectors corresponding to each of the n content items.

[0049] The user vector of a target user is used to characterize the user characteristics of that target user. In this embodiment, since the user vector of a target user is determined based on the content vector in the content browsing sequence of the target user, the user vector of the target user can characterize the characteristics of the target user in terms of content browsing, such as the relevance, order, or preference of the content being viewed.

[0050] Optionally, the user vector of the target user can be generated by summing and averaging or weighted summing the context-based content vectors corresponding to the n content items.

[0051] In summary, the technical solution provided in this application performs context-based feature extraction processing on the content viewing sequence of the target user through a semantic extraction model, obtaining context-based content vectors corresponding to each content in the content preview sequence. Then, user vectors of the target user are constructed from these content vectors. This ensures that the contextual information of the content in the content viewing sequence is fully considered during the extraction of user vectors, so that the final user vectors can more accurately reflect the user's characteristics in dimensions such as relevance, order, or preference among the viewed content, thereby improving the accuracy of user vectors.

[0052] Please refer to Figure 3 This illustrates a flowchart of a user feature extraction method according to another embodiment of this application. The execution entity for each step of this method can be a computer device, such as a server or terminal. The method may include the following steps (301-310):

[0053] Step 301: Obtain the target user's content viewing history, which includes the content viewed by the target user and the viewing information for each piece of content.

[0054] Content viewing information is used to indicate the viewing status of the content. Optionally, the content viewing information includes, but is not limited to, at least one of the following: timestamps for each viewing, duration of each viewing, etc.

[0055] Optionally, the target user's content browsing history can be obtained by recording the target user's content browsing behavior within a certain time period. This time period can be any historical time period, or a time period extending backward from the current moment for a set duration, thereby recording the target user's content browsing behavior in a recent period.

[0056] Step 302: Determine the statistical indicators corresponding to each content based on the viewing information of each content.

[0057] Statistical metrics are used to rank content. Optionally, statistical metrics include, but are not limited to, at least one of the following: number of views, total viewing time, etc. The number of views by a target user for a specific piece of content refers to how many times the target user has viewed the content; for example, the number of views could be 1, 2, 3, etc. In one possible implementation, the number of views can be determined based on the number of timestamps for each view recorded in the content's viewing information. The total viewing time by a target user for a specific piece of content refers to the cumulative time the target user spends viewing the content. For example, if a target user has viewed the content twice, once for 1 minute and once for 2 minutes, then the total viewing time for the target user on the target content is 3 minutes. In one possible implementation, the total viewing time can be obtained by summing the durations of each view recorded in the content's viewing information.

[0058] Step 303: Sort the content according to statistical indicators to generate the target user's content viewing sequence. The target user's content viewing sequence includes n pieces of content that the target user has viewed, where n is a positive integer.

[0059] When the number of statistical indicators is 1, the content can be sorted directly according to the statistical indicator from largest to smallest or from smallest to largest to generate the content viewing sequence for the target user.

[0060] When the number of statistical indicators is greater than one, for each piece of content, a ranking indicator can be determined based on multiple statistical indicators of that content. Then, the content is sorted according to the ranking indicator in descending or ascending order to generate a content viewing sequence for the target user. The ranking indicator of a content can be calculated based on multiple statistical indicators of that content, such as using weighted summation, etc., and this embodiment of the application does not limit this calculation.

[0061] Step 304: Obtain the original content vectors corresponding to the n contents in the content viewing sequence.

[0062] The original content vector refers to a feature vector representation that is independent of the contextual information of the content in the content viewing sequence. In this embodiment, the concept of the original content vector corresponds to the context-based content vector. As introduced above, the context-based content vector is a vector extracted by the semantic extraction model that can reflect the contextual information of the content. Here, the original content vector is a vector that cannot reflect the contextual information of the content. In this embodiment, the original content vectors of each content in the content viewing sequence are used as input to the semantic extraction model, which extracts the contextual relationships of each content in the sequence and generates a context-based content vector for each content.

[0063] Optionally, for the i-th content among the n content items mentioned above, obtain the token vector, category vector, and position vector of the i-th content; based on the token vector, category vector, and position vector of the i-th content, determine the original content vector of the i-th content, where i is a positive integer less than or equal to n. Here, the token vector (token embedding) refers to the vector representation corresponding to the content's token, the category vector (token class embedding) refers to the vector representation corresponding to the content's category, and the position vector (position embedding) refers to the vector representation corresponding to the content's position in the content viewing sequence.

[0064] The content's identifier is its ID (Identity), used to uniquely identify the content; different content pieces have different identifiers. The content's category refers to the classification to which the content belongs. For example, if the content is a book, the category could include science fiction, suspense, romance, urban fiction, etc. The content's position in the content viewing sequence is its sorted position within that sequence, such as its number from left to right.

[0065] In this embodiment, the original content vector of the i-th content is determined based on its identifier vector, category vector, and location vector. For example, the original content vector of the i-th content is obtained by performing a summation, averaging, or weighted summation operation on its identifier vector, category vector, and location vector. Thus, the original content vector records the content's identifier, category, and location information.

[0066] Step 305: Input the original content vectors corresponding to each of the n contents into the semantic extraction model.

[0067] Step 306: Perform context-based feature extraction processing through the semantic extraction model to obtain context-based content vectors corresponding to each of the n contents.

[0068] For example, such as Figure 4As shown, the semantic extraction model includes an input layer 41, an encoder layer 42, a classifier layer 43, and an output layer 44. The input layer 41 takes the original content vectors of each content item in the content viewing sequence as input. The encoder layer 42 performs context-based feature extraction on the original content vectors input from the input layer 41, obtaining context-based content vectors for each content item. The classifier layer 43 performs the mapping process from the context-based content vectors to the classification results. The output layer 44 outputs the classification results obtained by the classifier layer 43. A description of the classifier layer 43 can be found in the model training example below, and will not be repeated here. The encoder layer 42 may include a transformer-structured encoder that considers the contextual information of the content when performing feature vector extraction.

[0069] Step 307: Summate the context-based content vectors corresponding to each of the n content items to obtain a summation vector.

[0070] Step 308: Divide each element in the summation vector by n to obtain the user vector of the target user.

[0071] For example, the user vector U of the target user is calculated using the following formula:

[0072]

[0073] Among them, O i This represents the context-based content vector of the i-th content in the content viewing sequence, where n is the number of content items contained in the content viewing sequence.

[0074] Step 309: Based on the user vector of the target user and the user vectors of other users, determine the similar users of the target user.

[0075] For example, the similarity between the target user's user vector and the user vectors of other users can be calculated. If the similarity is greater than a threshold, the other user is determined to be a similar user to the target user; if the similarity is less than the threshold, the other user is determined not to be a similar user to the target user. The similarity between the two user vectors can be calculated using Euclidean distance, cosine distance, etc., and this application does not limit this method.

[0076] Step 310: Based on the content browsing history of similar users, determine the push content to be provided to the target user.

[0077] For example, based on the content browsing history of similar users, the content recently viewed by those similar users can be retrieved and then used as push content to target users. Since the similarity between two users is determined based on user vectors, which reflect characteristics such as the relevance, order, or preferences of users in the content they view, similar users often have the same or similar reading habits and preferences. Pushing content in this way can help improve the click-through rate of the pushed content.

[0078] In summary, the technical solution provided in this application determines the original content vector of the content by using the content-based identifier vector, category vector, and location vector. This ensures that the original content vector records the identifier, classification, and location information of the content, thereby providing more valuable information for feature extraction in the semantic extraction model and improving the accuracy and robustness of the extracted context-based content vector.

[0079] In addition, by summing and averaging the context-based content vectors of each content in the content viewing sequence, a user vector is obtained, which provides a simple and efficient way to calculate the user vector and saves the device's computing overhead.

[0080] The training process of the semantic extraction model will be described below through examples. In an exemplary embodiment, the training process of the semantic extraction model can be as follows:

[0081] 1. Obtain the content browsing sequence of the sample user, which includes at least one piece of content that the sample user has viewed.

[0082] Optionally, the content browsing records of the sample users are obtained. These records include the content viewed by the sample users and the browsing information for each piece of content. Based on the browsing information for each piece of content, statistical indicators corresponding to each piece of content are determined. The content is then sorted according to the statistical indicators to generate a content browsing sequence for the sample users. This process is the same as or similar to the method for obtaining the content browsing sequence of the target users described above. For details, please refer to the description in the above embodiments, which will not be repeated here.

[0083] In addition, there are usually multiple sample users. For each sample user, a content viewing sequence can be generated based on that user's content viewing history.

[0084] 2. Based on the content browsing sequences of sample users, construct training samples for the semantic extraction model.

[0085] In this embodiment of the application, the sample data of the training samples includes the content browsing sequence of the sample users after partially masking the content items, and the tag data of the training samples includes the masked content items in the content browsing sequence of the sample users.

[0086] A sample user's content browsing sequence may include multiple content items (i.e., content). By masking some of these content items (such as one or more), the sample user's content browsing sequence after partial content masking is obtained. For example, if a sample user's content browsing sequence includes 5 content items, one of these items can be masked, or p (where p is a positive integer less than 5) of these items can be masked. The semantic extraction model then predicts the masked content items based on the unmasked content items.

[0087] In addition, during the generation of training samples, a random masking method can be used to randomly select some content items in the content browsing sequence of sample users to mask, thereby generating richer and more comprehensive training samples.

[0088] 3. The semantic extraction model is trained using training samples.

[0089] For example, such as Figure 5 As shown, suppose a sample user's content viewing sequence is (t1, t2, t3, t4, t5), and the original content vectors of each item in the sample user's content viewing sequence are denoted as (w1, w2, w3, w4, w5). In a training sample of this sample user, the fourth content item is randomly masked. Figure 5 The input data to the model input layer 41 is represented by [mask], which includes (w1, w2, w3, [mask], w5). The encoder layer 42 performs context-based feature extraction on the original content vectors of each content input to the input layer 41, obtaining context-based content vectors corresponding to each content. The classifier layer 43 performs the mapping process from the context-based content vectors corresponding to each content to the classification results. The output layer 44 outputs the classification results obtained by the classifier layer 43. The classifier layer 43 may include fully connected layers, activation function layers, and normalization layers, etc.

[0090] During model training, the loss function of the semantic extraction model can be calculated based on the prediction results and label data corresponding to the training samples. The prediction results corresponding to the training samples refer to the predicted information of masked content items in the content browsing sequence of the sample user output by the semantic extraction model. Then, the parameters of the semantic extraction model are adjusted according to this loss function. For example, by adjusting the model parameters to minimize the value of the loss function, the model's performance can be improved.

[0091] In this embodiment, the semantic extraction model can employ the ALBERT model. The ALBERT model is a commonly used language model in the field of natural language processing. Compared to the BERT model, the ALBERT model offers the following improvements in natural language processing:

[0092] (1) Factorized Embedding Parameterization. The size (denoted as E) of the word embeddings (or word vectors) in the BERT vocabulary is equivalent to the number of hidden nodes (H) in the transformer layer, so E = H. However, the actual size of the vocabulary (corresponding to the number of viewed content in this application) is generally very large, which leads to a large number of model parameters. To solve these problems, ALBERT proposed a factorization-based method. This method does not directly map one-hot encoding to the hidden layer, but first maps one-hot encoding to a low-dimensional space, and then maps it to the hidden layer. This is actually similar to performing matrix factorization.

[0093] (2) Cross-layer parameter sharing. ALBERT proposed that parameters can be shared between different layers of the model, so that the number of parameters will not increase with the number of layers.

[0094] (3) Inter-sentence coherence loss. In BERT training, the next sentence prediction loss was proposed, which involves giving two sentence segments and having BERT predict their order. However, ALBERT pointed out that this approach is problematic and not very useful. ALBERT proposed sentence-order prediction loss (SOP), which predicts whether two sentences have been swapped based on topic association.

[0095] In this embodiment of the application, the training process of the ALBERT model has been further improved, including the following points:

[0096] (1) The loss function has been simplified, requiring only the calculation of the prediction loss for the masked content item. Furthermore, the weight information of the training samples is incorporated into the loss calculation. This loss function calculation method adopted in this application not only reduces the complexity of the model but also improves the final performance.

[0097] In an exemplary embodiment, the loss function of the semantic extraction model is calculated as follows:

[0098] (a) Determine the weights corresponding to the training samples;

[0099] Optionally, the viewing information of the masked content items in the training samples can be obtained, and the weights corresponding to the training samples can be determined based on the viewing information of the masked content items in the training samples. In this way, for the same user, different masked content items contribute differently to the loss.

[0100] (b) Calculate the loss function of the semantic extraction model based on the prediction results, label data and weights corresponding to the training samples;

[0101] For example, the formula for calculating the loss function Loss is as follows:

[0102]

[0103] Where N represents the number of training samples, L i w represents the loss for the i-th training sample. i Let y represent the weight corresponding to the i-th training sample, M represent the number of categories, and y represent the weight of the i-th training sample. ic Let p represent the label data of the i-th training sample. This label data can be a vector consisting of 0s and 1s (the total number of 0s and 1s is M), where 1 represents a masked item and 0 represents no prediction. ic This represents the prediction result for the i-th training sample, which can be a vector consisting of M probability values. The closer the vector corresponding to the label data is to the vector corresponding to the prediction result, the better the model's prediction performance.

[0104] (c) Adjust the parameters of the semantic extraction model according to the loss function.

[0105] For example, the model's performance can be improved by adjusting the model parameters to minimize the value of the loss function.

[0106] In this embodiment, when calculating the model loss function, the weights corresponding to each training sample are considered. Since the weights are determined based on the viewing information of the masked content items in the training samples, for the same user, the contribution of different masked content items to the loss is different. This allows for the selective strengthening of the influence of some content items on the loss and the weakening of the influence of other content items on the loss, thereby improving the flexibility and accuracy of loss calculation.

[0107] (2) Since the amount of content viewed is large (e.g., millions), directly sampling multi-class cross-entropy as the loss function will make the model's complexity difficult to train due to the excessive number of parameters in the last fully connected layer. Therefore, the loss function is optimized by using negative sampling, which can make the model converge quickly.

[0108] (3) In order to accelerate the training speed of the model, a multi-GPU (Graphics Processing Unit) training scheme is used, so that the training time of the model decreases as the number of GPUs increases.

[0109] In this embodiment, the ALBERT Transformer model architecture is used to construct a powerful semantic extraction model. Compared with traditional methods, this approach can predict a content item based on global context, resulting in a more robust user vector. Compared to BERT, ALBERT reduces model complexity by using parameter sharing and input matrix factorization. Furthermore, the training scheme is optimized by employing negative sampling to accelerate model training and ensure the model maintains periodic updates.

[0110] Taking book recommendations to users in reading apps as an example, there are numerous recommendation scenarios in reading apps, such as recommending new books to users who might be interested, recommending articles from WeChat official accounts based on personalization, and sharing articles based on likes. For these recommended books / articles, it is necessary to distribute them to other similar users based on the users who clicked or liked them, thereby increasing the number of recommendations and the accuracy of the target audience.

[0111] Experiments have shown that the similar user packages obtained using the technical solution of this application have significantly improved click-through rates, whether pushing new books to potentially interested users, pushing personalized articles from WeChat official accounts, or spreading the word through likes on WeChat official accounts.

[0112] The following are embodiments of the apparatus described in this application, which can be used to execute the embodiments of the method described in this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the method described in this application.

[0113] Please refer to Figure 6 This diagram illustrates a block diagram of a user feature extraction apparatus according to an embodiment of this application. The apparatus has the functionality to implement the method example described above; this functionality can be implemented in hardware or by hardware executing corresponding software. The apparatus can be the computer device described above, or it can be installed within a computer device. Figure 6 As shown, the device 600 includes: a reading sequence acquisition module 610, a content vector extraction module 620, and a user vector generation module 630.

[0114] The reading sequence acquisition module 610 is used to acquire the content reading sequence of the target user, wherein the content reading sequence of the target user includes n pieces of content that the target user has read, and n is a positive integer.

[0115] The content vector extraction module 620 is used to perform context-based feature extraction processing on the content viewing sequence through a semantic extraction model to obtain context-based content vectors corresponding to the n contents respectively; wherein, the context-based content vector refers to the feature vector representation that takes into account the context information of the contents in the content viewing sequence.

[0116] User vector generation module 630 is used to generate user vectors of the target user based on the context-based content vectors corresponding to the n contents respectively, and the user vectors of the target user are used to characterize the user features of the target user.

[0117] In an exemplary embodiment, such as Figure 7 As shown, the content vector extraction module 620 includes: a vector acquisition unit 621, a vector input unit 622, and a feature extraction unit 623.

[0118] The vector acquisition unit 621 is used to acquire the original content vectors corresponding to the n contents respectively. The original content vectors refer to the feature vector representations that are independent of the context information of the contents in the content viewing sequence.

[0119] The vector input unit 622 is used to input the original content vectors corresponding to the n contents into the semantic extraction model.

[0120] The feature extraction unit 623 is used to perform the context-based feature extraction process through the semantic extraction model to obtain the context-based content vectors corresponding to the n contents.

[0121] In an exemplary embodiment, the vector acquisition unit 621 is used for:

[0122] For the i-th content among the n content items, obtain the identifier vector, category vector, and position vector of the i-th content item; wherein, the identifier vector is the vector representation corresponding to the identifier of the content, the category vector is the vector representation corresponding to the category to which the content belongs, and the position vector is the vector representation corresponding to the position of the content in the content viewing sequence;

[0123] Based on the identifier vector, category vector, and position vector of the i-th content, the original content vector of the i-th content is determined, where i is a positive integer less than or equal to n.

[0124] In an exemplary embodiment, the user vector generation module 630 is configured to:

[0125] The context-based content vectors corresponding to the n contents are summed to obtain a summation vector;

[0126] Divide each element of the summation vector by n to obtain the user vector of the target user.

[0127] In an exemplary embodiment, the training process of the semantic extraction model is as follows:

[0128] Obtain the content browsing sequence of the sample user, wherein the content browsing sequence of the sample user includes at least one piece of content that the sample user has browsed;

[0129] Based on the content browsing sequence of the sample users, training samples for the semantic extraction model are constructed. The sample data of the training samples includes the content browsing sequence of the sample users after partial content item masking, and the tag data of the training samples includes the masked content items in the content browsing sequence of the sample users.

[0130] The semantic extraction model is trained using the training samples.

[0131] In an exemplary embodiment, training the semantic extraction model using the training samples includes:

[0132] Determine the weights corresponding to the training samples;

[0133] The loss function of the semantic extraction model is calculated based on the prediction results, label data, and weights corresponding to the training samples; wherein, the prediction results corresponding to the training samples refer to the prediction information of the masked content items in the content browsing sequence of the sample user output by the semantic extraction model.

[0134] The parameters of the semantic extraction model are adjusted based on the loss function.

[0135] In an exemplary embodiment, determining the weights corresponding to the training samples includes:

[0136] Obtain the viewing information of the masked content items in the training samples;

[0137] The weights corresponding to the training samples are determined based on the viewing information of the masked content items in the training samples.

[0138] In an exemplary embodiment, the reading sequence acquisition module 610 is configured to:

[0139] Obtain the content browsing history of the target user, the content browsing history including the content viewed by the target user and the browsing information of each content;

[0140] Based on the viewing information of each content, determine the statistical indicators corresponding to each content;

[0141] The content is sorted according to the statistical indicators to generate the content viewing sequence for the target user.

[0142] In an exemplary embodiment, such as Figure 7 As shown, the device 600 further includes: a similar user identification module 640 and a push content providing module 650.

[0143] The similar user determination module 640 is used to determine similar users of the target user based on the user vector of the target user and the user vectors of other users.

[0144] The push content providing module 650 is used to determine the push content to be provided to the target user based on the content browsing history of the similar users.

[0145] In summary, the technical solution provided in this application performs context-based feature extraction processing on the content viewing sequence of the target user through a semantic extraction model, obtaining context-based content vectors corresponding to each content in the content preview sequence. Then, user vectors of the target user are constructed from these content vectors. This ensures that the contextual information of the content in the content viewing sequence is fully considered during the extraction of user vectors, so that the final user vectors can more accurately reflect the user's characteristics in dimensions such as relevance, order, or preference among the viewed content, thereby improving the accuracy of user vectors.

[0146] It should be noted that the apparatus provided in the above embodiments is only illustrated by the division of the above functional modules when implementing its functions. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0147] Please refer to Figure 8 This diagram illustrates a structural block diagram of a computer device according to an embodiment of this application. This computer device can be used to implement the user feature extraction method provided in the above embodiments. Specifically:

[0148] The computer device 800 includes a processing unit (such as a CPU, GPU, and FPGA) 801, a system memory 804 including RAM (Random-Access Memory) 802 and ROM (Read-Only Memory) 803, and a system bus 805 connecting the system memory 804 and the CPU 801. The computer device 800 also includes a basic input / output system 806 to facilitate information transfer between various devices within the server, and a large-capacity storage device 807 for storing the operating system 813, application programs 814, and other program modules 815.

[0149] The basic input / output system 806 includes a display 808 for displaying information and an input device 809 for user input, such as a mouse or keyboard. Both the display 808 and the input device 809 are connected to the central processing unit 801 via an input / output controller 810 connected to the system bus 805. The basic input / output system 806 may also include the input / output controller 810 for receiving and processing input from multiple other devices such as a keyboard, mouse, or electronic stylus. Similarly, the input / output controller 810 also provides output to a display screen, printer, or other types of output devices.

[0150] The mass storage device 807 is connected to the central processing unit 801 via a mass storage controller (not shown) connected to the system bus 805. The mass storage device 807 and its associated computer-readable media provide non-volatile storage for the computer device 800. That is, the mass storage device 807 may include computer-readable media (not shown) such as a hard disk or a CD-ROM (Compact Disc Read-Only Memory) drive.

[0151] Without loss of generality, the computer-readable medium may include computer storage media and communication media. Computer storage media include volatile and non-volatile, removable and non-removable media implemented using any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include RAM, ROM, EPROM (Erasable Programmable Read-Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), flash memory or other solid-state storage technologies, CD-ROM, DVD (Digital Video Disc) or other optical storage, magnetic tape cassettes, magnetic tape, disk storage, or other magnetic storage devices. Of course, those skilled in the art will recognize that the computer storage medium is not limited to the above-mentioned types. The system memory 804 and mass storage device 807 described above can be collectively referred to as memory.

[0152] According to an embodiment of this application, the computer device 800 can also be connected to a remote computer on a network, such as the Internet, for operation. That is, the computer device 800 can be connected to a network 812 via a network interface unit 811 connected to the system bus 805, or it can also use the network interface unit 811 to connect to other types of networks or remote computer systems (not shown).

[0153] The memory also includes at least one instruction, at least one program, code set, or instruction set, which is stored in the memory and configured to be executed by one or more processors to implement the above-mentioned user feature extraction method.

[0154] In one exemplary embodiment, a computer-readable storage medium is also provided, the storage medium storing at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, the at least one program, the code set, or the instruction set, when executed by a processor, implements the above-described method for extracting user features.

[0155] Optionally, the computer-readable storage medium may include: ROM (Read-Only Memory), RAM (Random-Access Memory), SSD (Solid State Drives), or optical disc, etc. The random access memory may include ReRAM (Resistance Random Access Memory) and DRAM (Dynamic Random Access Memory).

[0156] In one exemplary embodiment, a computer program product or computer program is also provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the computer device to perform the aforementioned user feature extraction method.

[0157] It should be understood that "multiple" as used herein refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. Furthermore, the step numbers described herein are merely illustrative of one possible execution order. In some other embodiments, the steps may not be executed in numerical order, such as two steps with different numbers being executed simultaneously, or two steps with different numbers being executed in the reverse order of the illustration. This application does not limit this.

[0158] The above description is merely an exemplary embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A method of extracting user features, characterized by, The method includes: Obtain the content browsing history of the target user, wherein the content browsing history includes n pieces of content that the target user has viewed and the viewing information of each piece of content, where n is a positive integer; Based on the viewing information of the i-th content, multiple statistical indicators are determined, and the multiple statistical indicators are weighted and summed to obtain the ranking index of the i-th content, where i takes any positive integer from 1 to n. The n contents are sorted according to the sorting index to obtain the content viewing sequence of the target user; Based on the original content vector corresponding to the i-th content, a semantic extraction model is used to perform context-based feature extraction on the i-th content to obtain a context-based content vector corresponding to the i-th content. The original content vector corresponding to the i-th content is determined based on the identifier vector, category vector, and position vector of the i-th content. The context information of the i-th content includes other content in the content viewing sequence that is before and / or after the i-th content. Based on the context-based content vectors corresponding to the n content items, a user vector for the target user is generated, and the user vector for the target user is used to characterize the user features of the target user.

2. The method of claim 1, wherein, The step of performing context-based feature extraction on the i-th content based on the original content vector corresponding to the i-th content using a semantic extraction model to obtain the context-based content vector corresponding to the i-th content includes: Obtain the original content vectors corresponding to the n contents respectively. The original content vectors refer to the feature vector representations that are independent of the context information of the contents in the content viewing sequence. The original content vectors corresponding to the n contents are input into the semantic extraction model; The semantic extraction model performs context-based feature extraction on the i-th content based on the original content vector corresponding to the i-th content, thereby obtaining the context-based content vector corresponding to the i-th content.

3. The method of claim 2, wherein, The step of obtaining the original content vectors corresponding to the n contents includes: For the i-th content among the n content items, obtain the identifier vector, category vector, and position vector of the i-th content item; wherein, the identifier vector is the vector representation corresponding to the identifier of the content, the category vector is the vector representation corresponding to the category to which the content belongs, and the position vector is the vector representation corresponding to the position of the content in the content viewing sequence; Based on the identifier vector, category vector, and position vector of the i-th content, the original content vector of the i-th content is determined, where i is a positive integer less than or equal to n.

4. The method of claim 1, wherein, The step of generating the user vector of the target user based on the context-based content vectors corresponding to the n content items includes: The context-based content vectors corresponding to the n contents are summed to obtain a summation vector; Divide each element of the summation vector by n to obtain the user vector of the target user.

5. The method of claim 1, wherein, The training process of the semantic extraction model is as follows: Obtain the content browsing sequence of the sample user, wherein the content browsing sequence of the sample user includes at least one piece of content that the sample user has browsed; Based on the content browsing sequence of the sample users, training samples for the semantic extraction model are constructed. The sample data of the training samples includes the content browsing sequence of the sample users after partial content item masking, and the tag data of the training samples includes the masked content items in the content browsing sequence of the sample users. The semantic extraction model is trained using the training samples.

6. The method of claim 5, wherein, The step of training the semantic extraction model using the training samples includes: Determine the weights corresponding to the training samples; The loss function of the semantic extraction model is calculated based on the prediction results, label data, and weights corresponding to the training samples; wherein, the prediction results corresponding to the training samples refer to the prediction information of the masked content items in the content browsing sequence of the sample user output by the semantic extraction model. The parameters of the semantic extraction model are adjusted based on the loss function.

7. The method according to claim 6, characterized in that, Determining the weights corresponding to the training samples includes: Obtain the viewing information of the masked content items in the training samples; The weights corresponding to the training samples are determined based on the viewing information of the masked content items in the training samples.

8. The method according to any one of claims 1 to 7, characterized in that, After generating the user vector of the target user based on the context-based content vectors corresponding to the n content items, the method further includes: Based on the user vector of the target user and the user vectors of other users, determine the similar users of the target user; Based on the content browsing history of the similar users, the push content to be provided to the target user is determined.

9. A user feature extraction device, characterized in that, The device includes: The reading sequence acquisition module is used to acquire the content reading records of the target user, which include n pieces of content viewed by the target user and the reading information of each piece of content, where n is a positive integer; based on the reading information of the i-th piece of content, multiple statistical indicators are determined, and the multiple statistical indicators are weighted and summed to obtain the ranking index of the i-th piece of content, where i takes any positive integer from 1 to n; the n pieces of content are sorted according to the ranking index of the n pieces of content to obtain the content reading sequence of the target user; The content vector extraction module is used to perform context-based feature extraction processing on the i-th content based on the original content vector corresponding to the i-th content using a semantic extraction model, so as to obtain the context-based content vector corresponding to the i-th content; wherein, the original content vector corresponding to the i-th content is determined based on the identifier vector, category vector and position vector of the i-th content, and the context information of the i-th content includes other content located before and / or after the i-th content in the content viewing sequence; The user vector generation module is used to generate the user vector of the target user based on the context-based content vectors corresponding to the n contents, and the user vector of the target user is used to characterize the user features of the target user.

10. The apparatus according to claim 9, characterized in that, The content vector extraction module includes a vector acquisition unit, a vector input unit, and a feature extraction unit; The vector acquisition unit is used to acquire the original content vectors corresponding to the n contents respectively. The original content vectors refer to the feature vector representations that are independent of the context information of the contents in the content viewing sequence. The vector input unit is used to input the original content vectors corresponding to the n contents into the semantic extraction model; The feature extraction unit is used to perform the context-based feature extraction process on the i-th content based on the original content vector corresponding to the i-th content through the semantic extraction model, so as to obtain the context-based content vector corresponding to the i-th content.

11. The apparatus according to claim 10, characterized in that, The vector acquisition unit is used for: For the i-th content among the n content items, obtain the identifier vector, category vector, and position vector of the i-th content item; wherein, the identifier vector is the vector representation corresponding to the identifier of the content, the category vector is the vector representation corresponding to the category to which the content belongs, and the position vector is the vector representation corresponding to the position of the content in the content viewing sequence; Based on the identifier vector, category vector, and position vector of the i-th content, the original content vector of the i-th content is determined, where i is a positive integer less than or equal to n.

12. The apparatus according to claim 9, characterized in that, The user vector generation module is used for: The context-based content vectors corresponding to the n contents are summed to obtain a summation vector; Divide each element of the summation vector by n to obtain the user vector of the target user.

13. The apparatus according to claim 9, characterized in that, The training process of the semantic extraction model is as follows: Obtain the content browsing sequence of the sample user, wherein the content browsing sequence of the sample user includes at least one piece of content that the sample user has browsed; Based on the content browsing sequence of the sample users, training samples for the semantic extraction model are constructed. The sample data of the training samples includes the content browsing sequence of the sample users after partial content item masking, and the tag data of the training samples includes the masked content items in the content browsing sequence of the sample users. The semantic extraction model is trained using the training samples.

14. The apparatus according to claim 13, characterized in that, The step of training the semantic extraction model using the training samples includes: Determine the weights corresponding to the training samples; The loss function of the semantic extraction model is calculated based on the prediction results, label data, and weights corresponding to the training samples; wherein, the prediction results corresponding to the training samples refer to the prediction information of the masked content items in the content browsing sequence of the sample user output by the semantic extraction model. The parameters of the semantic extraction model are adjusted based on the loss function.

15. The apparatus according to claim 14, characterized in that, Determining the weights corresponding to the training samples includes: Obtain the viewing information of the masked content items in the training samples; The weights corresponding to the training samples are determined based on the viewing information of the masked content items in the training samples.

16. The apparatus according to any one of claims 9 to 15, characterized in that, The device also includes a similar user identification module and a push content provision module; The similar user determination module is used to determine similar users of the target user based on the user vector of the target user and the user vectors of other users; The push content providing module is used to determine the push content to be provided to the target user based on the content browsing history of the similar users.

17. A computer device, characterized in that, The computer device includes a processor and a memory, the memory storing at least one instruction, at least one program, a code set, or an instruction set, the at least one instruction, the at least one program, the code set, or the instruction set being loaded and executed by the processor to implement the user feature extraction method as described in any one of claims 1 to 8.

18. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the user feature extraction method as described in any one of claims 1 to 7.

19. A computer program product, characterized in that, The computer program product includes computer instructions stored in a computer-readable storage medium, a processor of a computer device reading the computer instructions from the computer-readable storage medium, and the processor executing the computer instructions to cause the computer device to perform the user feature extraction method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Method and device for training recommendation model, and recommendation system

    CN105589971A

  • Information push method based on internet-surfing log mining and user activity recognition

    CN105718579A