Method, device, equipment and medium for determining user characteristics and model training

By training the BERT model, the mask processes object identification that users are interested in but are not very popular. Combined with context information, the problem of inaccurate determination of user preference characteristics in the recommendation system is solved, and the access rate of recommended objects is improved.

CN112328778BActive Publication Date: 2025-08-12TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011209328.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-11-03
Publication Date
2025-08-12
Estimated Expiration
2040-11-03

AI Technical Summary

Technical Problem

Existing recommendation systems are difficult to accurately determine user preference characteristics, resulting in low access rates of recommended objects.

Method used

Using the BERT model training method, the object identification with low popularity but high interest to the user is processed by masking, and combining the context information of the user object sequence, the user's object feature vector is extracted, and the user's feature vector is then determined.

Benefits of technology

It improves the access rate of recommended objects and accurately reflects the user's preferences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112328778B_ABST
    Figure CN112328778B_ABST
Patent Text Reader

Abstract

The present application provides a method, apparatus, device and medium for determining user features and model training, wherein the method uses a trained BERT model as a language model for extracting object features of objects visited by users. In the process of training the BERT model, the probability of masking the object identifiers corresponding to objects that are not very popular but have a high degree of user interest will be increased, so that the trained BERT model can more accurately extract object features that can characterize user preferences from the user's object sequence. Moreover, based on the trained BERT model, the contextual information between the object identifiers in the object sequence visited by the user can be combined to determine the object features corresponding to the objects visited by the user, thereby more accurately extracting object features that can reflect the user's preferences. Therefore, the object features extracted based on BERT can more accurately determine user features that reflect user preferences.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of recommendation technology, and in particular to a method, apparatus, device, and medium for determining user characteristics and model training. Background Art

[0002] In content recommendation scenarios, recommendation systems can recommend content that users are interested in. For example, online reading platforms can recommend newly released books to interested users; or recommend books that a user has read to other users with similar preferences.

[0003] To more effectively recommend items to users, recommendation systems need to determine their preferences. Failure to accurately determine these preferences leads to inaccurate recommendations, which in turn impacts the accessibility of recommended items. Therefore, accurately determining user preferences within recommendation systems is a technical challenge facing those skilled in the art. Summary of the Invention

[0004] In view of this, the present application provides a method, device, equipment and medium for determining user features and model training, so as to train a feature extraction model that can more accurately extract object features that reflect user preferences, and enable the determined user features to more accurately reflect user preferences.

[0005] To achieve the above objectives, this application provides the following technical solutions:

[0006] In one aspect, the present application provides a method for determining user characteristics, comprising:

[0007] Obtaining an object sequence of a user to be analyzed, wherein the object sequence of the user includes: object identifiers of multiple objects visited by the user in the object recommendation system;

[0008] Determining an object feature vector for each object identifier in the user's object sequence using a language model combined with contextual information of each object identifier in the user's object sequence, wherein the language model is a transformer-based bidirectional encoding representation (BERT) model, and the BRET model is trained using masked object sequences corresponding to multiple sample users and taking the prediction of masked object identifiers in the masked object sequences as a training target; the masked object sequence of the sample user is an object sequence obtained after at least one object identifier in the sample user's object sequence is masked;

[0009] A user feature vector for characterizing the user's interest in an object in the object recommendation system is determined based on the object feature vector of each object identifier in the user's object sequence.

[0010] In a possible implementation of the above aspect, obtaining the object sequence of the user to be analyzed includes:

[0011] Obtaining object access information of a user to be analyzed, the object access information of the user including: object identifiers of multiple objects visited by the user in the object recommendation system, and access behavior characteristics of the user in accessing each of the multiple objects;

[0012] The object identifiers of the multiple objects visited by the user are sorted according to the interest levels represented by the access behavior characteristics of the objects to obtain the object sequence of the user.

[0013] In another aspect, the present application also provides a model training method, comprising:

[0014] Obtaining object sequences of respective multiple sample users, wherein the object sequences of the sample users include: object identifiers of multiple objects visited by the sample users in the object recommendation system;

[0015] For each sample user, combining importance information of each object identifier in the object sequence of the sample user, determining at least one object identifier to be masked from the object sequence of the sample user, and performing masking processing on the determined at least one object identifier to obtain a masked object sequence;

[0016] The importance information of the object identifier includes: one or both of an access behavior feature and an access popularity; the access behavior feature of the object identifier is a behavior feature that characterizes the degree of interest of the sample user in the object corresponding to the object identifier; the probability that the object identifier is determined to be the object identifier to be masked is positively correlated with the degree of interest characterized by the access behavior feature of the object identifier, and negatively correlated with the access popularity corresponding to the object identifier;

[0017] The training objective is to predict at least one masked object identifier in the masked object sequence of the sample user, and the masked object sequence of the sample user is used to train a BERT model to obtain a BERT model for extracting object features of each object identifier in the user's object sequence, wherein the object features are used to characterize the user's interest in the object.

[0018] In another aspect, the present application provides an apparatus for determining user characteristics, comprising:

[0019] A user sequence obtaining unit is configured to obtain an object sequence of a user to be analyzed, wherein the object sequence of the user includes object identifiers of multiple objects visited by the user in the object recommendation system;

[0020] a vector determination unit, utilizing a language model in combination with contextual information of each object identifier in the user's object sequence to determine an object feature vector for each object identifier in the user's object sequence, wherein the language model is a transformer-based bidirectional encoding representation (BERT) model, and the BRET model is trained using masked object sequences corresponding to multiple sample users and taking prediction of masked object identifiers in the masked object sequences as a training target; the masked object sequence of the sample user is an object sequence obtained after at least one object identifier in the sample user's object sequence is masked;

[0021] The feature determination unit determines a user feature vector for representing a feature of interest of the user to an object in the object recommendation system based on the object feature vector of each object identifier in the object sequence of the user.

[0022] In another aspect, the present application further provides a model training device, comprising:

[0023] a sample sequence obtaining unit, configured to obtain an object sequence of each of a plurality of sample users, wherein the object sequence of a sample user is composed of object identifiers of a plurality of objects visited by the sample user in the object recommendation system;

[0024] a mask processing unit, for each sample user, combining importance information of each object identifier in the object sequence of the sample user, determining at least one object identifier to be masked from the object sequence of the sample user, and performing mask processing on the determined at least one object identifier to obtain a masked object sequence;

[0025] The importance information of the object identifier includes: one or both of an access behavior feature and an access popularity; the access behavior feature of the object identifier is a behavior feature that characterizes the degree of interest of the sample user in the object corresponding to the object identifier; the probability that the object identifier is determined to be the object identifier to be masked is positively correlated with the degree of interest characterized by the access behavior feature of the object identifier, and negatively correlated with the access popularity corresponding to the object identifier;

[0026] A model training unit is configured to train a BERT model using the masked object sequence of the sample user, with the prediction of at least one masked object identifier in the masked object sequence of the sample user as a training target, to obtain a BERT model for extracting object features of each object identifier in the user's object sequence, wherein the object features are used to characterize the user's interest in the object.

[0027] In another aspect, the present application further provides a computer device, characterized in that it includes a memory and a processor;

[0028] Wherein, the memory is used to store programs;

[0029] The processor is used to execute the program, and when the program is executed, it is specifically used to implement the method for determining user characteristics as described in any one of the above or the model training method as described above.

[0030] On the other hand, the present application provides a storage medium for storing a program, which, when executed, is used to implement the method for determining user characteristics as described in any one of the above or the model training method as described above.

[0031] From the above content, it can be seen that the present application trains the BERT model as a feature extraction model for extracting object features of objects visited by users. In the process of training the BERT model, the probability of the object identifiers of objects that are not very popular but have a high degree of user interest being masked will be increased. Since objects that are not very popular but have a high degree of user interest can better reflect the user's preferences, in the process of training the BERT model, the object identifiers of objects that are not very popular but have a high degree of user interest are masked and trained in a focused manner, so that the trained BERT model can more accurately extract object features that can characterize the user's preferences from the user's object sequence.

[0032] Moreover, since the BERT model can combine the contextual information between the object identifiers in the object sequence visited by the user to determine the object features of the objects visited by the user, the BERT model can more accurately extract object features that can reflect user preferences. Therefore, the object features extracted based on BERT can more accurately determine the user features that reflect user preferences. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without any creative work.

[0034] Figure 1 A schematic diagram showing a composition architecture of a recommendation system to which this application is applicable is shown;

[0035] Figure 2 A flow chart showing an embodiment of a method for determining user characteristics of the present application is shown;

[0036] Figure 3 A schematic diagram of an implementation process of obtaining a user's object sequence in this application is shown;

[0037] Figure 4A flow chart of an embodiment of the model training method provided by the present application is shown;

[0038] Figure 5 A flow chart of another embodiment of the model training method provided by the present application is shown;

[0039] Figure 6 A schematic diagram of the structure of the BERT model in this application is shown;

[0040] Figure 7 A schematic diagram of the structure of the transformer model in the BERT model is shown;

[0041] Figure 8 A schematic diagram of the composition architecture of a specific application scenario to which the solution of the present application is applicable is shown;

[0042] Figure 9 A schematic diagram of a training process for training a BERT model in an application scenario of the present application is shown;

[0043] Figure 10 A schematic diagram of the implementation flow of the method for determining user characteristics provided by the present application in an application scenario is shown;

[0044] Figure 11 A schematic diagram showing the structure of an embodiment of a model training device provided by the present application is shown;

[0045] Figure 12 A schematic diagram showing a composition architecture of an apparatus for determining user characteristics provided by the present application is shown;

[0046] Figure 13 The figure shows a schematic diagram of the composition structure of a computer device to which the present application is applicable. DETAILED DESCRIPTION

[0047] The solution of this application is applicable to an object recommendation system, which is a service platform that can recommend objects to users. Different object recommendation platforms will recommend different objects to users. For example, the objects recommended by the object recommendation system can include books, articles, videos, applications, or images.

[0048] For example, the object recommendation system may be an online reading platform, through which users can read books available on the online reading platform. At the same time, the online reading platform may also specifically recommend books that the user is interested in.

[0049] For another example, the object recommendation system may be a multimedia service platform, which may return multimedia to the user terminal based on the media access request of the user terminal; at the same time, the multimedia service platform may also send recommended multimedia information to the user terminal.

[0050] In the embodiment of the present application, in order to determine the user characteristics of users in the object recommendation system, machine learning and other processing of the model will be performed based on artificial intelligence technology.

[0051] Artificial Intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also involves studying the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.

[0052] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0053] In the embodiments of the present application, at least the model will be trained based on machine learning. Among them, machine learning (ML) is a multi-disciplinary interdisciplinary subject involving multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specializes in how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all areas of artificial intelligence. Machine learning and deep learning generally include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and formula-based learning.

[0054] First, the object recommendation system to which the solution of this application is applicable is introduced.

[0055] like Figure 1 As shown, it shows a schematic diagram of the composition architecture of an object recommendation system applicable to this application.

[0056] Depend on Figure 1 It can be seen that the object recommendation system 100 may include at least one recommendation server 101 .

[0057] For example, the object recommendation system can consist of a single recommendation server, a cluster of multiple recommendation servers, a distributed system, or a cloud platform. A cloud platform, also known as a cloud computing platform, is a network platform built on cloud technology. Cloud technology refers to a hosting technology that unifies hardware, software, and network resources within a wide area network or local area network to enable data computing, storage, processing, and sharing.

[0058] The recommendation server of the object recommendation system may store multiple items for recommendation, such as the books and multimedia mentioned above.

[0059] Of course, the object recommendation system may also be configured to have a database outside of the recommendation server, and store relevant information about the objects that the object recommendation system needs to recommend in the database.

[0060] In the present application, the user's terminal 102 may establish a communication connection with the object recommendation system 100 via a network. The terminal 102 may request the object recommendation system to access a certain object.

[0061] At the same time, the object recommendation system will also determine the user's preferences based on the terminal user's historical access information to objects in the object recommendation system, and recommend objects to the user in a targeted manner.

[0062] In an embodiment of the present application, the object recommendation system may first obtain user features that can characterize user preferences, and recommend objects based on the user features.

[0063] In order to achieve more accurate object recommendations, this application will also train a Bidirectional Encoder Representations from Transformers (BERT) model based on the objects visited by the user to train a feature extraction model. The feature extraction model is used to extract the object features of the object identifiers corresponding to the objects visited by the user, and the object features extracted by the feature extraction model can reflect the user's preference characteristics for the objects in the object recommendation system.

[0064] For ease of understanding, the method for determining user characteristics provided by this application is first introduced below.

[0065] like Figure 2As shown, it shows a flow chart of an embodiment of a method for determining user characteristics of the present application. The method of this embodiment can be applied to a recommendation server or computer device of an object recommendation system. The method of this embodiment may include:

[0066] S201: Obtain an object sequence of a user to be analyzed.

[0067] The user's object sequence includes: object identifiers of multiple objects that the user has visited in the object recommendation system. The object identifiers of the multiple objects in the user's object sequence have a sequential order.

[0068] The object identifier of the object is used to uniquely identify the object in the object recommendation system.

[0069] For example, the object identifier of an object can be a unique number for the object.

[0070] For another example, the object identifier of the object may also be a number mapped from the object's identification information. For example, assuming that the object recommendation system includes 100 objects available for recommendation, the object identifiers of these 100 objects may be set to natural numbers from 1 to 100 in sequence.

[0071] S202 : Determine an object feature vector for each object identifier in the object sequence of the user by using a language model in combination with context information of each object identifier in the object sequence of the user.

[0072] In this application, the language model is the BERT model.

[0073] For example, the user's object sequence is input into the BERT model, and the BERT model can respectively analyze the contextual relationship between each object identifier in the object sequence, and determine the object feature vector of each object identifier based on the contextual relationship.

[0074] In this application, the BRET model is trained using masked object sequences corresponding to multiple sample users, with the goal of predicting the masked object identifiers in the masked object sequences. The masked object sequence for a sample user is the object sequence obtained after at least one object identifier in the sample user's object sequence has been masked.

[0075] Among them, after obtaining the masked object sequences of multiple sample users for training the BERT model and determining the training target, there can be many specific ways to train the BERT model, and this application does not impose any restrictions on this.

[0076] For ease of distinction, in this embodiment, the feature vector of each object identifier output by the BERT model is referred to as an object feature vector.

[0077] S203: Determine, based on the object feature vectors of the object identifiers in the object sequence of the user, a user feature vector for representing the user's interest in the object in the object recommendation system.

[0078] It can be understood that the object sequence is composed of object identifiers corresponding to objects visited by the user. Therefore, the object feature vector of each object identifier in the object sequence can reflect the user's interest in the objects in the object recommendation system. Therefore, based on the object features of each object identifier in the object sequence, a user feature vector used to characterize the user's interest features can be determined.

[0079] For example, in one possible implementation, the user feature vector can be determined by determining the average value of the object feature vectors of each object identifier in the object sequence to obtain an average value vector; and determining the obtained average value vector as the user feature vector of the user. That is, the user feature vector T can be calculated using the following formula 1:

[0080]

[0081] Among them, T i is the object feature vector of the i-th object identifier in the user's object sequence, i is a natural number from 1 to M, and M is the total number of object identifiers in the user's object sequence.

[0082] Of course, there may be other ways to determine the user characteristics, which are not limited thereto.

[0083] It is understandable that after determining the user feature vector of the user, the object recommendation system may further recommend at least one object to the user based on the user feature vector.

[0084] In one implementation, since a user's user feature vector can reflect the object features of the objects the user has visited, if the user feature vectors of different users are relatively similar, it means that the objects preferred by these two users are also relatively similar. Based on this, the present application can also determine at least one similar user with a similar user feature vector as the user, and recommend objects visited by the similar user to the user; or, recommend objects visited by the user to similar users.

[0085] It can be seen that after obtaining the object sequence visited by the user, the present application will accurately extract the object feature vector of each object identifier in the user's object sequence through the trained BERT model combined with the context information between the object identifiers of each object in the user's object sequence, so that the object feature vector of the object identifier can represent the context information between the object identifier and other object identifiers in the user's object sequence, and the context information between the object identifiers can reflect the characteristic relationship between the objects preferred by the user. Therefore, combining the object feature vectors corresponding to the object identifiers of multiple objects visited by the user to determine the user characteristics can enable the user characteristics to more accurately reflect the user's preference characteristics, thereby facilitating more reasonable object recommendations and improving the visit rate of recommended objects.

[0086] It is understandable that the ordering of the object identifiers of the multiple objects visited by the user in the user's object sequence can be set as needed.

[0087] In one possible implementation, in order to enable the BERT model to combine the contextual information between different object identifiers in the user's object sequence and to more accurately extract the object features represented by each object identifier in the user's object sequence, the present application can also sort the object identifiers of the multiple objects in combination with the user's access behavior characteristics to the multiple objects.

[0088] like Figure 3 As shown, it shows a schematic diagram of an implementation process of obtaining a user's object sequence in this application, which may include:

[0089] S301: Obtain user's object access information.

[0090] The user's object access information includes: object identifiers of multiple objects visited by the user, and access behavior characteristics of the user when accessing each of the multiple objects.

[0091] The user's access behavior characteristics for accessing the object can represent the user's access behavior toward the object. Specifically, the access behavior characteristics for the object identifier corresponding to the user are behavioral characteristics used to characterize the user's interest in the object corresponding to the object identifier. The access behavior characteristics for the object identifier corresponding to the user can be obtained from the user's access data for the object represented by the object identifier stored in the recommendation system.

[0092] For example, the access behavior characteristics of the object identifier corresponding to the user may include one or more of the following information: the access duration and the number of accesses of the user corresponding to the object identifier.

[0093] It is understood that the longer a user visits an object, the more likely they are to like or be interested in that object. For example, the longer a user spends reading a book, the more likely they are to like it. Similarly, the more times a user visits an object, the more likely they are to be interested in it.

[0094] Of course, the object identifier may also include the user's respective access times or the last access time, etc. For example, the last time a user accesses an object may represent the user's recent attention to the object, and thus reflect whether the user still likes the object recently.

[0095] S302 , sorting the object identifiers of multiple objects visited by the user according to the interest levels represented by the object access behavior characteristics to obtain the object sequence of the user.

[0096] It can be understood that since the access behavior characteristics of an object can represent the user's interest level in the object, sorting the object identifiers corresponding to the objects visited by the user according to the interest level represented by the access behavior characteristics of the object can make objects with similar user interests appear adjacent in the object sequence.

[0097] For example, the object identifiers of the multiple objects visited by the user may be sorted in order from high to low (or from low to high) according to the interest levels represented by the access behavior characteristics of the objects.

[0098] In one possible implementation, the object access behavior feature may include the object access duration. In this case, since the object access duration can represent the user's interest in the object, the object identifiers of multiple objects visited by the user are sorted from longest to shortest (or vice versa) according to the object access duration, thereby obtaining the user's object sequence.

[0099] In another possible implementation, the user's access behavior characteristics for objects include the duration and number of visits. In this case, the duration of the user's access to the object can be used as the primary sorting criterion, and the number of visits can be used as the secondary sorting criterion to sort the object identifiers of multiple objects visited by the user.

[0100] Among them, the access time is the main sorting basis, that is, the access time is used as the basis for sorting the object identifiers of the multiple objects. If the access time of two or more objects is the same, the object identifiers of the two or more objects can be sorted according to the number of visits to the two or more objects.

[0101] For example, the object identifiers of the user's multiple objects may be sorted in descending order of the object's access duration. If two or more objects have the same access duration, the object identifiers of the two or more objects may be sorted in descending order of the number of times the objects have been accessed.

[0102] It can be seen that in Figure 3 In an embodiment, the object identifiers of multiple objects visited by the user can be sorted according to the user's interest level represented by the object access behavior characteristics, so that objects with similar interest levels are adjacent in the object sequence of the sample user, which is more conducive to the BERT model combining the contextual information in the object sequence to more accurately extract the object features represented by each object identifier in the user's object sequence.

[0103] In order to make the determined user features more accurately reflect the user's preference features for objects in the object recommendation system, this application can combine the access behavior characteristics of sample users and the access popularity of objects to determine the object identifiers that need to be masked in the object sequence of the sample user during the training of the BERT model. Figure 4 Provide a detailed introduction.

[0104] like Figure 4 As shown, it shows a flow chart of an embodiment of a model training method of the present application. The method of this embodiment may include application to an object recommendation system, such as a server or computer device in a recommendation object system.

[0105] The method of this embodiment may include:

[0106] S401: Obtain object sequences of multiple sample users.

[0107] The object sequence of the sample user is composed of the object identifiers of multiple objects that the sample user has visited in the object recommendation system. It can be understood that the object identifiers of the multiple objects in the object sequence of the sample user have a sequential order.

[0108] In one possible implementation, object access information for multiple sample users can be obtained. The object access information for each sample user includes the object identifiers of multiple objects visited by the sample user and the sample user's access behavior characteristics for each of the multiple objects. For each sample user, the object identifiers of the multiple objects visited by the sample user are sorted according to the level of interest indicated by the object access behavior characteristics to obtain an object sequence for the sample user.

[0109] Through this implementation method, the object identifiers of objects with similar interests among sample users can be adjacent in the object sequence, which is beneficial to the training process of the BERT model. The contextual relationship between each object identifier in the object sequence can be analyzed more accurately, and the object features of the objects in the object sequence can be extracted more accurately.

[0110] It is understandable that, for ease of distinction, this application refers to the users to whom the object sequences required for model training belong as sample users.

[0111] The method for obtaining the object sequence of the sample user is the same as the specific implementation method for obtaining the object sequence of the user mentioned above. For details, please refer to the relevant introduction of the previous embodiment, which will not be repeated here.

[0112] S402 : For each sample user, based on the importance information of each object identifier in the object sequence of the sample user, determine at least one object identifier to be masked from the object sequence of the sample user, and perform masking processing on the determined at least one object identifier to obtain a masked object sequence.

[0113] The importance information of the object identifier includes one or both of the access behavior characteristics and access popularity.

[0114] The access behavior characteristics of the sample user accessing an object are similar to the information contained in the access behavior characteristics of the user accessing an object. For example, the access behavior characteristics of the object identifier corresponding to the sample user may include one or more of the following information: the duration and number of visits of the sample user to the object identifier.

[0115] The access popularity of an object identifier is the popularity of the multiple sample users accessing the object corresponding to the object identifier. The popularity of multiple sample users accessing an object can be represented by the total number of the multiple sample users who have visited the object. The more people who have visited the object, the higher the popularity of the object.

[0116] Of course, if the access popularity of different objects for all users in the recommendation platform can be counted, then the access popularity of an object identifier can be the total number of users who have visited the object corresponding to the object identifier in the object recommendation system.

[0117] Among them, for each sample user, the mask processing of the at least one object identifier refers to using set mask rules to replace or change part or all of the at least one object identifier so that the at least one object identifier after mask processing changes.

[0118] For example, 80% of the at least one object identifier may be replaced with a mask identifier, 10% may be replaced with other characters, and 10% may remain unchanged.

[0119] It is understandable that the purpose of masking some object identifiers in the sample user's object sequence during BERT model training is to enable the BERT model to predict the masked object identifiers in the object sequence. During BERT model training, the trained BERT model learns the contextual information of each object identifier in the object sequence.

[0120] It is understandable that if random extraction is used to determine the object identifiers to be masked from the user's object sequence, it is likely that the object identifiers corresponding to objects with high visit popularity will be extracted with a higher probability. However, objects with high visit popularity are actually objects that most users have seen and cannot accurately reflect the user's preference characteristics. At the same time, the use of random extraction may result in a lower probability of extracting object identifiers that the user likes but has low popularity. As a result, the object identifiers extracted as to be masked cannot accurately reflect the user's preferences, and the subsequently trained BERT model cannot accurately extract contextual information that can reflect the user's preference characteristics from the object identifiers in the object sequence.

[0121] In the present application, the probability of an object identifier being determined as an object identifier to be masked is positively correlated with the degree of interest represented by the access behavior characteristics of the object identifier, and negatively correlated with the access popularity corresponding to the object identifier. Therefore, the probability that an object identifier with a higher degree of user preference is extracted as an object identifier that needs to be masked is higher, and the probability of an object with a higher access popularity being extracted is lower, which is conducive to the subsequent trained BERT model to more accurately extract contextual information reflecting the user's preference characteristics from the user's object sequence.

[0122] In this application, for the sake of distinction, the object sequence obtained after masking the object sequence of the sample user is referred to as a masked object sequence.

[0123] S403, predicting at least one masked object identifier in the masked object sequence of the sample user as a training target, and using the masked object sequence corresponding to the sample user to train a BERT model, to obtain a BERT model for extracting object features of each object identifier in the user's object sequence.

[0124] Among them, the object features corresponding to each object identifier extracted by the trained BERT model from the user's object sequence can be used to represent the user's interest in the objects in the object recommendation system.

[0125] It is understandable that after masking some object identifiers in the object sequence of the sample user using the solution of the present application, there may be multiple specific implementation processes for training the BRET model, and the present application does not impose any restrictions.

[0126] For example, for each sample user, the masked object sequence of the sample user can be input into the BERT model to be trained to obtain the object feature vector of each object identifier in the masked object sequence output by the BERT model; then, based on the object feature vector of each object identifier in the masked object sequence corresponding to the sample user, at least one object identifier that is masked in the masked object sequence can be predicted.

[0127] Accordingly, the actual masked at least one object identifier in each sample user's masked object sequence and the predicted at least one masked object identifier are combined to determine whether the BERT model training meets the training requirements. If the training requirements are not met, the internal parameters of the BERT model are adjusted and the BERT model is retrained using the masked object sequences of each sample user until the training requirements are met.

[0128] Since the BERT model uses a multi-layer transformer to perform bidirectional learning on text (such as the object series in this application), it can more accurately learn the contextual relationship between each word in the text (the object identifier in this application), thereby accurately extracting the semantic features of each word in the text.

[0129] From the above content, it can be seen that in the process of training the BERT model, this application will control the probability of masking the object identifiers of objects that are not very popular but have a high degree of user interest. Since objects that are not very popular but have a high degree of user interest can better reflect the user's preferences, in the process of training the BERT model, the object identifiers of objects that are not very popular but have a high degree of user interest are masked and trained in a focused manner, so that the trained BERT model can more accurately extract object features that can represent the user's preferences from the user's object sequence, which is conducive to more accurate determination of user features based on the object features extracted by the BERT.

[0130] In order to more clearly understand the training method of the feature extraction model of this application, the following is an introduction using a training method as an example. Figure 5 , which shows a flow chart of another embodiment of a method for training a feature extraction model of the present application. The method of this embodiment may include:

[0131] S501: Obtain object access information of each of a plurality of sample users.

[0132] The object access information of the sample user includes: object identifiers of multiple objects visited by the user, and access behavior characteristics of the user in accessing each of the multiple objects.

[0133] S502 , for each sample user, sort the object identifiers of multiple objects visited by the sample user according to the interest level represented by the object access behavior characteristics to obtain the object sequence of the sample user.

[0134] This application uses S501 and S502 as an example to illustrate a method of obtaining the object sequence of a sample user. Specific implementation methods of obtaining the object sequence of a sample user by other methods are also applicable to this embodiment and are not limited thereto.

[0135] S503: For each sample user, obtain importance information of each object identifier in the object sequence of the sample user.

[0136] The importance information of the object identifier includes one or both of access behavior characteristics and access popularity.

[0137] S504 : For each sample user, determine the weight of each object identifier in the object sequence of the sample user based on the importance information of each object identifier in the object sequence of the sample user.

[0138] Among them, for each sample user, the access behavior characteristics corresponding to the object identifier represent that the higher the sample user's interest in the object corresponding to the object identifier, the higher the weight of the object identifier; and the higher the popularity corresponding to the object identifier, the lower the weight corresponding to the object identifier.

[0139] For example, for each sample user, the access behavior characteristic corresponding to an object identifier can be the duration of time the sample user accesses the object corresponding to that object identifier. The longer the access duration, the higher the weight of that object identifier in the sample user's object sequence. Similarly, the access behavior characteristic corresponding to an object identifier can be the number of times the sample user accesses the object corresponding to that object identifier. The higher the number of visits, the higher the weight of that object identifier.

[0140] In one possible implementation, if the importance information includes the access duration corresponding to an object identifier, then for any object identifier in the sample user's object sequence, the weight of the object identifier can be determined by the ratio of the access duration corresponding to that object identifier to the total access duration corresponding to all object identifiers in the sample user's object sequence. The total access duration corresponding to the sample user is the sum of the access durations corresponding to all object identifiers in the sample user's object sequence.

[0141] Similarly, if the importance information includes the number of visits corresponding to the object identifier, the ratio of the number of visits of the object identifier to the total number of visits corresponding to all object identifiers in the object sequence of the sample user can be determined as the weight of the object identifier.

[0142] In another possible implementation, for any object identifier corresponding to each sample user, the importance information may include: the access duration corresponding to the object identifier and the access popularity of the object identifier. In this case, the ratio of the access duration corresponding to the object identifier to the total cardinality can be determined as the weight of each object identifier in the object sequence of the sample user. The total cardinality is the product of the total access duration corresponding to the sample user and the derivative of the popularity corresponding to the object identifier. The total access duration is described above. It can be seen that for any object identifier of the sample user, the higher the popularity of the object identifier, the lower the weight of the object identifier; and the longer the access duration of the object identifier, the higher its weight.

[0143] It can be understood that the above is an introduction to the process of determining the corresponding weights of object identifiers of sample users using several cases as examples. In actual applications, as long as the weights of object identifiers with relatively long access times are relatively high for the sample users, and the weights of object identifiers with high access popularity are relatively low, this application does not impose any restrictions on the specific implementation method.

[0144] S505: For each sample user, based on the weight of each object identifier in the object sequence of the sample user, determine at least one object identifier to be masked from the object sequence of the sample user, and mask the at least one object identifier to be masked in the object sequence to obtain a masked object sequence corresponding to the sample user.

[0145] The higher the weight of the object identifier, the higher the possibility that the object identifier is determined to be the object identifier to be masked.

[0146] In a possible implementation, for each sample user, at least one object identifier to be masked may be determined from the object sequence of the sample user using a random weighting algorithm according to the weight of each object identifier in the object sequence of the sample user.

[0147] The process of performing masking on the object identifier to be masked can be referred to the relevant introduction of the previous embodiment, and will not be described in detail here.

[0148] It should be noted that this embodiment uses the example of first determining the weights of each object identifier in the sample user's object sequence and then determining the object identifier to be masked based on the object identifier weights. However, it is understood that other methods for determining the object identifier to be masked described in other embodiments are also applicable to this embodiment.

[0149] S506 : For each sample user, input the masked object sequence of the sample user into the BERT model to be trained, and obtain the object feature vector corresponding to each object identifier in the masked object sequence output by the BERT model.

[0150] After inputting the sample user's masked object sequence into the BERT model, the BERT model extracts semantic feature vectors for the object identifiers in the masked object sequence based on the contextual relationships between the object identifiers. The semantic feature vectors represent the contextual relationships between object identifiers and other object identifiers.

[0151] like Figure 6 As shown, it shows a schematic diagram of the composition structure of the BERT model.

[0152] BERT is a feature extraction model that includes a bidirectional transformer. The BERT model can obtain the initial vector representation of each object identifier in the object sequence, such as Figure 6 In the initial vector 1 to the initial vector M, M is the total number of object identifiers in the object sequence.

[0153] Among them, the initial vector of the object identifier can be the bottom layer of the BERT model (such as Figure 6 The vector encoding layer (the bottom layer) encodes each object identifier into a vector. Of course, it is also possible to use word vector encoding to encode the object identifier into a vector and then input it into the BERT model.

[0154] The initial vector of each object identifier in the object sequence passes through each layer of Transformer in turn, and finally a new feature representation of each object identifier can be extracted, that is, the object feature vector of the object identifier.

[0155] Among them, the composition structure of Transformer can be as follows Figure 7 As shown in Figure 2. Transformer is formed by stacking several encoders and decoders. Figure 7 The left part is the encoder, which consists of a multi-head attention and a full connection, and is used to convert the input object sequence into a feature vector. Figure 7 The right side of the figure is the decoder, which takes as input the encoder output and the predicted result. It consists of a masked multi-head attention, a multi-head attention, and a fully connected network to output the conditional probability of the final result.

[0156] S507 , for each sample user, inputting the object feature vector corresponding to each object identifier in the masked object sequence into the fully connected network model to be trained, and obtaining the mask probability of each object identifier in the masked object sequence predicted by the fully connected network model.

[0157] The mask probability of the object identifier is the probability that the object identifier belongs to the masked object identifier.

[0158] It is understandable that the object feature vectors of each object identifier in the masked object sequence extracted by the BERT model, and the object feature vectors of each object identifier can reflect the contextual relationship between the object identifiers. On this basis, this application requires the fully connected application network model to predict the probability that the object identifier belongs to the masked object identifier based on the object feature vector extracted by the BERT model. Therefore, this application needs to be continuously trained to make the object feature vector extracted by the BERT model more able to accurately reflect the relationship between the various objects in the object sequence, and at the same time, make the probability output by the fully connected network model more accurate.

[0159] For example, in the case where the object sequence is a sort of object identifiers of multiple objects accessed by a user according to the user's interest in the objects, the interest levels of adjacent objects in the object sequence are similar, and therefore, the features of adjacent objects will also have certain commonalities. On this basis, the user interest features represented by the semantic feature vectors (i.e., object feature vectors) corresponding to adjacent object identifiers in the object sequence are also similar. Correspondingly, if certain object identifiers in the object sequence are masked, the contextual relationship between the object identifier and the adjacent object identifier will also change. Therefore, it is necessary to train the BERT model to identify information such as the semantic contextual relationship of different object identifiers, and ultimately enable the fully connected network model to accurately determine the corresponding probability of each object identifier.

[0160] For ease of distinction, the probability that each object identity output by the fully connected network model belongs to the masked object identity is called the mask probability.

[0161] S508 : For each sample user, determine at least one object identifier in the masked object sequence corresponding to the sample user, whose masking probability is greater than a set value, as the predicted at least one masked object identifier.

[0162] The set value can be set as needed, for example, the set value can be 60%.

[0163] S509: Based on at least one object identifier actually masked in the masked object sequences corresponding to multiple sample users and at least one predicted masked object identifier, detect whether the BERT model and the fully connected network model meet the training requirements. If so, the training ends; if not, adjust the internal parameters of the BERT model and the fully connected network model, and return to execute S506 until the training requirements are met.

[0164] The BERT model and the fully connected network model meet the training requirements, indicating that the BERT model and the fully connected network can accurately predict the masked object identifiers in the masked object sequence. There are various possible training requirements that the BERT model and the fully connected network model need to meet.

[0165] For example, in one possible implementation, a loss function value can be calculated based on the set loss function, the actual masked object identifier of at least one object in each sample user's masked object sequence, and the predicted masked object identifier of at least one object. If the current loss function value converges, it means that the training requirements are met. If the loss function value has not converged, it indicates that further training of the BERT model and the fully connected network model is still necessary.

[0166] For example, the BERT model and the fully connected network model can meet the training requirements by calculating the loss function value based on the mask processing loss function (mask token loss function) in the BERT model, and determining that the training of the BERT model and the fully connected network model is completed when the loss function value converges.

[0167] Of course, the above S507 to S509 are only described by taking the training method as an example. In actual application, there may be other possible situations, which are not limited to this.

[0168] It is understandable that, when the model training is not completed and the training requirements are not met, the process may return to step S505 to perform masking on the object sequence of the sample user again.

[0169] For ease of understanding, the following describes the process of training the model and determining user characteristics in this application in conjunction with an application scenario.

[0170] Take the object recommendation system as an online reading platform (such as WeChat reading platform, etc.) as an example. Figure 8 , which shows a schematic diagram of the composition architecture of an application scenario to which the solution of this application is applicable.

[0171] Depend on Figure 8 It can be seen that the online reading platform 801 may include at least one server 802 .

[0172] The user terminal 803 can establish a communication connection with the server of the online reading platform 801 through the network.

[0173] like Figure 8 As shown, the user terminal 803 can send a book reading request to the server of the online reading platform 801, and the book reading request is used to request access to books in the online reading platform.

[0174] The server 802 of the online reading platform responds to the book reading request and can return the data information of the book to the user terminal, such as the data of the first page of the book and the data of the chapter introduction page.

[0175] Accordingly, the user terminal 803 can display the corresponding pages of the book based on the data information of the book, so that the user can read the book.

[0176] In this application, the online reading platform will also record information such as the identifier of the book read by the user, the time the book was read, the reading time of the book, and the number of times the user read the book.

[0177] The online reading platform can also determine the user characteristics of each user, which can represent the types of books the user is interested in. At the same time, the online reading platform can recommend a book list suitable for the user based on the user characteristics, and the book list can include information about at least one book.

[0178] Combine Figure 8 This article introduces the process of training the BERT model to extract book features of books that users are interested in based on the application scenario. Figure 9 As shown, it shows a schematic diagram of the process of applying the method of training feature extraction model of this application to the scenario of WeChat reading. This process may include:

[0179] S901: Obtain book sequence samples of respective sample users and access popularity and access duration corresponding to each book identifier in the book sequence samples.

[0180] Each sample user corresponds to a book sequence sample, which includes the book identifiers of multiple books that the sample user has read.

[0181] In order to distinguish the book sequences of the users whose user features are to be analyzed subsequently, in this embodiment, the book sequences of the sample users used to train the BERT model are referred to as book sequence samples.

[0182] For example, the book identifiers of multiple books read by the sample user can be sorted according to the reading time of each book read by the sample user on the online reading platform to obtain a book sequence sample of the sample user.

[0183] In practical applications, in order to ensure that the length of the book sequence samples of each sample user is the same, the present application can also select the book identifiers corresponding to the books in the pre-set position after sorting the multiple books read by the sample user to form the book sequence sample. The set position can be set as needed, for example, the books in the top 100 can be selected.

[0184] The access popularity corresponding to the book identifier is the access popularity of the book corresponding to the book identifier, such as the total number of people who have read the book in the multiple user samples, or the total number of users who have read the book on the online reading platform.

[0185] For each sample user, the access duration of a book ID is the total duration that the sample user spends reading the book corresponding to the book ID.

[0186] It is understandable that, for ease of understanding, this application uses the importance information of book identification as access popularity and access duration as an example for explanation, but the same applies to the case where the importance information includes the other information mentioned above.

[0187] S902 : For each sample user, determine the weight of each book identifier in the book sequence sample according to the access popularity and access duration corresponding to each book identifier in the book sequence sample.

[0188] Among them, the longer the access time corresponding to the book identifier is, the higher the weight of the book identifier is; and the higher the access popularity corresponding to the book identifier is, the lower the weight of the book identifier is.

[0189] S903: For each sample user, based on the weight of each book identifier in the book sequence sample, select at least one book identifier to be masked from the book sequence sample, and perform masking processing on the at least one book identifier to be masked in the book sequence sample to obtain a masked book sequence sample.

[0190] For the BERT model, the BERT model can use the masked language model MLM to mask the book identifiers in the book sequence samples. However, in this embodiment, the book identifiers that need to be masked are determined in combination with the weights of the book identifiers, which can increase the probability of book identifiers that users have read for a long time and are relatively unpopular being extracted and masked.

[0191] Since the weight of a book identifier can reflect the degree of influence of the book corresponding to the book identifier on the user's preference characteristics, masking the book identifiers in the book sequence samples in combination with the weight of the book identifier is beneficial for the trained BERT model to more accurately extract the book features of the book read by the user from the book identifier.

[0192] After determining at least one book identifier to mask, you can select a masking method as needed. For example, the BERT model can replace 80% of the selected at least one book identifier with a preset token and 10% with a random word; the remaining 10% can remain unchanged.

[0193] In order to facilitate distinction, the book sequence samples of the sample users obtained after masking are called masked book sequence samples.

[0194] S904 , for each sample user, inputting the masked book sequence sample of the sample user into the BERT model to be trained, and obtaining the book feature vector of each book identifier in the masked book sequence sample output by the BERT model.

[0195] In this embodiment, the feature vector corresponding to the book identifier extracted by BERT is called a book feature vector.

[0196] S905: For each sample user, the book feature vector of each book identifier in the masked book sequence sample corresponding to the sample user is input into the fully connected network model to be trained, and the mask probability corresponding to each book identifier in the masked book sequence sample predicted by the fully connected network model is obtained.

[0197] For each sample user, the mask probability of a book ID in the masked book sequence of the sample user is the probability that the book ID belongs to the masked book ID.

[0198] In practical applications, the object feature vector of each book identification in the masked book sequence sample can be combined with a negative sampling method, and the fully connected network model can be used to determine the mask probability of the book identification.

[0199] S906 : Determine at least one masked book identifier in the predicted masked book sequence sample based on the masking probability corresponding to each book identifier in the masked book sequence sample.

[0200] For example, a set number of book identifiers corresponding to relatively high masking probabilities may be determined as predicted masked book identifiers.

[0201] Of course, the masked language model can also be combined to determine the masked book identification according to the mask probability corresponding to each book identification output by the fully connected network. This application does not impose any restrictions on this.

[0202] S907 , based on at least one actually masked book identifier in the masked book sequence of each sample user among the multiple sample users and at least one predicted masked book identifier, detecting whether the BERT model meets the training requirements.

[0203] Among them, if the BERT model meets the training requirements, it will be combined with model training.

[0204] For example, the loss function can be calculated based on the MLM loss function used in BERT model pre-training, the actual masked book identifiers in the masked book sequence of each sample user, and the predicted masked book identifiers. If the loss function converges, the BERT model training is complete.

[0205] S908: If the BERT model has not yet met the training requirements, adjust the internal parameters of the BERT model and the fully connected network model, and return to the operation of S902 until the training requirements are met.

[0206] It is understandable that in the training of the BERT model in this application, only the task of predicting the masked book identification corresponding to the masked language model is trained (the object sequence based on this part has changed in order to extract the object identification that needs to be masked), and the training task of predicting the next sentence is not involved. Since the purpose of the next sentence prediction is to let the BERT model learn the relationship between sentences (book identifications), its essence is to determine whether two sentences are from the same topic. However, the sources of the various book identifications in the book sequence samples in this application are the same, and they all belong to books that have been read by the same user on the same book reading platform. Therefore, there is no need to predict the next sentence. In the process of training the BERT model in the embodiment of this application, the relevant training of the next sentence is removed, which can reduce the complexity of BERT model training and accelerate the model training speed.

[0207] In the process of training the BERT model, this application sets weights based on the access time and popularity corresponding to the book identifiers, so that the weights of book identifiers with higher popularity are relatively lower, while the weights of book identifiers with longer user access time are relatively higher. This is conducive to more reasonable selection of masked book identifiers from book sequence samples, and can also improve the accuracy of the BERT model in extracting features corresponding to book identifiers.

[0208] It is understandable that in Figure 8In the application scenario, after training the BERT model, the present application can determine the user characteristics of users in the online reading platform based on the BERT model, and the user characteristics can reflect the book characteristics of the books that the users are interested in. For example, see Figure 10 , which shows that the method for determining user characteristics of this application is applied to Figure 8 A flow chart of an application scenario, in this embodiment, may include:

[0209] S1001, obtaining a book sequence corresponding to the user to be analyzed in an online reading platform.

[0210] The book sequence includes: book identifiers of multiple books that the user has visited in the online reading platform, and the book identifiers of multiple books in the book sequence are sorted from longest to shortest according to the reading time of the user reading the books corresponding to the book identifiers.

[0211] In this embodiment, for ease of understanding, the book sequence is constructed according to the reading time of the books read by the user. However, other methods for obtaining the object sequence of the user mentioned above are also applicable to this embodiment and will not be repeated here.

[0212] In an alternative approach, if the length of the book sequence sample used to train the BERT model is the set number of positions, the book sequence can also be composed of graphic identifiers of the books read by the user in the set position. Of course, if the number of books read by the user is less than the set number of positions, the set image identifiers or character completion can be used.

[0213] S1002: Input the user's book sequence into the trained BERT model to obtain the book feature vector of each book identifier in the book sequence output by the BERT model.

[0214] S1003: Calculate the average value of the book feature vectors of the book identifiers in the book sequence of the user, and determine the obtained average value vector as the user feature vector of the user.

[0215] S1004: Based on the user feature vector, determine at least one similar user in the online reading platform who has similar features of interest to the user, and recommend books read by the similar user to the user.

[0216] It is understandable that step S1004 is an optional step. In actual applications, after determining the user feature vector for representing the books preferred by the user, the user feature vector can be simply saved so that the user feature vector can be called when needed to recommend books to the user.

[0217] It is understandable that different users' access characteristics, such as the books they read and the length of time they spend reading them, can reflect the characteristics of their interests. On this basis, since the user feature vector is determined based on the sequence of books read by the user, if the user feature vectors of different users are similar, then the books that these two users are interested in will also be relatively similar. On this basis, the present application can combine the user feature vectors of the user to determine similar users with similar preferences, so as to recommend books read by similar users to each other.

[0218] Corresponding to the method for determining user characteristics of the present application, the present application also provides an apparatus for determining user characteristics.

[0219] like Figure 11 , which shows a schematic diagram of the composition structure of an embodiment of a device for determining user characteristics of the present application. The device of this embodiment may include:

[0220] The user sequence obtaining unit 1101 is configured to obtain an object sequence of a user to be analyzed, wherein the object sequence of the user includes object identifiers of multiple objects that the user has visited in the object recommendation system;

[0221] A vector determination unit 1102 determines an object feature vector for each object identifier in the user's object sequence using a language model combined with contextual information of each object identifier in the user's object sequence, wherein the language model is a transformer-based bidirectional encoding representation (BERT) model, and the BRET model is trained using masked object sequences corresponding to multiple sample users and using prediction of masked object identifiers in the masked object sequences as a training target; the masked object sequence of the sample user is an object sequence obtained after at least one object identifier in the sample user's object sequence is masked;

[0222] The feature determination unit 1103 determines a user feature vector for representing the user's interest in the object in the object recommendation system based on the object feature vector of each object identifier in the user's object sequence.

[0223] In a possible implementation, the user sequence obtaining unit includes:

[0224] A user information obtaining unit, configured to obtain object access information of the user to be analyzed, the object access information of the user including: object identifiers of multiple objects visited by the user in the object recommendation system, and access behavior characteristics of the user in accessing each of the multiple objects;

[0225] The user sequence generating unit is used to sort the object identifiers of multiple objects visited by the user according to the interest level represented by the access behavior characteristics of the objects to obtain the object sequence of the user.

[0226] In an optional manner, the access behavior characteristics of the object visited by the user obtained by the user information obtaining unit include: the access duration and the number of visits of the user to the object;

[0227] The user sequence generating unit is specifically configured to use the duration of the user's access to the object as the primary sorting basis and the number of times the user accesses the object as the secondary sorting basis, sort the object identifiers of multiple objects visited by the user, and obtain the object sequence of the user.

[0228] In yet another possible implementation, the feature determination unit includes:

[0229] an average determination subunit, configured to determine an average value of the object feature vectors of the object identifiers in the object sequence of the user, and obtain an average value vector;

[0230] The vector determination subunit is configured to determine the average value vector as a user feature vector for characterizing the user's interest in the object in the object recommendation system.

[0231] Corresponding to the model training method in this application, this application also provides a model training device. Figure 12 As shown, it shows a schematic diagram of the composition structure of an embodiment of a model training device of the present application. The device of this embodiment may include:

[0232] The sample sequence obtaining unit 1201 is configured to obtain an object sequence of each of a plurality of sample users, wherein the object sequence of a sample user is composed of object identifiers of a plurality of objects that the sample user has visited in the object recommendation system;

[0233] The mask processing unit 1202 determines, for each sample user, at least one object identifier to be masked from the object sequence of the sample user based on the importance information of each object identifier in the object sequence of the sample user, and performs mask processing on the determined at least one object identifier to obtain a masked object sequence.

[0234] The importance information of the object identifier includes: one or both of an access behavior feature and access popularity. The access behavior feature of the object identifier is a behavior feature that characterizes the sample user's interest in the object corresponding to the object identifier. The probability that the object identifier is determined to be the object identifier to be masked is positively correlated with the interest level represented by the access behavior feature of the object identifier, and negatively correlated with the access popularity corresponding to the object identifier.

[0235] The model training unit 1203 is used to predict at least one masked object identifier in the masked object sequence of the sample user as a training target, and use the masked object sequence of the sample user to train the BERT model to obtain a BERT model for extracting object features of each object identifier in the user's object sequence, wherein the object features are used to characterize the user's interest in the object.

[0236] In a possible implementation, the sample sequence obtaining unit includes:

[0237] a sample information obtaining unit, configured to obtain object access information of each of a plurality of sample users, wherein the object access information of the sample users includes: object identifiers of a plurality of objects visited by the sample users, and access behavior characteristics of the sample users in accessing each of the plurality of objects;

[0238] The sample sequence generating unit is used to sort the object identifiers of multiple objects visited by each sample user according to the interest level represented by the object access behavior characteristics, so as to obtain the object sequence of the sample user.

[0239] In yet another possible implementation, the mask processing unit includes:

[0240] A weight determination subunit is configured to determine the weight of each object identifier in the object sequence of the sample user based on the importance information of each object identifier in the object sequence of the sample user, wherein the weight of the object identifier is positively correlated with the interest level represented by the access behavior characteristics of the object identifier and negatively correlated with the access popularity corresponding to the object identifier;

[0241] The mask processing subunit is configured to determine at least one object identifier to be masked from the object sequence of the sample user using a random weighting algorithm according to the weight of each object identifier in the object sequence of the sample user.

[0242] In another aspect, the present application further provides a computer device, which can be a recommendation server in an object recommendation system, or a computing device for data analysis and processing. Of course, it can also be a computer device independent of the object recommendation system.

[0243] like Figure 13 , which shows a schematic diagram of the composition architecture of the computer device provided by this application. Figure 13 In the embodiment, the computer device 1300 may include: a processor 1301 and a memory 1302.

[0244] Optionally, the computer device may further include: a communication interface 1303 , an input unit 1304 , a display 1305 , and a communication bus 1306 .

[0245] The processor 1301 , the memory 1302 , the communication interface 1303 , the input unit 1304 and the display 1305 all communicate with each other via a communication bus 1306 .

[0246] In the embodiment of the present application, the processor 1301 may be a central processing unit, an application-specific integrated circuit, etc.

[0247] The processor may call the program stored in the memory 1302 . Specifically, the processor may execute the operations executed by the object recommendation system in the above embodiment.

[0248] The memory 1302 is used to store one or more programs, which may include program codes, and the program codes include computer operating instructions. In an embodiment of the present application, the memory at least stores a program for implementing the model training method in any of the above embodiments, or a method for determining user characteristics.

[0249] In one possible implementation, the memory 1302 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, the programs mentioned above, etc.; the data storage area may store data created during the use of the computer device.

[0250] The communication interface 1303 may be an interface of a communication module.

[0251] The present application may further include an input unit 1304 , which may include a touch sensing unit, a keyboard, and the like.

[0252] The display 1305 includes a display panel, such as a touch display panel.

[0253] certainly, Figure 13 The computer device structure shown does not constitute a limitation on the computer device in the embodiment of the present application. In actual applications, the computer device may include Figure 13 More or fewer components than shown, or combinations of certain components.

[0254] On the other hand, the present application also provides a storage medium storing computer-executable instructions. When the computer-executable instructions are loaded and executed by a processor, the model training method or the method for determining user characteristics in any of the above embodiments is implemented.

[0255] This application also provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in various optional implementations of the aforementioned model training method, model training apparatus, method for determining user characteristics, or apparatus for determining user characteristics. The specific implementation process can be referred to the description of the corresponding embodiments above and is not described in detail here.

[0256] It should be noted that in the specific implementation of this application, data related to user characteristics, user information, etc. are involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent must be obtained, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.

[0257] Each embodiment in this specification is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. At the same time, the features described in the embodiments in this specification can be replaced or combined with each other, so that professionals in this field can implement or use this application. For device embodiments, since they are basically similar to method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0258] The above description of the disclosed embodiments will enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein, but is to be construed in the widest manner consistent with the principles and novel features disclosed herein.

[0259] The above are only preferred embodiments of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A method for determining user characteristics, characterized in that: include: Obtaining an object sequence of a user to be analyzed, wherein the object sequence of the user includes: object identifiers of multiple objects visited by the user in the object recommendation system; Determine the object feature vector of each object identifier in the user's object sequence by using a language model in combination with contextual information of each object identifier in the user's object sequence, wherein the language model is a transformer-based bidirectional encoding representation BERT model, and the BERT model is trained using masked object sequences corresponding to multiple sample users and taking the prediction of masked object identifiers in the masked object sequences as a training target; the masked object sequence of the sample user is an object sequence obtained after at least one object identifier in the sample user's object sequence is masked; the probability of an object identifier in the sample user's object sequence being masked is positively correlated with the degree of interest represented by the access behavior characteristics of the object identifier, and is negatively correlated with the access popularity corresponding to the object identifier; A user feature vector for characterizing the user's interest in an object in the object recommendation system is determined based on the object feature vector of each object identifier in the user's object sequence.

2. The method according to claim 1, characterized in that The obtaining of the object sequence of the user to be analyzed includes: Obtaining object access information of a user to be analyzed, the object access information of the user including: object identifiers of multiple objects visited by the user in the object recommendation system, and access behavior characteristics of the user in accessing each of the multiple objects; The object identifiers of the multiple objects visited by the user are sorted according to the interest levels represented by the access behavior characteristics of the objects to obtain the object sequence of the user.

3. The method according to claim 2, characterized in that The access behavior characteristics of the objects visited by the user include: the access duration and the number of visits of the user to the object; The step of sorting the object identifiers of the plurality of objects visited by the user according to the interest level represented by the object access behavior characteristics to obtain the object sequence of the user includes: The object IDs of multiple objects visited by the user are sorted by taking the duration of the user's visit to the object as the primary sorting basis and the number of times the user has visited the object as the secondary sorting basis to obtain the object sequence of the user.

4. The method according to claim 1, wherein The determining, based on the object feature vectors of the object identifiers of the user in the object sequence, a user feature vector for representing the user's interest in the object in the object recommendation system includes: Determining an average value of the object feature vectors of each object identifier in the object sequence of the user to obtain an average value vector; The average value vector is determined as a user feature vector for characterizing the user's interest in the object in the object recommendation system.

5. A model training method, characterized in that: include: Obtaining object sequences of respective multiple sample users, wherein the object sequences of the sample users include: object identifiers of multiple objects visited by the sample users in the object recommendation system; For each sample user, combining importance information of each object identifier in the object sequence of the sample user, determining at least one object identifier to be masked from the object sequence of the sample user, and performing masking processing on the determined at least one object identifier to obtain a masked object sequence; The importance information of the object identifier includes: one or both of an access behavior feature and an access popularity; the access behavior feature of the object identifier is a behavior feature that characterizes the degree of interest of the sample user in the object corresponding to the object identifier; the probability that the object identifier is determined to be the object identifier to be masked is positively correlated with the degree of interest characterized by the access behavior feature of the object identifier, and negatively correlated with the access popularity corresponding to the object identifier; The training objective is to predict at least one masked object identifier in the masked object sequence of the sample user, and the masked object sequence of the sample user is used to train a BERT model to obtain a BERT model for extracting object features of each object identifier in the user's object sequence, wherein the object features are used to characterize the user's interest in the object.

6. The method according to claim 5, characterized in that The obtaining of the object sequences of the plurality of sample users includes: Obtaining object access information of each of a plurality of sample users, wherein the object access information of the sample users includes: object identifiers of a plurality of objects visited by the sample users, and access behavior characteristics of the sample users in accessing each of the plurality of objects; For each sample user, object identifiers of multiple objects visited by the sample user are sorted according to the interest level represented by the object access behavior characteristics to obtain an object sequence of the sample user.

7. The method according to claim 5, characterized in that The determining at least one object identifier to be masked from the object sequence of the sample user in combination with the importance information of each object identifier in the object sequence of the sample user includes: Determining a weight for each object identifier in the object sequence of the sample user based on importance information of each object identifier in the object sequence of the sample user, wherein the weight of the object identifier is positively correlated with the interest level represented by the access behavior characteristics of the object identifier and negatively correlated with the access popularity corresponding to the object identifier; At least one object identifier to be masked is determined from the object sequence of the sample user using a random weighting algorithm according to the weight of each object identifier in the object sequence of the sample user.

8. A device for determining user characteristics, characterized in that: include: A user sequence obtaining unit is configured to obtain an object sequence of a user to be analyzed, wherein the object sequence of the user includes object identifiers of multiple objects visited by the user in the object recommendation system; A vector determination unit determines an object feature vector for each object identifier in the user's object sequence by using a language model in combination with context information of each object identifier in the user's object sequence, wherein the language model is a transformer-based bidirectional encoding representation BERT model, and the BERT model is trained using masked object sequences corresponding to multiple sample users and taking the prediction of masked object identifiers in the masked object sequences as a training target; the masked object sequence of the sample user is an object sequence obtained after at least one object identifier in the sample user's object sequence is masked; the probability of an object identifier in the sample user's object sequence being masked is positively correlated with the degree of interest represented by the access behavior characteristics of the object identifier, and is negatively correlated with the access popularity corresponding to the object identifier; The feature determination unit determines a user feature vector for representing a feature of interest of the user to an object in the object recommendation system based on the object feature vector of each object identifier in the object sequence of the user.

9. The device according to claim 8, characterized in that The user sequence obtaining unit includes: A user information obtaining unit, configured to obtain object access information of the user to be analyzed, wherein the object access information of the user includes: object identifiers of multiple objects visited by the user in the object recommendation system, and access behavior characteristics of the user when accessing each of the multiple objects; The user sequence generating unit is configured to sort the object identifiers of multiple objects visited by the user according to the interest levels represented by the access behavior characteristics of the objects, so as to obtain the object sequence of the user.

10. The device according to claim 9, characterized in that The access behavior characteristics of the user accessing each of the plurality of objects include: the access duration and the number of accesses of the user to the object; The user sequence generation unit is specifically configured to use the duration of the user's access to the object as a primary sorting criterion and the number of times the user has accessed the object as a secondary sorting criterion, to sort the object identifiers of multiple objects visited by the user, and obtain the user's object sequence.

11. The device according to claim 8, characterized in that The feature determination unit includes: an average determination subunit, configured to determine an average value of the object feature vectors of the object identifiers in the object sequence of the user to obtain an average value vector; The vector determination subunit is configured to determine the average value vector as a user feature vector for characterizing the user's interest in an object in the object recommendation system.

12. A model training device, characterized in that: include: a sample sequence obtaining unit, configured to obtain an object sequence of each of a plurality of sample users, wherein the object sequence of a sample user is composed of object identifiers of a plurality of objects visited by the sample user in the object recommendation system; a mask processing unit, for each sample user, combining importance information of each object identifier in the object sequence of the sample user, determining at least one object identifier to be masked from the object sequence of the sample user, and performing mask processing on the determined at least one object identifier to obtain a masked object sequence; The importance information of the object identifier includes: one or both of an access behavior feature and an access popularity; the access behavior feature of the object identifier is a behavior feature that characterizes the degree of interest of the sample user in the object corresponding to the object identifier; the probability that the object identifier is determined to be the object identifier to be masked is positively correlated with the degree of interest characterized by the access behavior feature of the object identifier, and negatively correlated with the access popularity corresponding to the object identifier; A model training unit is configured to train a BERT model using the masked object sequence of the sample user, with the prediction of at least one masked object identifier in the masked object sequence of the sample user as a training target, to obtain a BERT model for extracting object features of each object identifier in the user's object sequence, wherein the object features are used to characterize the user's interest in the object.

13. The device according to claim 12, characterized in that The sample sequence obtaining unit includes: a sample information obtaining unit, configured to obtain object access information of each of a plurality of sample users, wherein the object access information of the sample users includes: object identifiers of a plurality of objects visited by the sample users, and access behavior characteristics of the sample users in accessing each of the plurality of objects; The sample sequence generating unit is used to sort the object identifiers of multiple objects visited by each sample user according to the interest level represented by the object access behavior characteristics, so as to obtain the object sequence of the sample user.

14. The device according to claim 12, characterized in that The mask processing unit includes: a weight determination subunit, configured to determine a weight for each object identifier in the sample user's object sequence based on importance information of each object identifier in the sample user's object sequence, wherein the weight of the object identifier is positively correlated with the interest level represented by the access behavior characteristics of the object identifier and negatively correlated with the access popularity corresponding to the object identifier; The mask processing subunit is configured to determine at least one object identifier to be masked from the object sequence of the sample user by adopting a random weighting algorithm according to the weight of each object identifier in the object sequence of the sample user.

15. A computer device, characterized in that: including memory and processor; Wherein, the memory is used to store programs; The processor is used to execute the program, and when the program is executed, it is specifically used to implement the method for determining user characteristics as described in any one of claims 1 to 4 or the model training method as described in any one of claims 5 to 7.

16. A computer-readable storage medium, characterized in that Used to store a program, which, when executed by a processor, is used to implement the method for determining user characteristics as described in any one of claims 1 to 4 or the model training method as described in any one of claims 5 to 7.

17. A computer program product, characterized in that It includes computer instructions, which, when executed by a processor, implement the method for determining user characteristics as described in any one of claims 1 to 4 or the model training method as described in any one of claims 5 to 7.

Citation Information

Patent Citations

  • Object recommendation method and device, electronic equipment and readable storage medium

    CN111461812A

  • Data processing method, device and equipment and storage medium

    CN111831901A