User modeling method and system of universal user encoder based on Mindore

By adopting a general user encoder based on Mindspore in the recommendation system, integrating multimodal information, utilizing context information and adopting Hard-Negative negative sampling strategy, multiple shortcomings in the existing user modeling methods are solved, and more accurate and personalized user modeling and recommendation effects are achieved.

CN120144865APending Publication Date: 2025-06-13NORTHEASTERN UNIV CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510216373.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing user modeling methods have problems in the recommendation system such as single modal dependence, ignoring context information, unreasonable negative sampling, insufficient migration capabilities and insufficient feature extraction of interactive sequences, resulting in poor recommendation results.

Method used

The general user encoder based on Mindspore is adopted to build an encoder including data enhancement module, interactive sequence vector representation building module, user attribute information acquisition module, etc., integrate multimodal information, utilize context information, adopt Hard-Negative negative sampling strategy, enhance migration capabilities, and extract interactive sequence features.

Benefits of technology

It achieves a comprehensive understanding of user interests and preferences, improves the personalization and flexibility of the recommendation system, improves recommendation accuracy and migration capabilities, and enhances the discrimination ability of the user encoder and the accuracy and robustness of the recommendation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144865A_ABST
    Figure CN120144865A_ABST
Patent Text Reader

Abstract

The invention discloses a user modeling method and system of a universal user encoder based on Mindpoint, and relates to the technical field of recommendation systems. By integrating the multi-modal information contained in the interaction sequence of the user, comprehensive understanding of the user is realized, and the user encoder can generate rich user representations so as to capture interests and preferences of the user more accurately. By means of fusion of multi-modal information, the performance of the recommendation system in personalized recommendation tasks is improved. And a new Hard-Negative negative sampling strategy is adopted, so that the learning ability of a user encoder in distinguishing preferential and non-preferential items of the user is enhanced, and the discrimination ability of the user encoder is improved. According to the method, the historical interaction information and the attribute information of the user are utilized to form richer and multi-dimensional user representation, so that the recommendation accuracy can be improved, and the migration capability of a user encoder can also be improved. Various different data enhancement is performed on the interaction sequence, so that various features among the articles in the interaction sequence are extracted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of recommendation systems, and particularly relates to a user modeling method and system based on a general user encoder of Mindspore. Background Art

[0002] A recommendation system is a technology that automatically recommends items that a user may be interested in based on the user's behavior and preferences. With the development of information technology and the popularization of the Internet, we have entered an era of information overload. In this context, both consumers and producers face huge challenges. For consumers, it is a challenging question of how to select the data they need from an overwhelming amount of data, and for producers, it is also a challenging question of how to make their information stand out and be noticed in the vast ocean of information. The recommendation system is an important tool to solve this problem. The recommendation system plays a crucial role among consumers, producers, and data. On the one hand, it helps consumers select more interesting items, and on the other hand, it helps producers push items more accurately to users who may be interested, thus achieving a win-win situation.

[0003] In the current research of recommendation systems, user modeling is an important link to improve the recommendation effect. In the current field of recommendation systems, whether it is an online shopping website or a movie recommendation website, these systems help users select items that the user may be interested in from a large amount of information and then recommend them to the user. In this process, they will use various information to predict the user's preferences and behaviors, and represent the user explicitly or implicitly. This process is user modeling.

[0004] However, the existing user modeling methods have several significant defects in the recommendation system, which limit the performance and applicability of the recommendation system. Specifically:

[0005] 1. Single-modal dependence: With the improvement of computing power conditions, multi-modal data is widely used in recommendation systems. The new requirements have prompted researchers to explore more complex and accurate user modeling schemes. Traditional user modeling usually relies on the user's historical interaction data, especially encoded by user ID. This method fails to fully utilize the multi-modal information contained in user interactions, such as images, texts, and audios, resulting in an incomplete and inaccurate understanding of the user's interests.

[0006] 2. Ignoring context information: Existing user encoders often ignore the context information in the user interaction sequence and only focus on the item IDs of the interactions. This simplified processing method makes the user encoder unable to capture the preference changes of the user in different situations, limiting the personalization and flexibility of the recommendation.

[0007] 3. Unreasonable negative sampling: During the training of the user encoder, the selection of negative samples is crucial for the learning of the user encoder. Existing methods usually select negative samples by random sampling, which may lead to the user encoder failing to effectively distinguish between items preferred by users and non-preferred items, thus affecting the learning effect. The lack of a targeted negative sample strategy makes the user encoder insufficient in learning the subtle differences in user preferences.

[0008] 4. Insufficient transfer ability: Many existing user encoders perform well on specific datasets, but in cross-domain recommendation tasks, the transfer ability of the user encoder is limited. Due to the lack of full consideration of the multi-modal features of users during the training process, the user encoder often performs poorly when facing new datasets and cannot effectively adapt to the recommendation needs of different scenarios.

[0009] 5. Insufficient extraction of interaction sequence features: Interaction sequence data is the most important data form in the recommendation system, but in traditional recommendation systems, the mining of features between items in the interaction sequence is not sufficient. Summary of the Invention

[0010] Aiming at the deficiencies of the existing technology, the present invention provides a user modeling method and system based on a general user encoder of Mindspore. By constructing a flexible, efficient and accurate general user encoder, the needs of modern multi-modal recommendation systems are met, and the development of recommendation technology is promoted.

[0011] The first aspect of the present invention provides a user modeling method based on a general user encoder of Mindspore, including the following steps:

[0012] Step 1: Encode the information of each item using a general item encoder to obtain the item representation of each item, and further obtain a vector representation vocabulary including the item representations of all items;

[0013] Step 2: Obtain the user's historical interaction sequence;

[0014] The user's historical interaction sequence is a time series, including the numbers of items that have interacted with the user at different times, denoted as [i 0 , i 1 ,..., i q ,..., i his , where i q is the number of the item that interacted with the user at the q-th moment in the user's historical interaction sequence, q represents the moment, and his represents the length of the user's historical interaction sequence;

[0015] Step 3: Construct a general user encoder;

[0016] The general user encoder includes a data augmentation module, an interactive sequence vector representation construction module, a user attribute information acquisition module, a first splicing module, a padding module, a shared self-attention mechanism module, a first addition module, a self-attention mechanism module, a second addition module, four average pooling layers, a second splicing module, a multiplication module, and a fusion module;

[0017] The data augmentation module is used to perform data augmentation on the input interactive sequence of the user history through four data augmentation strategies respectively, and obtain the interactive sequence of the user history after data augmentation by the four data augmentation strategies as the input of the interactive sequence vector representation construction module;

[0018] The data augmentation strategies include:

[0019] a. Random masking strategy: randomly mask the numbers of items in the interactive sequence of the user history according to a set ratio;

[0020] b. Random shuffling strategy: shuffle the numbers of items in the interactive sequence of the user history in chronological order;

[0021] c. Truncation strategy: truncate the interactive sequence of the user history and remove the numbers of some items;

[0022] d. Substitution strategy: randomly replace the numbers of items in the interactive sequence of the user history with the numbers of implicitly similar items; the implicitly similar items are items whose similarity to a certain item is within a set interval;

[0023] The interactive sequence vector representation construction module is used to query the corresponding item representations from the vector representation vocabulary according to the interactive sequences of the user history after data augmentation by the four data augmentation strategies respectively to form the interactive sequence vector representations corresponding to each data-augmented user history as the input of the first splicing module or the padding module; query the corresponding item representations from the vector representation vocabulary according to the original interactive sequence of the user history to form the interactive sequence vector representation corresponding to the original user history as the input of the first splicing module or the padding module;

[0024] The interactive sequence vector representation is where, is the item representation obtained by querying the number of the item interacting with the user at the q-th moment from the vector representation vocabulary;

[0025] The user attribute information acquisition module is used to obtain user attributes and map the user attributes into vectors to obtain user attribute information as the input of the first splicing module when user attributes are required;

[0026] The user attribute information is expressed as wherein, is the z-th element in the user attribute information, and Z is the length of the user attribute information;

[0027] The first splicing module is configured to splice the user attribute information with the interaction sequence vector representations corresponding to each data-augmented user historical interaction sequence respectively to obtain multiple input sequences as the input of the padding module; splice the user attribute information with the interaction sequence vector representation corresponding to the original user historical interaction sequence to obtain the input sequence corresponding to the original user historical interaction sequence as the input of the padding module;

[0028] The padding module is configured to, when obtaining the user attribute, pad the input sequences output by the first splicing module in a left-padding manner to obtain the padded input sequences as the input of the shared self-attention mechanism module and the first addition module; when not obtaining the user attribute, use the interaction sequence vector representation output by the interaction sequence vector representation construction module as the input sequence and pad it in a left-padding manner to obtain the padded input sequences as the input of the shared self-attention mechanism module and the first addition module;

[0029] The shared self-attention mechanism module is configured to use the multi-head attention mechanism to calculate the vector embedding representations corresponding to each padded input sequence respectively as the input of the first addition module;

[0030] The first addition module is configured to add each vector embedding representation to its corresponding padded input sequence respectively to obtain multiple enhanced sequence representations as the input of the self-attention mechanism module, including the enhanced sequence representations corresponding to the data-augmented user historical interaction sequences and the enhanced sequence representations corresponding to the original user historical interaction sequences;

[0031] The self-attention mechanism module performs self-attention mechanism calculations on each enhanced sequence representation respectively to restore the enhanced sequence representations to obtain sequence representations containing subsequence context features, sequence representations containing position features, sequence representations containing subsequence order features, sequence representations containing distinctive features of similar items, and sequence representations containing transitional features between items in the interaction sequence;

[0032] The self-attention mechanism module includes a masking module, a random shuffling module, a cropping module, a substitution module, and a self module;

[0033] The masking module, random shuffling module, cropping module, and substitution module respectively perform self-attention mechanism calculations on the enhanced sequence representations obtained through the random masking strategy, the enhanced sequence representations obtained through the random shuffling strategy, the enhanced sequence representations obtained through the cropping strategy, and the enhanced sequence representations obtained through the substitution strategy, and respectively obtain sequence representations containing subsequence context features, sequence representations containing position features, sequence representations containing subsequence order features, and sequence representations containing distinctive features of similar items;

[0034] The self-module is used to perform a self-attention mechanism calculation on the enhanced sequence representation corresponding to the original user historical interaction sequence, and obtain a sequence representation containing transitional features between items in the interaction sequence as the input to the fusion module;

[0035] The second addition module is used to add the sequence representation containing subsequence context features, the sequence representation containing position features, the sequence representation containing subsequence order features, and the sequence representation containing distinctive features of similar items to their corresponding enhanced sequence representations respectively, and obtain intermediate vectors containing subsequence context features, intermediate vectors containing position features, intermediate vectors containing subsequence order features, and intermediate vectors containing distinctive features of similar items as the input to the average pooling layer;

[0036] The four average pooling layers are used to perform average pooling on the intermediate vectors containing subsequence context features, the intermediate vectors containing position features, the intermediate vectors containing subsequence order features, and the intermediate vectors containing distinctive features of similar items respectively, and obtain intermediate vectors containing subsequence context features, intermediate vectors containing position features, intermediate vectors containing subsequence order features, and intermediate vectors containing distinctive features of similar items after average pooling as the input to the second concatenation module and the multiplication module;

[0037] The second concatenation module is used to perform a concatenation operation on the intermediate vectors containing subsequence context features, the intermediate vectors containing position features, the intermediate vectors containing subsequence order features, and the intermediate vectors containing distinctive features of similar items after average pooling, and obtain a concatenated vector as the input to the multiplication module;

[0038] The multiplication module is used to multiply the intermediate vectors containing subsequence context features, the intermediate vectors containing position features, the intermediate vectors containing subsequence order features, and the intermediate vectors containing distinctive features of similar items with the concatenated vector respectively, and obtain four product vectors as the input to the fusion module;

[0039] The fusion module is used to fuse the four product vectors and the sequence representation containing transitional features between items in the interaction sequence to obtain a user representation.

[0040] Step 4: Train the general user encoder using the user's historical interaction sequences to obtain a trained general user encoder;

[0041] Step 5: Fine-tune the trained general user encoder to obtain a fine-tuned general user encoder;

[0042] Step 5.1: Divide each user's historical interaction sequence obtained using the leave-one-out strategy to obtain a training set, a validation set, and a test set;

[0043] Step 5.2: Perform negative sampling on the training set and the validation set using the Hard-Negative strategy to obtain the sampled training set and validation set;

[0044] Step 5.3: Apply the trained general user encoder to the recommendation system and use the sampled training set and validation set to train the entire recommendation system. The general user encoder in the trained recommendation system is the fine-tuned general user encoder;

[0045] The recommendation system includes a general user encoder, a general item encoder, and a downstream task module; the downstream task module is used to perform downstream tasks based on the item representation obtained from the general item encoder and the user representation obtained from the general user encoder;

[0046] Step 6: Obtain the user's historical interaction sequence to be modeled, and input the user's historical interaction sequence into the fine-tuned general user encoder to obtain a user representation.

[0047] The second aspect of the present invention provides a user modeling system for a general user encoder based on Mindspore, which is used to implement the user modeling method for the general user encoder based on Mindspore, including:

[0048] A general item encoder for encoding the information of each item to obtain an item representation of each item;

[0049] A vector representation vocabulary construction module that constructs a vector representation vocabulary from the item representations of each item;

[0050] A user historical interaction sequence acquisition module for acquiring the user's historical interaction sequence;

[0051] A general user encoder for encoding the user's historical interaction sequence based on the item representations of each item in the vector representation vocabulary to obtain a user representation.

[0052] A third aspect of the present invention provides an electronic device, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are executed by the processor, the steps of the user modeling method based on Mindspore's general user encoder are executed.

[0053] A fourth aspect of the present invention provides a computer-readable storage medium. A computer program is stored in the computer-readable storage medium. When the computer program is run by a processor, the steps of the user modeling method based on Mindspore's general user encoder are executed.

[0054] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0055] 1. Unified encoding of multimodal information: By integrating multimodal information (such as text, images, etc.) contained in the user's interaction sequence, a comprehensive understanding of the user is achieved. The user encoder can generate rich user representations to more accurately capture the user's interests and preferences.

[0056] 2. Improved recommendation accuracy: By leveraging the fusion of multimodal information, the performance of the recommendation system in personalized recommendation tasks is enhanced. Through in-depth mining of user behavior, it is ensured that the recommendation results better meet the actual needs of users, thereby improving user satisfaction.

[0057] 3. Better negative sampling strategy: Adopting the new Hard-Negative negative sampling strategy enhances the learning ability of the user encoder in distinguishing between user-preferred and non-preferred items. By specifically selecting negative samples, the discriminative ability of the user encoder is improved, which will contribute to enhancing the performance of the recommendation system in diverse and complex user behaviors, thereby improving the accuracy and robustness of recommendations.

[0058] 4. Enhanced transfer ability: Utilizing the user's historical interaction information and their attribute information to form a richer and more multi-dimensional user representation can not only improve the accuracy of recommendations but also enhance the transfer ability of the user encoder across different datasets and recommendation scenarios. This will enable the user encoder to adapt to diverse application requirements, thereby achieving higher recommendation efficiency and user satisfaction.

[0059] 5. Better data augmentation strategy: By performing various different data augmentations on the interaction sequence, various features between items in the interaction sequence are extracted. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 It is a training flowchart in an embodiment of the present invention;

[0061] Figure 2The structural diagram of the recommended large model in the embodiments of the present invention;

[0062] Figure 3 The processing flowchart of the general user encoder in the embodiments of the present invention. Detailed implementation manners

[0063] The present invention will be described in detail below with reference to the accompanying drawings and embodiments.

[0064] The present invention proposes a general user encoder that can accept multi-modal input information under the Mindspore framework. First, it is necessary to integrate the limited model resources under the Mindspore framework, and migrate, debug, and optimize models such as Transformer under the PyTorch framework to make them perform better.

[0065] The key point of the present invention is to propose a general user encoder based on multi-modal information, aiming to solve the deficiencies of existing user modeling methods. The general user encoder can effectively integrate the multi-modal interaction information of users, including text and images, etc., so as to generate rich user representations and achieve a comprehensive understanding of user interests and preferences. Adopting the two-stage training strategy shown in Figure 1 combines the pre-training and fine-tuning processes. First, preliminary training is carried out through various data augmentation strategies to enable the general user encoder to learn the representations of multi-modal information; subsequently, in the fine-tuning stage, specific recommendation tasks are combined for refinement adjustment to ensure the best performance of the general user encoder in downstream tasks. The low-rank matrix (LoRA) module is also introduced to reduce the training overhead by reducing the number of trainable parameters and enhance the adaptability and training efficiency of the general user encoder on different recommendation datasets. In addition, for the selection of negative samples, the present invention adopts a new Hard-Negative strategy. By calculating the similarity between candidate negative samples and user representations, low-similarity negative samples are eliminated, thereby improving the learning ability of the general user encoder in distinguishing user preferences from non-preferred items and enhancing its discriminative ability and recommendation accuracy. During the user encoding process, the present invention makes full use of context information. Through dynamic padding and positional encoding, it ensures that the general user encoder can capture the preference changes of users in different situations, thereby enhancing the personalization and flexibility of recommendations. At the same time, the general user encoder is designed with good cross-domain migration ability, enabling the general user encoder to quickly adapt to different datasets and recommendation scenarios and reducing performance degradation caused by data scarcity. Finally, the present invention has also optimized the computing performance to ensure that the general user encoder maintains high efficiency when processing large-scale datasets to support the requirements of real-time recommendations.

[0066] A user modeling method for a general user encoder based on Mindspore is provided in this embodiment, including the following steps:

[0067] Step 1: Use a general item encoder to encode the information of each item to obtain the item representation of each item, and further obtain a vector representation vocabulary (embedding_table) including the item representations of all items; the information of each item includes the text description, picture, name, etc. of the item.

[0068] In this embodiment, the general item encoder in the recommended large model as Figure 2 shown is adopted, and after pre-training, the pre-trained general item encoder is fine-tuned using the images and text descriptions of the items in the recommendation dataset so that it can adapt to the used recommendation dataset. In this fine-tuning step, a low-rank matrix (LoRA) module is added to the general item encoder, and its principle is:

[0069] W = BA

[0070] where W is the original weight matrix in the neural network, and BA is the low-rank matrix, which reduces the overhead by reducing the number of trainable parameters in the fine-tuning process;

[0071] Then, use the fine-tuned general item encoder to generate a vector representation vocabulary (embedding_table) corresponding to all items for use by the general user encoder and downstream tasks (such as recommendation tasks and similarity matching);

[0072] Step 2: Obtain the user's historical interaction sequence; the user's historical interaction sequence is a time series, including the numbers of the items that have interacted with the user at different times.

[0073] The user's historical interaction sequence is [i 0 , i 1 ,..., i q ,..., i his , where i q is the number of the item that interacts with the user at the q-th moment in the user's historical interaction sequence, q represents the moment, and his represents the length of the user's historical interaction sequence;

[0074] Step 3: Construct a general user encoder;

[0075] As Figure 3 shown, the general user encoder includes a data augmentation module, an interaction sequence vector representation construction module, a user attribute information acquisition module, a first splicing module, a padding module, a shared self-attention mechanism module, a first addition module, a self-attention mechanism module, a second addition module, four average pooling layers, a second splicing module, a multiplication module, and a fusion module;

[0076] The data augmentation module is used to perform data augmentation on the input user historical interaction sequence through four data augmentation strategies, and obtain the user historical interaction sequence after data augmentation by the four data augmentation strategies as the input of the interaction sequence vector representation construction module;

[0077] The data augmentation strategies include:

[0078] a. Random masking strategy (Mask): Randomly mask the item numbers in the user historical interaction sequence according to a set ratio, that is, [i 1 , i 2 , i 3 , i 4 → [i 1 , Mask, i 3 , Mask], where Mask represents the mask;

[0079] In this embodiment, the set ratio is 15%;

[0080] b. Random shuffling strategy (Shuffle): Shuffle the item numbers in the user historical interaction sequence in chronological order, that is, [i 1 , i 2 , i 3 , i 4 → [i 2 , i 1 , i 4 , i 3 ;

[0081] c. Cropping strategy (Crop): Crop the user historical interaction sequence and remove the item numbers of some items, that is, [i 1 , i 2 , i 3 , i 4 → [i 1 , i 2 ;

[0082] d. Substitution strategy (Substitute): Randomly replace the item numbers in the user historical interaction sequence with the numbers of implicitly similar items, that is where is the number of the implicitly similar item; The implicitly similar item is an item whose similarity to a certain item is within a set range;

[0083] In this embodiment, the set range of similarity is 40% - 70%;

[0084] The interactive sequence vector representation construction module is used to query the corresponding item representation from the vector representation vocabulary according to the interactive sequences of the user history after data augmentation by four data augmentation strategies respectively, to form the interactive sequence vector representation corresponding to each data-augmented user history, and use it as the input of the first splicing module or the padding module; query the corresponding item representation from the vector representation vocabulary according to the original user history's interactive sequence to form the interactive sequence vector representation corresponding to the original user history's interactive sequence, and use it as the input of the first splicing module or the padding module;

[0085] The interactive sequence vector representation is wherein, is the item representation obtained by querying the number of the item interacted with the user at the q-th moment from the vector representation vocabulary;

[0086] The user attribute information acquisition module is used to obtain user attributes and map the user attributes to vectors to obtain user attribute information as the input of the first splicing module when user attributes are required;

[0087] The user attribute information is represented as wherein, is the z-th element in the user attribute information, and Z is the length of the user attribute information;

[0088] In this embodiment, the user attributes are gender and age, etc., and the embedding layer is used to map the user attributes to vectors;

[0089] The first splicing module is used to splice the user attribute information with the interactive sequence vector representation corresponding to each data-augmented user history respectively to obtain multiple input sequences as the input of the padding module; splice the user attribute information with the interactive sequence vector representation corresponding to the original user history's interactive sequence to obtain the input sequence corresponding to the original user history's interactive sequence as the input of the padding module;

[0090] The input sequences obtained by splicing the user attribute information with the four interactive sequence vector representations respectively are:

[0091]

[0092] wherein, e in is the input sequence;

[0093] The padding module is used to pad the input sequence output by the first splicing module in a left-padding manner to obtain a padded input sequence as the input to the shared self-attention mechanism module and the first addition module when the user attribute is obtained; when the user attribute is not obtained, the interaction sequence vector representation output by the interaction sequence vector representation construction module is used as the input sequence and padded in a left-padding manner to obtain a padded input sequence as the input to the shared self-attention mechanism module and the first addition module;

[0094] The specific left-padding method is as follows: when the input sequence includes user attribute information, a sequence padding marker is added between the user attribute information and the interaction sequence vector representation; otherwise, a sequence padding marker is added before the interaction sequence vector representation;

[0095] In this embodiment, to ensure that the lengths of the input sequences are consistent, padding is performed on the input sequences with insufficient lengths <pad>The padding of the marker adopts the left-padding method, and the input sequence after padding is where e pad is the padding marker. Since position encoding will be added during the processing of the input sequence, right-padding will result in all the position encodings representing the most recent interactive items being empty;

[0096] The shared self-attention mechanism module is used to calculate the vector embedding representation corresponding to each padded input sequence respectively by using the multi-head attention mechanism as the input of the first addition module;

[0097] The first addition module is used to add each vector embedding representation to its corresponding padded input sequence to obtain multiple enhanced sequence representations respectively as the input of the self-attention mechanism module, including the enhanced sequence representation corresponding to the user historical interaction sequence after data enhancement and the enhanced sequence representation corresponding to the original user historical interaction sequence;

[0098] The self-attention mechanism module performs self-attention mechanism calculations on each enhanced sequence representation respectively to restore the enhanced sequence representation to obtain a sequence representation containing subsequence context features, a sequence representation containing position features, a sequence representation containing subsequence order features, a sequence representation containing the discriminative features of similar items, and a sequence representation containing the transitional features between interactive sequence items;

[0099] The self-attention mechanism module includes a masking module, a random shuffling module, a cropping module, a substitution module, and a self module;

[0100] The masking module, the random shuffling module, the cropping module, and the substitution module are used to perform self-attention mechanism calculations on the enhanced sequence representation obtained by the random masking strategy, the enhanced sequence representation obtained by the random shuffling strategy, the enhanced sequence representation obtained by the cropping strategy, and the enhanced sequence representation obtained by the substitution strategy respectively to obtain a sequence representation containing subsequence context features, a sequence representation containing position features, a sequence representation containing subsequence order features, and a sequence representation containing the discriminative features of similar items;

[0101] The self module is used to perform self-attention mechanism calculations on the enhanced sequence representation corresponding to the original user historical interaction sequence to obtain a sequence representation containing the transitional features between interactive sequence items as the input of the fusion module.

[0102] The second addition module is used to add the sequence representation containing subsequence context features, the sequence representation containing position features, the sequence representation containing subsequence order features, and the sequence representation containing the discriminative features of similar items to their corresponding enhanced sequence representations respectively, to obtain the intermediate vectors of the subsequence context features, the intermediate vectors of the position features, the intermediate vectors of the subsequence order features, and the intermediate vectors of the discriminative features of similar items as the input of the average pooling layer;

[0103] The four average pooling layers are used to perform average pooling on the intermediate vectors of the subsequence context features, the intermediate vectors of the position features, the intermediate vectors of the subsequence order features, and the intermediate vectors of the discriminative features of similar items respectively, to obtain the intermediate vectors of the subsequence context features, the intermediate vectors of the position features, the intermediate vectors of the subsequence order features, and the intermediate vectors of the discriminative features of similar items after average pooling as the input of the second concatenation module and the multiplication module;

[0104] The second concatenation module is used to perform a concatenation operation on the intermediate vectors of the subsequence context features, the intermediate vectors of the position features, the intermediate vectors of the subsequence order features, and the intermediate vectors of the discriminative features of similar items after average pooling, to obtain a concatenated vector as the input of the multiplication module;

[0105] The multiplication module is used to multiply the intermediate vectors of the subsequence context features, the intermediate vectors of the position features, the intermediate vectors of the subsequence order features, and the intermediate vectors of the discriminative features of similar items with the concatenated vector respectively, to obtain four product vectors as the input of the fusion module;

[0106] The fusion module is used to fuse the four product vectors and the sequence representation containing the transitional features between interactive sequence items to obtain a user representation.

[0107] In this embodiment, the fusion is implemented by a multi-layer perceptron (MLP);

[0108] Step 4: Use the user's historical interaction sequence to train the general user encoder to obtain a trained general user encoder;

[0109] Step 5: Fine-tune the trained general user encoder to obtain a fine-tuned general user encoder;

[0110] Step 5.1: Adopt the leave-one-out strategy to divide each user's historical interaction sequence obtained to obtain a training set, a validation set, and a test set;

[0111] In this embodiment, the last sample of each user's historical interaction sequence is used as the test set, the penultimate sample is used as the validation set, and the previous samples are used as the training set; the sample is the number of an item in the user's historical interaction sequence.

[0112] Step 5.2: Perform negative sampling on the training set and the validation set using the Hard-Negative strategy to obtain the sampled training set and validation set.

[0113] Specifically: For a sample in a user's historical interaction sequence, the sample itself is a positive sample, and the corresponding position samples in other users' historical interaction sequences are negative samples. First, calculate the similarity between each negative sample in the negative sample candidate set and the positive sample, and eliminate the negative samples with similarity lower than the set value. Such samples are too easy and cannot train the representation effect of the general user encoder well. Then, eliminate the negative samples with similarity higher than the set value, because such samples are likely to be items that the user has not interacted with but is interested in, so it is unreasonable to use them as negative samples to train the general user encoder.

[0114] Step 5.3: Apply the trained general user encoder to the recommendation system, and use the sampled training set and validation set to train the entire recommendation system. The general user encoder in the trained recommendation system is the fine-tuned general user encoder.

[0115] The recommendation system includes a general user encoder, a general item encoder, and a downstream task module; the downstream task module is used to perform downstream tasks (recommendation tasks, similarity matching, etc.) based on the item representation obtained by the general item encoder and the user representation obtained by the general user encoder.

[0116] Step 6: Obtain the user's historical interaction sequence to be modeled, input the user's historical interaction sequence into the fine-tuned general user encoder to obtain the user representation, and the subsequent downstream task module can use the user representation obtained by the fine-tuned general user encoder to implement the recommendation task.

[0117] In this embodiment, a user modeling system based on the general user encoder of Mindspore is provided, which is used to implement the user modeling method based on the general user encoder of Mindspore, including:

[0118] A general item encoder, which is used to encode the information of each item to obtain the item representation of each item.

[0119] A vector representation vocabulary construction module, which constructs a vector representation vocabulary from the item representations of each item.

[0120] A user historical interaction sequence acquisition module, which is used to acquire the user's historical interaction sequence.

[0121] A general user encoder, which is used to encode the user's historical interaction sequence based on the item representation that represents each item in the vocabulary by vectors, so as to obtain the user representation;

[0122] In this embodiment, an electronic device is provided, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are executed by the processor, the steps of the user modeling method based on the general user encoder of Mindspore are executed;

[0123] This embodiment provides a computer-readable storage medium, in which a computer program is stored. When the computer program is run by a processor, the steps of the user modeling method based on the general user encoder of Mindspore are executed.< / pad>

Claims

1. A user modeling method based on a universal user encoder of Mindspore, characterized in that: The following steps are involved: Step 1: Use a universal item encoder to encode the information of each item to obtain the item representation of each item, and then obtain a vector representation vocabulary including the item representations of all items; Step 2: Obtain the user's historical interaction sequence; Step 3: Build a universal user encoder; Step 4: Use the user's historical interaction sequence to train the universal user encoder to obtain a trained universal user encoder; Step 5: Fine-tune the trained universal user encoder to obtain a fine-tuned universal user encoder; Step 6: Obtain the interaction sequence of the user history to be modeled, and input the interaction sequence of the user history into the fine-tuned universal user encoder to obtain the user representation.

2. The user modeling method based on the universal user encoder of Mindspore according to claim 1, characterized in that: The user's historical interaction sequence in step 2 is a time series, including the numbers of items that the user interacted with at different times, represented as [i0,i1,...,i q ,...,i his ],i q is the number of the item that the user interacted with at the qth moment in the user's historical interaction sequence, q represents the moment, and his represents the length of the user's historical interaction sequence.

3. The user modeling method based on the Mindspore universal user encoder according to claim 2, characterized in that: The universal user encoder in step 3 includes a data enhancement module, an interaction sequence vector representation construction module, a user attribute information acquisition module, a first splicing module, a completion module, a shared self-attention mechanism module, a first addition module, a self-attention mechanism module, a second addition module, four average pooling layers, a second splicing module, a multiplication module and a fusion module; The data enhancement module is used to perform data enhancement on the input user history interaction sequence respectively through four data enhancement strategies, and obtain the user history interaction sequence after data enhancement through the four data enhancement strategies as the input of the interaction sequence vector representation construction module; The interaction sequence vector representation construction module is used to query the corresponding item representation from the vector representation vocabulary according to the interaction sequence of the user history after data enhancement by the four data enhancement strategies to form the interaction sequence vector representation corresponding to each interaction sequence of the user history after data enhancement, as the input of the first splicing module or the completion module; query the corresponding item representation from the vector representation vocabulary according to the original interaction sequence of the user history to form the interaction sequence vector representation corresponding to the original interaction sequence of the user history, as the input of the first splicing module or the completion module; The interaction sequence vector is represented as in, The item representation obtained by searching the vector representation vocabulary for the item number that the user interacted with at the qth moment; The user attribute information acquisition module is used to acquire user attributes and map the user attributes into vectors when user attributes are needed, so as to obtain user attribute information as input of the first splicing module; The user attribute information is represented as in, is the zth element in the user attribute information, and Z is the length of the user attribute information; The first concatenation module is used to concatenate the user attribute information with the interaction sequence vector representation corresponding to each data-enhanced user history interaction sequence, to obtain multiple input sequences as inputs of the completion module; concatenate the user attribute information with the interaction sequence vector representation corresponding to the original user history interaction sequence, to obtain the input sequence corresponding to the original user history interaction sequence as input of the completion module; The padding module is used to pad the input sequence output by the first concatenation module by left padding when obtaining user attributes, and obtain the padding input sequence as the input of the shared self-attention mechanism module and the first addition module; when the user attributes are not obtained, the interaction sequence vector representation output by the interaction sequence vector representation construction module is used as the input sequence and padding is performed by left padding, and the padding input sequence is used as the input of the shared self-attention mechanism module and the first addition module; The shared self-attention mechanism module is used to use the multi-head attention mechanism to calculate the vector embedding representation corresponding to each padded input sequence as the input of the first addition module; The first adding module is used to add each vector embedding representation to its corresponding padded input sequence to obtain multiple enhanced sequence representations as inputs of the self-attention mechanism module, including the enhanced sequence representation corresponding to the interaction sequence of the user history after data enhancement and the enhanced sequence representation corresponding to the original interaction sequence of the user history; The self-attention mechanism module performs self-attention mechanism calculation on each enhanced sequence representation, and restores the enhanced sequence representation to obtain a sequence representation containing subsequence context features, a sequence representation containing position features, a sequence representation containing subsequence order features, a sequence representation containing distinctive features of similar items, and a sequence representation containing transition features between interactive sequence items; The second adding module is used to add the sequence representation containing subsequence context features, the sequence representation containing position features, the sequence representation containing subsequence order features, and the sequence representation containing the distinctive features of similar items to their corresponding enhanced sequence representations, and obtain the intermediate vector containing subsequence context features, the intermediate vector containing position features, the intermediate vector containing subsequence order features, and the intermediate vector containing the distinctive features of similar items as the input of the average pooling layer; The four average pooling layers are used to respectively average pool the intermediate vectors containing subsequence context features, the intermediate vectors containing position features, the intermediate vectors containing subsequence order features, and the intermediate vectors containing the distinctive features of similar items, and obtain the intermediate vectors containing subsequence context features, the intermediate vectors containing position features, the intermediate vectors containing subsequence order features, and the intermediate vectors containing the distinctive features of similar items after average pooling as inputs of the second concatenation module and the multiplication module; The second concatenation module is used to concatenate the intermediate vector containing the subsequence context feature, the intermediate vector containing the position feature, the intermediate vector containing the subsequence order feature, and the intermediate vector containing the distinctive feature of similar items after average pooling, and obtain a concatenated vector as an input of the multiplication module; The multiplication module is used to respectively multiply the intermediate vector containing the subsequence context feature, the intermediate vector containing the position feature, the intermediate vector containing the subsequence order feature and the intermediate vector containing the distinctive feature of similar items with the concatenation vector to obtain four product vectors as inputs of the fusion module; The fusion module is used to fuse the four product vectors with the sequence representation containing the transitional features between the items in the interaction sequence to obtain the user representation.

4. The user modeling method based on the Mindspore universal user encoder according to claim 3, characterized in that: The data enhancement strategy includes: a. Random masking strategy: Randomly mask the item numbers in the user's historical interaction sequence according to a set ratio; b. Random shuffling strategy: shuffle the item numbers in the user's historical interaction sequence in chronological order; c. Clipping strategy: Clip the user's historical interaction sequence and remove the numbers of some items; d. Replacement strategy: randomly replace the numbers of items in the user's historical interaction sequence with the numbers of implicitly similar items; the implicitly similar items are items whose similarity to a certain item is within a set range.

5. The user modeling method based on the Mindspore universal user encoder according to claim 4, characterized in that: The self-attention mechanism module includes a masking module, a random shuffling module, a cropping module, a replacement module and a self module; The masking module, random shuffling module, cropping module and replacement module respectively perform self-attention mechanism calculation on the enhanced sequence representation obtained by the random masking strategy, the enhanced sequence representation obtained by the random shuffling strategy, the enhanced sequence representation obtained by the cropping strategy and the enhanced sequence representation obtained by the replacement strategy, and respectively obtain a sequence representation containing subsequence context features, a sequence representation containing position features, a sequence representation containing subsequence order features and a sequence representation containing distinctive features of similar items; The self-module is used to perform self-attention mechanism calculation on the enhanced sequence representation corresponding to the original user history interaction sequence, and obtain the sequence representation containing the transitional features between items in the interaction sequence as the input of the fusion module.

6. The user modeling method based on the Mindspore universal user encoder according to claim 1, characterized in that: Step 5 specifically includes: Step 5.1: Use the leave-one-out strategy to divide the interaction sequence of each user's history into training set, validation set and test set; Step 5.2: Use the Hard-Negative strategy to perform negative sampling on the training set and validation set to obtain the sampled training set and validation set; Step 5.3: Apply the trained universal user encoder to the recommendation system, and use the sampled training set and validation set to train the entire recommendation system. The trained universal user encoder in the recommendation system is the fine-tuned universal user encoder. The recommendation system includes a universal user encoder, a universal item encoder and a downstream task module; the downstream task module is used to perform downstream tasks according to the item representation obtained by the universal item encoder and the user representation obtained by the universal user encoder.

7. A user modeling system based on a universal user encoder of Mindspore, used to implement the user modeling method based on a universal user encoder of Mindspore according to any one of claims 1 to 6, characterized in that: include: A universal item encoder, used to encode the information of each item to obtain an item representation of each item; A vector representation vocabulary building module that constructs a vector representation vocabulary from the item representation of each item; A user history interaction sequence acquisition module is used to acquire the user history interaction sequence; A universal user encoder is used to encode the user's historical interaction sequence based on the item representation of each item in the vector representation vocabulary to obtain the user representation.

8. An electronic device, characterized in that: include: A processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor and the memory communicate via the bus, and when the machine-readable instructions are executed by the processor, the steps of the user modeling method based on the universal user encoder of Mindspore according to any one of claims 1 to 6 are performed.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, executes the steps of the user modeling method based on the universal user encoder based on Mindspore as claimed in any one of claims 1 to 6.