A text-based recommendation method, device, equipment and storage medium

Through the text encoder and semantic meaning encoder combined with the cross attention mechanism, the problem of insufficient user interaction data is solved, and the accuracy and effectiveness of the recommendation model are improved.

CN119807541BActive Publication Date: 2025-07-08SHENZHEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510301917.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-07-08
Estimated Expiration
2045-03-14

AI Technical Summary

Technical Problem

The existing recommendation model is difficult to accurately learn user interest characteristics when user interaction data is insufficient, resulting in poor recommendation results.

Method used

The text-based recommendation method is adopted to obtain the user's semantic intention and behavioral intention through text encoder and semantic intention encoder, and cross-domain learning is carried out in combination with the cross-attention mechanism to generate recommended items.

Benefits of technology

Even in the case of insufficient interaction sequences in the target domain, the accuracy and effectiveness of the recommendation model are improved through cross-domain learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119807541B_ABST
    Figure CN119807541B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of item recommendation, and specifically, to a text-based recommendation method, device, equipment, and storage medium. The present invention applies a text encoder and a semantic intention encoder to an interaction sequence generated by a user in a target domain to obtain a semantic intention; the interaction sequence is further applied with an item encoder and a behavior intention encoder to obtain a behavior intention. Finally, according to the semantic intention and the behavior intention, as well as the item number and text semantics of the candidate items, items are screened out from the candidate items and recommended to the user. Since a cross-attention mechanism for assisting training is introduced during training, the cross-attention mechanism enables the semantic intention encoder in the target domain to learn based on both the target domain and the source domain. Therefore, even if the number of interaction sequences in the target domain is insufficient, the semantic intention encoder can fully learn the user's semantic intention, and ultimately improves the recommendation effect of the recommendation model based on the semantic intention encoder of the present invention.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of item recommendation, and in particular, to a text-based recommendation method, device, equipment, and storage medium. Background Art

[0002] Recommendation systems are a class of automated tools widely used in various online platforms, aiming to help users discover products or content that match their preferences by modeling users' interest intentions, thereby effectively alleviating the problem of information overload. Sequential recommendation algorithms consider a series of interaction behaviors between users and items to capture the changing patterns of users' preferences over time and predict the products or content that users are most likely to be interested in in the future. However, most existing sequential recommendation models often rely on rich user interaction data to learn accurate interest representations. When the user interaction records are insufficient, such models may be difficult to fully capture the interest characteristics of users, resulting in poor recommendation effects. That is, when the user interaction record data is less, using less interaction record data makes it difficult for the recommendation model to accurately learn users' interests.

[0003] In summary, the existing recommendation models reduce the recommendation effect in the case of less interaction data of users.

[0004] Therefore, the existing technology still needs to be improved. Summary of the Invention

[0005] To solve the above technical problems, the present invention provides a text-based recommendation method, device, equipment, and storage medium, which solves the problem that the existing recommendation models reduce the recommendation effect in the case of less interaction data of users.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] In a first aspect, the present invention provides a text-based recommendation method, which includes:

[0008] Obtain a target domain interaction sequence of a user containing interaction items, apply a text encoder to the target domain interaction sequence to obtain a target text semantic representation matrix; apply a semantic intention encoder to the target text semantic representation matrix to obtain the semantic intention of the user;

[0009] Apply an item encoder to the target domain interaction sequence to obtain an item encoding matrix; apply a behavior intention encoder to the item encoding matrix to obtain the user's behavior intention, where the semantic intention encoder, the text encoder, the item encoder, and the behavior intention encoder are all pre-trained encoders, and a cross-attention mechanism for auxiliary training is introduced during the training process. The cross-attention mechanism is used to enable the semantic intention encoder to perform cross-learning between the source domain and the target domain, so that the semantic intention encoder becomes a cross-domain encoder;

[0010] Obtain candidate items and the item numbers of the candidate items, and generate a text semantic representation matrix of the candidate items. Based on the text semantic representation matrix of the candidate items, the item numbers, the user's semantic intention, and the behavior intention, filter out the recommended items to be recommended to the user from the candidate items.

[0011] In one implementation, filtering out the recommended items to be recommended to the user from the candidate items based on the text semantic representation matrix of the candidate items, the item numbers, the user's semantic intention, and the behavior intention includes:

[0012] Generate an item vector of the candidate item based on the text semantic representation matrix and the item number of the candidate item;

[0013] Obtain a preference vector of the user based on the semantic intention and the behavior intention, where the preference vector is used to represent the preference degree of the user for each interaction item;

[0014] Filter out the recommended items to be recommended to the user from the candidate items based on the preference vector of the user and the item vector of each candidate item.

[0015] In one implementation, the semantic intention encoder, the text encoder, the item encoder, and the behavior intention encoder adopt a joint training method, where the joint training method includes:

[0016] Set up a text auxiliary encoder and an item auxiliary encoder. The text auxiliary encoder and the item auxiliary encoder are two encoders set up for auxiliary training of the semantic intention encoder, the text encoder, the item encoder, and the behavior intention encoder;

[0017] Obtain the interaction sample sequence of the same user in the source domain and the interaction sample sequence in the target domain;

[0018] Apply the text auxiliary encoder and the item auxiliary encoder to the interaction sample sequence in the source domain respectively to obtain a source sample text semantic representation matrix and a source sample item encoding matrix;

[0019] Apply the text encoder and the item encoder to the interaction sample sequence in the target domain respectively to obtain a target sample text semantic representation matrix and a target sample item encoding matrix;

[0020] Through the cross-attention mechanism, based on the target sample text semantic representation matrix, the source sample text semantic representation matrix, and the source sample item encoding matrix, obtain the cross-domain preference as a positive sample;

[0021] Through the semantic intention encoder, based on the target sample text semantic representation matrix and the neighborhood shared attention matrix composed of the source sample text semantic representation matrix and the target sample text semantic representation matrix, obtain the target sample semantic intention;

[0022] Through the behavior intention encoder, based on the target sample item encoding matrix, obtain the target sample behavior intention;

[0023] Based on the target sample semantic intention and the target sample behavior intention, obtain the in-domain preference as a positive sample;

[0024] Obtain the cross-domain preference of other users and use this cross-domain preference as the cross-domain preference of the negative sample;

[0025] Based on the cross-domain preference as a positive sample, the in-domain preference as a positive sample, and the cross-domain preference as a negative sample, obtain the contrastive loss;

[0026] Based on the contrastive loss, construct the total loss and perform iterative training based on the total loss.

[0027] In one implementation, constructing the total loss based on the contrastive loss includes:

[0028] Aggregate the cross-domain preference and the in-domain preference of the same user to obtain the final preference of the sample of the same user;

[0029] Based on the text semantic vector and item number of the candidate item and the final preference of the sample, obtain the cross-entropy loss;

[0030] Based on the contrastive loss and the cross-entropy loss, construct the total loss.

[0031] In one implementation, through the cross-attention mechanism, based on the target sample text semantic representation matrix, the source sample text semantic representation matrix, and the source sample item encoding matrix, obtaining the cross-domain preference as a positive sample includes:

[0032] By using the target sample text semantic representation matrix as the query vector of the cross-attention mechanism, the source sample text semantic representation matrix as the key vector of the cross-attention mechanism, and the source sample item encoding matrix as the value vector of the cross-attention mechanism, an inter-domain preference as a positive sample is obtained.

[0033] In one implementation, through the semantic intention encoder, according to the target sample text semantic representation matrix and the neighborhood shared attention matrix composed of the source sample text semantic representation matrix and the target sample text semantic representation matrix, the target sample semantic intention is obtained, including:

[0034] Using the target sample text semantic representation matrix as the query vector and the key vector of the attention mechanism inside the semantic intention encoder respectively, a target domain attention matrix is obtained;

[0035] Combining the neighborhood shared attention matrix and the target domain attention matrix to obtain a shared semantic association attention matrix;

[0036] By using the target sample text semantic representation matrix as the value vector, according to the shared semantic association attention matrix, the target sample semantic intention is obtained.

[0037] In one implementation, the semantic intention and the behavior intention of the user respectively represent the semantic intention and the behavior intention reflected by the user's last interaction.

[0038] In a second aspect, an embodiment of the present invention further provides a text-based recommendation device, where the device includes the following components:

[0039] A semantic intention prediction module, configured to obtain a target domain interaction sequence of a user including interaction items, apply a text encoder to the target domain interaction sequence to obtain a target text semantic representation matrix; apply a semantic intention encoder to the target text semantic representation matrix to obtain the semantic intention of the user;

[0040] A behavior intention prediction module, configured to apply an item encoder to the target domain interaction sequence to obtain an item encoding matrix; apply a behavior intention encoder to the item encoding matrix to obtain the behavior intention of the user; the semantic intention encoder, the text encoder, the item encoder, and the behavior intention encoder are all pre-trained encoders, and a cross-attention mechanism for auxiliary training is introduced during the training process, and the cross-attention mechanism is used to enable the semantic intention encoder to perform cross-learning in the source domain and the target domain, so that the semantic intention encoder becomes a cross-domain encoder;

[0041] A recommendation module, configured to obtain candidate items and the item numbers of the candidate items, generate a text semantic representation matrix of the candidate items, and screen out recommended items to be recommended to the user from the candidate items according to the text semantic representation matrix of the candidate items, the item numbers, the semantic intention and the behavior intention of the user.

[0042] In a third aspect, an embodiment of the present invention further provides a terminal device, where the terminal device includes a memory, a processor, and a text-based recommendation program stored in the memory and executable on the processor. When the processor executes the text-based recommendation program, the steps of the above-mentioned text-based recommendation method are implemented.

[0043] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which a text-based recommendation program is stored. When the text-based recommendation program is executed by a processor, the steps of the above-mentioned text-based recommendation method are implemented.

[0044] Beneficial effects: The present invention applies a text encoder and a semantic intention encoder to the interaction sequence generated by the user in the target domain to obtain the semantic intention of the user; the item encoder and the behavior intention encoder are further applied to the interaction sequence to obtain the behavior intention of the user. Finally, according to the semantic intention and behavior intention of the user, the item numbers and text semantics of the candidate items, recommended items are screened out from the candidate items and recommended to the user. Since the present invention introduces a cross-attention mechanism for auxiliary training when training the above four encoders, the cross-attention mechanism enables the semantic intention encoder in the target domain to learn based on both the interaction sequence in the target domain and the interaction sequence in the source domain. Therefore, even if the number of interaction sequences in the target domain is insufficient, since the semantic intention encoder also learns in the source domain, the semantic intention encoder can fully learn the user's semantic intention, and finally improves the recommendation effect of the recommendation model based on the semantic intention encoder of the present invention. Description of the Drawings

[0045] Figure 1 is the overall flowchart of the present invention;

[0046] Figure 2 is the prediction flowchart in the embodiment of the present invention;

[0047] Figure 3 is the training flowchart in the embodiment of the present invention;

[0048] Figure 4 is the schematic diagram of positive and negative sample pairs in the embodiment of the present invention;

[0049] Figure 5 is the schematic diagram of the influence of the contrast loss weight in the embodiment of the present invention;

[0050] Figure 6 Structural diagram of the text-based recommendation device provided by the present invention;

[0051] Figure 7 Internal structure principle block diagram of the terminal device provided by the embodiment of the present invention. Detailed implementation manners

[0052] The following combines the embodiments and the accompanying drawings of the specification to clearly and completely describe the technical solutions in the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.

[0053] It has been found through research that the recommendation system is a type of automated tool widely used in various online platforms, aiming to help users discover products or content that match their preferences by modeling the user's interest intentions, thereby effectively alleviating the problem of information overload. The sequential recommendation algorithm considers a series of interaction behaviors between users and items to capture the changing patterns of user preferences over time and predict the products or content that the user is most likely to be interested in in the future. However, most existing sequential recommendation models often rely on rich user interaction data to learn accurate interest representations. When the user interaction records are insufficient, such models may be difficult to fully capture the user's interest characteristics, resulting in poor recommendation effects. That is, when the user interaction record data is less, using less interaction record data makes it difficult for the recommendation model to accurately learn the user's interests and hobbies.

[0054] To solve the above technical problems, the present invention provides a text-based recommendation method, device, equipment and storage medium, which solves the problem that the existing recommendation model reduces the recommendation effect in the case of less user interaction data.

[0055] The text-based recommendation method of this embodiment can be applied to a terminal device, and the terminal device can be a terminal product with data processing functions, such as a computer, etc. In this embodiment, as Figure 1 shown in, the text-based recommendation method specifically includes the following steps:

[0056] S100, obtain the target domain interaction sequence of the user containing interaction items, apply a text encoder to the target domain interaction sequence to obtain a target text semantic representation matrix; apply a semantic intention encoder to the target text semantic representation matrix to obtain the semantic intention of the user;

[0057] S200. Apply the item encoder to the target domain interaction sequence to obtain an item encoding matrix; apply the behavior intention encoder to the item encoding matrix to obtain the user's behavior intention. Among them, the semantic intention encoder, the text encoder, the item encoder, and the behavior intention encoder are all pre-trained encoders. During the training process, a cross-attention mechanism for auxiliary training is introduced. The cross-attention mechanism is used to enable the semantic intention encoder to perform cross-learning between the source domain and the target domain, so that the semantic intention encoder becomes a cross-domain encoder.

[0058] S300. Obtain the candidate items and the item numbers of the candidate items, and generate a text semantic representation matrix of the candidate items. Based on the text semantic representation matrix of the candidate items, the item numbers, the user's semantic intention and behavior intention, screen out the recommended items to be recommended to the user from the candidate items.

[0059] The above recommendation method of the present invention can be used to recommend movies to users on the mobile phone side. The specific process is as follows:

[0060] First, train the recommendation model. The recommendation model is Figure 2 The text encoder located on the text encoding layer, the item encoder located on the item encoding layer, the cross-domain semantic intention encoder, the behavior intention encoder in the target domain, two mapping layers, the text encoding layer and the item encoder layer on the candidate item side, where the text encoder located on the text encoding layer, the item encoder located on the item encoding layer, the cross-domain semantic intention encoder, and the behavior intention encoder in the target domain are the encoders to be trained.

[0061] When training the recommendation model, as Figure 3 shown, it is necessary to introduce a text auxiliary encoder (located on the text encoding layer), an item auxiliary encoder (located on the item encoding layer), and a neighborhood-sharing cross-attention mechanism for processing the source domain interaction sequence. That is, when training the text encoder, item encoder, semantic intention encoder, and behavior intention encoder in the target domain, it is necessary to jointly train the text auxiliary encoder, item auxiliary encoder, and cross-attention mechanism.

[0062] During training, arrange the relevant information of the movies (movies are items) watched by the user in the order of viewing time to form a target domain interaction sequence (this interaction sequence is the interaction sample sequence in the target domain); arrange the books related to movies read by the user in the order of reading time to form a source domain interaction sequence (this interaction sequence is the interaction sample sequence in the source domain, Figure 3 in , , , For each interaction information in the interaction sequence of the source domain). As Figure 3 shown, the interaction sequence of the target domain is respectively input into the text encoder and the item encoder, and the interaction sequence of the source domain is also respectively input into the text encoder and the item encoder. After Figure 3 processing, the final user preference, semantic intention and behavioral intention are obtained. Then, according to the final user preference, candidate items that the user may prefer are predicted. Finally, according to whether the candidate items that may be preferred are the items that the user actually likes, semantic intention and behavioral intention, the final loss function is obtained, and the text encoder, item encoder, semantic intention encoder, behavioral intention encoder, training text auxiliary encoder, item auxiliary encoder and cross-attention mechanism are adjusted by backpropagation according to the final loss function until the training is completed.

[0063] After the training is completed, when a movie needs to be recommended to a user, only the relevant information of the movies that the user has watched needs to be formed into an interaction sequence of the target domain, and then the interaction sequence of the target domain is input into Figure 2 the recommendation model, and the movies recommended to the user can be obtained. Figure 2 in , , , represent each interaction information in the interaction sequence of the target domain.

[0064] Embodiment 1 provides a training method for jointly training a semantic intention encoder, a text encoder, an item encoder, and a behavioral intention encoder, including the following specific steps S01 to S015:

[0065] S01, set a text auxiliary encoder and an item auxiliary encoder, and the text auxiliary encoder and the item auxiliary encoder are two encoders set for assisting in training the semantic intention encoder, the text encoder, the item encoder, and the behavioral intention encoder.

[0066] The text auxiliary encoder is the encoder on the text encoding layer located Figure 3 above, and the item auxiliary encoder is the encoder on the item encoding layer located Figure 3 above. Both the text auxiliary encoder and the text encoder use the pre-trained language model BERT, and the semantic intention encoder uses the causal Transformer model (i.e., a deep learning model based on the attention mechanism).

[0067] S02, obtain the interaction sample sequence of the same user in the source domain and the interaction sample sequence in the target domain.

[0068] During training, due to the insufficient interaction sequences of the user in the target domain, the interaction sequences of the user in the source domain are jointly used as the training dataset to make up for the deficiency of the interaction sequences in the target domain and provide sufficient datasets for training.

[0069] S03, apply the text auxiliary encoder and the item auxiliary encoder to the interaction sample sequences in the source domain respectively to obtain the source sample text semantic representation matrix and the source sample item encoding matrix .

[0070]

[0071]

[0072] ;

[0073] ;

[0074] ;

[0075] In the formula, represents the source domain, 1, 2, n are the serial numbers of the items, represents the original semantics corresponding to the target sample text, and the information of the text includes comment information and meta information, where the meta information is used to represent the category of the item and the information describing the item, and the comment information is used to represent the user's evaluation of the item, represents the position of the item (this position is used to characterize the interaction time of the item), is the sample text in the source domain (the sample text included in the interaction sample sequence of this sample source domain) records the meta information text corresponding to the item , represents the meta information, is the sample text in the source domain (the sample text included in the interaction sample sequence of this sample source domain) records the comment information text corresponding to the item , represents the user 's comment semantic representation vector for the item (the item is the item with the interaction serial number ), represents the dimension of the output vector, is the set composed of all users , represents the meta information semantic representation vector of the item , represents the item ID number, represents a pre-trained language model, projection refers to the linear mapping layer, which is mainly implemented by a simple fully-connected layer. Its purpose is to perform a linear transformation on the semantic representation matrix. represents the normalization process, represents the averaging process.

[0076] S04, apply the text encoder and the item encoder to the interaction sample sequence of the target domain respectively to obtain the target sample text semantic representation matrix and the target sample item encoding matrix ;

[0077] The calculation method of is the same as that of , as long as the sample text of the source domain is replaced with the sample text of the target domain y. Similarly, The calculation method of is the same as that of

[0078] S05, as Figure 3 shown, by using the target sample text semantic representation matrix as the query vector Q of the cross-attention mechanism, using the source sample text semantic representation matrix as the key vector K of the cross-attention mechanism, and using the source sample item encoding matrix as the value vector V of the cross-attention mechanism, the inter-domain preference as a positive sample is obtained.

[0079] Since there are high-level semantic connections of the same user in different domains, these information reflect the long-term stable interest tendency of the user and can assist user representation modeling during cross-domain recommendation. Therefore, a domain-shared cross-attention mechanism can be designed. Its core idea is to calculate the semantic connection between the text representations of the source domain and the target domain in the attention map and construct a cross-domain preference representation by combining the user's behavior preferences in the source domain.

[0080] ;

[0081] is the preference vector output by the cross-attention mechanism. Taking the th (the th, that is, the last one) preference through a perception mapping based on a fully-connected layer, the inter-domain preference can be obtained.

[0082] ;

[0083] In the formula, , represents an encoder, and the encoder in the formula is used to encode the sequence, (i.e., Feed - forward Networ) is a feed - forward neural network, and T is the transpose of the matrix, represents normalization processing, represents the latent dimension of the vector, represents a mapping based on a fully - connected layer, is a learnable weight matrix, and the domain - shared attention matrix learns the shared semantic association between the source domain and the target domain.

[0084] S06, use the target sample text semantic representation matrix as the query vector and the key vector of the attention mechanism inside the semantic intention encoder respectively, to obtain the target domain attention matrix .

[0085] ;

[0086] Among them, is a learnable weight matrix.

[0087] S07, combine the neighborhood - shared attention matrix and the target domain attention matrix to obtain the attention matrix of the shared semantic association:

[0088] ;

[0089] In the formula, is a learnable mapping matrix.

[0090] In this embodiment, S06 and S07 use a causal Transformer as an independent encoder (Transformer is the semantic intention encoder), and introduce a re - weighted attention graph mechanism based on a Semantic Alignment Operator (SAO). By adaptively fusing the semantic attention graph of the target domain and the semantic attention graph of the shared semantic association, the semantic intention learning of the target domain is enhanced, so as to more accurately capture the cross - domain interest preferences of users.

[0091] S08, as Figure 3 shown, by using the target sample text semantic representation matrix As a value vector, according to the attention matrix of shared semantic association , the semantic intention of the target sample is obtained .

[0092] is the vector output by the cross-domain semantic intention encoder The semantic intention of the last n (i.e., the nth) in is defined by the following formula

[0093] ;

[0094] represents the encoder, and the encoder in the formula is a cross-domain semantic intention encoder represents the feed-forward neural network is a learnable weight matrix

[0095] When the target domain is the movie domain, the semantic intention is the user's preference for the movie theme, that is, the semantic intention represents the movie theme preferred by the user

[0096] S09, as Figure 3 shown, through the behavioral intention encoder, according to the target sample item coding matrix, the target sample behavioral intention is obtained .

[0097] is the vector output by the behavioral intention encoder of the target domain The last behavioral intention in is defined by the following formula

[0098] ;

[0099] represents the encoder, which is the behavioral intention encoder here

[0100] In this embodiment, considering that the behavioral intention and semantic intention of the user in the target domain are similar, therefore, in order to promote the alignment of intention representation, this embodiment uses a Transformer that shares weights with the cross-domain intention encoder to model the target domain item ID representation matrix , and obtains the preference vector of the user in the target domain .

[0101] When the target domain is the movie domain, the behavioral intention represents the actor or director or screenwriter, etc. of the movie preferred by the user

[0102] S010, according to the target sample semantic intention and the target sample behavioral intention, the in-domain preference as a positive sample is obtained :

[0103] ;

[0104] In the formula, represents Figure 3 the merging in that is, aggregating the coarse-grained cross-domain semantic intention representation and the fine-grained behavior intention representation to obtain the in-domain preference

[0105] S011. Obtain the inter-domain preference of other users, and use this inter-domain preference as the inter-domain preference of the negative sample .

[0106] Through steps S02, S03, S04, and S05, obtain the in-domain preference between the source domain and the target domain of the same user (this user is the Figure 4 shared user in ), and use the inter-domain preference corresponding to this user as the positive sample. Replace the user in steps S02, S03, S04, and S05 with another user (this other user is other users), and use the same processing method for the interaction sequence of other users in the source domain and the interaction in the target domain as in steps S02, S03, S04, and S05, then the inter-domain preference as the negative sample Figure 4 can be obtained. As shown in the inter-domain preference and the in-domain preference constitute a positive sample pair, while the inter-domain preference

[0107] S012. Based on the inter-domain preference as the positive sample, the in-domain preference as the positive sample, and the inter-domain preference as the negative sample, .

[0108] Although users have different distributed behavior patterns in the source domain and the target domain, their long-term preferences are consistent. Therefore, it is necessary to design effective supervision signals to align the user representations in different domains. This embodiment proposes a semantic-behavior hybrid enhanced contrast learning mechanism. Consider using the in-domain preference and the inter-domain preference of the same user as the positive sample pair, while using the inter-domain preference As negative samples. This contrastive learning maximizes the preference representations of overlapping users across domains and within domains, and minimizes the cross-domain preference representations with other users. Since each vector representation in the positive sample pairs and negative sample pairs aggregates semantic intentions and behavioral intentions, it can fully promote knowledge transfer across domains. The contrastive loss in this embodiment is calculated as follows:

[0109] ;

[0110] In the formula, (·) represents the sigmoid function (the sigmoid function is the binary function), (·) represents the dot product used to calculate the similarity of the representation vectors. Among them, represents the -th time step in the user-item interaction sequence, represents the total number of time steps.

[0111] S013 aggregates the cross-domain preferences and intra-domain preferences of the same user to obtain the final sample preference of the same user . In this embodiment, the cross-domain preference of the same user uses the cross-domain preference as the positive sample .

[0112] ;

[0113] In the training stage, in order to promote cross-domain knowledge transfer, this embodiment uses aggregated preferences to calculate the loss, enabling the model to fully utilize the comprehensive information of the source domain and the target domain and optimize the alignment of cross-domain user representations.

[0114] S014 obtains the cross-entropy loss according to the text semantic vector and item number of the candidate item and the final sample preference :

[0115] ;

[0116] ;

[0117] ;

[0118] is the text semantic vector of the candidate item in the target domain , and the item number is the ID encoding vector of the candidate item in the target domain . Denotes negative samples sampled from the target domain. (·) denotes the activation function, for the target domain set of candidate items, for the target domain set of user behavior sequences. Denotes any user , denotes the set of all users, Denotes the th time step in the user-item interaction sequence (a total of time steps).

[0119] S015, based on the contrastive loss and the cross-entropy loss , constructs the total loss , and performs iterative training based on the total loss.

[0120] ;

[0121] where is a hyperparameter that controls the semantic-behavior hybrid enhanced contrastive loss.

[0122] Example 2, in this example, each encoder in Example 1 is defined as a Transformer-based sequence encoder , and the encoding process of the sequence encoder is as follows:

[0123] Given an input representation matrix , first convert it through three linear projection layers into the query , key and value in the attention mechanism. Among them, the attention map can be defined as:

[0124]

[0125] where is a scaling factor, is the potential vector dimension, is the learnable weight matrix. Finally, the output of the attention can be defined as:

[0126] ;

[0127] where is the output potential representation, is a learnable weight matrix. The latent representation output by the attention is passed through a pointwise feed-forward network (FNN), and its formula is as follows:

[0128]

[0129] where, is a learnable weight matrix, is a learnable bias vector, is the final output of the encoder.

[0130] In Embodiment 3, after the training is completed, based on Figure 2 a recommended item can be selected from the candidate items and recommended to the user, including the following specific steps:

[0131] S301, generating an item vector of the candidate item according to the text semantic representation matrix and item number of the candidate item;

[0132] S302, obtaining a preference vector of the user according to the semantic intention and the behavioral intention, where the preference vector is used to characterize the preference degree of the user for each interactive item;

[0133] S303, screening out a recommended item to be recommended to the user from the candidate items according to the preference vector of the user and the item vector of each candidate item.

[0134] The effectiveness of the recommendation method of the present invention is demonstrated by the following experiments:

[0135] The present invention uses three data sets from a shopping website, namely Movie (Movie means movie), CD (CD means record), and Book (Book means book). The above data sets contain the interaction behaviors of overlapping users in different fields and are suitable for cross-domain sequential recommendation research.

[0136] The present invention performs the following data preprocessing operations on the above three data sets: i) only retain users and items with at least five interactions; ii) delete duplicate users and delete duplicate items; iii) sort the interactions of each user in chronological order; iv) only retain users who have interactions in all three fields.

[0137] To evaluate the recommendation performance of the present invention, the present invention adopts two common metrics, including the hit rate metric (the hit rate metric is HR@K) and the normalized discounted cumulative gain metric (the normalized discounted cumulative gain metric is NDCG@K), where K {5, 10}, and K is the number of items recommended to the user. The recommendation algorithm generally selects the K items that the user may like the most from the entire candidate item set and recommends them to the user.

[0138] The present invention uses the leave-one-out method for evaluation, that is, the last item in each sequence is used as the test set, the penultimate item is used as the validation set, and the previous items are used as the training set. To reduce the computational complexity, the present invention adopts a sampling-based evaluation strategy, that is, 100 negative samples and 1 true positive sample are randomly selected from the entire set of candidate items according to the popularity of the items for evaluation.

[0139] To verify the effectiveness of the recommendation method proposed by the present invention, the present invention compares four types of baseline models: the sequence recommendation model based on item ID (the item ID is the item number), the sequence recommendation model enhanced by additional information, the sequence recommendation model based on item ID and text enhancement, and the cross-domain sequence recommendation model.

[0140] Among them, the sequence recommendation model based on item ID includes the following recommendation methods:

[0141] GRU4Rec: A sequence recommendation method based on recurrent neural network, using gated neural network to model user preferences.

[0142] SASRec: A sequence recommendation method based on causal Transformer, using Transformer to capture users' dynamic preferences.

[0143] The sequence recommendation model enhanced by additional information includes the following recommendation methods:

[0144] DIF: A sequence recommendation method based on decoupled attention matrix for enhanced fusion of additional information (category information).

[0145] MSSR: An additional information enhanced sequence recommendation model that adaptively aggregates through multi-sequence attention.

[0146] The sequence recommendation model based on item ID and text enhancement includes the following recommendation methods:

[0147] DIF++: An extended version of DIF that fuses and enhances item review information and meta information as additional information.

[0148] SASRec++: An extended version of SASRec that uses two identical attention modules to model the item ID sequence and the text sequence.

[0149] CCA: A sequence recommendation model that uses cascaded cross-attention to aggregate review information.

[0150] The cross-domain sequence recommendation model includes the following recommendation methods:

[0151] CD-SASRec: A cross-domain sequential recommendation model based on Transformer, which conducts cross-domain knowledge fusion in the attention mechanism.

[0152] MGCL: A cross-domain sequential method based on multi-view graph neural network, and adopts multi-view contrastive learning for cross-domain knowledge transfer.

[0153] UniSRec: A cross-domain sequential recommendation method based on text semantic contrast, which learns general semantic representations from text information in multiple domains and fine-tunes in combination with item ID representations in downstream domains for cross-domain recommendation.

[0154] VQRec: A semantic representation transfer recommendation method based on vector quantization. This method uses a cross-domain fine-tuning strategy based on differentiable permutation network in downstream domains to enhance the self-adaptability of semantic transfer to the recommendation space.

[0155] The experimental results are shown in Table 1. Among them, the best experimental results are marked in bold black (the experimental results marked in bold black are the data in the CT2CSR row in Table 1); the sub-optimal experimental results are marked with an underline (the experimental results marked with an underline are the data in the UniSRec(T+ID) row in Table 1). Among them, N@K represents NDCG@K, H@K represents Hit@K, N@K includes N@5 and N@10 in Table 1, N@5 is the normalized discounted cumulative gain with parameter 5, and N@10 is the normalized discounted cumulative gain with parameter 10; H@5 is the hit rate index with parameter 5, and H@10 is the hit rate index with parameter 10.

[0156] In Table 1, the present invention has the following conclusions: (i) In single-row sequence recommendation, SASRec is superior to GRU4Rec, indicating that the attention-based method is superior to the recurrent neural network in capturing users' dynamic preferences. Compared with SASRec, DIF and MSSR perform better, indicating that the method enhanced by additional information effectively enhances user representation. (ii) Both SASRec++ and CCA are superior to SASRec, indicating that text information, as an effective auxiliary signal, can enhance the learning of user preferences. The performance of DIF++ is better than that of SASRec but worse than that of SASRec++, which indicates that the fusion method of this method may not be applicable to text semantic representation. (iii) The text-based cross-domain recommendation methods (e.g., VQRec(T), where T is text, UniSRec(T)) perform better than the item-ID-based cross-domain recommendation methods (e.g., MGCL, CD-SASRec), indicating that using text information as a cross-domain bridge can effectively enhance the effect of cross-domain recommendation. The performance of UniSRec(T+ID) is second only to the CT2CSR model proposed by the present invention, where T+ID is text plus number. (iv) The present invention (the present invention is CT2CSR in Table 1, and CT2CSR is cross-domain sequence recommendation based on contrastive text-enhanced attention mechanism) outperforms all baseline methods on three datasets. Compared with the best-performing baseline method, improvements of 5.51%, 5.86%, 6.11%, and 6.68% are achieved in N@5, H@5, N@10, and H@10, respectively.

[0157] To verify the effectiveness of each component in the CT2CSR model proposed by the present invention, ablation experiments were conducted on three datasets. Each ablated component is as follows:

[0158] (1) w / o TCA, remove the domain-shared cross-attention mechanism and also remove the inter-domain user preferences in the prediction layer.

[0159] (2) w / o SPO: Remove the semantic alignment operator in the cross-domain semantic intent encoder and only consider the attention matrix map of the target domain text representation.

[0160] (3) w / o BIE: Remove the behavior intent encoder in the target domain and also remove the ID representation of the candidate item in the prediction layer.

[0161] (4) w / o SIE: Remove the cross-domain semantic intent encoder and also remove the semantic representation of the candidate item in the prediction layer.

[0162] (5) w / o SBCL: Remove the contrastive learning task of the semantic-behavior dual view.

[0163] The present invention reports the results of ablation experiments in Table 2. The present invention has the following conclusions: (1) Removing any one component in CT2CSR leads to a decline in model performance, demonstrating the effectiveness of each module in the model; (2) Removing the cross-domain semantic intent encoder brings the greatest degree of performance decline, indicating that the cross-domain semantic intent based on text sequences effectively improves user representation learning; (3) Removing the behavior intent encoding in the target domain brings the second greatest degree of performance decline, indicating that the user's fine-grained interaction behavior intent is the key to capturing user preferences; (4) Removing the domain-shared cross-attention mechanism reduces the model performance, indicating that the cross-domain semantic connection has a positive impact on cross-domain representation learning; (5) Removing the semantic alignment operator and the semantic-behavior hybrid contrast learning task both weaken the recommendation performance, indicating that these two components are irreplaceable.

[0164] Table 1

[0165]

[0166] Table 2

[0167]

[0168] To further verify the robustness of the model, the present invention specifically studies three key parameters: the weight of the semantic-behavior hybrid enhanced contrast learning loss, the source domain sequence length, and the target domain sequence length. First, the hyperparameter of the contrast loss is selected from {0, 0.25, 0.5, 0.75, 1} , and the corresponding results are shown in Figure 5 ( Figure 5 the abscissa of which is the hyperparameter ). Figure 5 NDCG@10 in Figure 5 is N@10, N@10 is the normalized discounted cumulative gain with parameter 10, HR@10 is H@10, and H@10 is the hit rate metric with parameter 10. Figure 5 In the left graph, the middle graph, and the right graph in , the upper broken line represents NDCG@10 and the lower broken line represents H@10. Observing

[0169] ​In summary, the present invention first designs a cross-attention network based on domain sharing, effectively captures the text semantic relationships shared between domains by calculating the semantic attention matrix graphs of the source domain and the target domain, and designs a novel cross-domain intent encoder to adaptively learn the semantic intents between different domains. Secondly, the present invention models user preferences from the dual perspectives of in-domain preference learning and cross-domain preference learning. In cross-domain preference learning, CT2CSR (CT2CSR is the recommendation method of the present invention) uses the semantic relationships shared between domains as a bridge, combines the behavioral preferences of users in the source domain, and constructs a cross-domain preference representation. In in-domain preference learning, CT2CSR decouples the fine-grained user intent preferences and the coarse-grained text semantic preferences, and aggregates these preferences into in-domain preferences. Finally, CT2CSR aggregates the in-domain preferences and cross-domain preferences as the final user preference representation. Finally, the present method proposes a contrast learning mechanism based on semantic-behavior hybrid enhancement, which promotes cross-domain knowledge transfer by aligning the cross-domain preference representation and in-domain preference representation of the same user.

[0170] This embodiment also provides a text-based recommendation device, as Figure 6 shown, the device includes the following components:

[0171] A semantic intent prediction module 01, configured to obtain a target domain interaction sequence of a user including an interactive item, apply a text encoder to the target domain interaction sequence to obtain a target text semantic representation matrix; apply a semantic intent encoder to the target text semantic representation matrix to obtain the semantic intent of the user;

[0172] A behavioral intent prediction module 02, configured to apply an item encoder to the target domain interaction sequence to obtain an item encoding matrix; apply a behavioral intent encoder to the item encoding matrix to obtain the behavioral intent of the user; the semantic intent encoder, the text encoder, the item encoder, and the behavioral intent encoder are all pre-trained encoders, and a cross-attention mechanism for auxiliary training is introduced during the training process, and the cross-attention mechanism is used to enable the semantic intent encoder to perform cross-learning in the source domain and the target domain, so that the semantic intent encoder becomes a cross-domain encoder;

[0173] A recommendation module 03, configured to obtain candidate items and the item numbers of the candidate items, generate a text semantic representation matrix of the candidate items, and screen out the recommended items to be recommended to the user from the candidate items according to the text semantic representation matrix and item numbers of the candidate items and the semantic intent and behavioral intent of the user.

[0174] Based on the above embodiment, the present invention also provides a terminal device, and its principle block diagram can be as Figure 7As shown in the figure. The terminal device includes a processor, a memory, a network interface, and a display screen connected through a system bus. Among them, the processor of the terminal device is used to provide computing and control capabilities. The memory of the terminal device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface of the terminal device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a text-based recommendation method. The display screen of the terminal device can be a liquid crystal display screen or an electronic ink display screen.

[0175] Those skilled in the art can understand that Figure 7 the block diagram of the principle shown in the figure is only the block diagram of the partial structure related to the solution of the present invention, and does not constitute a limitation on the terminal device to which the solution of the present invention is applied. The specific terminal device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0176] In one embodiment, a terminal device is provided. The terminal device includes a memory, a processor, and a text-based recommendation program stored in the memory and executable on the processor. When the processor executes the text-based recommendation program, the following operation instructions are implemented:

[0177] Obtain the target domain interaction sequence of the user containing interaction items, apply a text encoder to the target domain interaction sequence to obtain text semantics; apply a semantic intention encoder to the text semantics to obtain the semantic intention of the user;

[0178] Apply an item encoder to the target domain interaction sequence to obtain an item coding matrix; apply a behavior intention encoder to the item coding matrix to obtain the behavior intention of the user; the semantic intention encoder, the text encoder, the item encoder, and the behavior intention encoder are all trained encoders. During the training process, a cross-attention mechanism for assisting training is introduced. The cross-attention mechanism is used to enable the semantic intention encoder to perform cross-learning between the source domain and the target domain, so that the semantic intention encoder becomes a cross-domain encoder;

[0179] Obtain candidate items and the item numbers of the candidate items, generate the text semantics of the candidate items, and screen out the recommended items to be recommended to the user from the candidate items according to the text semantics and item numbers of the candidate items and the semantic intention and behavior intention of the user.

[0180] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided by the present invention can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0181] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or equivalently replace some of the technical features. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A text-based recommendation method, characterized in that, Including: Obtain the target domain interaction sequence of a user including interactive items, apply a text encoder to the target domain interaction sequence to obtain a target text semantic representation matrix; Apply a semantic intention encoder to the target text semantic representation matrix to obtain the semantic intention of the user; Apply an item encoder to the target domain interaction sequence to obtain an item encoding matrix, and apply a behavior intention encoder to the item encoding matrix to obtain the behavior intention of the user. Among them, the semantic intention encoder, the text encoder, the item encoder, and the behavior intention encoder are all pre-trained encoders. During the training process, a cross-attention mechanism for auxiliary training is introduced. The cross-attention mechanism is used to enable the semantic intention encoder to perform cross-learning in the source domain and the target domain, so that the semantic intention encoder becomes a cross-domain encoder; Obtain candidate items and the item numbers of the candidate items, and generate a text semantic representation matrix of the candidate items. Based on the text semantic representation matrix of the candidate items, the item numbers, the semantic intention and the behavior intention of the user, screen out the recommended items to be recommended to the user from the candidate items; The semantic intention encoder, the text encoder, the item encoder, and the behavior intention encoder adopt a joint training method, where the joint training method includes: Set a text auxiliary encoder and an item auxiliary encoder. The text auxiliary encoder and the item auxiliary encoder are two encoders set for assisting in training the semantic intention encoder, the text encoder, the item encoder, and the behavior intention encoder; Obtain the interaction sample sequence of the same user in the source domain and the interaction sample sequence in the target domain; Apply the text auxiliary encoder and the item auxiliary encoder to the interaction sample sequence in the source domain respectively to obtain a source sample text semantic representation matrix and a source sample item encoding matrix; Apply the text encoder and the item encoder to the interaction sample sequence in the target domain respectively to obtain a target sample text semantic representation matrix and a target sample item encoding matrix; Through the cross-attention mechanism, obtain the cross-domain preference as a positive sample based on the target sample text semantic representation matrix, the source sample text semantic representation matrix, and the source sample item encoding matrix; Through the semantic intention encoder, obtain the target sample semantic intention based on the target sample text semantic representation matrix and the neighborhood sharing attention matrix composed of the source sample text semantic representation matrix and the target sample text semantic representation matrix; Through the behavior intention encoder, obtain the target sample behavior intention based on the target sample item encoding matrix; Based on the target sample semantic intention and the target sample behavior intention, obtain the in-domain preference as a positive sample; Obtain the cross-domain preference of other users and use this cross-domain preference as the cross-domain preference of the negative sample; Based on the cross-domain preference as a positive sample, the in-domain preference as a positive sample, and the cross-domain preference as a negative sample, obtain a contrastive loss; Based on the contrastive loss, construct a total loss and perform iterative training based on the total loss.

2. The text-based recommendation method according to claim 1, wherein Filter out the recommended items to be recommended to the user from the candidate items according to the text semantic representation matrix and item numbers of the candidate items, as well as the semantic intention and the behavioral intention of the user, including: Generate the item vectors of the candidate items according to the text semantic representation matrix and item numbers of the candidate items; Obtain the preference vector of the user according to the semantic intention and the behavioral intention, where the preference vector is used to characterize the preference degree of the user for each of the interactive items; Filter out the recommended items to be recommended to the user from the candidate items according to the preference vector of the user and the item vectors of each of the candidate items.

3. The text-based recommendation method according to claim 1, wherein Construct the total loss according to the contrastive loss, including: Aggregate the cross-domain preference and in-domain preference of the same user to obtain the final sample preference of the same user; Obtain the cross-entropy loss according to the text semantic vector and item number of the candidate item and the final sample preference; Construct the total loss according to the contrastive loss and the cross-entropy loss.

4. The text-based recommendation method according to claim 1, wherein Through the cross-attention mechanism, obtain the cross-domain preference as the positive sample according to the target sample text semantic representation matrix, the source sample text semantic representation matrix, and the source sample item encoding matrix, including: Obtain the cross-domain preference as the positive sample by using the target sample text semantic representation matrix as the query vector of the cross-attention mechanism, the source sample text semantic representation matrix as the key vector of the cross-attention mechanism, and the source sample item encoding matrix as the value vector of the cross-attention mechanism.

5. The text-based recommendation method according to claim 1, wherein Through the semantic intention encoder, obtain the target sample semantic intention according to the target sample text semantic representation matrix and the neighborhood shared attention matrix composed of the source sample text semantic representation matrix and the target sample text semantic representation matrix, including: Use the target sample text semantic representation matrix as the query vector and the key vector of the attention mechanism inside the semantic intention encoder respectively to obtain the target domain attention matrix; Merge the neighborhood shared attention matrix and the target domain attention matrix to obtain the attention matrix with shared semantic association; Obtain the target sample semantic intention by using the target sample text semantic representation matrix as the value vector according to the attention matrix with shared semantic association.

6. The text-based recommendation method according to any one of claims 1-5, characterized in that, The semantic intention and the behavioral intention of the user respectively represent the semantic intention and the behavioral intention reflected by the user's last interaction.

7. A text-based recommendation device, characterized in that, The device includes the following components: A semantic intention prediction module, configured to obtain the target domain interaction sequence of the user including interactive items, apply a text encoder to the target domain interaction sequence to obtain a target text semantic representation matrix; Apply a semantic intention encoder to the target text semantic representation matrix to obtain the semantic intention of the user; A behavior intention prediction module, which is used to apply an item encoder to the target domain interaction sequence to obtain an item encoding matrix, and apply a behavior intention encoder to the item encoding matrix to obtain the user's behavior intention. Among them, the semantic intention encoder, the text encoder, the item encoder, and the behavior intention encoder are all trained encoders. During the training process, a cross-attention mechanism for auxiliary training is introduced. The cross-attention mechanism is used to enable the semantic intention encoder to perform cross-learning between the source domain and the target domain, so that the semantic intention encoder becomes a cross-domain encoder; A recommendation module, which is used to obtain candidate items and the item numbers of the candidate items, and generate a text semantic representation matrix of the candidate items. According to the text semantic representation matrix of the candidate items, the item numbers, and the semantic intention and behavior intention of the user, the recommended items to be recommended to the user are screened out from the candidate items; The semantic intention encoder, the text encoder, the item encoder, and the behavior intention encoder adopt a joint training method, where the joint training method includes: Set a text auxiliary encoder and an item auxiliary encoder. The text auxiliary encoder and the item auxiliary encoder are two encoders set for auxiliary training of the semantic intention encoder, the text encoder, the item encoder, and the behavior intention encoder; Obtain the interaction sample sequence of the same user in the source domain and the interaction sample sequence in the target domain; Apply the text auxiliary encoder and the item auxiliary encoder to the interaction sample sequence in the source domain respectively to obtain a source sample text semantic representation matrix and a source sample item encoding matrix; Apply the text encoder and the item encoder to the interaction sample sequence in the target domain respectively to obtain a target sample text semantic representation matrix and a target sample item encoding matrix; Through the cross-attention mechanism, according to the target sample text semantic representation matrix, the source sample text semantic representation matrix, and the source sample item encoding matrix, obtain the cross-domain preference as a positive sample; Through the semantic intention encoder, according to the target sample text semantic representation matrix and the neighborhood sharing attention matrix composed of the source sample text semantic representation matrix and the target sample text semantic representation matrix, obtain the target sample semantic intention; Through the behavior intention encoder, according to the target sample item encoding matrix, obtain the target sample behavior intention; According to the target sample semantic intention and the target sample behavior intention, obtain the in-domain preference as a positive sample; Obtain the cross-domain preference of other users and use this cross-domain preference as the cross-domain preference of the negative sample; According to the cross-domain preference as a positive sample, the in-domain preference as a positive sample, and the cross-domain preference as a negative sample, obtain the contrast loss; According to the contrast loss, construct a total loss and perform iterative training according to the total loss.

8. A terminal device, characterized in that, The terminal device includes a memory, a processor, and a text-based recommendation program stored in the memory and executable on the processor. When the processor executes the text-based recommendation program, the steps of the text-based recommendation method according to any one of claims 1-6 are implemented.

9. A computer-readable storage medium, characterized in that, A text-based recommendation program is stored on the computer-readable storage medium. When the text-based recommendation program is executed by a processor, the steps of the text-based recommendation method according to any one of claims 1-6 are implemented.

Citation Information

Patent Citations

  • Proxy-aware cross-domain sequence recommendation method and device, medium and product

    CN116881548A

  • API sequence recommendation method, storage medium and electronic equipment

    CN118861432A