Book recommendation method based on cross-domain learning

By adopting cross-domain learning methods in the recommendation system, high-order representations of books-category and user-books are solved, and the recommendation accuracy and adaptability of the recommendation system are improved.

CN120011633AActive Publication Date: 2025-05-16ZHENGZHOU UNIVERSITY OF LIGHT INDUSTRY
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510091536.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-16
Estimated Expiration
2045-01-21

AI Technical Summary

Technical Problem

When existing recommendation systems face data sparsity and cold start problems, it is difficult to effectively recommend new books or content that new users are interested in, resulting in a decrease in recommendation accuracy.

Method used

A book recommendation method based on cross-domain learning is adopted, by obtaining user and book information from the source and target book domains, building inter-domain book-in-domain relationship diagrams and intra-domain books-user interaction diagrams, supervised training is carried out to obtain advanced representations, and achieving dual-domain unified representations of users and books.

Benefits of technology

It improves the recommendation accuracy of sparse domains, enhances adaptability to new users and new books, and can more accurately recommend books that meet user interests.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011633A_ABST
    Figure CN120011633A_ABST
Patent Text Reader

Abstract

The invention provides a book recommendation method based on cross-domain learning, and the method comprises the steps: obtaining user and book information from a source book domain and at least one target book domain, and dividing the user and book information into intra-domain book-user interaction data and cross-domain book-category data; constructing an inter-domain book-category relation graph based on the cross-domain book-category data; constructing an intra-domain book-user interaction graph based on the intra-domain book-user interaction data; supervised training is carried out on the inter-domain book-category relation graph, and inter-domain book-book category high-order representation is obtained; performing supervision training on the intra-domain book-user interaction graph of each book domain in combination with inter-domain book-book category high-order characterization, and obtaining double-domain unified user high-order characterization and unified book high-order characterization; and predicting scores between the users in the source book domain and the non-interactive books in the target book domain, and recommending potential interested books to the users in the source book domain in sequence. According to the method, the recommendation accuracy of the high sparse domain and the adaptability to new data are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of deep learning cross-domain recommendation, and in particular to a book recommendation method. Background Art

[0002] With the widespread use of the internet and the development of information technology, people are faced with a vast amount of information and content, leading to information overload. In this situation, traditional advertising push or generalized information push cannot meet the interests of individual users. Users need personalized recommendation services to help them filter content that suits their interests from a vast library of books. Recommender systems are technologies that use algorithms to recommend books that may be of interest to users. They help users quickly find books that match their preferences within a vast library, providing personalized recommendations. Recommender systems typically operate based on one or a combination of two approaches: content-based recommendations and collaborative filtering. These systems record user preferences based on their browsing history within a single data domain. The systems then filter the data to find books that are closest to or similar to the user's preferences and recommend them to the user. This significantly improves users' ability to find the books they need online. Recommender systems enable platforms to better understand user needs, provide more targeted services, and thus increase commercial value. However, in most real-world applications, few users are able to provide ratings or reviews for many books (i.e., data sparsity). This reduces the recommendation accuracy of CF-based models. Almost all existing CF-based recommendation systems suffer to some extent from the long-standing data sparsity problem, especially for new books or new users (the cold start problem). This problem can lead to overfitting when training CF-based models, significantly reducing recommendation accuracy. Therefore, to address data sparsity and cold start issues in the data domain, cross-domain recommendation (CDR) has become a new research topic. Summary of the Invention

[0003] Aiming at the data sparsity and cold start problems existing in existing recommendation systems, this paper proposes a book recommendation method based on cross-domain learning to improve the recommendation accuracy in sparse domains and the adaptability to new data.

[0004] In order to achieve the above object, the technical solution of the present invention is achieved as follows:

[0005] A book recommendation method based on cross-domain learning, comprising the following steps:

[0006] Step 1: Obtain user and book information from a source book domain and at least one target book domain, and divide the obtained user and book information into intra-domain book-user interaction data and cross-domain book-category data;

[0007] Step 2: Construct an inter-domain book-category relationship graph based on cross-domain book-category data; construct an intra-domain book-user interaction graph based on intra-domain book-user interaction data;

[0008] Step 3: Using cross-domain book-category data as supervision data, supervised training is performed on the inter-domain book-category relationship graph to obtain high-level representations of inter-domain books and book categories;

[0009] Step 4: Using the intra-domain book-user interaction data as supervision data and combining it with the inter-domain book-book category high-level representation, supervised training is performed on the intra-domain book-user interaction graph of each book domain to obtain unified user high-level representations and unified book high-level representations for both domains.

[0010] Step 5: Repeat steps 3 to 4 until convergence or the set number of training rounds is reached to obtain the final unified user high-level representation and unified book high-level representation;

[0011] Step 6: Using the final unified user high-level representation and unified book high-level representation, predict the recommendation ratings between the source book domain users and the target book domain books that have not been interacted with, and recommend the target book domain books of potential interest to the source book domain users in a ranked manner.

[0012] Furthermore, the domain book-user interaction data includes the domain book-user interaction data of the source book domain A. and the book-user interaction data within the target book domain B in, and m represents the book ID, u represents the user ID, r mu represents the rating of book m by user u;

[0013] The cross-domain book-category data is in, and They represent the book title, book introduction and book category name respectively, and c represents the book category ID.

[0014] Furthermore, the method for performing supervised training on the inter-domain book-category relationship graph is as follows: obtaining original feature vectors of books and original feature vectors of book categories through a large language model based on cross-domain book-category data;

[0015] Map and concatenate the original feature vectors of the book and the original feature vectors of the book category to obtain the zero-order representation matrix of the book-book category;

[0016] Perform feature aggregation on the inter-domain book-category relationship graph to obtain the inter-domain book-book category high-level representation;

[0017] The inter-domain book-category relationship loss function is calculated based on the inter-domain book-book category high-order representation, and the network parameters are updated using the back gradient propagation algorithm according to the loss function calculation result.

[0018] Furthermore, the method for performing supervised training on the book-user interaction graph within each book domain is as follows:

[0019] Based on the book-user interaction data in the domain The user ID and book ID are mapped, embedded, and concatenated respectively to obtain the zero-order representation matrix of book ID-user ID. Feature aggregation is performed on the book-user interaction graph within the source book domain A to obtain the high-order representation H of the book ID-user ID within the source book domain A. A ;

[0020] Based on the source book domain A, the domain book ID-user ID high-level representation H A The high-order prediction representation of the target book domain user ID obtained by the overlapping user interest mapper of the source book domain;

[0021] Based on the source book domain A, the domain book ID-user ID high-level representation H A , Inter-domain book-book category high-level representation H C and the target book domain user ID high-order prediction representation calculation domain book-user interaction loss function Update network parameters using the reverse gradient propagation algorithm based on the loss function calculation results;

[0022] Based on the book-user interaction data in the domain Calculate the high-level representation H of the book ID-user ID in the target book domain B B and

[0023] Calculate the cross-domain loss of user interest representation based on the high-order prediction representation of the user ID in the target book domain Backward gradient propagation is used to update the parameters of the overlapping user interest mapper for the source book domain.

[0024] Furthermore, the calculation method of the original feature vector of the book is:

[0025]

[0026] in, is the original feature vector of the book, and is the original feature vector matrix of all books, LLM(·) represents the large language model feature extraction function;

[0027] The calculation method of the original feature vector of the book category is:

[0028]

[0029] in, is the original feature vector of the category name of the book, is the original eigenvector matrix of all book categories.

[0030] Furthermore, the calculation formula of the zero-order representation matrix of book-book category is:

[0031]

[0032] in, Both represent the original eigenvector matrix mapping parameters, represents the zero-order representation matrix of the book, The zero-order representation matrix of book categories, is the zero-order representation matrix of book-book category, stack(·) represents the concatenation of matrices with columns unchanged;

[0033] The calculation formula for the inter-domain book-book category high-order representation is:

[0034]

[0035] Among them, H C represents the final aggregated high-level representation of inter-domain books and book categories, L represents the number of graph aggregation layers, and A C is the adjacency matrix of the domain book-category relationship graph, A C The adjacency matrix A C The calculated graph neural network normalized adjacency matrix, Δ (l) is the noise matrix, and the i-th row vector Δ of each noise matrix (l) i is random noise, random noise Δ (l) i Satisfy the two norm

[0036] ||Δ (l) i ||2=∈ and is a parameter that obeys a uniform distribution on the interval [0,1], ⊙ represents the vector outer product, sign(·) is the sign function, Represents the i-th order representation of the inter-domain book-book category.

[0037] Furthermore, the method for calculating the inter-domain book-category relationship loss function based on the inter-domain book-book category high-order representation is:

[0038]

[0039]

[0040] in, Represents the total loss function of the book-category relationship graph between domains, V C represents the node set in the inter-domain book-category relationship graph, represents the book-book category edge loss function, represents the graph contrastive learning loss function, and λ represents the weight of the graph contrastive learning loss function;

[0041] ln(·) is the logarithmic function, σ(·) is the sigmoid activation function, exp(·) is the exponential function, Represents the high-level representation H from the domain book-book category C The high-level representation of books obtained by indexing, Represents the high-level representation H from the domain book-book category C The high-order representation of the book category obtained by the index, τ is the contrastive learning temperature coefficient;

[0042] (m,c)∈E C Indicates that in the inter-domain book-category relationship graph, there is a book m belonging to category c, E C The edge set representing the book-category relationship graph between domains; in Indicates that book m does not belong to category c′ in the inter-domain book-category relationship graph.

[0043] Furthermore, the calculation process for obtaining the zero-order representation matrix of book ID-user ID is as follows:

[0044]

[0045] Among them, Embedding(·) represents the mapping embedding operation, represents the zero-order representation matrix of book ID, represents the zero-order representation matrix of user ID, represents the zero-order representation matrix of book ID-user ID;

[0046] The calculation formula of the high-level representation of the domain book ID-user ID of the source book domain A is:

[0047]

[0048] Among them, H A The final aggregated high-level representation of the source book domain A, the book ID-user ID, A A Represents the book-user interaction graph within the source book domain A, A A The adjacency matrix A AThe calculated normalized adjacency matrix of the graph neural network.

[0049] Furthermore, the calculation method of the target book domain user ID high-order prediction representation is:

[0050]

[0051] in, Represents the high-level representation H of the book ID-user ID in the source book domain A A The high-level representation of the user ID obtained by the index, projector A→B (·) represents the overlapping user interest mapping function of the source book domain A, Represents the high-order prediction representation of the user ID of the target book domain B;

[0052] The computational domain book-user interaction loss function The method is:

[0053]

[0054] Among them, h u represents the unified user high-level representation, h m Represents a unified high-level representation of books, V A The node set representing the book-user interaction graph within the source book domain A, Represents the high-level representation H of the book ID-user ID in the source book domain A A The high-level representation of the book ID obtained by the index, Represents the high-level representation H from the domain book-book category C High-level representation of books obtained from the index;

[0055] (u,m)∈E A , representing the book-user interaction graph within the domain There is an interaction pair (r um ≥1), E A represents the edge set of the book-user interaction graph within the source book domain A, in Represents the book-user interaction diagram in the domain There exists an interaction pair (r um′ <1).

[0056] Furthermore, the cross-domain loss of user interest representation is calculated based on the high-order prediction representation of the target book domain user ID The method is to freeze all parameters of the overlapping user interest mapper projector except the source book domain A, and calculate the cross-domain loss of user interest representation

[0057]

[0058] Among them, Wasserstein(·) represents the function of finding Wasserstein distance, is the user set in source book domain A, is the user set in the target book domain B, Represents the high-level representation H of the book ID-user ID in the target book domain B B The high-level representation of the user ID obtained by the index.

[0059] The beneficial effects of the present invention are:

[0060] The present invention aims to solve the data sparsity and cold start problems that exist in traditional recommendation systems. Since not all users will leave interaction data (ratings, comments, etc.) after browsing books, the training prediction data of the recommendation system is insufficient, resulting in the problem of data sparsity. In order to improve the accuracy of recommendations in sparse domains, the interaction data of rich domains is learned and mapped to the sparse domains. In addition, every time a new user or book is added to the data domain, the traditional recommendation system cannot adapt to the new data in a timely manner. Therefore, in order to solve this problem, the present invention proposes a learning method based on inter-domain mapping, which realizes rapid matching of new users or new books by historical interaction data of users and books in the data domain, and then displays the N books that best meet the user's interests as recommended books to the user for selection. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0062] Figure 1 This is a flowchart of the book recommendation method based on cross-domain learning of the present invention. DETAILED DESCRIPTION

[0063] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without creative work are within the scope of protection of the present invention.

[0064] A book recommendation method based on cross-domain learning, such as Figure 1 As shown, the steps include:

[0065] Step 1: Obtain user and book information from a source book domain and at least one target book domain, and divide the obtained user and book information into intra-domain book-user interaction data and cross-domain book-category data.

[0066] In this embodiment, user and book information comes from two book domains, namely book domain A and book domain B. Book domain A is the source book domain and book domain B is the target book domain. Book domain A is Douban Reading and book domain B is Qidian Reading. The data of the two book domains may share and overlap to a certain extent. Users can have different accounts on different platforms, so the data of one data domain can be used to recommend the results of another data domain.

[0067] Furthermore, the Pandas library is used to process and format user and book information to obtain the domain book-user interaction data. and cross-domain book-category data

[0068] Furthermore, the domain book-user interaction data Includes book-user interaction data within source book domain A and the book-user interaction data within the target book domain B And there is Among them, m represents the book ID, u represents the user ID, r mu represents the rating of book m by user u.

[0069] Furthermore, for cross-domain book-category data have and They represent the book title, book introduction and book category name respectively, and c represents the book category ID.

[0070] Examples of datasets are shown in Tables 1 and 2.

[0071] Table 1

[0072]

[0073]

[0074] Table 2

[0075]

[0076] Step 2: Construct an inter-domain book-category relationship graph based on cross-domain book-category data; construct an intra-domain book-user interaction graph based on intra-domain book-user interaction data.

[0077] Further, based on the book-user interaction data in the source book domain A Get the book collection of source book domain A User collection of source book domain A And the book-user relationship set, with books and users as nodes, and book-user relationships as edges, construct the book-user interaction graph in the source book domain A. Based on the book-user interaction graph in the source book domain A Construct the adjacency matrix A A ;in, Represents a node set, E A ={(m,u)} represents the edge set, that is, the book-user relationship set, and the element A in the adjacency matrix A [m][u]=r mu Represents user u's rating of book m as r mu , element A in the adjacency matrix A [m][u]=0 means that user u has not rated book m.

[0078] Further, based on the book-user interaction data in the target book domain B Get the book collection of target book domain B User collection of target book domain B And the book-user relationship set, with books and users as nodes, and book-user relationships as edges, build the book-user interaction graph in the target book domain B. Based on the book-user interaction diagram in the target book domain B Construct the adjacency matrix A B ;in, Represents a node set, E B ={(m,u)} represents the edge set, that is, the book-user relationship set, and the element A in the adjacency matrix B [m][u]=r mu Represents the user's rating of book m as r mu , element A in the adjacency matrix B [m][u]=0 means the user has no rating.

[0079] Further, obtain the book category collection based on the cross-domain book-category data According to the book collection of source book domain A and the book collection of target book domain B Get the book collection of two domains Using the books and book categories of the two domains as nodes and the book-user relationship as edges, we construct a book-category relationship graph between domains. According to the domain book-category relationship diagram Construct the adjacency matrix A C ;in, Represents a node set, E C ={(m,c)} represents the edge set, A C [m][c]=1 means that book m is related to book category c, otherwise A C [m][c]=0 Book m is irrelevant to book category c.

[0080] Step 3: Using cross-domain book-category data as supervision data, supervised training is performed on the inter-domain book-category relationship graph to obtain high-level representations of inter-domain books and book categories.

[0081] The method for supervised training on the inter-domain book-category relationship graph is:

[0082] Furthermore, based on the cross-domain book-category data, the original feature vectors of the books and the original feature vectors of the book categories are obtained through a large language model.

[0083] The calculation method of the original feature vector of the book is:

[0084]

[0085] in, is the original feature vector of the book, and is the original feature vector matrix of all books, LLM(·) represents the large language model feature extraction function, and the model used is bge-base-zh-v1.5.

[0086] The calculation method of the original feature vector of the book category is:

[0087]

[0088] in, is the original feature vector of the category name of the book, is the original eigenvector matrix of all book categories.

[0089] Furthermore, the original eigenvector matrix of the books and the original eigenvector matrix of the book categories are mapped and concatenated to obtain the zero-order representation matrix of the book-book category.

[0090]

[0091] in, Both represent the original eigenvector matrix mapping parameters, represents the zero-order representation matrix of the book, The zero-order representation matrix of book categories, is the zero-order representation matrix of book-book category, d is the embedding space dimension, d=64, |·| represents the number of elements in the set, and stack(·) represents concatenating the matrices without changing the columns. The number of rows in the concatenated matrix is ​​equal to the sum of the number of rows in the original matrices.

[0092] Furthermore, in the domain book-category relationship diagram Perform feature aggregation on the domain to obtain the high-level representation of book-book categories between domains;

[0093]

[0094] Among them, H C It represents the final aggregated domain book-book category high-level representation, L represents the set number of graph aggregation layers, and L = 3 is used in the implementation process. A C The adjacency matrix A C The calculated graph neural network normalized adjacency matrix, set The elements of the noise matrix are the i-th row vector Δ (l) i is random noise, random noise Δ (l) i Satisfies the second norm ||Δ (l) i ||2=∈ and is a parameter that obeys the uniform distribution on the interval [0,1], ⊙ represents the vector outer product, sign(·) is the sign function, which takes 1 when the input scalar is greater than 0, takes 0 when it is equal to 0, and takes -1 when it is less than 0; ∈ represents the upper bound of the noise threshold, Represents the i-th order representation of the inter-domain book-book category.

[0095] Furthermore, based on the inter-domain book-book category high-level representation H C Calculate the inter-domain book-book category relationship loss function, and use the reverse gradient propagation algorithm to update the network parameters according to the loss function calculation results, including the original feature vector matrix mapping parameters and

[0096] The method for calculating the inter-domain book-category relationship loss function is:

[0097]

[0098] in, represents the total loss function of the inter-domain book-category relationship graph, represents the book-book category edge loss function, represents the graph contrastive learning loss function, and λ represents the weight of the graph contrastive learning loss function; Used to learn the edges existing on the book-category relationship graph between domains, Used to further distinguish the book-book category positive-negative sample pairs, obtained by summing the weights λ Unify gradient propagation and model parameter updates.

[0099] ln(·) is the logarithmic function, σ(·) is the sigmoid activation function, exp(·) is the exponential function, represents transpose, Represents the high-level representation H from the domain book-book category C The high-level representation of books obtained by indexing, Represents the high-level representation H from the domain book-book category C The high-order representation of the book category obtained by the index in , τ represents the contrastive learning temperature coefficient.

[0100] (m,c)∈E C , is a positive sample pair, representing the book-category relationship graph between domains There is a book m in category c; in is a negative sample pair, representing the book-category relationship graph between domains The book m does not belong to category c′. The loss is minimized by maximizing the number of positive sample pairs and minimizing the number of negative sample pairs.

[0101] Step 4: Using the intra-domain book-user interaction data as supervision data and combining it with the inter-domain book-book category high-level representation, supervised training is performed on the intra-domain book-user interaction graph of each book domain to obtain unified user high-level representations and unified book high-level representations for both domains.

[0102] The method for supervised training on the in-domain book-user interaction graph of each book domain is:

[0103] In this embodiment, the book domain A is used as the source domain and the book domain B is used as the target domain for explanation; the supervised training can also be performed with the book domain A as the target domain and the book domain B as the source domain, using the same method, such as Figure 1 shown.

[0104] Furthermore, based on the domain book-user interaction data The user ID and book ID are mapped, embedded, and concatenated to obtain the zero-order representation matrix of book ID-user ID. The calculation process is:

[0105]

[0106] Among them, Embedding(·) represents the mapping embedding operation, represents the zero-order representation matrix of book ID, represents the zero-order representation matrix of user ID, Represents the zero-order representation matrix of book ID-user ID.

[0107] Furthermore, in the domain book-user interaction diagram Perform feature aggregation on the source book domain A to obtain the high-level representation of the book ID-user ID in the domain A. The calculation formula is:

[0108]

[0109] Among them, H A It represents the final aggregated high-level representation of the source book domain A’s book ID-user ID. L represents the number of graph aggregation layers. A A The adjacency matrix A A The calculated normalized adjacency matrix of the graph neural network.

[0110] Furthermore, based on the domain book ID-user ID high-level representation H of the source book domain A A The high-order prediction representation of the target book domain user ID obtained by the overlapping user interest mapper of the source book domain:

[0111]

[0112] in, Represents the high-level representation H of the book ID-user ID in the source book domain A A The high-level representation of the user ID obtained by the index, projector A→B (·) represents the overlapping user interest mapping function of the source book domain A, Represents the high-order prediction representation of the user ID of the target book domain B.

[0113] Furthermore, based on the domain book ID-user ID high-level representation H of the source book domain A A , Inter-domain book-book category high-level representation H C and the target book domain user ID high-order prediction representation calculation domain book-user interaction loss function According to the calculation results of the loss function, the back gradient propagation algorithm is used to update the network parameters, specifically the parameters of the embedding mapping module.

[0114] Computational domain book-user interaction loss function The method is:

[0115]

[0116] in,

[0117] represents the book-user interaction loss, which is used to fit the existing user-book interaction behavior. ln(·) is the logarithmic function, and σ(·) is the sigmoid activation function.

[0118] h u represents the unified user high-level representation, h m represents a unified high-level representation of books, Represents the high-level representation H of the book ID-user ID in the source book domain A A The high-level representation of the book ID obtained by the index, Represents the high-level representation H from the domain book-book category C High-level representation of books obtained from the index;

[0119] (u,m)∈E A , is a positive sample pair, representing the book-user interaction graph in the domain There is an interaction pair (r um ≥1), in is a negative sample pair, representing the book-user interaction graph in the domain There exists an interaction pair (r um′ <1).

[0120] Furthermore, based on the domain book-user interaction data Calculate the high-level representation H of the book ID-user ID in the target book domain B B The calculation method is the same as that of source book domain A and will not be repeated here.

[0121] Furthermore, all parameters of the overlapping user interest mapper projector except the source book domain A are frozen, and the cross-domain loss of user interest representation is calculated.

[0122]

[0123] in, Indicates finding vector to vector The function of Wasserstein distance, Represents the high-level representation H of the book ID-user ID in the target book domain BB The high-level representation of the user ID obtained by the index.

[0124] Furthermore, back gradient propagation is used to update the parameters of the overlapping user interest mapper for the source book domain A.

[0125] Step 5: Repeat steps 3 to 4 until convergence or the set number of training rounds is reached to obtain the final unified user high-level representation and unified book high-level representation.

[0126] Step 6: Using the final unified user high-level representation and unified book high-level representation, predict the recommendation scores between the source book domain users and the target book domain books that have not been interacted with, and recommend the target book domain books of potential interest to the source book domain users in a ranked manner.

[0127] Predict the recommendation score between users in the source book domain and books in the target book domain that have not been interacted with. The formula for calculating the recommendation score is:

[0128] p um =σ(h u h m T );

[0129] Among them, p um Indicates the recommendation score.

[0130] Furthermore, based on the recommendation scores between the user and the books with which he / she has not interacted, N books that the user may be interested in in the target book domain are recommended to the user in the source book domain according to the order of scores from high to low.

[0131] This invention breaks away from the traditional single-objective recommendation system model and utilizes a dual-objective cross-domain recommendation method based on inter-domain mapping. By learning the rich user-book interaction data in the source domain, it maps the user's target book domain. This can solve the problems of data sparsity and cold start in the traditional model, thereby adapting to the needs of rapid social development and iterative updates. It can give full play to the role of deep learning and quickly recommend books that suit users' preferences, thereby improving production efficiency and having good economic and social benefits.

[0132] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A book recommendation method based on cross-domain learning, characterized in that: Includes steps: Step 1: Obtain user and book information from a source book domain and at least one target book domain, and divide the obtained user and book information into intra-domain book-user interaction data and cross-domain book-category data; Step 2: construct an inter-domain book-category relationship graph based on cross-domain book-category data; construct an intra-domain book-user interaction graph based on intra-domain book-user interaction data; Step 3: Using cross-domain book-category data as supervision data, supervised training is performed on the inter-domain book-category relationship graph to obtain inter-domain book-book category high-level representations; Step 4: Using the intra-domain book-user interaction data as supervision data and combining the inter-domain book-book category high-level representation, supervised training is performed on the intra-domain book-user interaction graph of each book domain to obtain unified user high-level representations and unified book high-level representations for both domains. Step 5: Repeat steps 3 to 4 until convergence or the set number of training rounds is reached to obtain the final unified user high-level representation and unified book high-level representation; Step 6: Use the final unified user high-level representation and unified book high-level representation to predict the recommendation scores between the source book domain users and the target book domain uninteracted books, and recommend the target book domain books of potential interest to the source book domain users in a ranked manner.

2. The book recommendation method based on cross-domain learning according to claim 1 is characterized in that: The domain book-user interaction data includes the domain book-user interaction data of the source book domain A. and the book-user interaction data in the target book domain B in, and m represents the book ID, u represents the user ID, r mu represents the rating of book m by user u; The cross-domain book-category data is in, and They represent the book name, book introduction and book category name respectively, and c represents the book category ID.

3. The book recommendation method based on cross-domain learning according to claim 2 is characterized in that: The method for performing supervised training on the inter-domain book-category relationship graph is as follows: obtaining the original feature vector of the book and the original feature vector of the book category through a large language model based on the cross-domain book-category data; Map and concatenate the original feature vectors of the books and the original feature vectors of the book categories to obtain a zero-order representation matrix of the book-book category; Perform feature aggregation on the inter-domain book-category relationship graph to obtain the inter-domain book-book category high-level representation; The inter-domain book-category relationship loss function is calculated based on the inter-domain book-book category high-order representation, and the network parameters are updated using the back gradient propagation algorithm according to the loss function calculation result.

4. The book recommendation method based on cross-domain learning according to claim 2 is characterized in that: The method for performing supervised training on the intra-domain book-user interaction graph of each book domain is: Based on the book-user interaction data in the domain The user ID and book ID are mapped, embedded and concatenated respectively to obtain the zero-order representation matrix of book ID-user ID; feature aggregation is performed on the book-user interaction graph in the source book domain A. Get the high-level representation H of the book ID-user ID in the source book domain A A ; Based on the source book domain A, the domain book ID-user ID high-level representation H A The high-order prediction representation of the user ID in the target book domain obtained by the overlapping user interest mapper in the source book domain; Based on the source book domain A, the domain book ID-user ID high-level representation H A , Inter-domain book-book category high-level representation H C and the target book domain user ID high-order prediction representation calculation domain book-user interaction loss function The network parameters are updated using the reverse gradient propagation algorithm based on the loss function calculation results; Based on the book-user interaction data in the domain Calculate the high-level representation H of the book ID-user ID in the target book domain B B And calculate the cross-domain loss of user interest representation based on the high-order prediction representation of the target book domain user ID Backward gradient propagation is used to update the parameters of the overlapping user interest mapper for the source book domain.

5. The book recommendation method based on cross-domain learning according to claim 3 is characterized in that: The calculation method of the original feature vector of the book is: in, is the original feature vector of the book, and is the original eigenvector matrix of all books, LLM(·) represents the large language model feature extraction function; The calculation method of the original feature vector of the book category is: in, is the original feature vector of the book category name, is the original feature vector matrix of all book categories.

6. The book recommendation method based on cross-domain learning according to claim 5 is characterized in that: The calculation formula of the book-book category zero-order representation matrix is: in, All represent the original eigenvector matrix mapping parameters, represents the zero-order representation matrix of the book, The zero-order representation matrix of book categories, is the zero-order representation matrix of book-book category, stack(·) means concatenating the matrices without changing the columns; The calculation formula of the inter-domain book-book category high-order representation is: Among them, H C represents the final aggregated domain book-book category high-level representation, L represents the number of graph aggregation layers, A C is the adjacency matrix of the domain book-category relationship graph, A C The adjacency matrix A C The calculated graph neural network normalized adjacency matrix, Δ (l) is the noise matrix, and the i-th row vector of each noise matrix is (l) i is random noise, random noise Δ (l) i Satisfy the two norm is a parameter that follows a uniform distribution on the interval [0,1], ⊙ represents the vector outer product, sign(·) is the sign function, Represents the i-th order representation of the domain book-book category.

7. The book recommendation method based on cross-domain learning according to claim 6 is characterized in that: The method for calculating the inter-domain book-category relationship loss function based on the inter-domain book-book category high-order representation is: in, represents the total loss function of the book-category relationship graph between domains, V C represents the node set in the domain book-category relationship graph, represents the book-book category edge loss function, represents the graph contrastive learning loss function, λ represents the weight of the graph contrastive learning loss function; ln(·) is the logarithmic function, σ(·) is the sigmoid activation function, exp(·) is the exponential function, Represents the high-level representation H from the domain book-book category C The high-level representation of books obtained by indexing in Represents the high-level representation H from the domain book-book category C The high-order representation of the book category obtained by the index in , τ is the contrastive learning temperature coefficient; (m,c)∈E C Indicates that in the inter-domain book-category relationship graph, there is a book m belonging to category c, E C The edge set representing the book-category relationship graph between domains; In It means that book m does not belong to category c′ in the inter-domain book-category relationship graph.

8. The book recommendation method based on cross-domain learning according to any one of claims 2 to 7, characterized in that: The calculation process of obtaining the zero-order representation matrix of book ID-user ID is as follows: Among them, Embedding(·) represents the mapping embedding operation, represents the zero-order representation matrix of book ID, represents the zero-order representation matrix of user ID, represents the zero-order representation matrix of book ID-user ID; The calculation formula of the high-level representation of the book ID-user ID in the source book domain A is: Among them, H A The final aggregated high-level representation of the source book domain A, i.e., the book ID-user ID, is represented by A. A Represents the book-user interaction graph in the source book domain A, A A The adjacency matrix A A The calculated normalized adjacency matrix of the graph neural network.

9. The book recommendation method based on cross-domain learning according to claim 8 is characterized in that: The calculation method of the target book domain user ID high-order prediction representation is: in, Represents the high-level representation H of the book ID-user ID in the source book domain A A The high-level representation of the user ID obtained by the index, projector A→B (·) represents the overlapping user interest mapping function of source book domain A, Represents the high-order prediction representation of the user ID of the target book domain B; The computational domain book-user interaction loss function The method is: Among them, h u represents the unified user high-level representation, h m represents a unified high-level representation of books, V A The node set representing the book-user interaction graph in the source book domain A, Represents the high-level representation H of the book ID-user ID in the source book domain A A The high-level representation of the book ID obtained by the index, Represents the high-level representation H from the domain book-book category C High-level representation of books obtained from the index; (u,m)∈E A , representing the book-user interaction graph in the domain There exists an interaction pair (r um ≥1), E A represents the edge set of the book-user interaction graph in the source book domain A, In Represents a book-user interaction graph in the domain There exists an interaction pair (r um′ <1).

10. The book recommendation method based on cross-domain learning according to claim 9, characterized in that: The cross-domain loss of calculating the user interest representation based on the target book domain user ID high-order prediction representation The method is: freeze all parameters of the overlapping user interest mapper projector except the source book domain A, and calculate the cross-domain loss of user interest representation Among them, Wasserstein(·) represents the function of finding Wasserstein distance, A is the user set in source book domain A, B is the user set in the target book domain B, Represents the high-level representation H of the book ID-user ID in the target book domain B B The high-level representation of the user ID obtained by the index.

Citation Information

Patent Citations

  • Cross-domain recommendation method based on multi-view knowledge representation

    CN112541132A

  • Double-target cross-domain recommendation method and system based on knowledge graph and multi-head attention

    CN115760279A

  • Cross-domain recommendation via contrastive learning of user behaviors in attentive sequence models

    US20240161165A1

  • Model training method and apparatus, electronic device, computer readable medium, and computer program product

    WO2024114263A1