Self-supervised personalized recommendation method based on space-time aggregation and adaptive graph prompt learning
By introducing self-supervised personalized recommendation methods for spatiotemporal aggregation and adaptive graph prompt learning in the recommendation system, the problems of excessive attention to local information and neglecting time information in the prior art are solved, and a more accurate and personalized recommendation effect is achieved.
Patent Information
- Application Number
- CN202510213496.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-13
AI Technical Summary
Existing recommendation methods based on graph neural networks pay too much attention to local neighborhood information and ignore the temporal information of interactive data, making it difficult for the model to accurately capture the evolution of user preferences and changes in project popularity, affecting the accuracy and timeliness of the recommendation system.
A self-supervised personalized recommendation method based on spatiotemporal aggregation and adaptive graph prompt learning is proposed. By serializing user project interaction data, dynamic graphs and dynamic adjacency matrix are constructed, and graph convolution networks and global graph encoders are used for encoding. Combining adaptive spatiotemporal prompts and multi-view fusion, self-supervised tasks are generated to improve the generalization ability of the model.
By comprehensively capturing the spatial and temporal knowledge in interactive data, the accuracy and generalization of the recommendation system can be improved, and the accuracy and personalization of the recommendation results can be significantly improved, and data sparsity and noise problems can be alleviated.
Smart Images

Figure CN120144864A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a personalized graph recommendation method, and more precisely, to a self-supervised graph recommendation method based on spatio-temporal aggregation and adaptive graph prompt learning. Background Art
[0002] As an important information filtering technology, recommendation systems are widely used in e-commerce, video and music streaming, financial services, and other fields. Personalized recommendation is an important task in the field of recommendation systems, aiming to meet the personalized needs of users and improve the user experience. Early recommendation systems mainly adopted shallow models such as collaborative filtering, matrix factorization, etc. With the rise of deep learning technology, various neural networks such as deep neural networks (DNNs), recurrent neural networks (RNNs), convolutional neural networks (CNNs), etc. have been applied to recommendation systems to capture the complex relationships between users and items.
[0003] However, the research objects of recommendation systems mainly consist of the interactions between users and items or the historical behaviors of users. These interactions and behaviors with users and items usually present as graph data with a non-Euclidean structure. For example, in the commodity recommendation scenario, the click behavior data between users and commodities can be modeled as a user-item bipartite graph. Traditional deep learning technologies are difficult to model the complex relationships in graph data, resulting in poor recommendation performance. In recent years, graph neural networks (GNNs) have received extensive attention in the research of recommendation systems due to their powerful non-Euclidean embedding learning ability. It captures the high-order representation of the central node by iteratively aggregating neighborhood information, thereby effectively modeling user preferences for more accurate recommendations.
[0004] Most of the existing GNN-based recommendation methods focus on static scenarios and pay less attention to the recommendation problems in dynamic scenarios. In real-world application scenarios, users' interests and preferences are not static but constantly evolve over time and with the change of the environment. Dynamic interaction information can capture the time-series characteristics of user behavior, helping the model understand users' different interests at different time periods, thereby providing more accurate personalized recommendations. In addition, the popularity of items (such as movies, commodities, news, etc.) is not constant. New items may suddenly become popular due to certain events or trends, while old items may gradually be forgotten by users. Dynamic interaction information can reflect the changing trend of item popularity, helping the model adjust the recommendation strategy in real time. Therefore, in such dynamic scenarios, the existing personalized recommendation methods based on graph neural networks still have the following limitations:
[0005] 1) Existing aggregation methods based on graph neural networks usually over - focus on the local structural information of central nodes, often ignoring the important impact of dynamic interaction information between nodes on the recommendation results, thus causing the model to be unable to accurately capture the evolution of user preferences and the changes in item popularity, and affecting the accuracy and timeliness of the recommendation system.
[0006] 2) Most existing personalized graph recommendation models adopt a supervised data - driven training paradigm, while the observable data usually only accounts for a small part of the entire interaction space, which limits the ability of the recommendation model to learn more comprehensive representations, thus affecting the recommendation performance of the model.
[0007] 3) Most of the interaction data used in the recommendation system is implicit user feedback, such as browsing, clicking, etc., resulting in the data usually being doped with noise of different degrees and following different distributions, leading to a decline in the recommendation performance of the model. Summary of the Invention
[0008] The technical problem to be solved by the present invention is that existing recommendation methods based on graph neural networks usually over - focus on local neighborhood information and insufficiently explore the impact of time information of interaction data, resulting in the model being difficult to make accurate recommendations. In addition, since the implicit feedback information used in the recommendation task is usually sparse and noisy, the generalization ability of the recommendation model decreases, affecting the recommendation performance of the model. Therefore, to solve the above problems, the present invention proposes a self - supervised personalized recommendation model based on spatio - temporal aggregation and adaptive graph prompt learning, designs a spatio - temporal aggregation method based on graph prompt learning to more comprehensively capture user interests and dynamic behavior patterns, and designs self - supervised tasks to alleviate the deficiency of the generalization ability of the recommendation system, thereby improving the recommendation performance of the model.
[0009] To achieve the above objectives, the technical solution of the self - supervised personalized recommendation method based on spatio - temporal aggregation and adaptive graph prompt learning proposed by the present invention includes the following steps:
[0010] Step 1: Serialize the user - item interaction data based on the interaction timestamp to obtain a user - item interaction sequence, and construct a user - item interaction dynamic graph and the corresponding adjacency matrix.
[0011] (1) Given and represent the user set and the item set respectively, E = {<u, i, t>|u ∈ U, i ∈ I} represents the user - item interaction set, t represents the timestamp corresponding to the interaction. Based on E, the interaction set is represented in the form of a user - item bipartite graph G = {U, I, E}.
[0012] (2) Use the timestamp information to convert a specific user - item interaction set into a user - item interaction sequence. Taking user u as an example, its interaction sequence is:
[0013] Su =[i 1 ,…,i L ]
[0014] Among them, S u is the interaction sequence of user u, i n For user u at time t n With Project I n interaction, L is the length of the interaction sequence.
[0015] (3) S u Refactoring into interactive dynamic graph in t n Momentary dynamic graph time slice, express The node set of express A collection of links.
[0016] (4) Constructing dynamic graph time slices The adjacency matrix of Integrate the interaction adjacency matrices at different times to obtain the final interaction dynamic graph G u The dynamic adjacency matrix is defined as follows:
[0017]
[0018] in Represents a collection of interactions.
[0019] (5) The processing of item i also follows the operations (1)(2)(3)(4), and the item-user interaction sequence S is obtained. v , interactive dynamic graph G v and the dynamic adjacency matrix A v .
[0020] Step 2: Perform data enhancement on the interaction sequence to obtain the enhancer sequence and its corresponding dynamic interaction graph and adjacency matrix.
[0021] (1) According to a certain ratio, the processed user interaction sequence S u Interaction sequence S v Perform random masking to generate enhancer sequences S u For example, the corresponding enhancer sequence The definition is as follows:
[0022]
[0023] where f mask (·) indicates a random mask operation.
[0024] (2) Reconstruct the enhancer sequence into a dynamic interaction graph Taking user u as an example at time t n the dynamic graph time slice at time t is defined as follows:
[0025]
[0026] where represents the node set, represents the link set.
[0027] (3) Construct the adjacency matrix of the dynamic graph time slice Integrate the interaction adjacency matrices at different times to obtain the dynamic adjacency matrix of the final interaction dynamic graph which is defined as follows:
[0028]
[0029] where represents the interaction set.
[0030] Step 3: The multi-modal encoding module encodes the original interaction sequence and the enhancer interaction sequence respectively to obtain the spatio-temporal aware interaction sequence representations of the original sequence and the subsequence.
[0031] (1) Use the graph convolutional network GCN to encode the local structural information of the interaction graph at different time slices, obtain the local structural encoding embedding representations of the nodes in the interaction graph, pool the node embeddings at different time slices, and aggregate them through learnable weights to obtain the structural embedding representation of the final central node. Taking the interaction sequence S u of user u as an example, its encoding process is defined as follows:
[0032]
[0033]
[0034] where and are respectively 's feature matrix and adjacency matrix, h u ∈H is the embedding of node u in the user-item interaction bipartite graph G. The embedding h v of the item node is obtained in the same way.
[0035] (2) Use the global graph encoder to encode the user-item interaction graph again to further improve the global expression ability of its nodes, which is defined as follows:
[0036]
[0037] Among them, H and A are the node embedding matrix and the adjacency matrix in the bipartite graph G of user-item interaction respectively.
[0038] (3) Construct a set of shared learnable vectors T for all user interaction sequences p As time position cues, then, fuse the user node embeddings and item node embeddings obtained in (2) with the p corresponding time position cues in T to obtain the spatio-temporal aware interaction sequence representation S u of the interaction sequence, which is defined as follows:
[0039]
[0040] Among them, is the embedding representation of user u, is the embedding representation of the nth item i interacted with user u n p n ∈ T p is the time position cue corresponding to the nth interaction.
[0041] Step Four: Perform multi-view fusion on the spatio-temporal aware interaction sequence representation to obtain more expressive user-item embeddings, and generate adaptive spatio-temporal cues to guide the message passing process of the model.
[0042] (1) Map the spatio-temporal aware interaction sequence encoding of any interaction sequence in the user-item interaction set into different subspaces to obtain representations of different views Then input it into the multi-head attention mechanism to capture and fuse different view features in the data. Taking the interaction sequence S u as an example, its process is defined as follows:
[0043]
[0044] Among them, is the encoding mapping function of the ith subspace, ω i is its learnable parameter, is the query vector, is the key vector, is the value vector, d is the dimension of the mapping encoding
[0045] (2) Concatenate the weighted representations obtained from different subspaces to get the multi-view serialized embedding Z u , and then use the pooling operation to fuse Z u to obtain the spatio-temporal embedding representation h of the central node u , which is defined as follows:
[0046]
[0047] h u = pooling(Z u ), h v = pooling(Z v )
[0048] Among them, MLP ω (·) is a fusion function used to fuse the weighted representations of different subspaces, ω is its learnable weight, pooling(·) is a pooling function. Similarly, for any sequence S v 、 the above operations are also performed to obtain the embedding representation of the corresponding central node.
[0049] (3) Use the spatio-temporal embeddings h u and h v to generate an adaptive spatio-temporal cue P, which is defined as follows:
[0050]
[0051] Among them, f ψ is a mapping function, ψ is a learnable parameter, is a weight parameter used to scalarize the concatenated vector, is the interaction set of user u.
[0052] (4) Combine the adaptive spatio-temporal cue P with the adjacency matrix A to guide the message passing process of the model, which is defined as follows:
[0053]
[0054] Among them, h l+1 is the representation of the central node at the (l + 1)th layer, is the degree diagonal matrix, and any of its elements
[0055] (5) Aggregate the output vectors of each layer through weighted pooling to obtain the final node representations e u and e v of the user and the item. The pooling process is defined as follows:
[0056]
[0057] Among them, is the weight coefficient, which is set to 1 / (L * + 1) in this paper, L * is the number of network layers, and are the representations of the user and the item at the lth layer respectively.
[0058] Step 5: Construct a self-supervised task, jointly optimize the recommendation task and the self-supervised auxiliary task, ensure the recommendation goal, and train the model by generating supervision signals from the data itself, thereby improving the model's ability to understand the data.
[0059] (1) Construct the optimization objective of the main recommendation task, and its specific definition is as follows:
[0060]
[0061] Among them, is the set of first-order neighbors of node u, is the set of items that have not interacted with node u, Θ is the learnable parameter, β is the regularization coefficient, and σ(·) is the sigmoid function.
[0062] (2) Use the embedding representations of the obtained enhanced interaction data and the original interaction data to construct a contrastive loss as the optimization objective of the self-supervised auxiliary task, and its definition is as follows:
[0063]
[0064] Among them, τ is the hyperparameter, and sim(·) is the similarity function.
[0065] (3) Adopt a jointly optimized training method to construct a joint loss optimization function, and its definition is as follows:
[0066]
[0067] Among them, is the target recommendation loss, is the self-supervised contrastive loss, λ is the hyperparameter, and controls 's contribution degree to the overall optimization objective.
[0068] Through the technical solution proposed by the present invention, the following remarkable effects can be achieved:
[0069] Existing recommendation methods based on graph neural networks usually over - focus on local neighborhood information when aggregating neighbors, and insufficiently explore the impact of temporal information in interaction data, resulting in the difficulty of the model to make accurate recommendations. In addition, since the implicit feedback information used in the recommendation task is usually sparse and noisy, the generalization ability of the recommendation model is greatly reduced. Therefore, the present invention proposes a self - supervised personalized recommendation method STGPrec based on spatio - temporal aggregation and adaptive graph prompt learning to comprehensively capture spatio - temporal knowledge in interaction data and improve the accuracy and generalization of the recommendation system. First, we serialize the user - item interaction information and obtain its encodings in different modes. Second, we use the spatio - temporal aggregation module to fuse the user - item interaction encodings in different modes, enabling the model to more comprehensively understand the user's interest evolution and behavior patterns. At the same time, an adaptive graph prompt is introduced to guide the learning process of the model, and the implicit knowledge learned by the model is introduced into the spatio - temporal aggregation process, so that the model can make accurate personalized recommendations. Then, to enhance the robustness and generalization ability of the model, we use the data augmentation module to generate positive and negative samples to construct auxiliary tasks to increase the diversity of the dataset, and enable the model to capture user interests during the training process, thus helping the model learn more generalized feature representations. Finally, to further improve the recommendation performance, we construct a joint loss optimization function for the self - supervised task and the recommendation task, and train the model by generating supervision signals from the data itself, thereby improving the model's understanding ability of the data. We conducted experiments on 5 datasets and verified the effectiveness of the model.
[0070] In summary, the present invention proposes a self - supervised personalized recommendation method STGPrec based on spatio - temporal aggregation and adaptive graph prompt learning. This method fuses temporal information and the structural information captured by graph neural networks to more comprehensively capture the user's interests and behavior patterns. At the same time, an adaptive graph prompt is introduced to guide the learning process of the model, and the implicit knowledge learned by the model is integrated into the spatio - temporal aggregation process, significantly improving the accuracy and personalization of the recommendation results. In addition, to alleviate the problems of data sparsity and noise, this method designs self - supervised contrast tasks, fully utilizes the effective information of the data to mine the true preferences of users, improves the generalization performance of the recommendation model, and finally realizes high - quality personalized recommendations. Brief Description of the Drawings
[0071] Figure 1 It is a schematic diagram of the functions and connection relationships of the self - supervised personalized recommendation model based on spatio - temporal aggregation and adaptive graph prompt learning described in the present invention and its constituent modules.
[0072] Figure 2 It is a flowchart of the self - supervised personalized recommendation method based on spatio - temporal aggregation and adaptive graph prompt learning described in the present invention. Detailed Embodiments
[0073] The technical solution of the present invention will be elaborated in detail below with reference to the accompanying drawings. It should be noted that the embodiments listed herein are intended to help understand the core idea of the present invention, rather than limiting the protection scope of the present invention.
[0074] As Figure 1 shown, we innovatively propose a self-supervised personalized recommendation method (STGPrec), which uses key technologies such as spatio-temporal fusion, adaptive prompting, and self-supervised learning to improve the generalization and personalized recommendation capabilities of the recommendation model. Specifically, STGPrec mainly consists of three parts: a multi-modal encoding module, a multi-view spatio-temporal embedding fusion module, and a joint optimization module. First, the user-item interaction information is serialized and its encodings in different modes are obtained. Secondly, a spatio-temporal aggregation module is used to fuse the user-item interaction encodings in different modes, enabling the model to more comprehensively understand the user's interest evolution and behavior patterns. At the same time, an adaptive graph prompt is introduced to guide the learning process of the model, and the implicit knowledge learned by the model is introduced into the spatio-temporal aggregation process, so that the model can make accurate personalized recommendations. Finally, to enhance the robustness and generalization ability of the model, a data augmentation module is used to generate positive and negative samples to construct auxiliary tasks to increase the diversity of the dataset, and enable the model to capture user interests during the training process, thereby helping the model learn more general feature representations. To further improve the recommendation performance, we construct a joint loss optimization function for the self-supervised task and the recommendation task, and train the model by generating supervision signals from the data itself, thereby improving the model's understanding ability of the data.
[0075] As Figure 2 shown, the steps of the self-supervised personalized recommendation method based on spatio-temporal aggregation and adaptive graph prompt learning according to the present invention are as follows:
[0076] step1: Given and represent the user set and the item set respectively, E = {<u, i, t>|u ∈ U, i ∈ I} represents the user-item interaction set, t represents the timestamp corresponding to the interaction. Based on E, the interaction set is represented in the form of a user-item bipartite graph G = {U, I, E}.
[0077] step2: Use the timestamp information to convert the specific user-item interaction set into a user-item interaction sequence S u / v . Taking user u as an example, its interaction sequence is:
[0078] S u = [i 1 , …, i L
[0079] where S u is the interaction sequence of user u, in The interaction of user u at time t n with project i n where L is the length of the interaction sequence.
[0080] step3: Reconstruct the interaction sequence into an interaction dynamic graph. Taking S u as an example, reconstruct S u into an interaction dynamic graph where is the dynamic graph time slice at time t n , denotes the set of nodes, denotes the set of links.
[0081] step4: Generate a dynamic adjacency matrix. Taking S u as an example, construct the adjacency matrix of the dynamic graph time slice Integrate the interaction adjacency matrices at different times to obtain the dynamic adjacency matrix of the final interaction dynamic graph G which is defined as follows: u
[0082]
[0083] where denotes the interaction set.
[0084]
[0084] step5: Randomly mask the processed user interaction sequence S u / v in a certain proportion to generate an enhanced subsequence Taking S u as an example, its corresponding enhanced subsequence is defined as follows:
[0085]
[0086] where f mask (·) represents the random masking operation, is the randomly generated mask index sequence, and L mask denotes the mask length.
[0087] step6: Reconstruct the enhanced subsequence into a dynamic interaction graph Taking user u as an example, the dynamic graph time slice at time t n is defined as follows:
[0088]
[0089] where denotes the set of nodes, Represents a set of links.
[0090] Step 7: Generate a dynamic adjacency matrix. Taking as an example, construct the adjacency matrix of the dynamic graph time slice of Integrate the interaction adjacency matrices at different times to obtain the dynamic adjacency matrix of the final interaction dynamic graph which is defined as follows:
[0091]
[0092] where represents the interaction set.
[0093] Step 8: Use the graph convolutional network GCN to encode the local structural information of the interaction graphs at different time slices to obtain the local structural encoding embedding representation of the nodes in the interaction graph. Taking the interaction sequence S u of user u as an example, its encoding process is defined as follows:
[0094]
[0095] where, and are respectively the feature matrix and the adjacency matrix of
[0096] Step 9: Pool the node embeddings at different time slices, and aggregate them through learnable weights to obtain the structural embedding representation of the final central node. Taking the interaction sequence S u of user u as an example, its pooling process is defined as follows:
[0097]
[0098] where, α t is the learnable weight, h u ∈H is the embedding of node u in the user-item interaction bipartite graph G, and the embedding h v of the item node is obtained in the same way.
[0099] Step 10: Use the global graph encoder to encode the user-item interaction graph again to further improve the global expression ability of its nodes, which is defined as follows:
[0100]
[0101] where, H is the node embedding matrix of the user-item interaction bipartite graph G obtained in step 9, and A is the adjacency matrix.
[0102] step11: Construct a set of shared learnable vectors T for all user-item interaction sequences p As the time position hint, then, fuse the user node embedding and item node embedding obtained in step10 with the p corresponding time position hint in T to obtain the spatio-temporal aware interaction sequence representation of the interaction sequence. Taking S u as an example, its process is defined as follows:
[0103]
[0104] Among them, is the embedding representation of user u, is the embedding representation of the nth item i interacted with user u n , p n ∈T p is the time position hint corresponding to the nth interaction.
[0105] step12: Map the spatio-temporal aware interaction sequence encoding of any interaction sequence in the user-item interaction set into different subspaces to obtain representations of different views Then input it into the dot product attention mechanism to capture and fuse the different view features in the data. Taking the interaction sequence S u as an example, its process is defined as follows:
[0106]
[0107] Among them, is the encoding mapping function of the u-th subspace, ω i is its learnable parameter, is the query vector, is the key vector, is the value vector, d is the dimension of the mapping encoding .
[0108] step13: Concatenate the weighted representations obtained from different subspaces to get the multi-view serialized embedding Z u , then use the pooling operation to fuse Z u to obtain the spatio-temporal embedding representation h u of the central node, which is defined as follows:
[0109]
[0110] h u =pooling(Z u ), h v =pooling(Z v )
[0111] Among them, MLP ω (·) is a fusion function used to fuse the weighted representations of different subspaces, ω is its learnable weight, pooling(·) is a pooling function. Similarly, for any sequence S v 、 also performs the above operations to obtain the embedding representation of the corresponding central node.
[0112] step14: Generate an adaptive spatio-temporal cue P using the spatio-temporal embeddings h u and h v , and its definition is as follows:
[0113]
[0114] Among them, f ψ is a mapping function, ψ is a learnable parameter, is a weight parameter used to scalarize the concatenated vector, is the interaction set of user u.
[0115] step15: Combine the adaptive spatio-temporal cue P with the adjacency matrix A to guide the model message passing process, and its definition is as follows:
[0116]
[0117] Among them, h l+1 is the representation of the central node at layer (l + 1), is the degree diagonal matrix, and any of its elements
[0118] step16: Aggregate the output vectors of each layer through weighted pooling to obtain the final node representations e u and e v of the user and the item, and the pooling process is defined as follows:
[0119]
[0120] Among them, is the weight coefficient, which is set to 1 / (L * +1) in this paper, L * is the number of network layers, and are the representations of the user and the item at layer l respectively.
[0121] step17: Construct the optimization objective of the main recommendation task. Taking user u as an example, its specific definition is as follows:
[0122]
[0123] Among them, is the set of first-order neighbors of node u. is the set of items that have not interacted with node u, Θ is a learnable parameter, β is a regularization coefficient, and σ(·) is the sigmoid function.
[0124] Step 18: Use the embedding representations of the obtained enhanced interaction data and the original interaction data to construct a contrastive loss as the optimization objective of the self-supervised auxiliary task, which is defined as follows:
[0125]
[0126] where τ is a hyperparameter and sim(·) is a similarity function.
[0127] Step 19: Adopt a joint optimization training method to construct a joint loss optimization function, which is defined as follows:
[0128]
[0129] where is the target recommendation loss, is the self-supervised contrastive loss, λ is a hyperparameter that controls the contribution degree to the overall optimization objective.
[0130] Step 20: Judge whether the model reaches the preset early stopping condition. If it is satisfied, enter Step 22; if not, enter Step 8 to continue execution.
[0131] Step 21: Terminate the training process and save the trained model parameters. Subsequently, use the feature representations generated by this model to complete the recommendation task and comprehensively evaluate the recommendation performance of the model.
Claims
1. A self-supervised personalized recommendation method based on spatiotemporal aggregation and adaptive graph cue learning, characterized by: The following steps are involved: Step 1: Serialize the user-item interaction data based on the interaction timestamp to obtain the user-item interaction sequence, and construct the user-item interaction dynamic graph and the corresponding adjacency matrix; Step 2: Perform data enhancement on the interaction sequence to obtain the enhancer sequence and its corresponding dynamic interaction graph and adjacency matrix; Step 3: The multimodal encoding module encodes the original interaction sequence and the enhancer interaction sequence respectively to obtain the spatiotemporal perception interaction sequence representation of the original sequence and the subsequence; Step 4: Perform multi-view fusion on the spatiotemporal-aware interaction sequence representation to obtain more expressive user-item embeddings, and generate adaptive spatiotemporal cues to guide the model message passing process; Step 5: Construct a self-supervised task to jointly optimize the recommendation task and the self-supervised auxiliary task. While ensuring the recommendation goal, the model is trained by generating supervisory signals from the data itself, thereby improving the model's ability to understand the data.
2. The self-supervised personalized recommendation method based on spatiotemporal aggregation and adaptive graph prompt learning according to claim 1, characterized in that: The step one comprises: (1) Given and Represents the user set and the project set respectively, E={<u,i,t> |u∈U,i∈I} represents the user-item interaction set, t represents the timestamp corresponding to the interaction, based on E, the interaction set is represented as a user-item bipartite graph in the form of G={U,I,E}; (2) Use timestamp information to convert a specific user-item interaction set into a user-item interaction sequence. Taking user u as an example, its interaction sequence is: S u =[i1,…,i L ] Among them, S u is the interaction sequence of user u, i n For user u at time t n With Project I n interaction, L is the length of the interaction sequence; (3) S u Refactoring into interactive dynamic graph in t n Momentary dynamic graph time slice, express The node set of express A collection of links; (4) Constructing dynamic graph time slices The adjacency matrix of Integrate the interaction adjacency matrices at different times to obtain the final interaction dynamic graph G u The dynamic adjacency matrix is defined as follows: in Represents a collection of interactions; (5) The processing of item i also follows the operations (1)(2)(3)(4), and the item-user interaction sequence S is obtained. v , interactive dynamic graph G v and the dynamic adjacency matrix A u .
3. The self-supervised personalized recommendation method based on spatiotemporal aggregation and adaptive graph prompt learning according to claim 1, characterized in that: The second step comprises: (1) According to a certain ratio, the processed user interaction sequence S u Interaction sequence S v Perform random masking to generate enhancer sequences S u For example, the corresponding enhancer sequence The definition is as follows: where f mask (·) indicates a random mask operation; (2) The enhancer sequence Refactoring into a dynamic interaction diagram Take user u as an example, t n The time slice of the dynamic graph at a moment is defined as follows: in, Represents a collection of nodes. Represents a collection of links; (3) Constructing dynamic graph time slices The adjacency matrix of Integrate the interaction adjacency matrices at different times to obtain the final interaction dynamic graph The dynamic adjacency matrix is defined as follows: in Represents a collection of interactions.
4. The self-supervised personalized recommendation method based on spatiotemporal aggregation and adaptive graph prompt learning according to claim 1, characterized in that: The step three comprises: (1) The graph convolutional network (GCN) is used to encode the local structural information of the interaction graph at different time slices, and the local structural encoding embedding representation of the nodes in the interaction graph is obtained. The node embeddings of different time slices are pooled and aggregated through learnable weights to obtain the structural embedding representation of the final central node. The interaction sequence S of user u is used as the u For example, the encoding process is defined as follows: in, and They are The characteristic matrix and adjacency matrix of u ∈H is the embedding of node u in the user-item interaction bipartite graph G. The embedding h of the project node is obtained in the same way. v ; (2) Using a global graph encoder The user-item interaction graph is encoded again to further improve the global expression ability of its nodes, which is defined as follows: Among them, H and A are the node embedding matrix and adjacency matrix in the user-item interaction bipartite graph G, respectively; (3) Construct a set of shared learnable vectors T for all user interaction sequences p As a temporal position hint, the user node embedding and item node embedding obtained in (2) are then combined with T p The corresponding temporal position cues in the u The spatiotemporal-aware interaction sequence representation is defined as follows: in, is the embedding representation of user u, is the nth item i interacted with by user u n The embedding representation of p n ∈T p It is the time position prompt corresponding to the nth interaction.
5. The self-supervised personalized recommendation method based on spatiotemporal aggregation and adaptive graph prompt learning according to claim 1, characterized in that: The step 4 comprises: (1) Map the spatiotemporal-aware interaction sequence representation of any interaction sequence in the user-item interaction set into different subspaces to obtain representations of different views. It is then fed into a multi-head attention mechanism to capture and fuse different view features in the data to form an interactive sequence S u For example, the process definition is as follows: in, is the encoding mapping function of the i-th subspace, ω i For its learnable parameters, is the query vector, is the key vector, is the value vector, d is the mapping code Dimensions; (2) The weighted representations obtained from different subspaces Concatenate to get multi-view serialization embedding Z u , and then use the pooling operation to fuse Z u , and obtain the spatiotemporal embedding representation h of the central node u , which is defined as follows: h u =pooling(Z u ),h v =pooling(Z v ) Among them, MLP ω (·) is the fusion function used to fuse the weighted representations of different subspaces, ω is its learnable weight, and pooling(·) is the pooling function. Similarly, for any sequence The above operation is also performed to obtain the embedding representation of the corresponding central node; (3) Using spatiotemporal embedding h u and h v Generate an adaptive spatiotemporal cue P, which is defined as follows: Among them, f ψ is the mapping function, ψ is a learnable parameter, is the weight parameter used to scalarize the concatenated vector, is the interaction set of user u; (4) The adaptive spatiotemporal hint P is combined with the adjacency matrix A to guide the model message passing process, which is defined as follows: Among them, h l+1 is the representation of the central node of the (l+1) layer, is a diagonal matrix of degree, any element of which (5) Aggregate the output vectors of each layer through weighted pooling to obtain the node representation of the final user and project e u and e v , the pooling process is defined as follows: in, is the weight coefficient, which is set to 1 / (L * +1), L * is the number of network layers, and are the representations of users and items at the lth layer, respectively.
6. The self-supervised personalized recommendation method based on spatiotemporal aggregation and adaptive graph prompt learning according to claim 1, characterized in that: The step five comprises: (1) Construct the optimization objective of the main recommendation task, which is specifically defined as follows: in, is the first-order neighbor set of node u, is the set of items that have not interacted with node u, Θ is a learnable parameter, β is a regularization coefficient, and σ(·) is a sigmoid function; (2) Using the embedded representation of the enhanced interaction data and the original interaction data, we construct a contrast loss as the optimization objective of the self-supervised auxiliary task, which is defined as follows: Among them, τ is a hyperparameter, sim(·) is the similarity function; (3) Using the joint optimization training method, a joint loss optimization function is constructed, which is defined as follows: in, is the target recommendation loss, is the self-supervised contrast loss, λ is a hyperparameter that controls The degree of contribution to the overall optimization goal.