A Spatiotemporal Data Augmentation Method for Recommendation Systems

Through a large model, the space-time characteristics of users and objects are explicitly fused, and positive and negative sample pairs are generated, which solves the problems of insufficient feature fusion and low user feedback data in the space-time recommendation system, and improves recommendation performance and stability.

CN120179926BActive Publication Date: 2025-07-22ZHEJIANG UNIV OF SCI & TECH +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510641967.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-07-22
Estimated Expiration
2045-05-19

AI Technical Summary

Technical Problem

The existing space-time recommendation systems have problems with insufficient feature fusion and low-quality user feedback data, resulting in insufficient recommendation performance.

Method used

The large model is used to explicitly summarize the spatiotemporal characteristics of users and objects, and generate positive and negative sample pairs through the large model, expand the training sample set, optimize the recommended model with linear layer and BPR loss function, and use sample pruning strategy to reduce noise.

Benefits of technology

Explicitly fuse spatiotemporal features, expand the training sample set, reduce noise, improve the performance and stability of the recommended model, and is suitable for a variety of recommended models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179926B_ABST
    Figure CN120179926B_ABST
Patent Text Reader

Abstract

The present invention discloses a spatio-temporal data augmentation method for a recommendation system. The present invention explicitly summarizes the spatio-temporal characteristics of users and items, and uses a large model as an encoder to unify the embeddings in a single vector space; then uses a linear layer to align the feature dimensions for fusion into the recommendation model, solving the problem of insufficient feature fusion. The present invention uses the large model to understand the user's interaction history and candidate set, infer the user's preferences and generate positive and negative sample pairs, and combines the generated sample set with the original sample set as the final training set. Thanks to the excellent reasoning ability and natural language understanding ability of the large model, this not only expands the training sample set but also alleviates the noise problem.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of information recommendation, and particularly relates to a spatio-temporal data enhancement method for a recommendation system. Background Art

[0002] Spatio-temporal recommendation systems greatly expand the application boundary of recommendation systems by integrating users' geographical location and time information. Such systems not only consider users' historical behaviors and preferences, but also use various techniques to process time features and spatial features.

[0003] To improve recommendation performance, many efforts have been made in existing work. Some traditional recommendation models improve recommendation performance by modeling the interaction relationship between users and items. They mainly rely on collaborative filtering technology, have limited ability to process complex features, and cannot use spatio-temporal attributes to assist in recommendation. In addition, they also face the problem of insufficient supervision signals (feedback data). Other models use spatio-temporal sequence information to assist in modeling. However, they do not explicitly encode users' preferences and item features and effectively integrate them into the model, which leads to a shallow understanding of users' preferences and item features.

[0004] Data enhancement is to generate new training samples by performing various transformations on the original data. For spatio-temporal data, previous studies mostly used traditional methods for enhancement. They have made certain progress, but still have not solved the inherent problems of recommendation systems. Specifically, they do not directly enhance user-item interactions and still have the problem of low-quality user feedback data.

[0005] In recent years, large models have been developed vigorously and have achieved remarkable success in many tasks and fields. Initially, large models were developed as pre-trained language-based models to solve different natural language tasks. However, as time goes by, the number of parameters of large models has increased sharply, which enables them to show powerful natural language understanding capabilities, learn complex semantics and knowledge representations from large-scale text corpora, and complete various reasoning tasks. Some recent studies have tried to use large models to assist in recommendation. But they only enhance ordinary features and do not consider spatio-temporal features.

[0006] Based on the deficiencies of existing research, it can be seen that current recommendation systems mainly face two challenges. The first is the lack of explicitly summarizing spatio-temporal features and encoding them into the recommendation model. Due to the limitations of privacy protection in collecting user location information and behavior time, there may be insufficient data to mine users' spatio-temporal preferences, summarize them in text form, and finally encode and integrate them into the recommendation system. The second is the low-quality user feedback data, including data sparsity and noise problems. In terms of data sparsity, due to time and space limitations, users can often only interact with a small part of the items in the system, which leads to sparse feedback data. For example, in a food recommendation system, limited by working hours, users may only be active on holidays. In terms of the noise problem, users may interact with items influenced by trends or social circles, which does not represent their true preferences. For instance, a user may choose a restaurant because of its high rating in the recommendation system but actually dislike it.

[0007] In summary, current research on spatio-temporal recommendation systems has problems such as insufficient feature fusion and low-quality user feedback data. At this time, it is urgent to find a solution that can solve the above problems and improve the performance of the spatio-temporal recommendation model. Summary of the Invention

[0008] In view of the deficiencies of the prior art, the present invention proposes a spatio-temporal enhancement recommendation method based on a large model. By using the excellent natural language processing and reasoning capabilities of the large model, the large model summarizes the features of users and items and encodes and integrates them into the recommendation system to solve the problem of insufficient feature fusion in spatio-temporal recommendation systems.

[0009] In addition, the present invention uses the large model to generate positive and negative sample pairs for each user, expands the number of training samples, and reduces the noise caused by random sampling of the Bayesian Personalized Ranking (BPR) algorithm to solve the problem of low-quality user feedback data.

[0010] The present invention provides a spatio-temporal data enhancement method for a recommendation system, including the following steps:

[0011] Step 1: Use the large model to enhance the time features of items to obtain enhanced item time features;

[0012] Step 2: Use the reasoning ability of the large model to summarize the general features, time features, and space features of users based on the user's interaction history to obtain enhanced user features;

[0013] Step 3: Use the large model to understand the user's interaction history and candidate set, and infer the items that the user is interested in and not interested in as positive and negative sample pairs to expand the training sample set;

[0014] Step 4: Input the enhanced item time features, the original item general features and spatial features in the dataset, and the enhanced user features into the encoding layer to generate high-dimensional vectors, and then use a linear layer for dimensionality reduction;

[0015] Step 5: Input the dimensionality-reduced feature vectors and the id embedding vectors generated by the basic recommendation model into the feature fusion layer, and obtain a comprehensive representation through weighted fusion;

[0016] Step 6: Train the recommendation model using the BPR loss function. Input the comprehensive representation into the BPR loss function without sample pruning, and merge the enhanced samples and the original training set as the final training set to obtain the first loss function;

[0017] Step 7: For the time features and spatial features in the dimensionality-reduced item and user features, input them into the BPR layer with pruning, and adopt a sample pruning strategy to obtain the second loss function;

[0018] Step 8: Fuse the first loss function and the second loss function according to certain weights to obtain the total loss function, and use this total loss function to optimize the model training.

[0019] The present invention also provides a recommendation system integrated with the above spatio-temporal data enhancement method.

[0020] Compared with the prior art, the beneficial technical effects of the present invention are as follows:

[0021] 1. The present invention uses a large model to enhance spatio-temporal data and applies it to the recommendation system. Compared with traditional spatio-temporal recommendation methods, the present invention explicitly summarizes the spatio-temporal features of users and items, uses a large model as the encoding layer to unify the embeddings in a single vector space, and then uses a linear layer to align the feature dimensions for fusion into the recommendation model, solving the problem of insufficient feature fusion.

[0022] 2. The sparsity and noise of user feedback are inherent problems in the recommendation system. The present invention uses a large model to understand the user's interaction history and candidate set, infer the user's preferences and generate positive and negative sample pairs, and merge the generated sample set and the original sample set as the final training set. Thanks to the excellent reasoning ability and natural language understanding ability of the large model, this not only expands the training sample set but also alleviates the noise problem.

[0023] 3. In order to emphasize relevant supervision signals and prevent incorrect gradient descent, the present invention adopts a pruning strategy. Using spatio-temporal features for sample pruning, removing a part of easily distinguishable training sample pairs, making the optimization more stable and effective.

[0024] 4. The spatio-temporal data enhancement method proposed by the present invention is pluggable and can be integrated into various recommendation models to improve their recommendation effects. Brief Description of the Drawings

[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art.

[0026] Figure 1 It is a schematic diagram of the framework of the offline stage of a large model-based spatio-temporal enhanced recommendation technology according to an embodiment of the present application.

[0027] Figure 2 It is a schematic diagram of the framework of the online stage of a large model-based spatio-temporal enhanced recommendation technology according to an embodiment of the present application.

[0028] Figure 3 It is the structure of the BPR layer with pruning in the embodiment of the present application.

[0029] Figure 4 It is a schematic diagram of three prompt templates in the embodiment of the present application. Detailed Embodiment

[0030] The following will illustrate the specific embodiments of the present invention in conjunction with the drawings and embodiments. However, the following embodiments are only used to illustrate the present invention in detail and do not limit the scope of the present invention in any way.

[0031] Overall concept of the present application: The present application can be divided into two stages, namely the offline stage of using a large model to enhance data and the online stage of training a recommendation model using the enhanced data. Among them, the offline stage can be further divided into three modules: item enhancement, user enhancement, and user-item enhancement. In the item enhancement module, the large model is used to transform the time features of items. In the user enhancement module, the large model infers the preference information of users. In the user-item enhancement module, the large model processes the interaction history of users and candidate sets to generate positive and negative sample pairs. In the online stage, in order to fuse the enhanced features into the recommendation model, these text features are encoded and dimensionality-reduced. Finally, the present application uses the BPR loss function to optimize the model and uses spatio-temporal features to prune the samples.

[0032] An embodiment of the present application discloses a large model-based spatio-temporal data enhancement method, Figure 1 and Figure 2 respectively gives the schematic diagrams of the frameworks of the offline stage and the online stage of this embodiment. The following will describe the present application in detail in conjunction with the embodiments, specifically including the following steps:

[0033] Step 1: Use a large model to enhance the time features (business hours) of items.

[0034] Concatenate the specific business hours of the merchant (item) with the item prompt template. The item prompt template includes three parts: task description, specific business hours, and output format description. See Figure 4, input the concatenated prompt into the large model to generate enhanced time features.

[0035] Furthermore, the features of merchants can be divided into general features C i (category), spatial features S i (location), and time features T i (business hours). They all exist in the original dataset and are stored using the nested dictionary items. The keys in the first layer are the three types of features, and the second-layer dictionary stores the corresponding relationship between merchant IDs and text features. Among them, the business hours of merchants are too specific and have a fine granularity, which may lead to large errors in the user time features inferred by the large model and make it difficult to capture the user-merchant collaborative signal. Therefore, this step is used to enhance the time features of merchants. Specifically, concatenate the specific business hours of merchants with the item prompt template and then input it into the large model , to obtain the enhanced time feature T i , and replace the corresponding content in the dictionary items. The formula is as follows:

[0036]

[0037] Among them, P i is the item prompt template, and T i is the feature in text form; after enhancement, in this implementation, these three types of features are input into the encoding layer for further processing.

[0038] Step 2: Utilize the inference ability of the large model to summarize the general features (preferred categories), time features (preferred dining times), and spatial features (preferred dining locations) of the user based on the user's interaction history.

[0039] First, replace the specific business hours of each merchant in the user's interaction history with the enhanced result of Step 1, then concatenate the interaction history with the user prompt template, and enhance it through the large model. The user prompt template includes three parts: task description, user interaction history, and output format description. See Figure 4 .

[0040] Then call the split function of python to split the enhanced result to obtain the general feature C u 、time feature T u and spatial feature S u of each user. Store them using the nested dictionary users. The keys in the first layer are the three types of features, and the second-layer dictionary stores the corresponding relationship between user IDs and text features. The formula is as follows:

[0041]

[0042] Among them, P u is the user prompt template, and C u, S u , T u are text-form features. After enhancement, this embodiment inputs these three features into the encoding layer for further processing.

[0043] Step 3: Since BPR random sampling can lead to problems of false positive samples (noise) and false negative samples (uninteracted), it is necessary to use a large model to understand the user's interaction history and candidate set, and infer the items that the user is interested in and not interested in as positive and negative sample pairs, so as to achieve the purpose of expanding the training sample set and reducing noise.

[0044] Specifically, the user's interaction history and candidate set are concatenated with the user-item prompt template and input into the large model. Since the large model has a token number limit when processing input, this embodiment scores all items with a basic recommendation model and ranks them in descending order, and selects a small part of them as the positive sample candidate set and the negative sample candidate set (five each). This embodiment selects the top five items as the positive sample candidate set.

[0045] The user-item prompt template includes four parts: task description, interaction history, candidate set, and output format description, as shown in Figure 4 . The candidate set is obtained by screening the entire item set using a basic recommendation model.

[0046] Selecting the negative sample candidate set is relatively difficult. On the one hand, to avoid being a positive sample, their rankings cannot be too high. On the other hand, to stimulate and improve the effect of the recommendation model, these negative sample candidate sets should have a certain correlation with the positive samples. Therefore, their rankings cannot be too low either.

[0047] After many experiments, it is found that screening five items from the items ranked in the top 10% as the negative sample candidate set can achieve the best results. The large model outputs enhanced positive and negative sample pairs <p, n>, and the formula is as follows:

[0048]

[0049] where P ui is the user-item prompt template, and p and n represent the positive sample and negative sample generated by the large model respectively.

[0050] The enhanced implicit feedback set S A consists of paired training triples <u, p, n>, where u represents a certain user. This embodiment combines S A with the original training set S in a certain proportion as the final training set.

[0051] Step 4: Input the enhanced item time features generated in Step 1, the original item general features (categories), spatial features (locations) in the dataset, and the enhanced user features generated in Step 2 into the encoding layer to generate high-dimensional vectors, and then use a linear layer to reduce the dimension.

[0052] The item features and user features output by the large model are in text form. To be able to fuse with traditional recommendation models, these texts need to be encoded into vector representations in this implementation. Since the large model has excellent context awareness ability and can learn rich semantic information, the large model is used as the encoding layer in this implementation. It is represented formulaically as follows:

[0053]

[0054]

[0055] Among them, c i , s i , t i represent the three categories of encoded item features, and c u , s u , t u represent the three categories of encoded user features, which are embedding vectors of shape 1×h.

[0056] Since the vectors encoded by the large model have too high a dimension and cannot be directly fused with the recommendation model. Therefore, this implementation uses a linear layer with dropout to reduce its dimension. While capturing key information, it can also reduce the risk of overfitting. The formula is as follows:

[0057]

[0058]

[0059] Among them, , , represent the three categories of item features after dimension reduction, represents the three categories of user features after dimension reduction. They are embedding vectors of shape 1×l, where l << h.

[0060] Step 5: Input the dimension-reduced feature vectors obtained in Step 4 and the id embedding vectors generated by the basic recommendation model into the feature fusion layer, and obtain a comprehensive representation through weighted fusion.

[0061] Traditional recommendation models, such as NGCF and GCMC, use collaborative filtering mechanisms and multi-layer neural networks to aggregate information. Through multi-layer propagation mechanisms, the id embedding vector Id i of the item and the id embedding vector Id u of the user can be obtained.

[0062] This implementation inputs the item features, user features, and Id after dimensionality reduction i and Id u into the feature fusion layer together, and obtains the comprehensive representation through weighted fusion. The specific calculation formula is as follows:

[0063]

[0064]

[0065] wherein, respectively represent the comprehensive item representation and user representation. This implementation uses to adjust the fusion weight of spatio-temporal features, and uses to adjust the fusion weight of ordinary features. In addition, this implementation uses L2 normalization to process the low-dimensional feature vectors to eliminate the scale differences between different features and improve numerical stability.

[0066] Step 6: Train the recommendation model using the BPR loss function. Since not all the enhanced sample pairs generated in Step 3 are reliable interactions and may contain noise, this implementation uses a parameter to control its fusion ratio with the original training set to generate the final training set.

[0067] Input the comprehensive representation generated in Step 5 into the BPR loss function without sample pruning, and merge the enhanced samples generated in Step 3 and the original training set as the final training set. This step will generate Loss1, and the specific calculation formula is as follows:

[0068]

[0069] wherein, the training triple <u, p, n> is selected from S∪S A . Since the samples generated by the user-item enhancement module are not all reliable interactions and may still contain noise, using all of them for model training will damage the performance. Therefore, this implementation needs to control the number of samples in S A .

[0070] Specifically, . Wherein, represents the enhanced sample fusion rate, which controls the proportion of enhanced samples in the final BPR training set, B represents the batch size, which controls the number of samples used to train the model in each iteration. . Wherein, is the comprehensive user representation calculated in Step 5, is the comprehensive item representation of the positive sample, is the comprehensive item representation of the negative sample, respectively represent the scores of the user for the positive sample and the negative sample, The non-linear activation function is sigmoid. To prevent overfitting, this implementation uses to control the L2 regularization strength .

[0071] Step 7: For the dimension-reduced item and user features generated in Step 4, select the time features and spatial features among them, input them into the BPR layer, and adopt a sample pruning strategy to emphasize relevant supervision signals, as shown in Figure 3 .

[0072] Specifically, to calculate Loss2, this implementation takes the time feature of the item , the spatial feature , the time feature of the user , and the spatial feature and inputs them into a BPR layer with sample pruning. It contains four BPR loss functions, which are respectively used to receive four combinations of spatio-temporal features.

[0073] To emphasize relevant supervision signals, this implementation uses spatio-temporal features to prune samples, removing some easily distinguishable training triples, which can make the optimization more stable and effective. The calculation formula is as follows:

[0074]

[0075] Among them, the function sorts the losses of each training triple in ascending order and selects the top M items. is the pruning rate. When ; When ; When When .

[0076] Each BPR loss function calculates the loss of a spatio-temporal combination, and Loss2 is obtained by adding the losses of the four BPR loss functions. The specific calculation formula is:

[0077]

[0078] Step 8: Fuse the loss calculated in Step 6 and the loss calculated in Step 7 according to a certain weight to obtain the final total loss . Use this loss function to optimize the model training. The specific calculation formula is:

[0079]

[0080] Among them For control In The proportion in

[0081] Verification example:

[0082] The dataset used in this example is the Yelp dataset provided by Yelp Inc. Since the original dataset is too large, this example selects the interaction data in 2018, filters out merchants and users with less than 6 interactions, and regards all user-merchant interactions as implicit feedback.

[0083] This example adopts three commonly used evaluation metrics for recommendation algorithms. Among them, R@20 is the proportion of actually relevant results among the top 20 recommended results in all relevant results in the test set, P@20 is the proportion of actually relevant results among the top 20 recommended results in the top 20 recommended results, and N@20 measures the ranking quality of the top 20 results of the recommended results.

[0084] Table 1 shows the performance improvement after combining two traditional recommendation models with the enhanced method of this embodiment.

[0085] Table 1 Overall recommendation performance improvement after different models are combined with this embodiment

[0086]

[0087] This embodiment improves the performance of these two models, proving the effectiveness and applicability of this embodiment. The improvement of each of their indicators exceeds 20%. Among them, GCN is better, and the improvement of each indicator exceeds 25%.

[0088] Table 2 shows the impact of the ranking of the negative sample candidate set in the user-enhanced module on the performance of the recommendation model. The recommendation model here takes (After being enhanced using this embodiment ) as an example for illustration. The ranking refers to the ranking of the negative sample candidate set in the entire item set. For example, if the dataset has 9591 items, ranking = 0.1% means selecting 5 items around the 10th place.

[0089] Table 2 Analysis of the ranking of the negative sample candidate set

[0090]

[0091] It can be seen from Table 2 that the results with ranking = 0.1% are the worst because their positions are too forward, which will lead to bias in the model learning. As the ranking increases, the recommendation performance first rises and then falls. Ranking = 10.0% achieves the best performance. Too large a ranking will lead to a performance decline because their positions are too backward, and the similarity with the positive samples is low, making it difficult to promote model optimization.

[0092] The embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings, but the present invention is not limited to the described embodiments. For those skilled in the art, without departing from the principles and spirit of the present invention, various changes, modifications, substitutions, and variations to these embodiments, including components, still fall within the protection scope of the present invention.

Claims

1. A spatio-temporal data augmentation method for a recommendation system, characterized in that It includes the following steps: Step 1: Use a large model to enhance the time features of items, obtaining enhanced item time features; Step 2: Use the inference ability of the large model to summarize the general features, time features, and spatial features of the user based on the user's interaction history, obtaining enhanced user features; Step 3: Use the large model to understand the user's interaction history and candidate set, and infer the items that the user is interested in and not interested in as positive and negative sample pairs to expand the training sample set; Step 4: Input the enhanced item time features, the original item general features and spatial features in the dataset, and the enhanced user features into the encoding layer to generate high-dimensional vectors, and then use a linear layer for dimensionality reduction; Step 5: Input the dimensionality-reduced feature vectors and the id embedding vectors generated by the basic recommendation model into the feature fusion layer, and obtain a comprehensive representation through weighted fusion; Step 6: Use the BPR loss function to train the recommendation model, input the comprehensive representation into the BPR loss function without sample pruning, and merge the enhanced samples and the original training set as the final training set to obtain the first loss function; Step 7: For the time features and spatial features in the dimensionality-reduced item and user features, input them into the BPR layer with pruning, and adopt a sample pruning strategy to obtain the second loss function; Step 8: Fuse the first loss function and the second loss function according to a certain weight to obtain the total loss function, and use this total loss function to optimize the model training.

2. The spatio-temporal data augmentation method for a recommendation system according to claim 1, wherein, In the above Step 1, enhancing the item time features specifically includes: Concatenate the specific business hours of the item with the item prompt template, where the item prompt template includes a task description, specific business hours, and output format description; Input the concatenated prompt into the large model to generate enhanced time features, and replace the time features of the corresponding items in the original dataset.

3. The spatio-temporal data augmentation method for a recommendation system according to claim 1 or 2, characterized in that In the above Step 2, summarizing the user features specifically includes: Replace the specific business hours of each item in the user interaction history with the enhanced time features in Step 1; Concatenate the replaced interaction history with the user prompt template, where the user prompt template includes a task description, user interaction history, and output format description; Input the concatenated prompt into the large model to generate enhanced general features, time features, and spatial features.

4. The spatio-temporal data augmentation method for a recommendation system according to claim 1, characterized in that In the above Step 3, generating positive and negative sample pairs specifically includes: Concatenate the user's interaction history and candidate set with the user-item prompt template, where the user-item prompt template includes a task description, interaction history, candidate set, and output format description; Use the basic recommendation model to score all items and rank them in descending order, select the top five items as the positive sample candidate set, and select several items from the items ranked in the top 10% as the negative sample candidate set; Input the concatenated prompt into the large model to generate positive and negative sample pairs.

5. The spatio-temporal data augmentation method for a recommendation system according to claim 1, characterized in that In the above Step 4, encoding and dimensionality reduction specifically include: Use the large model as the encoding layer to encode the enhanced text features into high-dimensional vectors; Use a linear layer with dropout to perform dimensionality reduction on the high-dimensional vectors to obtain low-dimensional feature vectors.

6. The spatio-temporal data augmentation method for a recommendation system according to claim 1, wherein In the above Step 5, it also includes using L2 normalization to process the low-dimensional feature vectors to eliminate the scale differences between different features.

7. The spatio-temporal data augmentation method for a recommendation system according to claim 6, wherein In step 6, the specific use of the BPR loss function is as follows: Input the comprehensive representation generated in step 5 into the BPR loss function without sample pruning; Merge the enhanced samples and the original training set in a certain proportion as the final training set; Use to control the proportion of enhanced samples in the training set to prevent the influence of noise on the training of the recommendation model.

8. The spatio-temporal data augmentation method for a recommendation system according to claim 1, characterized in that In step 7, the specific sample pruning strategy is as follows: Input the time and space features of the items and users after dimensionality reduction into the BPR layer with pruning; Use the time and space features to prune the samples and remove some easily distinguishable training triples; Process four different combinations of time and space features through four BPR loss functions respectively. After each BPR loss function processes, adopt the sample pruning strategy to obtain the second loss function.

9. The spatio-temporal data augmentation method for a recommendation system according to claim 1, wherein In step 8, the specific fusion loss function is as follows: Fuse the first loss function calculated in step 6 and the second loss function calculated in step 7 according to certain weights to obtain the total loss function; Use the total loss function to optimize the training of the recommendation model to improve the performance of the recommendation system.

10. A recommendation system, characterized in that, Integrate the spatio-temporal data augmentation method described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Distributed intelligent retrieval method and system for spatio-temporal data

    CN118838937A

  • Interactive data analysis method based on large model

    CN118886415A