Multi-view comparative learning recommendation method and system fusing semantic perception

By integrating a semantically aware multi-view contrastive learning method, the problem of insufficient utilization of semantic information in recommendation systems is solved, achieving deep integration of structure and semantics, and improving the performance of recommendation systems in scenarios with sparse data and cold start.

CN120873272APending Publication Date: 2025-10-31TIANJIN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510735804.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing recommendation systems suffer from insufficient utilization of semantic information, poor multi-view fusion performance, weak cold start performance, imprecise selection of negative samples, and a lack of hierarchical semantic enhancement mechanisms, resulting in performance limitations in complex recommendation scenarios.

Method used

A systematic semantic enhancement mechanism is constructed. Semantic features are extracted through the Sentence-BERT model, and enhanced structures and feature views are generated by combining multi-head attention mechanism and progressive fusion strategy. The view is processed by the parameter-sharing LightGCN encoder, and negative samples are filtered by semantic similarity threshold. The model representation is optimized by combining InfoNCE variant loss function and BPR recommendation loss.

Benefits of technology

It achieves deep fusion of structural and semantic information, improves the recommendation performance of the model in data-sparse and cold-start scenarios, and enhances the accuracy and efficiency of the recommendation system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120873272A_ABST
    Figure CN120873272A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-view comparative learning recommendation method fusing semantic perception. The multi-view comparative learning recommendation method comprises the steps of semantic feature preprocessing, semantic perception view generation and semantic optimization comparative learning. The system mainly comprises four core modules: a semantic feature preprocessing module, a semantic perception view generator, a semantic optimization comparison learning module and a collaborative optimization recommendation generation module. The method is widely applied to Internet platforms and application systems requiring personalized recommendation services, such as e-commerce, online videos, social networks, news pushing and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence and data mining technology, specifically to the cross-application of recommendation systems, graph neural networks, and contrastive learning, and more specifically to a multi-view contrastive learning recommendation method and system that integrates semantic awareness. Background Technology

[0002] In today's digital age, recommender systems, as a key technology for filtering and personalizing information, have become a core component of e-commerce, social media, and content platforms. The development of recommender systems has evolved from content-based filtering to collaborative filtering, and then to deep learning. Among deep learning recommender models, graph neural network-based methods have attracted considerable attention due to their ability to naturally model user-item interaction networks, effectively aggregating neighborhood information through message passing mechanisms to obtain richer structural features.

[0003] Graph neural network recommendation methods treat users and items as nodes and interactions as edges, forming a bipartite graph structure. They then utilize graph algorithms to uncover implicit user preferences and item relevance. LightGCN, as a representative of lightweight graph convolutional networks, significantly reduces computational complexity while maintaining powerful expressive capabilities by removing feature transformations and nonlinear activation operations from traditional GCNs, and has become the mainstream baseline for current graph neural network-based recommendation systems.

[0004] Meanwhile, contrastive learning, as an important paradigm of self-supervised learning, has achieved significant success in computer vision and natural language processing. By constructing pairs of positive and negative samples, contrastive learning enables models to learn to distinguish between similar and dissimilar entities, acquiring more discriminative feature representations. Introducing contrastive learning into recommender systems can effectively alleviate data sparsity and cold-start problems without relying on additional labeled data.

[0005] Furthermore, the development of pre-trained language models, especially sentence-level representation models such as Sentence-BERT, has provided powerful tools for the accurate extraction of semantic information in recommender systems. These models can extract rich semantic features from textual data such as user reviews and item descriptions, providing recommender systems with deep knowledge that structured data cannot express.

[0006] Although graph neural networks and contrastive learning have brought new opportunities for the development of recommender systems, existing research still has the following main limitations:

[0007] (1) Insufficient integration of structural and semantic information. Most graph neural network recommendation methods only focus on the topological structure of the interaction graph, failing to fully utilize the rich semantic information contained in the text. Although existing research has begun to explore integrating text features into recommendation models, most adopt simple splicing or fusion strategies, lacking a deep interaction and collaborative mechanism between semantics and graph structure. Traditional graph neural networks mainly rely on the topological structure constructed from explicit interaction data, failing to fully utilize semantic content such as user reviews and item descriptions that contain important user preferences and item attribute information.

[0008] (2) Contrastive learning view construction lacks semantic guidance. Existing recommendation methods based on contrastive learning mostly rely on random data augmentation (such as edge deletion and feature masking) to generate multiple views. This approach is not only highly random and poorly interpretable, but also fails to introduce domain knowledge to guide the sample augmentation process, easily introducing noise and spurious negative examples. Traditional graph contrastive learning methods mainly rely on random data augmentation (such as random edge deletion and random masking) to generate different views. This random augmentation strategy lacks semantic guidance and is difficult to retain key semantic information.

[0009] (3) The negative sample selection strategy is not refined enough. Existing methods either select negative sample pairs completely randomly or only based on graph structure characteristics when constructing negative sample pairs, failing to fully consider the semantic similarity between samples, resulting in the neglect of some high-quality negative samples. High-quality negative samples are crucial for the model to learn discriminative feature representations, but traditional random sampling may misclassify semantically similar items as negative samples.

[0010] (4) Lack of hierarchical semantic enhancement mechanisms. Existing methods often introduce semantic information only at a single level (such as data preprocessing or model training), lacking a systematic semantic enhancement mechanism from data to model. Semantic information can enhance the recommendation system at multiple levels, including graph structure enhancement at the data layer, representation enhancement at the feature layer, and training strategy optimization at the model layer, but existing methods lack this hierarchical design.

[0011] These shortcomings severely limit the performance of graph neural network recommendation methods in complex recommendation scenarios, especially under cold start conditions. More effective semantic enhancement mechanisms and contrastive learning strategies need to be designed to achieve deep integration of structural and semantic information. Summary of the Invention

[0012] This invention aims to address the technical problems in existing recommendation systems, such as insufficient utilization of semantic information, poor multi-view fusion performance, and weak cold-start performance. It provides an efficient multi-view comparative learning recommendation method and system that integrates semantic awareness, with the following specific objectives:

[0013] 1) Construct a systematic semantic enhancement mechanism to achieve deep integration of semantic information and graph structure;

[0014] 2) Design an effective multi-view contrast learning framework to improve the model's representation learning ability;

[0015] 3) Addressing the performance bottlenecks of traditional recommendation methods in scenarios with sparse data and cold starts;

[0016] 4) Provides hierarchical semantic enhancement strategies, optimizing the entire process from data preprocessing to model training.

[0017] The technical solution of this invention is a multi-view comparison learning recommendation method that integrates semantic awareness, comprising the following steps:

[0018] 1) Preprocess the raw interaction data, retaining active users with at least 3 interactions and popular items with at least 2 interactions, and retain at least 20% of cold start users and items in the test set;

[0019] 2) For different datasets, adopt corresponding semantic information extraction strategies, use the Sentence-BERT model to extract features from the processed text, generate sentence vector representations, and calculate the semantic similarity matrix between users and between items;

[0020] 3) Through the structural view enhancement algorithm, the original interaction graph is weighted and fused with the semantic similarity matrix, the multi-head attention mechanism is applied to optimize the graph structure, and a hybrid pruning strategy is used to construct the enhanced structural view;

[0021] 4) Through the feature view enhancement algorithm, ID embedding and semantic embedding are mapped to a unified space, and an enhanced feature view is generated by adopting a progressive fusion strategy;

[0022] 5) The LightGCN encoder with shared parameters is used to process the structure enhancement view and the feature enhancement view respectively, to construct cross-view positive sample pairs, and to filter negative samples based on semantic similarity threshold;

[0023] 6) The InfoNCE variant loss function is used to optimize the consistency and discriminativeness of multi-view representations, and the weights are dynamically adjusted by combining BPR recommendation loss and contrastive learning loss with a cosine decay strategy.

[0024] 7) Generate the final node embedding representation through weighted fusion, calculate the user's predicted score for the item based on the fused embedding, sort the uninterrupted items in descending order of the predicted score, and select the Top-K as the final recommendation list.

[0025] Furthermore, in the semantic information processing step, for the MovieLens-1M dataset, the structured features are converted into natural language sentences; for the Amazon-Electronics and Yelp2018 datasets, user reviews are used as the semantic information source, and a head truncation strategy is applied to reviews with more than 256 words.

[0026] Furthermore, in the structure view enhancement algorithm, progressive weight α is calculated so that the model prioritizes learning the original interaction structure in the early stage of training and gradually enhances semantic associations in the middle and later stages; a multi-head attention mechanism is applied to optimize the graph structure, and attention is applied directly at the graph structure level rather than the node embedding level, which more effectively captures the implicit semantic relationships between all nodes.

[0027] Furthermore, in the feature view enhancement algorithm, ID embedding and semantic embedding are mapped to a unified 64-dimensional space through channel projection technology, and an enhanced feature view is generated by a progressive fusion strategy.

[0028] Furthermore, in the semantic optimization contrastive learning step, a parameter-shared LightGCN encoder is used to process the structure-enhanced view and the feature-enhanced view respectively, construct cross-view positive sample pairs, and filter negative samples based on semantic similarity thresholds to avoid the problem of false negative samples.

[0029] Furthermore, in the dynamic joint optimization step, the weights are dynamically adjusted using a cosine decay strategy, combining the BPR recommendation loss and the contrastive learning loss, where the initial weight λ0 = 0.1 and the regularization coefficient β = 0.001.

[0030] Furthermore, in the final recommendation generation step, the final node embedding representation is generated by weighted fusion, the user's predicted score for the item is calculated based on the fused embedding, and the uninteracted items are sorted in descending order of the predicted score. The Top-K items are selected as the final recommendation list.

[0031] Another technical solution proposed in this invention is a multi-view comparison learning recommendation system that integrates semantic awareness, including: a semantic feature preprocessing module, used to preprocess the original interaction data and extract semantic features;

[0032] A semantically aware view generator is used to generate enhanced structure views and enhanced feature views;

[0033] The semantic optimization contrastive learning module is used to optimize the consistency and discriminability of multi-view representations;

[0034] The collaborative optimization recommendation generation module is used to generate the final recommendation list.

[0035] The semantic feature preprocessing module includes:

[0036] The data preprocessing unit is used to retain active users and popular items, and to retain cold-start users and items in the test set;

[0037] The semantic information processing unit is used to extract semantic features and generate sentence vector representations;

[0038] The semantic similarity calculation unit is used to calculate the semantic similarity matrix between users and between items.

[0039] The semantically aware view generator includes:

[0040] Structural view enhancement unit, used to generate enhanced structural views;

[0041] Feature view enhancement unit, used to generate enhanced feature views.

[0042] The semantic optimization contrastive learning module includes:

[0043] A dual-path encoder unit is used to process structurally enhanced views and feature-enhanced views;

[0044] Sample building unit, used to build positive sample pairs across views and filter negative samples;

[0045] The comparison loss calculation unit is used to optimize the consistency and distinguishability of multi-view representations;

[0046] The dynamic joint optimization unit is used to dynamically adjust weights and optimize the model.

[0047] The collaborative optimization recommendation generation module includes:

[0048] A multi-view embedding fusion unit is used to generate the final node embedding representation;

[0049] The recommendation generation unit is used to generate the final recommendation list.

[0050] Beneficial effects: Deep learning-based intelligent recommendation technology involves key technologies such as graph structure data processing, multi-view representation learning, semantic feature extraction, and contrastive learning optimization. It is widely used in internet platforms and application systems that require personalized recommendation services, such as e-commerce, online video, social networks, and news push.

[0051] like Figure 2 As shown, five experiments were conducted on the datasets MovieLens-1M, Amazon Reviews, and Yelp2018, and the average results were taken. The evaluation metrics were Recall@20 and NDCG@20. Compared with the baseline methods LightGCN, SGL, and SimGCL, the results showed performance improvements. Attached Figure Description

[0052] Figure 1 A framework for recommending methods that integrate semantic-aware multi-view contrastive learning;

[0053] Figure 2 Performance comparison chart;

[0054] Figure 3 Here is the pseudocode for Algorithms 1-3;

[0055] Figure 4 This is the pseudocode for Algorithms 4-6. Detailed Implementation

[0056] The present invention will be further described below with reference to the accompanying drawings.

[0057] The multi-view contrastive learning recommendation method of this invention, which integrates semantic awareness, adopts a hierarchical semantic enhancement architecture, such as... Figure 1 As shown, the system mainly consists of four core modules: a semantic feature preprocessing module, a semantic-aware view generator, a semantic optimization and comparative learning module, and a collaborative optimization recommendation generation module.

[0058] This method specifically includes the following steps:

[0059] Semantic feature preprocessing

[0060] Data preprocessing and cold start filtering: First, the raw interaction data is preprocessed. In the training set, active users with at least 3 interactions and popular items with at least 2 interactions are retained to ensure that the model can learn sufficient interaction patterns. At the same time, in the test set, at least 20% of cold start users (interaction count ≤ 2) and items (interaction count ≤ 1) are intentionally retained to evaluate the model's performance in cold start scenarios.

[0061] Semantic information processing: Appropriate semantic information extraction strategies are adopted for different datasets. For the MovieLens-1M dataset, structured features (such as "Action|Comedy") are converted into natural language sentences (such as "This movie combines Action and Comedy genres"). For the Amazon-Electronics and Yelp2018 datasets, user reviews are used as the source of semantic information, and a head truncation strategy is applied to reviews with more than 256 words.

[0062] Semantic embedding generation: The Sentence-BERT model (all-MiniLM-L6-v2 version) is used to extract features from the processed text, generating 384-dimensional sentence vector representations. Based on these semantic embeddings, semantic similarity matrices between users and between items are calculated, using cosine similarity as the metric.

[0063]

[0064] The similarity matrices of users and items are combined into a complete semantic similarity matrix. The complete algorithm pseudocode is shown in Algorithm 1.

[0065] Semantic-aware view generation

[0066] Structural view enhancement algorithm: First, progressive weights α are calculated, allowing the model to prioritize learning the original interaction structure in the early stages of training, and gradually enhancing semantic associations in the later stages. Then, the original interaction graph A_raw is weighted and fused with the semantic similarity matrix S_sbert.

[0067] A base = (1-α)·A raw +α·S sbert

[0068]

[0069] Next, a multi-head attention mechanism is applied to optimize the graph structure. Attention is applied directly at the graph structure level rather than the node embedding level, which more effectively captures the implicit semantic relationships between all nodes.

[0070]

[0071] Where H = 3 represents the number of attention heads, and each attention head is calculated as follows:

[0072] Finally, a hybrid pruning strategy is applied. First, candidate edges are filtered by a static threshold. Then, the Top-K edges with the highest similarity are selected from the candidate edges to construct the final enhanced structural view. The complete algorithm pseudocode can be found in Algorithm 2.

[0073] Feature view augmentation algorithm: Maps ID embeddings (obtained directly from the LightGCN model) and semantic embeddings (obtained from the Sentence-BERT pre-trained model) to a unified 64-dimensional space using a multi-channel projection technique.

[0074]

[0075] Then, a progressive fusion strategy is used to generate an enhanced feature view, where λ employs a progressive weighting strategy similar to that used for structural enhancement.

[0076] The complete pseudocode for the algorithm can be found in Algorithm 3.

[0077] Semantic optimization comparative learning

[0078] Dual-path encoder design: A parameter-sharing LightGCN encoder is used to process the structure-enhanced view and the feature-enhanced view separately. The structure-view path inputs the enhanced graph structure A_augmented and the original feature X_raw into LightGCN to obtain the structure-view embedding Z_struct. The feature-view path inputs the original graph structure A_raw and the enhanced feature X_augmented into LightGCN to obtain the feature-view embedding Z_feat.

[0079] Z struct =LightGCN(Aaug (X)

[0080] Z feat =LightGCN(A, X) aug )

[0081] The complete pseudocode for the algorithm can be found in Algorithm 4.

[0082] Semantic-guided sample construction: Constructing cross-view positive sample pairs by combining the representations of the same node in two views into positive sample pairs; For negative sample selection, filtering is based on semantic similarity thresholds to avoid the problem of false negative samples.

[0083]

[0084] Here, θ is set to the 70th percentile value of the semantic similarity matrix. The complete pseudocode for the algorithm can be found in Algorithm 5.

[0085]

[0086] Contrast Loss Calculation: The InfoNCE variant loss function is used to optimize the consistency and discriminability of multi-view representations.

[0087] Where s_i,j represents the cosine similarity, τ is the temperature parameter, and the overall contrast loss is the average of the contrast losses of all nodes:

[0088]

[0089] Dynamic Joint Optimization: Combining BPR recommendation loss and contrastive learning loss, a cosine decay strategy is used to dynamically adjust the weights: where λ_0 = 0.1 is the initial weight, and β = 0.001 is the regularization coefficient.

[0090] Final recommendation generation

[0091] z final =α·Z struct +(1-α)·Z feat

[0092]

[0093] Multi-view embedding fusion: Generating the final node embedding representation through weighted fusion.

[0094] Where α is the fusion weight. The predicted score for user u on item i is calculated based on the fusion embedding:

[0095] For each user, uninterrupted items are sorted in descending order of predicted score, and the Top-K items are selected as the final recommendation list. See Algorithm 6 for the complete pseudocode.

Claims

1. A multi-view comparison learning recommendation method integrating semantic awareness, characterized in that, Includes the following steps: 1) Preprocess the raw interaction data, retaining active users with at least 3 interactions and popular items with at least 2 interactions, and retain at least 20% of cold start users and items in the test set; 2) For different datasets, adopt corresponding semantic information extraction strategies, use the Sentence-BERT model to extract features from the processed text, generate sentence vector representations, and calculate the semantic similarity matrix between users and between items; 3) Through the structural view enhancement algorithm, the original interaction graph is weighted and fused with the semantic similarity matrix, the multi-head attention mechanism is applied to optimize the graph structure, and a hybrid pruning strategy is used to construct the enhanced structural view; 4) Through the feature view enhancement algorithm, ID embedding and semantic embedding are mapped to a unified space, and an enhanced feature view is generated by adopting a progressive fusion strategy; 5) The LightGCN encoder with shared parameters is used to process the structure enhancement view and the feature enhancement view respectively, to construct cross-view positive sample pairs, and to filter negative samples based on semantic similarity threshold; 6) The InfoNCE variant loss function is used to optimize the consistency and discriminativeness of multi-view representations, and the weights are dynamically adjusted by combining BPR recommendation loss and contrastive learning loss with a cosine decay strategy. 7) Generate the final node embedding representation through weighted fusion, calculate the user's predicted score for the item based on the fused embedding, sort the uninterrupted items in descending order of the predicted score, and select the Top-K as the final recommendation list.

2. The multi-view comparison learning recommendation method with fused semantic awareness according to claim 1, characterized in that, In the semantic information processing step, for the MovieLens-1M dataset, the structured features are converted into natural language sentences; for the Amazon-Electronics and Yelp2018 datasets, user reviews are used as the semantic information source, and a head truncation strategy is applied to reviews with more than 256 words.

3. The multi-view comparison learning recommendation method with fused semantic awareness according to claim 1, characterized in that, In the structure view enhancement algorithm, progressive weights α are calculated so that the model prioritizes learning the original interaction structure in the early stage of training and gradually enhances semantic associations in the middle and later stages. By applying a multi-head attention mechanism to optimize the graph structure, attention is applied directly at the graph structure level rather than the node embedding level, which more effectively captures the implicit semantic relationships between all nodes.

4. The multi-view comparison learning recommendation method with fused semantic awareness according to claim 1, characterized in that, In the feature view enhancement algorithm, ID embedding and semantic embedding are mapped to a unified 64-dimensional space through channel projection technology, and an enhanced feature view is generated by a progressive fusion strategy.

5. The multi-view comparison learning recommendation method with fused semantic awareness according to claim 1, characterized in that, In the semantic optimization contrastive learning step, the parameter-shared LightGCN encoder processes the structure enhancement view and feature enhancement view respectively, constructs cross-view positive sample pairs, and filters negative samples based on semantic similarity threshold to avoid the problem of false negative samples.

6. The multi-view comparison learning recommendation method with fused semantic awareness according to claim 1, characterized in that, In the dynamic joint optimization step, the weights are dynamically adjusted using a cosine decay strategy, combining the BPR recommendation loss and the contrastive learning loss, where the initial weight λ0 = 0.1 and the regularization coefficient β = 0.

001.

7. The multi-view comparison learning recommendation method with fused semantic awareness according to claim 1, characterized in that, In the final recommendation generation step, the final node embedding representation is generated by weighted fusion, the user's predicted score for the item is calculated based on the fused embedding, and the uninterrupted items are sorted in descending order of the predicted score. The Top-K items are selected as the final recommendation list.

8. A multi-view comparison learning recommendation system integrating semantic awareness, characterized in that, include: The semantic feature preprocessing module is used to preprocess the raw interaction data and extract semantic features; A semantically aware view generator is used to generate enhanced structure views and enhanced feature views; The semantic optimization contrastive learning module is used to optimize the consistency and discriminability of multi-view representations; The collaborative optimization recommendation generation module is used to generate the final recommendation list.

9. The multi-view comparison learning recommendation system with fused semantic awareness according to claim 8, characterized in that, The semantic feature preprocessing module includes: The data preprocessing unit is used to retain active users and popular items, and to retain cold-start users and items in the test set; The semantic information processing unit is used to extract semantic features and generate sentence vector representations; - Semantic similarity calculation unit, used to calculate the semantic similarity matrix between users and between items.

10. The multi-view comparison learning recommendation system with fused semantic awareness according to claim 8, characterized in that, The semantically aware view generator includes: Structural view enhancement unit, used to generate enhanced structural views; Feature view enhancement unit, used to generate enhanced feature views; The semantic optimization contrastive learning module includes: A dual-path encoder unit is used to process structurally enhanced views and feature-enhanced views; Sample building unit, used to build positive sample pairs across views and filter negative samples; The comparison loss calculation unit is used to optimize the consistency and distinguishability of multi-view representations; Dynamic joint optimization unit, used to dynamically adjust weights and optimize the model; The collaborative optimization recommendation generation module includes: A multi-view embedding fusion unit is used to generate the final node embedding representation; The recommendation generation unit is used to generate the final recommendation list.