Graph mixed learning friend recommendation method for sparse social network

By generating diverse and denoised contrast views for comparative learning, combining graph convolutional models to update feature embeddings and filter high-confidence pseudo links, the problem of inaccurate recommendations in sparse social networks is solved, achieving more accurate friend recommendations.

CN120655448AActive Publication Date: 2025-09-16CHENGDU UNIV OF INFORMATION TECH

Patent Information

Application Number
CN202510775387.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-09-16
Estimated Expiration
2045-06-11

AI Technical Summary

Technical Problem

Existing technologies have difficulty in fully exploring potential friend relationships in sparse social networks, resulting in inaccurate recommendation results.

Method used

By generating diverse social network comparison views and denoised comparison views for comparative learning, the feature embedding is updated in combination with the graph convolutional model, a set of pseudo links with high confidence is screened out, and the model parameters are iteratively optimized.

Benefits of technology

It improves the accuracy of friend recommendations in sparse social network environments, effectively mines users' social and interest characteristics, and alleviates the problem of sparse social networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120655448A_ABST
    Figure CN120655448A_ABST
Patent Text Reader

Abstract

The invention discloses a sparse social network-oriented graph mixed learning friend recommendation method, which comprises the following steps of: establishing an original graph structure data set, and respectively generating a diversified social network comparison view and a de-noising comparison view; performing comparative learning to obtain an initial feature embedding representation; updating initial feature embedding; fusing the updated feature embedding representation, and establishing a user node link probability prediction model; performing friend prediction, generating a plurality of user-user pseudo links, and obtaining a candidate pseudo link set; performing reliability screening to obtain a final pseudo-link set; obtaining a new graph structure data set; and replacing the original graph structure data set with the new graph structure data set, and repeating the steps until a new user-user pseudo link cannot be generated. The method is used for solving the problems that in the prior art, potential friend relations are difficult to fully excavate in the face of sparse social networks, and recommendation results are not accurate, and the purpose of better adapting to friend recommendation requirements in the sparse social network environment is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of friend recommendation in social networks, and in particular to a graph hybrid learning friend recommendation method for sparse social networks. Background Art

[0002] In today's digital age, social networks have become a crucial part of people's lives, serving as a vital platform for information dissemination, interpersonal communication, and social interaction. As a key component in the formation and evolution of social networks, friend recommendation is becoming increasingly important. It not only helps users discover potential friends, expand their social circles, and access more information resources, but also helps social platforms increase user engagement and retention.

[0003] In the real world, social networks are often highly sparse. When constructing a social graph with users as nodes and their friendships as edges, the result is often a sparse graph with only a few links. Furthermore, some users also have very limited interactions with items. While existing hybrid friend recommendation methods have achieved some success, they still struggle to fully tap into potential friendships, resulting in inaccurate recommendations. Summary of the Invention

[0004] The present invention provides a graph hybrid learning friend recommendation method for sparse social networks to solve the problems in the existing technology that it is difficult to fully explore potential friend relationships and the recommendation results are inaccurate when facing sparse social networks, so as to better adapt to the friend recommendation needs in sparse social network environments.

[0005] The present invention is achieved through the following technical solutions:

[0006] A graph hybrid learning friend recommendation method for sparse social networks includes the following steps:

[0007] S1. Establishing an original graph structure dataset, wherein the original graph structure dataset includes user-user relationships and user-item relationships;

[0008] S2. Based on the original graph structure dataset, generating a diversified social network comparison view and a denoised comparison view respectively;

[0009] S3. Perform comparative learning based on the diversified social network comparative view and the denoised comparative view to obtain an initial feature embedding representation; the initial feature embedding representation includes a social embedding representation and an interest embedding representation;

[0010] S4. Updating the initial feature embedding through a graph convolutional model to obtain an updated feature embedding representation;

[0011] S5. Fusing the updated feature embedding representation to establish a user node link probability prediction model;

[0012] S6. Use the user node link probability prediction model to perform friend prediction on all users in the original graph structure data set to generate a number of user-user pseudo links; remove existing user-user relationships to obtain a candidate pseudo link set;

[0013] S7, screening the candidate pseudo link set for reliability to obtain a final pseudo link set;

[0014] S8. Add the final pseudo link set to the original graph structure dataset to obtain a new graph structure dataset;

[0015] S9. Replace the original graph structure dataset with the new graph structure dataset, and return to step S1 until no new user-user pseudo links can be generated; and perform friend recommendations using the current graph structure dataset.

[0016] During their research, the inventors discovered that existing technologies often struggle to fully exploit their advantages when dealing with sparse social networks. They may even introduce unnecessary noise, which can affect the learning performance of recommendation models and hinder the full exploitation of potential friend relationships, leading to inaccurate recommendation results. To overcome these issues, the present invention proposes a graph hybrid learning friend recommendation method for sparse social networks. This method first establishes an original graph structure dataset consisting of user-user and user-item relationships, then generates two different comparative views: a diversified social network comparative view and a denoised comparative view. Based on comparative learning techniques, an initial feature embedding representation is obtained, comprising a social embedding representation and an interest embedding representation. Subsequently, the user's social embedding and interest embedding are propagated and updated through a graph convolutional layer, respectively, to obtain an updated feature embedding representation. The original graph structure dataset is then used to train friend recommendation for the user, yielding an initial model. Subsequently, friend prediction is performed on the user, i.e., the probability of a link existing between user nodes is predicted, generating a number of user-user pseudo links. Since not all pseudo links are reliable, this application also performs reliability screening on these pseudo links to obtain a final set of pseudo links. If the final pseudo-link set is empty, the graph structure dataset at this time will be used as the friend recommendation dataset; otherwise, the final pseudo-link set will be added to the original graph structure dataset to form a new user-user relationship and user-item relationship graph, and the iterative training of the model will continue.

[0017] Compared with the existing technology, this application: by generating two different comparative views and conducting comparative learning based on these two views, it can better enable friend recommendations; in the feature embedding part, it fully considers the interaction of multi-dimensional features between users and between users and items; in the continuous "training-prediction-pseudo-link selection" iterative process, the model adds pseudo-links with higher confidence to the training data, thereby continuously optimizing the model parameters. This has been experimentally verified by the inventor's team. Compared with the existing technology in sparse social network environments, this application has better recommendation accuracy.

[0018] Furthermore, step S2 specifically includes:

[0019] S201: Extracting a dataset containing user-user relationships from the original graph structure dataset, defined as a dataset ;

[0020] S202, the data set Input into the graph generator to obtain a diverse comparative view of social networks;

[0021] S203, the data set Input to the image denoiser to get the denoised comparison view.

[0022] This solution can be understood as employing a graph generator and a graph denoiser to form a comparative view generator, thereby generating the desired comparative view. The graph generator and graph denoiser can be implemented using existing technologies and are not specifically limited here. The graph generator is responsible for generating a comparative view of a diverse social network structure using a specific augmentation strategy to enrich the data representation and enhance the model's generalization capabilities. The graph denoiser is responsible for removing noisy edges through learned parameters to generate a cleaner comparative view.

[0023] Furthermore, the graph generator uses the following formula to determine the loss function:

[0024]

[0025]

[0026]

[0027] Where: is the loss function of the graph generator; is the first BPR loss; To repair losses; stands for weighted sum; is the ratio coefficient used to control the number of retained samples; is the regularization term weight; represents the regularization term; represents the dataset enhanced by the graph generator; u represents the user; v represents the item; 、 Represent positive sample nodes and negative sample nodes respectively; function It means sorting and selecting the first X values; Represents the sigmoid activation function; Representative User Positive sample node 's prediction score; Representative User For negative sample nodes 's prediction score; Representative dataset The node set of represent A subset of represents the original enhancement feature; Represents the enhanced features after mask operation; is the scaling factor.

[0028] Those skilled in the art should understand that the positive sample nodes refer to other nodes that have similar structures, attributes or semantic relationships with the current node, and the negative sample nodes are nodes that are significantly different from the current node in structure and semantics. This solution clearly defines the loss function of the graph generator and uses enhanced supervisory signals. Optimizing the loss can improve the model's ability to understand user preferences and the accuracy of recommendation results while preventing overfitting.

[0029] Furthermore, the image denoiser uses the following formula to determine the loss function:

[0030]

[0031]

[0032]

[0033] Where: is the loss function of the image denoiser; is the denoising loss; is the second BPR loss; Represents the regularization strength used to control weight decay; represents the regularization term; Representation dataset The number of layers in the middle graph; l represents the node in the lth layer; Representation dataset midpoint and connected edges; express The set of middle edges; Indicates the Nodes in the layer and the score of the edge between; represent The cumulative distribution function of represents the sigmoid activation function, Representative Model parameters corresponding to the layer; Represents a triplet describing a user For items and items preference relationship; 、 Represents users For items 's rating.

[0034] It should be understood by those skilled in the art that the model parameters These are the parameters used by the model at each layer to calculate the likelihood of user-item interactions. Model parameters are used to evaluate the quality of model predictions and adjust the model for better learning. This solution explicitly defines the loss function for the graph denoiser, enabling it to better combat noisy data and provide high-quality training signals.

[0035] Furthermore, in the process of comparative learning, the objective function is determined by the following formula:

[0036]

[0037]

[0038]

[0039] Where: represents the objective function of contrastive learning; represents the contrast loss of the user node; Represents the contrast loss of item nodes; Represents the i-th user node; represent The user collection in ; represents the cosine similarity function; 、 All are temperature parameters; Represents the user positive sample pair; Represents user negative sample pairs; Represents the i-th item node; represent The collection of items in ; Represents the positive sample pair of the item; represents an item-negative sample pair.

[0040] This approach combines the contrastive loss of user nodes with the contrastive loss of item nodes to derive the modeling objective function for contrastive learning. The temperature parameter is used to control the flatness of the function structure distribution to avoid vanishing or exploding gradients.

[0041] Furthermore, step S4 specifically includes:

[0042] S401, input the initial feature embedding into the graph convolution model and perform normalization processing;

[0043] S402, inputting the normalized features into the Bi-Li+ module to obtain a first feature vector;

[0044] S403: Input the normalized features into the SENet+ module, and perform compression, excitation, reweighting, and fusion operations in sequence to obtain a second feature vector;

[0045] S404, concatenating the first eigenvector and the second eigenvector;

[0046] S405: The concatenated feature vector is further propagated through several graph convolutional layers to obtain an updated feature embedding representation.

[0047] In the process of propagating user social embedding and interest embedding, this solution allows neighbor nodes to undergo feature crossover when aggregating, which can more effectively capture high-order interaction information between user and item features. Specifically:

[0048] Normalize the input feature embedding to enhance the stability of model training.

[0049] Then, the eigenvector passes through the Bi-Li+ module to obtain the first eigenvector ; The Bi-Li+ module replaces the Hadamard product with the inner product, reducing the parameters of feature interaction from a d-dimensional vector to 1 bit, thereby reducing the model size.

[0050] The SENet+ module is used to dynamically learn the importance of features. After four stages of compression, excitation, reweighting and fusion, the output is obtained. as the second feature vector. The compression operation provides a foundation for subsequent feature weight learning. The excitation operation refines domain-level attention to a more fine-grained bit-level attention. Using two fully connected layers to learn weights, the model relearns more detailed weights for each dimension of the feature, thereby more accurately determining feature importance. The reweighting operation increases the weights of important features and decreases the weights of unimportant features. The fusion operation provides richer information for subsequent feature interactions and model predictions.

[0051] Finally, the outputs of Bi-Li+ and SENet+ are concatenated to obtain the updated feature embedding representation q u ,Right now .

[0052] This solution can fully consider the interaction of multi-dimensional features between users and between users and items, which is conducive to improving the accuracy of friend recommendations.

[0053] Furthermore, the user node link probability prediction model is:

[0054]

[0055] Where: a and b represent users ,user ; Representative User and The probability that a link exists between them; Represents the sigmoid activation function; stands for weighted sum; represents the dot product; Representative User Social embedding representation of Representative User Interest embedding representation; Representative User Social embedding representation of Representative User Interest embedding representation.

[0056] After updating the social embedding representation and interest embedding representation in the initial feature embedding in step S4, updated social embedding representations and updated interest embedding representations are obtained. To predict the link probability between users, these two representations need to be effectively fused. After experimentally comparing various fusion methods, the weighted summation strategy demonstrated the best results, resulting in the user-node link probability prediction model defined in this solution.

[0057] Furthermore, step S7 specifically includes:

[0058] S701: Perform intra-instance link selection on users in the candidate pseudo link set to obtain an intra-instance candidate pseudo link set;

[0059] S702: Perform inter-instance link selection on users in the candidate pseudo link set to obtain an inter-instance candidate pseudo link set;

[0060] S703 : Obtain the final pseudo link set based on the intra-instance candidate pseudo link set and the inter-instance candidate pseudo link set.

[0061] This scheme proposes an adaptive strategy for reliability screening of candidate pseudo-link sets. Through two steps, intra-instance link selection and inter-instance link selection, high-quality pseudo-links are gradually selected. Its beneficial effects include: 1. Ensuring the inclusion of true links: Intra-instance link selection prioritizes potential links from high to low based on the predicted confidence of each unlabeled instance until the cumulative confidence exceeds a threshold. This ensures that, for each instance, links with higher confidence scores are prioritized, increasing the likelihood that true links will be included in the candidate pseudo-link set. 2. Improving link estimation accuracy: Intra-instance link selection ensures that the links in the candidate set have nearly equal confidence scores, ensuring comparable link estimation accuracy for each sample, reducing the number of low-confidence links in the candidate pseudo-link set, and improving overall link quality. 3. Improving category balance: Inter-instance link selection ranks the confidence scores of all instances in each category from a category perspective and selects the top percentage of instances as candidate pseudo-links for that category. This ensures that a sufficient number of samples from each category are selected, preventing overselection of certain categories due to excessively high confidence and underestimation of others. ④ Enhance model learning capabilities: By combining intra-instance link selection with inter-instance link selection, unlabeled data is filtered and utilized from different perspectives, providing the model with richer and more diverse learning samples, enabling the model to learn more comprehensive feature representations and semantic information from multiple dimensions. ⑤ Achieve diverse and fair learning: This approach ensures a diverse and fair distribution of candidate pseudo-links, preventing the model from being overly influenced by certain dominant categories or high-confidence links during learning, thereby enabling the model to learn balanced knowledge from a wider range of data.

[0062] Furthermore, the candidate pseudo link set in the instance is calculated using the following formula:

[0063]

[0064] Where: is the set of candidate pseudo links within the instance; Represents a function that returns the minimum size set; j is the candidate user number; Represents a user For users Confidence of interest; is a threshold used to determine which links have high enough confidence, which can be adaptively updated as the model is trained;

[0065] The candidate pseudo link set between instances is calculated using the following formula:

[0066]

[0067] Where: is the set of candidate pseudo links between instances; Represents the function used to return a given quantile vector-valued function at ; Represents a function that sorts a vector into ascending order; A vector representing the confidence scores of all instances.

[0068] Furthermore, the final pseudo link set is obtained by the following formula:

[0069]

[0070] Where: is the final pseudo link set; is the intersection operator.

[0071] Compared with the prior art, the present invention has at least the following advantages and beneficial effects:

[0072] 1. The present invention provides a graph hybrid learning friend recommendation method for sparse social networks. By generating two different comparative views and performing comparative learning based on the two views, it can better enable friend recommendation.

[0073] 2. The present invention proposes a graph hybrid learning friend recommendation method for sparse social networks, which fully considers the interaction of multi-dimensional features between users and between users and items in the feature embedding part.

[0074] 3. The present invention proposes a graph hybrid learning friend recommendation method for sparse social networks. In the continuous iterative process of "training-prediction-pseudo-link selection", the model adds pseudo-links with higher confidence to the training data, thereby continuously optimizing the model parameters.

[0075] 4. This invention proposes a graph hybrid learning friend recommendation method for sparse social networks, which can effectively mine users' social and interest characteristics, improve the accuracy of friend recommendations, and effectively alleviate the problem of sparse social networks. BRIEF DESCRIPTION OF THE DRAWINGS

[0076] The drawings described herein are used to provide a further understanding of the embodiments of the present invention, constitute a part of this application, and do not constitute a limitation of the embodiments of the present invention. In the drawings:

[0077] Figure 1 It is a flowchart of a specific embodiment of the present invention;

[0078] Figure 2 A schematic diagram of an original graph structure dataset in a specific embodiment of the present invention;

[0079] Figure 3 This is a schematic diagram of a new graph structure dataset in a specific embodiment of the present invention. DETAILED DESCRIPTION

[0080] In order to make the objects, technical solutions and advantages of the present invention more clear, the present invention is further described in detail below in conjunction with the examples and drawings. The schematic embodiments of the present invention and their description are only used to explain the present invention and are not intended to limit the present invention. In the description of this application, it should be understood that the orientations or positional relationships indicated by terms such as "front", "back", "left", "right", "up", "down", "vertical", "horizontal", "high", "low", "inside", "outside", etc. are based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the scope of protection of this application.

[0081] Example 1:

[0082] like Figure 1 The graph hybrid learning friend recommendation method for sparse social networks shown in FIG includes the following steps:

[0083] Step 1: Create an original graph structure dataset, which includes user-user relationships and user-item relationships; Figure 2 shown.

[0084] Step 2: Based on the original graph structure dataset, a comparison view generator is used to generate a diversified social network comparison view and a denoised comparison view.

[0085] Specifically:

[0086] S201: Extracting a dataset containing user-user relationships from the original graph structure dataset, defined as a dataset ;

[0087] S202, the data set Input to the graph generator to obtain the diverse social network comparison view View1;

[0088] S203, the data set Input to the image denoiser to obtain the denoised comparison view View2.

[0089] The comparative view generator in this embodiment consists of a graph generator and a graph denoiser. The graph generator uses a specific augmentation strategy to generate a comparative view of a diverse social network structure, enriching the data representation and enhancing the model's generalization capabilities. The graph denoiser removes noisy edges by learning parameters, generating a cleaner comparative view.

[0090] The graph generator in this embodiment adopts the user profile and item enhancement strategy of the LLMRec model to improve the model's ability to understand user preferences and the accuracy of recommendation results. The dataset enhanced by the LLMRec model is represented as , using enhanced supervisory signals To optimize the loss function:

[0091]

[0092]

[0093]

[0094] Where: is the loss function of the graph generator; is the first BPR loss; To repair losses; stands for weighted sum; is the ratio coefficient used to control the number of retained samples; is the regularization term weight; represents the regularization term; represents the dataset enhanced by the graph generator; u represents the user; v represents the item; 、 Represent positive sample nodes and negative sample nodes respectively. Positive sample nodes refer to other nodes with similar structures, attributes or semantic relationships with the current node, and negative sample nodes refer to nodes with large differences in structure and semantics from the current node. Function It means sorting and selecting the first X values; Represents the sigmoid activation function; Representative User Positive sample node 's prediction score; Representative User For negative sample nodes 's prediction score; Representative dataset The node set of represent A subset of represents the original enhancement feature; Represents the enhanced features after mask operation; is the scaling factor.

[0095] The image denoiser in this embodiment uses the AdaGCL model and the following loss function:

[0096]

[0097]

[0098]

[0099] Where: is the loss function of the image denoiser; is the denoising loss; is the second BPR loss; Represents the regularization strength used to control weight decay; represents the regularization term; Representing a dataset The number of layers in the middle graph; l represents the node in the lth layer; Representing a dataset midpoint and connected edges; express The set of middle edges; Indicates the Nodes in the layer and the score of the edge between; represent The cumulative distribution function of represents the sigmoid activation function, Representative Model parameters corresponding to the layer; Represents a triplet describing a user For items and items preference relationship; 、 Represents users For items 's rating.

[0100] Step 3: Based on the diversified social network comparison view and the denoised comparison view, comparative learning is performed to obtain an initial feature embedding representation; the initial feature embedding representation includes a social embedding representation and an interest embedding representation.

[0101] In the process of contrastive learning, the objective function is:

[0102]

[0103]

[0104]

[0105] Where: represents the objective function of contrastive learning; represents the contrast loss of the user node; Represents the contrast loss of item nodes; Represents the i-th user node; represent The user collection in ; represents the cosine similarity function; 、 All are temperature parameters; Represents the user positive sample pair; Represents user negative sample pairs; Represents the i-th item node; represent The collection of items in ; Represents the positive sample pair of the item; represents an item-negative sample pair.

[0106] Step 4: Update the initial feature embedding through the graph convolution model to obtain an updated feature embedding representation; the updated feature embedding representation includes an updated social embedding representation and an updated interest embedding representation.

[0107] Specifically, for social embedding representation and interest embedding representation, the following steps are performed respectively:

[0108] S401. Input the initial feature embedding into the graph convolutional model and perform normalization to enhance the stability of model training; use layer normalization for numerical features; use batch normalization for categorical features.

[0109] S402: Input the normalized features into the Bi-Li+ module to obtain the first feature vector. At the same time, a low-rank layer (Low Rank Layer) can be introduced to project feature interactions from the sparse space to the low-rank space, further significantly reducing storage requirements.

[0110] S403. Input the normalized features into the SENet+ module, and perform compression, excitation, reweighting, and fusion operations in sequence to obtain the second feature vector; in the compression stage, each normalized feature embedding is divided into several groups, and the maximum value and average value of each group are selected as "summary statistics" to capture the global information of the features in different dimensions, providing a basis for subsequent learning of feature weights. In the excitation stage, the domain-level attention is improved to a more fine-grained bit-level attention, and two fully connected layers are used to learn weights. The model performs more detailed weight learning on each dimension of the feature, thereby more accurately determining the importance of the feature. In the reweighting stage, element-level multiplication operations are performed to reweight the features according to the learned weights, so that the weights of important features are increased and the weights of unimportant features are reduced. In the fusion stage, the reweighted features are fused with the outputs of other modules to provide richer information for subsequent feature interactions and model predictions.

[0111] S404: Concatenate the first eigenvector and the second eigenvector.

[0112] S405: The concatenated feature vector is further propagated through several graph convolutional layers to obtain an updated feature embedding representation.

[0113] Step 5: Fuse the updated feature embedding representation to establish a user node link probability prediction model:

[0114]

[0115] Where: a and b represent users ,user ; Representative User and The probability that a link exists between them; Represents the sigmoid activation function; stands for weighted sum; represents the dot product; Representative User Social embedding representation of Representative User Interest embedding representation; Representative User Social embedding representation of Representative User Interest embedding representation.

[0116] Step 6: Use the user node link probability prediction model to predict friends for all users in the original graph structure dataset, generate several user-user pseudo links, and eliminate existing user-user relationships to obtain a candidate pseudo link set.

[0117] Step 7: Reliability screening of the candidate pseudo link set to obtain a final pseudo link set; specifically, the following steps are performed:

[0118] S701: Perform intra-instance link selection on users in the candidate pseudo link set to obtain an intra-instance candidate pseudo link set:

[0119]

[0120] Where: is the set of candidate pseudo links within the instance; Represents a function that returns the minimum size set; j is the candidate user number; Represents a user For users Confidence of interest; is the threshold.

[0121] S702: Perform inter-instance link selection on the users in the candidate pseudo link set to obtain an inter-instance candidate pseudo link set:

[0122]

[0123] Where: is the set of candidate pseudo links between instances; Represents the function used to return a given quantile vector-valued function at ; Represents a function that sorts a vector into ascending order; A vector representing the confidence scores of all instances.

[0124] S703: Obtain the final pseudo link set based on the intra-instance candidate pseudo link set and the inter-instance candidate pseudo link set:

[0125]

[0126] Where: is the final pseudo link set; is the intersection operator.

[0127] Step 8: Add the final pseudo link set to the original graph structure dataset to obtain a new graph structure dataset, such as Figure 3 As shown;

[0128] Step 9: Replace the original graph structure dataset with the new graph structure dataset, and return to step S1 until no new user-user pseudo links can be generated; and perform friend recommendations using the current graph structure dataset.

[0129] It should be noted that in Figure 2 and Figure 3 In the figure, circles represent “users” and squares represent “items”.

[0130] Example 2:

[0131] In order to evaluate the effectiveness of the present application method in different recommendation scenarios, this embodiment selected three very representative public datasets in three scenarios for experimental evaluation. The three public datasets are: Ciao, Epinions and Last.FM. Among them, the Ciao dataset records the interactive behavior of users on the e-commerce platform, covering detailed data such as purchases and browsing, as well as friendships between users. The Epinions dataset contains users' evaluations and comments on various products and services, as well as the trust relationship network between users. The Last.FM dataset focuses on the music platform, including users' listening records, tag information and social relationships.

[0132] For all datasets used, this embodiment adopts a unified ratio of training set, validation set and test set, where 70% of the data is used for training, 10% of the data is used to verify the performance of the model and adjust the hyperparameters, and the remaining 20% ​​of the data is used to finally evaluate the generalization ability of the model.

[0133] This example selects a variety of representative existing friend recommendation methods for comparison, including traditional matrix decomposition methods, DNN-based recommendation models, GNN-based recommendation models, etc. The evaluation indicators are recall rate ( ), hit rate( ) and normalized discounted cumulative gain ( The top 10 (i.e., k=10) users are selected as recommended users. The comparative experimental results are shown in Table 1, where the best performance result is marked in bold and the second-best performance result is underlined:

[0134] Table 1 Comparative experimental results

[0135]

[0136] As can be seen from Table 1, our friend recommendation method outperforms other baseline models on all three datasets. On the Ciao dataset, our Recall@10, HR@10, and NDCG@10 performance improves by 2.5%, 2.5%, and 3.3%, respectively, over the next best model. On the Epinions dataset, our Recall@10, HR@10, and NDCG@10 performance improves by 2.2%, 2.4%, and 2.6%, respectively, over the next best model. On the Last.FM dataset, our Recall@10, HR@10, and NDCG@10 performance improve by 3.1%, 3.1%, and 2.6%, respectively, over the next best model. This fully demonstrates the effectiveness and superiority of our friend recommendation method in sparse social network environments.

[0137] Example 3:

[0138] To verify the effectiveness of each core step in the friend recommendation method of this application and its impact on performance, this example conducts an ablation experiment on the complete method of this application and the following three variants to compare performance:

[0139] (1) w / o CVG: Disable the contrast view generator, that is, do not use steps 2 and 3.

[0140] (2) w / o SM: disable the self-training module, that is, do not implement step 7 and do not perform reliability screening on the candidate pseudo-link set.

[0141] (3) w / o SLGN: The aggregation process of neighbor nodes in the graph convolutional model is handled normally, that is, the initial feature embedding is not updated in step 4.

[0142] The ablation experiment results are shown in Table 2:

[0143] Table 2 Ablation experiment results

[0144]

[0145] As can be seen from Table 2, the friend recommendation method of this application outperforms its variants on all data sets. Among them, after disabling the self-training module (w / o SM), the performance of the model drops significantly, which shows that the reliability screening of the candidate pseudo-link set plays a vital role in the method of this application. It can find possible UU links through continuous iterative optimization, thereby expanding the social scope. The addition of the contrast view generator can effectively alleviate the problem of sparse social networks. By enhancing the contrast view and denoising the contrast view, the user's consistency features can be learned more accurately, and the Recall@10, HR@10 and NDCG@10 are improved by 10.6%~13.9%, 11.1%~14.1% and 11.5%~14.7% respectively. After optimizing the aggregation method of neighbor nodes in the graph convolutional layer, Recall@10, HR@10, and NDCG@10 were improved by 2.4%~6.4%, 4.4%~6.2%, and 4.9%~6%, respectively. This shows that after updating the initial feature embedding, this application can learn the features of the node more deeply.

[0146] Example 4:

[0147] This application uses a specific method to aggregate features of neighboring nodes through a graph convolution model, thereby updating the initial feature embedding. To verify the effectiveness of the graph convolution model, this embodiment uses a variety of existing feature cross-strategies to carry out performance comparisons. The experimental results are shown in Table 3:

[0148] Table 3 Performance comparison of different methods for updating initial feature embedding

[0149]

[0150] As can be seen from Table 3, when using the graph convolutional model of this application, it is possible to more effectively capture high-order feature interactions in social networks, and its performance is better than other existing methods.

[0151] In this embodiment, the hyperparameter of the graph convolution model, that is, the number of graph convolution layers, is set to 4.

[0152] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

[0153] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprises", or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In addition, the term "connected" as used in this document, unless otherwise specified, may refer to a direct connection or an indirect connection via other components.

Claims

1. A graph hybrid learning friend recommendation method for sparse social networks, characterized by: The following steps are involved: S1. Establishing an original graph structure dataset, wherein the original graph structure dataset includes user-user relationships and user-item relationships; S2. Based on the original graph structure dataset, generating a diversified social network comparison view and a denoised comparison view respectively; S3. Performing comparative learning based on the diversified social network comparative view and the denoised comparative view to obtain an initial feature embedding representation; The initial feature embedding representation includes a social embedding representation and an interest embedding representation; S4. Updating the initial feature embedding through a graph convolutional model to obtain an updated feature embedding representation; S5. Fusing the updated feature embedding representation to establish a user node link probability prediction model; S6. Use the user node link probability prediction model to perform friend prediction on all users in the original graph structure data set to generate a number of user-user pseudo links; remove existing user-user relationships to obtain a candidate pseudo link set; S7, screening the candidate pseudo link set for reliability to obtain a final pseudo link set; S8. Add the final pseudo link set to the original graph structure dataset to obtain a new graph structure dataset; S9, replacing the original graph structure dataset with the new graph structure dataset, and returning to step S1 until no new user-user pseudo links can be generated; Recommend friends based on the current graph structure dataset.

2. The graph hybrid learning friend recommendation method for sparse social networks according to claim 1, characterized in that: Step S2 specifically includes: S201: Extracting a dataset containing user-user relationships from the original graph structure dataset, defined as a dataset ; S202, the data set Input into the graph generator to obtain a diverse comparative view of social networks; S203, the data set Input to the image denoiser to get the denoised comparison view.

3. The graph hybrid learning friend recommendation method for sparse social networks according to claim 2, characterized in that: The graph generator uses the following formula to determine the loss function: Where: is the loss function of the graph generator; is the first BPR loss; To repair losses; stands for weighted sum; is the ratio coefficient used to control the number of retained samples; is the regularization term weight; represents the regularization term; represents the dataset enhanced by the graph generator; u represents the user; v stands for item; 、 Represent positive sample nodes and negative sample nodes respectively. Positive sample nodes refer to other nodes with similar structures, attributes or semantic relationships with the current node, and negative sample nodes refer to nodes with large differences in structure and semantics from the current node. Function It means sorting and selecting the first X values; Represents the sigmoid activation function; Representative User Positive sample node 's prediction score; Representative User For negative sample nodes 's prediction score; Representative dataset The node set of represent A subset of represents the original enhancement feature; Represents the enhanced features after mask operation; is the scaling factor.

4. The graph hybrid learning friend recommendation method for sparse social networks according to claim 2, characterized in that: The image denoiser uses the following formula to determine the loss function: Where: is the loss function of the image denoiser; is the denoising loss; is the second BPR loss; Represents the regularization strength used to control weight decay; represents the regularization term; Representation dataset The number of layers in the middle graph; l represents the node in the lth layer; Representation dataset midpoint and connected edges; express The set of middle edges; Indicates the Nodes in the layer and the score of the edge between; represent The cumulative distribution function of represents the sigmoid activation function, Representative Model parameters corresponding to the layer; Represents a triplet describing a user For items and items preference relationship; 、 Represents users For items 's rating.

5. The graph hybrid learning friend recommendation method for sparse social networks according to claim 1, characterized in that: In the process of contrastive learning, the objective function is determined by the following formula: Where: represents the objective function of contrastive learning; Represents the contrast loss of the user node; Represents the contrast loss of item nodes; Represents the i-th user node; represent The user collection in ; represents the cosine similarity function; 、 All are temperature parameters; Represents the user positive sample pair; Represents user negative sample pairs; Represents the i-th item node; represent The collection of items in ; Represents the positive sample pair of the item; represents an item-negative sample pair.

6. The graph hybrid learning friend recommendation method for sparse social networks according to claim 1, characterized in that: Step S4 specifically includes: S401, input the initial feature embedding into the graph convolution model and perform normalization processing; S402, inputting the normalized features into the Bi-Li+ module to obtain a first feature vector; S403: Input the normalized features into the SENet+ module, and perform compression, excitation, reweighting, and fusion operations in sequence to obtain a second feature vector; S404, concatenating the first eigenvector and the second eigenvector; S405: The concatenated feature vector is further propagated through several graph convolutional layers to obtain an updated feature embedding representation.

7. The graph hybrid learning friend recommendation method for sparse social networks according to claim 1, characterized in that: The user node link probability prediction model is: Where: a and b represent users ,user ; Representative User and The probability that a link exists between them; Represents the sigmoid activation function; stands for weighted sum; represents the dot product; Representative User Social embedding representation of Representative User Interest embedding representation; Representative User Social embedding representation of Representative User Interest embedding representation.

8. The graph hybrid learning friend recommendation method for sparse social networks according to claim 1, characterized in that: Step S7 specifically includes: S701: Perform intra-instance link selection on users in the candidate pseudo link set to obtain an intra-instance candidate pseudo link set; S702: Perform inter-instance link selection on users in the candidate pseudo link set to obtain an inter-instance candidate pseudo link set; S703 : Obtain the final pseudo link set based on the intra-instance candidate pseudo link set and the inter-instance candidate pseudo link set.

9. The graph hybrid learning friend recommendation method for sparse social networks according to claim 8, characterized in that: The candidate pseudo link set in the instance is calculated using the following formula: Where: is the set of candidate pseudo links within the instance; Represents a function that returns the minimum size set; j is the candidate user number; Represents a user For users Confidence of interest; is the threshold; The candidate pseudo link set between instances is calculated using the following formula: Where: is the set of candidate pseudo links between instances; Represents the function used to return a given quantile vector-valued function at ; Represents a function that sorts a vector into ascending order; A vector representing the confidence scores of all instances.

10. The graph hybrid learning friend recommendation method for sparse social networks according to claim 9, characterized in that: The final pseudo link set is obtained by the following formula: Where: is the final pseudo link set; is the intersection operator.

Citation Information

Patent Citations

  • LBSN supernetwork link prediction method based on time-space relationship

    CN107784124A

  • Social recommendation method of heterogeneous graph convolutional network combining social contact and interest information

    CN111428147A

  • Social network service platform friend recommendation method based on global attention mechanism representation learning

    CN112100514A

  • Session recommendation method and system based on enhanced graph neural network, and medium

    CN114116995A

  • Local information retention depth contrast clustering method applied to graph structure data clustering

    CN117710716A

Cited By

  • Multi-network unified link prediction method based on social network

    CN122112355A