Social network group tendency discrimination method and system

By constructing the MuLa-GCN algorithm in the social network, integrating the tag-label and node-label features and pruning optimization, the weak representation ability and overfitting of the node classification algorithm in the social network are solved, and more accurate group tendency discrimination and information transmission effects are achieved.

CN120277211APending Publication Date: 2025-07-08中科天玑数据科技股份有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510350943.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The node classification algorithm in existing social networks lacks rich representation information, resulting in weak representation ability and overfitting and oversmoothing problems, making it difficult to effectively determine group tendencies.

Method used

By constructing the MuLa-GCN algorithm, integrating the feature information of tag-label and node-label, combining pruning strategies to optimize the graph convolution network, improving the loss function and classifier, and using weighted fusion or voting method to generate the optimal algorithm to improve the performance of multi-label node classification.

Benefits of technology

提高了社交网络中群体倾向性的判别准确性,增强了推荐效果和广告投放的有效性,提升了信息传递和分发效果。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277211A_ABST
    Figure CN120277211A_ABST
Patent Text Reader

Abstract

The invention discloses a social network group tendency discrimination method and system, and the method comprises the steps: collecting data in a social network category, and constructing node and structure information; preprocessing the data, and constructing features; a Baseline algorithm is constructed, and a Baseline algorithm is constructed; the cross entropy loss based on the GCN algorithm is changed into'binary cross entropy loss', and the classifier is changed from Softmax into Sigmod to construct a Bi-GCN algorithm; adding label-label feature information based on a GCN algorithm, and fusing representation of an original node-label to construct a MuLa-GCN algorithm; optimizing the MuLa-GCN algorithm and the Bi-GCN algorithm by adopting a pruning strategy so as to generate a Drop-Bi-GCN algorithm and a Drop-MuLa-GCN algorithm; selecting one or more fusion generation optimal algorithms from the algorithms, and classifying the multi-label nodes; the characteristics of the graph convolution network in processing a multi-label node classification task are fully exerted, the performance of the graph convolution network is improved through feature expansion of the graph convolution and integration of a pruning mechanism, and the effectiveness of social network advertisement putting is improved; and the information transmission and distribution effect of the social network is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multi-label scenario classification, and particularly relates to a method and system for discriminating group tendency in a social network. Background Art

[0002] With the rapid development of Internet technology, social networks have gradually become an important social tool. "Groups" play an important role in social networks, and the interactions between different groups make the complex network more valuable for research. In this paper, "group" is defined as the topic label of node behavior, and "tendency" corresponds to the research on the possibility of a node participating in multiple topics, that is, the multi-label node classification problem. At present, great progress has been made in the research of this problem, but it still faces many problems and challenges, such as: the diversity of node attributes, data sparsity, and the conventional way of representing nodes only through attribute and structure information, lacking label representation information in the multi-label scenario.

[0003] Therefore, in view of the problem of attribute diversity, targeted selection of attribute information with strong correlation to the classification task is proposed. For the problem of data sparsity, nodes with rich labels are refined through methods such as data cleaning and label analysis. For the problem of insufficient representation, label embedding is considered to enhance the node representation ability based only on attributes and structures, and the algorithm is optimized through a pruning mechanism to achieve multi-label node classification, and further achieve the purpose of discriminating group tendency.

[0004] 1. Existing algorithms unilaterally consider the attributes or structures of nodes, lacking rich representation, resulting in weak representation information ability.

[0005] 2. Problems such as overfitting and oversmoothing are more obvious.

[0006] Summary of the Invention

[0007] Aiming at the deficiencies of the existing technology, the purpose of the present invention is to provide a method and system for discriminating group tendency in a social network.

[0008] To achieve the above purpose, the present invention provides the following technical solutions:

[0009] A method for discriminating group tendency in a social network, comprising the steps of:

[0010] Collect data within the scope of the social network and construct node and structure information;

[0011] Preprocess the data,

[0012] Construct features, and construct the representation of the attributes and relevant speech text information of the nodes, as well as the adjacency matrix and feature matrix between nodes;

[0013] Construct the Baseline algorithm based on LINE, DeepWalk, and Node2vec;

[0014] Change the cross-entropy loss based on the GCN algorithm to "binary cross-entropy loss", and change the classifier from Softmax to Sigmod to construct the Bi-GCN algorithm;

[0015] Add label-label feature information based on the GCN algorithm, and integrate the original node-label representation to construct the MuLa-GCN algorithm;

[0016] Optimize the MuLa-GCN algorithm and the Bi-GCN algorithm using a pruning strategy to generate the Drop-Bi-GCN algorithm and the Drop-MuLa-GCN algorithm;

[0017] Select one or more from the Baseline algorithm, Bi-GCN algorithm, MuLa-GCN algorithm, Drop-Bi-GCN algorithm, and Drop-MuLa-GCN algorithm to fuse and generate the optimal algorithm for classifying multi-label nodes.

[0018] In the present invention, preferably, the data within the scope of the social network includes entity data and relationship data.

[0019] In the present invention, preferably, the construction of the Bi-GCN algorithm specifically includes:

[0020] Based on the original GCN:

[0021]

[0022] Among them, H (l+1) represents the node feature representation (node embedding) of the (l + 1)-th layer, σ() is a non-linear activation function, is the normalized adjacency matrix, H l is the node feature representation of the l-th layer, W l is the learnable weight matrix of the l-th layer;

[0023] Construct the loss function:

[0024]

[0025]

[0026] Among them, l n is the loss function, x n and y n are the features and labels of the corresponding samples, w i is the learnable weight.

[0027] In the present invention, preferably, the construction of the MuLa - GCN algorithm is specifically as follows:

[0028] Construct Label - label features and Node - label features,

[0029] Co - optimize the Label - label features and Node - label features.

[0030] In the present invention, preferably, maximize the co - occurrence probability between labels:

[0031]

[0032] Maximize the occurrence probability of nodes and corresponding labels:

[0033]

[0034] When the number of nodes is too large, use negative sampling to reduce the computational complexity of node - label and label - label:

[0035]

[0036] where l i and respectively represent different labels, and h i is the vector corresponding to x i .

[0037] In the present invention, preferably, the pruning strategy includes sparsifying the adjacency matrix:

[0038] A drop = A - A ′ ,

[0039] and normalizing:

[0040]

[0041] where A is the adjacency matrix, A ′ is the pruned adjacency matrix, and A drop and are respectively the results after pruning and normalizing.

[0042] In the present invention, preferably, generate the optimal algorithm by combining business selection, and the fusion methods include weighted fusion method and voting method.

[0043] In the present invention, preferably, the weighted fusion method is:

[0044] P final = w1·P1 + w2·P2+…+w n ·P n ,

[0045] Wherein, w1, w2, w3…, w n are the weights corresponding to each algorithm M1, M2, M3,…, M n and P1, P2, P3…, P n are the outputs corresponding to each algorithm M1, M2, M3,…, M n .

[0046] In the present invention, preferably, the voting method adopts the predicted probability output by each algorithm, and takes the average value of all probabilities.

[0047] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0048] The method of the present invention gives full play to the characteristics of the graph convolutional network in processing multi-label node classification tasks, and improves its performance by integrating the feature expansion and pruning mechanism of graph convolution. The application directions of this method are numerous. For example, through the analysis of group tendencies in the network, understanding the characteristics of target users' interests, activities, etc. can effectively improve the recommendation effect; improve the effectiveness of social network advertising placement; improve the information transmission and distribution effect of social networks. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 is a schematic flowchart of a method for discriminating group tendencies in a social network according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0050] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0051] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field of the present invention. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The term “and / or” used herein includes any and all combinations of one or more of the related listed items.

[0052] Please refer to Figure 1, a preferred embodiment of the present invention provides a method for judging the group tendency of a social network. By defining the group tendency problem as a multi-label node classification task, first, the Loss function of the original GCN for the single-label scenario is adjusted to adapt to the multi-label scenario as the basic ability of the graph convolutional network to process the node classification task. At the same time, combined with the characteristics of the task, the label-label features are integrated into the original node-label features for collaborative optimization to improve the overall performance of the algorithm. Existing algorithms only focus on shallow GCNs and rarely perform deep extensions. However, if the number of layers of the graph convolutional neural network is too small, it is difficult to learn rich hidden information. Due to the limitations of GCN itself, such as overfitting and oversmoothing, which are particularly obvious after deepening the number of layers, adding a pruning strategy can alleviate this problem, making the GCN algorithm after deepening the number of layers have better performance and being able to better handle the classification task of multi-label nodes.

[0053] The specific steps of this method are as follows:

[0054] Collect data within the scope of the social network, and construct node and structure information. The data within the scope of the social network includes entity data and relationship data;

[0055] Preprocess the data. The preprocessing includes data cleaning, removing redundant label content, labeling and screening the original text data according to the required theme. Through statistical analysis, isolated nodes are removed.

[0056] Construct features, and construct the representation of the node's attributes and related speech text information, as well as the adjacency matrix and feature matrix between nodes. The adjacency matrix is constructed through static attention and user mapping, and is used to describe the connection relationship between nodes. The theme is determined through user mapping, original pushed content, and forwarding behavior, and a feature matrix is constructed. The feature matrix stores the feature information of the nodes.

[0057] Construct a Baseline algorithm based on LINE, DeepWalk, and Node2vec, that is, use LINE, DeepWalk, and Node2vec as baseline models to construct a Baseline version of the algorithm, in order to verify the improvement effect of the GCN algorithm in local structure preservation and global context modeling in the future.

[0058] Modify the cross-entropy loss in the GCN algorithm to "binary cross-entropy loss", and change the classifier from Softmax to Sigmod, thus constructing the Bi-GCN algorithm. Specifically, change the cross-entropy loss of the original GCN applied to the single-label multi-classification scenario to "binary cross-entropy loss" (Binary cross entropy), where the classifier is changed from Softmax to Sigmod to complete the transition from the single-label to the multi-label scenario. That is, the softmax output represents a multinomial distribution, while for multi-classification, multi-label classification needs to be divided into multiple binary classifications, and the Sigmoid output is a binomial distribution. The formula of the original GCN is as follows:

[0059]

[0060] Among them, H (l+1) represents the node feature representation (node embedding) of the (l + 1)-th layer, σ(·) is a non-linear activation function, is the normalized adjacency matrix, H l is the node feature representation of the l-th layer, W l is the learnable weight matrix of the l-th layer.

[0061] The modification of the model is to change the cross-entropy loss of the original GCN to the "binary cross-entropy loss" in formula (4). Among them, the classifier is changed from softmax to sigmod. The data D(x, y) contains N samples, and each sample may have M labels.

[0062]

[0063]

[0064] Among them, l n is the loss function, x n and y n are the features and labels of the corresponding samples, w i is the learnable weight.

[0065] Based on the GCN algorithm, add label-label feature information and integrate the original node-label representation to construct the MuLa-GCN algorithm. Specifically, add label-label feature information on the basis of the graph convolutional network of the original GCN algorithm, integrate it into the original node-label representation, and perform collaborative optimization to improve the overall performance of the algorithm. Construct Label-label features and Node-label features, and perform collaborative optimization through Node-label features and Label-label features.

[0066] When constructing label-label and node-label features, the correlation matrices between labels and between nodes and labels are constructed as feature matrices respectively. For label-label representation, common methods include explicit modeling based on probabilistic graphical models or recurrent neural networks, and implicit modeling of label correlations through attention mechanisms. This paper adopts a data-driven approach, using graph convolutional networks (GCNs) to mine the co-occurrence patterns of labels in the dataset, thereby learning the correlations between labels and constructing a correlation coefficient matrix. For node-label representation, drawing on the idea of word embedding, nodes are analogized to words and labels to context words, and the idea of the Skip gram model is adopted to predict the possible surrounding labels through the central node. For example, given node x1 and its label set y1, y2... yn, multiple node-label pairs are formed between the node and each label, forming a sequence y1, x1, y2... yn.

[0067] In the joint optimization process, the representations of label-label and node-label share the mapping parameters from word embedding to the classifier, such that the gradients of all classifiers affect the classifier function based on GCN, thereby implicitly modeling label correlations. To reduce the computational complexity, when the number of nodes is too large, negative sampling techniques are used to reduce the computational amount of node-label and label-label pairs. In this way, the model can effectively capture the correlations between labels and the associations between nodes and labels, improving the performance of the multi-label node classification task.

[0068] Maximize the probability of co-occurrence between labels, maximize the probability of the occurrence of nodes and their corresponding labels. When the number of nodes is too large, use negative sampling to reduce the computational complexity of node-label and label-label.

[0069] Among them, to maximize the probability of co-occurrence between labels:

[0070]

[0071] To maximize the probability of the occurrence of nodes and their corresponding labels:

[0072]

[0073] When the number of nodes is too large, use negative sampling to reduce the computational complexity of node-label and label-label:

[0074]

[0075] where, l i and represent different labels respectively, and h i is the vector corresponding to x i respectively.

[0076] The pruning strategy is adopted to optimize the MuLa-GCN algorithm and the Bi-GCN algorithm to generate the Drop-Bi-GCN algorithm and the Drop-MuLa-GCN algorithm. Most previous work only focused on shallow GCNs and rarely performed deeper extensions. However, if the number of layers of the graph convolutional neural network is too small, it is difficult to learn rich hidden information. But after deepening the network, the generalization ability on small data decreases, which will lead to overfitting. More and more layers make the network complex, isolate the input and output, and problems such as over-smoothing occur. By randomly performing the sparsification and renormalization operations on the adjacency matrix, as shown in formulas (8) and (9). It increases the randomness and richness of the input data and effectively prevents the overfitting problem. The pruning strategy can be regarded as a message passing attenuator, deleting some edges to make the graph sparse, and to a certain extent avoiding the over-smoothing problem of deep GCNs. Modifying the MuLa-GCN algorithm and the Bi-GCN algorithm according to the pruning strategy can obtain higher accuracy and better results.

[0077] The pruning strategy includes performing the sparsification of the adjacency matrix:

[0078] A drop = A - A ′ ,

[0079] and renormalization:

[0080]

[0081] where A is the adjacency matrix, A ′ is the pruned adjacency matrix, A drop and are respectively the results after pruning and after renormalization.

[0082] Select one or more from the Baseline algorithm, Bi-GCN algorithm, MuLa-GCN algorithm, Drop-Bi-GCN algorithm and Drop-MuLa-GCN algorithm for fusion to generate the optimal algorithm for classifying multi-label nodes. According to the business scenario, if it is a social network domain scenario, the method of the present invention can be used, or in other fields and a network structure dataset of nodes and relationships can be constructed, and it can also be adopted. Use the test dataset to test the above algorithms according to the Micro-f1 and Macro-f1 evaluation metrics to select the one with the best effect. After selecting the optimal algorithm, the fusion methods include the weighted fusion method and the voting method. In scenarios with obvious business intervention, the weighted average method can be adopted, and the more preferred fusion methods for algorithms are the weighted fusion method and the voting method. Whether to specifically adopt a single algorithm or a multi-algorithm fusion is determined according to the specific actual scenario test situation.

[0083] The weighted fusion method is to assign different weights to the results of different algorithms, and perform weighted averaging or summation on the prediction results of multiple models according to their performance. The weighted fusion method is as follows:

[0084] P final = w1·P1 + w2·P2 + … + w n ·P n ,

[0085] where w1, w2, w3…, w n are the weights corresponding to each algorithm M1, M2, M3, …, M n , and P1, P2, P3…, P n are the outputs corresponding to each algorithm M1, M2, M3, …, M n .

[0086] The voting method uses the predicted probabilities output by each algorithm and takes the average of all probabilities. For example, predicting multiple classes P1 = (label 1_1 , label 2_1 ), P n = (label n_1 , label n_2 ). Then P final = (avg(label 1_1 , … label n_1 ), …, avg(label 2_1 , … label n_2 ) (12)

[0087] The whole process of the present invention from data collection to application mainly includes steps such as data acquisition, analysis, and modeling. First, the original data of relevant topics is obtained through data acquisition, and after processing, in-depth analysis and modeling are carried out, and finally the model is applied to the corresponding scenarios.

[0088] The research is carried out from three perspectives: individual, key users, and groups. The individual user analysis module focuses on the basic information, tendencies, interaction records, and speech situations of users to help comprehensively understand user behavior. The key user analysis module targets influential figures in specific regions and evaluates their potential influence in the social network by statistically analyzing their attribute information (such as educational experience) and various relationship networks (such as social relationships, social platform relationships). The group analysis module focuses on group tendencies, behavior characteristics, emotional changes, and attribute portraits. By visually displaying the interaction degree and tendencies between different groups, the possibility of individuals participating in specific groups is analyzed. The group behavior characteristic analysis reveals the behavior differences of users in different fields in the theme dissemination, the emotional portrait shows the emotional preferences of the group for the theme, and the group attribute portrait reflects the differences in attributes such as age, gender, and religion.

[0089] The comparative experiment data proves that:

[0090] The data sets used by the algorithm involved in the present invention are BlogCatalog, MicroBlog_5, MicroBlog_30, and Facebook_23, where MicroBlog_5 represents screening 5 topics and Facebook_23 represents screening 23 topics. The effects of GCN-Bin, MuLa-GCN, Drop-GCN-Bin, and Drop-MuLa GCN are verified. Among them, the pruning mechanism significantly improves the model performance and is effective for both GCN-Bin and MuLa-GCN. In the experiment, a 2-layer pruned GCN is selected with a pruning rate of 0.3 (retention rate of 0.7). Due to the high proportion of single labels in the MicroBlog_5 data set and the simple classification task, the indicators are significantly better than other data sets, but the performance of GCN-based models is not significant. As shown in Table 1, the experiment shows that the performance of MuLa-GCN based on label embedding is better than that of GCN-Bin without label embedding, and the algorithm with the Drop Edge mechanism has better performance.

[0091] Table 1 Comparison of Macro-f1(%) indicators in different algorithm results

[0092]

[0093] In some other preferred embodiments of the present invention, a computer-readable storage medium is provided, storing a computer program, which when executed by a processor, causes the processor to execute the steps of the method as described in the above embodiments.

[0094] If the said function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. And the aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs and other various media that can store program codes.

[0095] The above description is a detailed description of the preferred and feasible embodiments of the present invention. However, the embodiments are not intended to limit the scope of the patent application of the present invention. Any equivalent changes or modifications made under the technical spirit disclosed by the present invention shall fall within the scope of the patent covered by the present invention.

Claims

1. A method for judging the tendency of a social network group, characterized in that, Including steps: Collect data within the scope of social networks, and construct node and structural information; Preprocess the data, Construct features, characterize the attributes of nodes and the relevant speech text information, and construct the adjacency matrix and feature matrix between nodes; Construct the Baseline algorithm based on LINE, DeepWalk, and Node2vec; Change the cross-entropy loss based on the GCN algorithm to "binary cross-entropy loss", and change the classifier from Softmax to Sigmod to construct the Bi-GCN algorithm; Add label-label feature information based on the GCN algorithm, and integrate the original node-label representation to construct the MuLa-GCN algorithm; Optimize the MuLa-GCN algorithm and the Bi-GCN algorithm using pruning strategies to generate the Drop-Bi-GCN algorithm and the Drop-MuLa-GCN algorithm; Select one or more from the Baseline algorithm, Bi-GCN algorithm, MuLa-GCN algorithm, Drop-Bi-GCN algorithm, and Drop-MuLa-GCN algorithm to fuse and generate the optimal algorithm for classifying multi-label nodes.

2. The discriminant method for the social network group tendency according to claim 1, characterized in that The data within the scope of the social network includes entity data and relationship data.

3. A method for discriminating the group tendency of a social network according to claim 2, characterized in that, The specific construction of the Bi-GCN algorithm includes: Based on the original GCN algorithm: Among them, H (l+1) represents the node feature representation (node embedding) of the (l + 1)-th layer, σ(·) is a non-linear activation function, is the normalized adjacency matrix, H l is the node feature representation of the l-th layer, and W l is the learnable weight matrix of the l-th layer; Construct a loss function: where l n is the loss function, x n and y n are the features and labels of the corresponding samples, and w i is the learnable weight.

4. A method for discriminating the group tendency of a social network according to claim 1, characterized in that, The specific construction of the MuLa-GCN algorithm is: Construct Label-label features and Node-label features, Co-optimize the Label-label features and Node-label features.

5. A method for judging the group tendency of a social network according to claim 4, characterized in that Maximize the co-occurrence probability between labels: Maximize the occurrence probability of nodes and corresponding labels: When the number of nodes is too large, use negative sampling to reduce the computational complexity of node-label and label-label: where l i and represent different labels respectively, and h i is the vector corresponding to x i .

6. A method for judging the group tendency of a social network according to claim 1, characterized in that, The pruning strategy includes sparsifying the adjacency matrix: A drop = A - A', And normalization: Among them, A is the adjacency matrix, A' is the pruned adjacency matrix, and A drop and are the results after pruning and normalization respectively.

7. A method for discriminating the group tendency of a social network according to claim 1, characterized in that, Combine business selection to generate the optimal algorithm, and the fusion methods include weighted fusion method and voting method.

8. A method for discriminating the group tendency of a social network according to claim 1, characterized in that, The weighted fusion method is: P final = w1·P1 + w2·P2 + … + w n ·P n , Wherein, w1, w2, w3…, w n are the weights corresponding to respective algorithms M1, M2, M3,…, M n , and P1, P2, P3…, P n are the outputs corresponding to respective algorithms M1, M2, M3,…, M n .

9. A method for discriminating the tendency of a social network group according to claim 1, characterized in that The voting method uses the predicted probability output by each algorithm and takes the average of all probabilities.

10. A social network group tendency discrimination system for implementing a social network group tendency discrimination method according to any one of claims 1-9, characterized in that, Including Baseline model, Bi-GCN model, MuLa-GCN model, Drop-Bi-GCN model, Drop-MuLa-GCN model and data preprocessing model, and the data preprocessing model is used for the received data.