Content-oriented group opinion prediction method and system

By employing techniques such as BERT, forgetting curves, and graph attention networks, the problems of user self-drive and timeliness in group opinion prediction were solved, thereby improving the accuracy of the model.

CN115577288BActive Publication Date: 2026-01-23SOUTHEAST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211309757.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-25
Publication Date
2026-01-23
Estimated Expiration
2042-10-25

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively capture user self-motivation, consider the timeliness of activities, and the order in which users join in predicting group opinions, resulting in insufficient model accuracy.

Method used

We use BERT for text feature extraction, combine forgetting curves, user self-driven representations, domain representations and attention mechanisms, and use LSTM and graph attention networks to fuse group features to build a group opinion prediction model.

Benefits of technology

It improves the accuracy of predicting group opinions, better captures the timeliness of users' spontaneous behavior and activities, and enhances the model's predictive capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115577288B_ABST
    Figure CN115577288B_ABST
Patent Text Reader

Abstract

The application discloses a content-oriented group opinion prediction method and system, and the method steps of the application are as follows: firstly, a BERT model is used to pre-train activity description text features to obtain user initial representation; secondly, a cooperation network is constructed based on user cooperation relations to extract self-driven representation of the user; then, the user is clustered according to the interest and hobby label field of the user itself to obtain field-specific feature representation of the user; group features are obtained by fusing the user initialization representation, the self-driven representation and the field-specific representation at the individual level; finally, the attitude of the group to a target activity is predicted through a group opinion prediction model. The system realizes visual display of the description generation result by using web interaction technology. The application can effectively predict the attitude of the group in an interest activity community to whether an activity is held, and provides effective technical support for platform management and related activity recommendation.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to a method for predicting group opinion, and particularly relates to a method and system for predicting group opinion based on text. BACKGROUND

[0002] With the development of the Internet, social networks have gradually accumulated a large number of users, and the huge user group can freely and fully share and exchange their own opinions. Social network platforms not only provide a convenient interaction method for individuals, but also provide a full exchange for a group of people with similar interests and backgrounds. In some specific interest community, the initiator of the interest activity often initiates the activity in a certain way in a group, and whether the activity can be held depends on the opinion of the whole group. Therefore, group opinion prediction is particularly important for community management platforms.

[0003] Group opinion prediction can be regarded as an extension task of text-based stance detection or group recommendation, but it is different from these two technologies. The text-based stance detection task is from the perspective of text, and considers the stance of a single or multiple targets on the text. Group recommendation is to recommend different items to a group. On the one hand, text-based stance detection needs to clearly distinguish the stance, that is, to distinguish the positive, negative and other stances. The difference between group opinion and stance detection is that it needs to predict the group opinion, and the target in the group only holds a positive opinion, which also shows that the formation of the group has a purpose. On the other hand, the research focus of group recommendation is how to mine the common preference characteristics of members in the group, and needs to balance the differences between members in the group to alleviate the preference conflict between members.

[0004] Text-based stance detection can be divided into single-objective text stance detection and multi-objective text stance detection. The task of single-objective stance detection is to determine the attitude and opinion of a given target towards the current text, i.e., to find the mapping relationship between the text and the target's stance, given a single target and text content. Early work mainly used rule-based and machine learning approaches. SVM (Single Object Vector Machine) dominated research using feature engineering. With the development of deep learning, more and more work in the field of stance detection is adopting deep learning methods. Isablelle et al. used RNNs to encode the target and text, using the output layer of the target encoding module as the initial value of the text encoding module, meaning the text encoding module needs to wait for the output of the target encoding module. Vijayaraghavan et al. used convolutional neural networks to train features at two levels: word level and character level, and performed stance detection and analysis by fusing these two levels of features. Compared with single-objective stance detection, multi-objective stance detection changes the research object from a single target to multiple targets; given n targets and text content, it needs to determine the stance inclination of multiple targets towards the text. Furthermore, multi-objective stance detection involves the propagation and opposition of stances, meaning that differences in roles between targets lead to different stances. Sobhani argues that single-objective stance detection treats each individual equally, ignoring potential influence and conflicting interests between them. He proposes a multi-objective stance detection task and releases a dataset for this purpose. The authors also propose an attention-based multi-objective stance detection method, leveraging the advantages of attention to more effectively adjust the weights of text information when determining the stance of each object. Wei et al. proposed a dynamic memory enhancement network that uses two bidirectional long short-term memory neural networks in the text encoding module, fuses features using an attention mechanism, and then extracts the association information between multiple objects and their stances using shared dynamic memory units. Siddiqua et al. proposed an ensemble model based on neural networks. This model concatenates the object vector and text vector to obtain input features, then uses multiple convolutional kernels to convolve the input features, feeding them into a densely connected bidirectional long short-term memory network and a nested long short-term memory network. The outputs of these two networks are then concatenated again to obtain the final features used to determine the probability of the object's stance.

[0005] In research on group opinion prediction based on group recommendation, the research object changes from individuals to groups, representing a leap from individual to group. Therefore, the fusion of individual and group preferences becomes a core issue in group recommendation, and preference fusion has become a key step in the field. From the perspective of preference fusion, it can be divided into two aspects: implicit preference fusion methods and explicit preference fusion methods. The difference between these two methods is that implicit fusion methods fuse group opinions without obtaining the explicit preferences of individuals; individual preferences are represented by features. Explicit fusion methods, on the other hand, require obtaining the opinions or preferences of individuals in advance before fusing them into group opinions or preferences. Early work on implicit preference fusion was mainly based on probabilistic models and the idea of ​​information aggregation. Seko et al. proposed a model that integrates content classification with group decision-making, arguing that item classification influences group decisions. Liu et al. proposed a personalized topic model, which assumes that the most influential users represent groups and have a significant impact on group decisions. However, this method is only applicable to predefined groups, i.e., groups with relatively stable relationships. For sporadically formed groups, this method has significant limitations. With the development of deep learning, more and more work is using deep learning methods for research. Cao et al. were the first to use an attention mechanism to aggregate individual preferences to obtain a representation of group preferences, but they only aggregated individual preferences without considering social influence. Building on this work, the authors considered social influence by applying attention to each user's neighbors once to represent the influence of those neighbors, and then applying attention again to the group, obtaining a group feature representation through a hierarchical attention mechanism. He et al. proposed the GAME model, applicable to incidental groups. This paper models the relationships between users, groups, and content from multiple perspectives. User preference features are obtained from both user-content and individual-group perspectives. Furthermore, the authors argue that a group is formed under the influence of a topic, thus the topic's influence on individual opinions also has an impact. Since incidental groups lack historical behavioral data, this work proposes using the fusion of individual features within the group as a group representation.

[0006] Opinion dynamics uses mathematics, physics, and computer science, particularly agent-based modeling and simulation methods, to study the evolutionary processes and rules governing the convergence or clustering of group opinions. The research scope of opinion dynamics is very broad, encompassing various social phenomena such as individual opinion evolution, group decision-making, consensus achievement, and the survival of minority opinions. An opinion is an individual's view, choice, or inclination towards an event. Based on the way opinions are described, opinion dynamics models can be divided into discrete models and continuous models; this patent will introduce the current state of research on opinion dynamics from the perspective of opinion description methods. Discrete models use binary values ​​or other discrete integer values ​​to model opinions, such as 0 and 1, just like buy and sell, left and right, neutral, support and opposition in the real world. These include Ising models, voter models, and local majority models and their extended models. The Ising model was originally proposed in physics to explain the phase transition properties of ferromagnetic materials. The phase transition properties of ferromagnetic materials share many similarities with the evolutionary properties of group viewpoints in sociology. Therefore, some have proposed using the Ising model to characterize viewpoint conflicts within a group. This method replaces the polarity of ferromagnetic materials with the polarity of individual viewpoints and replaces the system energy with the group's viewpoint. If two adjacent nodes have opposing viewpoints, the total system energy decreases by one; otherwise, it increases by one. Conversely, if two nodes have the same viewpoint, the total system energy decreases by one. The Voter model, proposed by Clifford and Sudbury, posits that individuals always refer to the viewpoints of their neighbors and are unaffected by external information. Furthermore, an individual's viewpoint originates from only one neighbor, while the majority of neighbors do not directly influence the individual. As the model evolves, individuals with similar viewpoints begin to form clusters within the group network. Since this model is equivalent to a random walk, the system often converges to a specific viewpoint, but which viewpoint it converges to is unpredictable. Galam improved upon the Ising model by proposing the local majority model. This model considers the herding effect prevalent in sociology. Its evolutionary rule is as follows: in a group of n individuals, each holds either a +1 or -1 opinion. The quantification of the group's opinion is reflected in the model selecting the opinion with the most prevalent members as the group's opinion. The Sznazjd model, proposed by sznazjd, posits that opinions are always influenced by pairs of individuals. An individual's opinion is affected by their neighbors within two hops, exhibiting significant "outflow" of information, and is therefore more often used to simulate the spread of opinions in society. Summary of the Invention

[0007] Purpose of the invention: In order to overcome the shortcomings of the prior art, the present invention provides a content-oriented method and system for predicting group opinions.

[0008] Technical solution: To achieve the above-mentioned objectives, the present invention adopts the following technical solution:

[0009] The present invention provides a content-oriented method for predicting group opinions, comprising the following steps:

[0010] (1) Text feature extraction

[0011] The activity text is preprocessed and pre-trained using BERT, which is then used for classification training according to different domains of the activity text to obtain the feature representation of the activity text.

[0012] (2) User initialization representation

[0013] A forgetting curve f(t) is constructed to represent the change in the importance of the active text over time. The corresponding forgetting curve value and the active text features obtained in step (1) are then combined. Multiply and sum to obtain the user feature representation u. self ′; then subtract the negative text u after average pooling nay Get the user-initialized representation u self ;

[0014] (3) User self-driving representation

[0015] A topological graph G1 of user relationships is established using the relationships between users, and then a user self-driving representation u is obtained using a two-layer convolutional neural network. effected The initial features of the convolutional neural network user are the user initialization representation u obtained in step (2). self ;

[0016] (4) User domain representation

[0017] Based on each user's different domain, the GMM algorithm is used to perform overlapping clustering of the user's domains and construct a domain graph G2. Then, GAT is used to fuse the features, and finally obtain the features of the user's domain, i.e., the user's domain representation u. group ;

[0018] (5) Integration of group characteristics

[0019] Based on the outputs of steps (2), (3), and (4), a weighted summation is performed using an attention mechanism to obtain the user representation S in the group. Then, the group features are obtained using LSTM and an attention mechanism.

[0020] (6) Group opinion prediction

[0021] Represent the active text features obtained in step (1) and the group features obtained in step (5) The data is then concatenated and input into a classifier composed of multilayer perceptrons for classification, ultimately yielding the prediction result.

[0022] Furthermore, the method of the present invention also includes a system function display step, that is, the results obtained in step (6) are visualized and analyzed on a web page, and the accuracy of this method compared with other methods is given.

[0023] Furthermore, the activity text mentioned in step (1) includes the text title of the activity and a brief textual description of the activity, and the word count is required to be no more than 160 words;

[0024] In step (1), BERT is used to encode the active text. The feature representation of an active text, i.e. the feature representation of the active text, has a dimension of 1*768. The sentence length processed by BERT is set to 160. In the pre-training process of the active text, four methods are used for training: feature concatenation of the last four layers of BERT, maximum pooling of the last four layers of features, feature of the last layer, and the output of the last layer plus LSTM.

[0025] When representing the active text supported by the group, for the split words Average pooling is used as its encoding representation, as shown in Formula 1.

[0026]

[0027] in, The output of words in the BERT vocabulary, n w This represents the number of words that have corresponding output in this BERT table. This indicates something that is not in the vocabulary.

[0028] Furthermore, the specific steps of step (2) are as follows:

[0029] The forgetting curve f(t) is constructed to represent the change in the importance of the active text over time, as shown in Formula 2.

[0030]

[0031] Where f(t) represents the importance of the activity over time, k0, c, and t0 are constants, and t represents time. For an activity proposed by the activity initiator, the activity text features are represented... Sort the data by time, multiply it by the forgetting function designed above, and sum the results to obtain the user feature representation u. self Furthermore, since the opposing text contradicts the user's viewpoint, average pooling was used to obtain the opposing text u. nay and from u self Subtract u from ′nay , obtain the user's initialization representation u self .

[0032] Furthermore, step (3) specifically includes:

[0033] First, we need to establish the topological relationship between users based on the relationship between the initiator and the co-initiator, and obtain the topological relationship graph G1 between users. Based on G1, we obtain the propagation path of influence. In step (2), we obtain the user initialization representation u. self That is, there are initial features. The initial representation for graph convolution computation is the output of the user-initialized representation, as shown in Equation 3.

[0034]

[0035] Self-driving ability is represented by Formula 4.

[0036]

[0037] in This represents the output after l1+1 convolution operations, taking the output u of the last layer of the network. effected Let σ(·) represent user self-driven behavior, and let σ(·) represent the activation function. Where A is the adjacency matrix of topological relation G1, and I is the identity matrix. for The degree matrix, These are the parameters that the model needs to learn.

[0038] Furthermore, the specific method for calculating the domain representation of the user in step (4) is as follows:

[0039] A Gaussian Mixture Model (GMM) is used to cluster user domains. GMM can classify a user into multiple domains. The specific algorithm is shown in Equation 5.

[0040]

[0041] Where p(x) represents the distribution of the Gaussian mixture model, k cluster Represents the number of categories. The representative observation data belongs to the i-th cluster The mixing coefficient of each category Let be the probability density function of a random vector x that follows a Gaussian distribution, where The mean vector represents the data. The optimization formula of the GMM clustering algorithm cannot be directly solved analytically; therefore, the EM (Expectation Maximization Algorithm) algorithm is often used for iterative optimization.

[0042]

[0043] in, This represents the GMM optimization objective, where N represents the number of users in the network;

[0044] Through the clustering process described above, users are divided into different interest domains. However, the influence of these domains has not yet spread across different domains. Therefore, the divided domains are abstracted as nodes in a graph, and a graph G2 with domains as nodes is constructed. The specific construction process is as follows: For users in multiple domains, they are regarded as central nodes, connecting two or more domains. The self-driving nature of users indicates that, under the action of the GMM algorithm, they will cluster into different domains. Then, the domains are abstracted as nodes in the graph, and users spanning multiple domains are used as anchor points to connect two or more nodes in G2.

[0045] Before using a graph attention network, the parameters of the nodes need to be initialized. The initialization method for each node in the constructed graph G2 is as follows: an attention mechanism is used to fuse user features within an interest domain as the representation of the current domain. As shown in Formula 7,

[0046]

[0047] in This represents the user in the current domain, and Attention(·) represents the attention mechanism.

[0048] Subsequently, a graph attention network is used in the constructed G2 to obtain a representation of the group influence on the user. The calculation process of the attention coefficient is shown in Equation 8.

[0049]

[0050] in Representing a group For groups Attention coefficient Representatives and groups i g Connected groups, k g Representing the groups within it, Representation and group Connected groups j g =1,2,...,n pairs The resulting impact weight;

[0051] This weight needs to be obtained by linearly transforming the features of all nodes adjacent to a given node, and then applying the LeakeyReLU activation function, as shown in equations 9 and 10.

[0052]

[0053]

[0054] in This represents the characteristics of a group after a linear transformation. and Let l1 be the matrix of linear transformation, and l2 and l3 be the network layer numbers, respectively.

[0055] The calculation process for GAT is shown in Formula 11.

[0056]

[0057] The influence on the target domain comes from the spread of influence from neighboring domains. Representatives and groups i cluster Connected groups, j cluster Representing the groups within it, Indicate j cluster The influence coefficient of the current domain on the current domain, l4 is the number of layers in the network, and after passing through the activation function σ, the output of layer l4+1 is obtained. The final layer output is the feature after being influenced by the domain.

[0058] To overlay the obtained domain features onto each user's feature representation, a user-domain relationship map is used. ug To determine a user's domain, for users belonging to only one domain, the obtained domain features are used as part of their overall characteristics. For users belonging to multiple domains (i.e., possessing multiple domain features), average pooling is used to incorporate these multiple domain features as part of the user's domain influence, resulting in the user's domain-specific representation, u. group As shown in Formula 12,

[0059]

[0060] Where N cluster This indicates the number of domains a user belongs to.

[0061] Furthermore, the group feature fusion described in step (5) specifically involves using an attention mechanism to fuse the three features obtained in steps (2), (3), and (4) respectively: the user initialization representation, the user self-driven representation, and the user domain representation, to obtain the final representation S of user i. i As shown in Formula 13.

[0062]

[0063] in This indicates that user i, based on their own interests, proactively proposed some activity plans, which is the initialization representation; For user i's self-driven characteristic, after the user initiates an activity, they spontaneously seek support from other users, which means the direct impact of actively initiating cooperation with other users; Let i be the domain representation of user i, indicating that the activity proposed by the user belongs to a certain domain. There must be other users in this domain of interest, and the proposal of the activity plan will also be influenced by other users in the domain.

[0064] Features are fused using a Long Short-Term Memory (LSTM) network and an attention mechanism to reflect the temporal characteristics of co-initiators joining the group.

[0065] The output h is obtained by using a Long Short-Term Memory network that satisfies the temporal characteristics. lstm Once the hidden states of the LSTM are output, an attention mechanism can be used to fuse the features, thereby obtaining group features. As shown in formulas 14-16,

[0066] h u =ReLU(W sg h lstm (14)

[0067] e i =W gh h u (15)

[0068]

[0069] Where h u Represents h lstm The result after linear transformation and ReLU, e i Let α represent the attention coefficient for user i. i Let W represent the attention weight after normalization for user i, τ(i) represent the user set of the group to which user i belongs, and W... sg W gh Indicates the parameters that the model needs to learn;

[0070] Through α i Hidden state features corresponding to LSTM By performing multiplication, the group features can be obtained, as shown in Formula 17.

[0071]

[0072] Its h lstmThe features of users in the group are represented by LSTM encoding, and N represents the number of users.

[0073] Furthermore, the specific method for predicting group opinions in step (6) is as follows:

[0074] The group characteristics obtained in step (5) and the active text feature representation obtained in step (1) By splicing the data and feeding it into a classifier composed of a multilayer perceptron, the probability p that the group supports initiating the activity can be obtained.

[0075]

[0076] Among them W gb The parameter to be learned is the current group's opinion on the corresponding activity text, whether it holds a positive or negative attitude. If the opinion is positive, the model believes that the current group will act as a co-initiator and propose the activity; otherwise, the model believes that the current group will not initiate the activity.

[0077] Furthermore, the system's functional demonstration includes the display of supplementary data, the display of model performance comparison results, and the display of group opinion prediction results. In the data supplementation display section, the platform can upload group information, including user domain information and information on previously initiated or rejected activities, to supplement the model training dataset. The model performance comparison result display section provides performance analysis results of the model proposed in this patent compared with other relevant models. In the group opinion prediction display section, the platform can select a group and upload a new activity text, and the system can automatically determine whether the group will support the initiation of the activity.

[0078] This invention also provides a content-oriented group opinion prediction system, including a data management and storage module, a data preprocessing module, a model training module, and a user interaction module. The data management and storage module is responsible for data supplementation and storage; the data supplementation function uploads new data through the platform to supplement the original dataset; the data storage function is responsible for storing relevant original data, preprocessed data, related datasets, and the finally trained model; the data preprocessing module preprocesses the original dataset for subsequent model training; the model training module builds and trains the model, including model parameter initialization, iterative input, and parameter updates; finally, the user interaction module is mainly responsible for receiving and processing user requests and visualizing the description results.

[0079] After extensive experimental testing, it was finally proven that the invention has higher accuracy than other group recommendation technologies and opinion prediction technologies.

[0080] Beneficial effects: Compared with the prior art, the present invention adopts the above technical solution and has the following advantages:

[0081] (1) The self-driven characteristics of users are modeled. This invention designs a method to represent user self-driven characteristics, which can capture the process of activity initiators spontaneously seeking other users with similar interests.

[0082] (2) Taking into account the timeliness of the proposed activities, the present invention relates to a forgetting function that can effectively reflect the importance of the timeliness of the activities and will have a greater impact on recently proposed activities;

[0083] (3) The order in which users join the group is considered, and the model is modeled and integrated with the group’s perspective using LSTM and attention mechanism, which improves the accuracy of the model. Attached Figure Description

[0084] Figure 1 This is the overall framework diagram of the present invention;

[0085] Figure 2 This is a diagram illustrating user relationship building. Figure 2 (a) represents the bipartite graph relationship between the user and the activity text. Figure 2 (b) represents the topology construction diagram of influence among users. Figure 2 (c) indicates the influence propagation relationship between users;

[0086] Figure 3 This is a diagram illustrating the construction of a domain network. Figure 3 (a) represents the initial state where users have not yet been clustered. Figure 3 (b) represents domain clustering based on user self-driven behavior. Figure 3 (c) indicates the established domain topology;

[0087] Figure 4 This is a diagram illustrating the spread of influence in a specific domain. Figure 4 (a) represents the constructed domain topology. Figure 4 (b) indicates that GAT is used for convolution in the domain topology to obtain the user's domain features. Figure 4 (c) indicates that the user is affected by the domain;

[0088] Figure 5 This is the system result display interface of the present invention. Detailed Implementation

[0089] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0090] The following is only one embodiment of the present invention. The present invention has many other embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art can make various corresponding changes and modifications according to the present invention. All such corresponding changes and modifications should fall within the protection scope of the appended claims.

[0091] The method of this invention mainly consists of three modules: an individual feature extraction module, a group feature extraction module, and a group opinion prediction module. The individual feature extraction module's main function is to extract relevant user features at the individual level. It can be further divided into three sub-modules: an activity description text feature extraction module, a user self-driven feature extraction module, and a user domain feature extraction module. The activity description text feature extraction module processes the activity text to obtain activity text features and forms an initial user representation. The user self-driven feature extraction module extracts user self-driven representations based on the constructed cooperation network. The user domain feature extraction module extracts user domain representations based on a domain clustering graph. The group feature extraction module integrates the three features extracted by the individual feature extraction module into group features, which are then fed into the group opinion prediction module for group opinion prediction. A detailed flowchart is shown below. Figure 1 As shown, the detailed description method includes the following steps:

[0092] 1. Feature extraction from activity description text

[0093] (1) Text feature extraction

[0094] In the application scenario of this invention, the activity text content supported by a group is crucial. Therefore, this invention treats the domain to which the activity belongs as a label, and the text describing the activity as training content. The classification task and training involve classifying the activity text according to its different categories, i.e., preprocessing the activity text. This stage uses the BERT model for pre-training. First, the activity content is segmented into words, resulting in a word list. Then, through word mapping, the ID of each word is obtained, and the word ID is passed into BERT to complete the encoding of the activity text.

[0095] When representing the activity text supported by a group, some proper nouns may appear in the activity text that do not exist in the BERT vocabulary. By default, BERT will split non-existent words into existing words, which can affect the accuracy of the judgment. For example, after the outbreak of the novel coronavirus, discussions about the coronavirus and nucleic acid testing increased in groups, so the activities supported by the group would naturally involve the coronavirus. However, "Coronavirus" does not exist in the BERT model, and BERT will default to treating it as an unrecognized word. Breaking down into the superposition of the root words of multiple words, that is Once a word is split, its corresponding semantics will be affected. In this invention, for the split words... Average pooling is used as its encoding representation, as shown in Formula 1, where, The output of words in the BERT vocabulary. This indicates something that is not in the vocabulary.

[0096]

[0097] This task is a multi-class classification task, therefore this paper adopts the multi-class cross-entropy loss function, where Indicates the i-th b Each activity text has a corresponding activity text feature representation as follows: It is the output of BERT, i.e., the feature representation of the active text; y represents the probability of the current activity's predicted category; ic W represents the true category. b b0 are the parameters that the model needs to learn. Represents the loss function, N b M represents the number of text samples during training. i This indicates the amount of activity text related to the user.

[0098]

[0099]

[0100]

[0101] (2) User initialization representation

[0102] There are two types of interaction between users and the event text: supporting edges between users and the event text, and opposing edges between dissenting users and the event text. The edges between users and the event text can be further refined into edges between the event initiator and co-initiator and the event text. This invention uses data from user-initiated and dissenting texts to initialize user representations. Since activities popular with a group are time-sensitive, activities supported by the group more recently are more important, while activities proposed earlier have less impact on the group and users. This invention designs and implements a time-influence function based on the forgetting curve, as shown in Equation 2.

[0103]

[0104] Where f(t) represents the importance of the activity over time, k0, c, and t0 are constants, and t represents time. For an activity proposed by the activity initiator, the text features are represented as follows: Sort the data by time, multiply it by the forgetting function designed above, and sum the results to obtain the user feature representation u. self Furthermore, since the opposing text contradicts the user's viewpoint, average pooling was used to obtain the opposing text u. nay and from u self Subtract u from ′ nay , obtain the user's initialization representation u self .

[0105] 2. User self-driven feature extraction

[0106] First, a topological relationship graph G1 between users needs to be established based on the relationship between the initiator and co-initiators. The establishment process is as follows: Figure 2 As shown. Figure 2 (a) Represents the bipartite relationship between a user and the activity text, where circles represent text and dashed lines indicate that the user is the initiator of the activity. i A solid line indicates that this user is a co-initiator of the activity. k ,u k+1 ,..., Where n u The ellipse represents the number of co-sponsors; it is a group consisting of the initiator and co-sponsors. Figure 2 (b) Representing the user influence topology construction graph by jointly proposing activity texts As a bridge, the event initiator is the starting point, and the co-initiators are the ending point, building a relationship for the spread of influence among users, and the effect is as follows: Figure 2 As shown in (c), in each group, an ego network is formed with the initiator at the center, and the edges in the network are unidirectional edges from the initiator to the co-initiator.

[0107] Based on the topological relationships between users, the propagation path of influence can be obtained. In step 1, we obtained the user initialization representation u. self That is, there are initial features. The initial representation for graph convolution computation is the output of the user-initialized representation, as shown in Equation 3.

[0108]

[0109] Self-driving ability is represented by Formula 4.

[0110]

[0111] in This represents the output after l1+1 convolution operations, taking the output u of the last layer of the network. effected Let σ(·) represent user self-driven behavior, and let σ(·) represent the activation function. Where A is the adjacency matrix of topological relation G1, and I is the identity matrix. for The degree matrix, These are the parameters that the model needs to learn. The specific modeling process for user self-drive is shown in Table 1.

[0112] Table 1 shows the pseudocode for the user self-driven representation algorithm.

[0113]

[0114]

[0115] 3. User domain feature extraction

[0116] The users involved in the application scenarios of this invention do not belong to just one domain; therefore, traditional clustering methods are not applicable. This invention employs a Gaussian Mixture Model (GMM) to cluster user domains. GMM can group a user into multiple domains, and the specific algorithm is shown in Equation 5.

[0117]

[0118] Where p(x) represents the distribution of the Gaussian mixture model, k cluster Represents the number of categories. The representative observation data belongs to the i-th cluster The mixing coefficient of each category Let be the probability density function of a random vector x that follows a Gaussian distribution, where The mean vector represents the data. The optimization formula of the GMM clustering algorithm cannot be directly solved analytically; therefore, the EM (Expectation Maximization Algorithm) algorithm is often used for iterative optimization.

[0119]

[0120] in, This represents the GMM optimization objective, where N represents the number of users in the network;

[0121] Through the clustering process described above, users can be divided into different interest domains. However, the influence of these domains has not yet spread across different domains. Therefore, this invention abstracts the divided domains as nodes in a graph, constructing a graph G2 with domains as nodes. The specific construction process is as follows: For users in multiple domains, they are considered as central nodes, connecting two or more domains, such as... Figure 3 As shown, where, Figure 3 (a) represents the initial state where users have not yet been clustered; Figure 3(b) The self-driven nature of users indicates that under the action of the GMM algorithm, they will cluster into different categories (domains). In order for the clustered categories to form another graph, this invention abstracts the domains into nodes in the graph and connects two or more nodes in G2 by using users who span multiple categories as anchor points. Figure 3 (c) represents the established domain topology.

[0122] Before using a graph attention network, the parameters of the network nodes need to be initialized. The initialization method for each node in the constructed graph G2 is as follows: an attention mechanism is used to fuse user features within an interest domain as the representation of the current domain. As shown in Formula 7,

[0123]

[0124] in This represents the user in the current domain, and Attention(·) represents the attention mechanism.

[0125] Subsequently, a graph attention network is used in the constructed G2 to obtain a representation of the group influence on the user. The calculation process of the attention coefficient is shown in Equation 8.

[0126]

[0127] in Representing a group For groups Attention coefficient Representatives and groups i g Connected groups, k g Representing the groups within it, Representation and group Connected groups j g =1,2,...,n pairs The resulting impact weight;

[0128] This weight needs to be obtained by linearly transforming the features of all nodes adjacent to a given node, and then applying the LeakeyReLU activation function, as shown in equations 9 and 10.

[0129]

[0130]

[0131] in This represents the features of a group after a linear transformation. and Let l1 be the matrix of linear transformation, and l2 and l3 be the network layer numbers, respectively.

[0132] The calculation process for GAT is shown in Formula 11.

[0133]

[0134] The influence on the target domain comes from the spread of influence from neighboring domains. Representatives and groups i cluster Connected groups, j cluster Representing the groups within it, Indicate j cluster The influence coefficient of the current domain on the current domain, l4 is the number of layers in the network, and after passing through the activation function σ, the output of layer l4+1 is obtained. The final layer output is the feature after being influenced by the domain.

[0135] exist Figure 4 (c) indicates that the user is affected by the domain and needs to be... Figure 4 The domain features obtained in (b) are superimposed on the feature representation of each user. To achieve this, the present invention designs and implements a Back function, which maps the relationship between users and the domain. ug To determine a user's domain, for users belonging to only one domain, the domain features are used as part of their overall characteristics. For users belonging to multiple domains (i.e., possessing multiple domain features), average pooling is used to incorporate these multiple domain features as part of the user's domain influence, as shown in Formula 12, where N... cluster This indicates the number of domains a user belongs to.

[0136]

[0137] The overall implementation process based on clustering for domain partitioning and group influence calculation is shown in Table 2.

[0138] Table 2 shows the pseudocode for the user domain representation algorithm.

[0139]

[0140] 4. Integration of group viewpoints

[0141] Steps 1, 2, and 3 respectively introduce the four parts of feature extraction from the active text, user initialization representation, user self-drive, and user domain modeling. Step 4 uses the three user features obtained above as the user representation. In order to dynamically adjust the weights of these three features, this invention uses an attention mechanism to fuse these three features, such as the public...

[0142] As shown in Equation 13.

[0143]

[0144] in This indicates that the user proactively proposes some activity plans based on their own interests; this is the user initialization representation.

[0145] For user i's self-driven characteristic, after the user initiates an activity, they spontaneously seek support from other users, which means the direct impact of actively initiating cooperation with other users; Let i be the domain-specific representation of user i, indicating that the activity proposed by the user belongs to a certain domain. Within this domain of interest, other users inevitably exist, and the proposed activity plan will be influenced by other users within the domain. This invention will construct the final group opinion prediction model based on the above three aspects of user characteristics and activity text features. The role of the group feature part is to obtain group characteristics. To obtain group characteristics, it is necessary to move from individuals to the group. The modeling of individual users has been completed in the first three steps, that is, fusing individual opinion features to obtain group opinion features. This invention uses a Long Short-Term Memory (LSTM) network and an attention mechanism to fuse features to reflect the temporal characteristics of co-initiators joining the group.

[0146] The reason for using Long Short-Term Memory (LSTM) networks before attention is that when an active text is presented, the initiating user may find some other users in the short term. Even if the campaign text hasn't been approved by the sponsors or community platform, the initiating user can still persuade other users to join as initiators. Therefore, the order in which users join exists, so a Long Short-Term Memory (LSTM) network that satisfies temporal characteristics is used for computation to obtain the output h. lstm Once the hidden states of the LSTM are output, an attention mechanism can be used to fuse the features, thereby obtaining group features. As shown in formulas 14-16, where e i α represents the attention coefficient. i W represents the normalized attention weights. i This represents the parameters that the model needs to learn.

[0147] h u =ReLU(W sg h lstm (14)

[0148] e i =W gh h u (15)

[0149]

[0150] Where h u Represents h lstm The result after linear transformation and ReLU, e i Let α represent the attention coefficient for user i. i Let W represent the attention weight after normalization for user i, τ(i) represent the user set of the group to which user i belongs, and W... sg W gh Indicates the parameters that the model needs to learn;

[0151] Through α i Hidden state features corresponding to LSTM By performing multiplication, as shown in Formula 17, the group features can be obtained.

[0152]

[0153] Its h lstm The features of users in the group are represented by LSTM encoding, and N represents the number of users.

[0154] 5. Group opinion prediction

[0155] The group features obtained in step 4 and the active text features obtained in step 1 By splicing the data and feeding it into a classifier composed of a multilayer perceptron, the probability p that the group supports initiating the activity can be obtained.

[0156]

[0157] Among them W gb The parameter to be learned is the current group's opinion on the corresponding activity text, whether it holds a positive or negative attitude. If the opinion is positive, the model believes that the current group will act as a co-initiator and propose the activity; otherwise, the model believes that the current group will not initiate the activity.

[0158] After extensive experimental testing, it was finally proven that the invention has higher accuracy than other group recommendation technologies and opinion prediction technologies.

[0159] 6. System Function Demonstration

[0160] The system functionality demonstration includes displays of data supplementation, model performance comparison results, and group opinion prediction results. Details are as follows:

[0161] (1) Data supplementation and display: The platform can upload group information (including user domain information, etc.) and information on activities that have been initiated or rejected in the past to supplement the model training dataset.

[0162] (2) Display of model performance comparison results, providing performance analysis results of the model proposed in this patent and other relevant comparative models;

[0163] (3) Group opinion prediction display: The platform can select a group and upload a new activity copy information. The system can automatically determine whether the group will support the launch of the activity.

[0164] The group opinion prediction system of the present invention includes a data management and storage module, a data preprocessing module, a model training module, and a user interaction module, as detailed below.

[0165] (1) Data Management and Storage Module: This module is mainly responsible for data supplementation and data storage. The data supplementation function uploads new data through the platform to supplement the original dataset; the data storage function is responsible for storing the relevant original data, preprocessed data, related datasets, and the finally trained model.

[0166] (2) Data preprocessing module, which preprocesses the original data of the dataset for subsequent model training.

[0167] (3) In the model training module, the model is built and trained, including model parameter initialization, iterative input and parameter update.

[0168] (4) User interaction module, which is mainly responsible for receiving and processing user requests and visualizing the description results.

Claims

1. A content-oriented method for predicting group opinions, comprising the following steps: (1) Text feature extraction The activity text is preprocessed and pre-trained using BERT, which is then used for classification training according to different domains of the activity text to obtain the feature representation of the activity text. (2) User initialization representation A forgetting curve f(t) is constructed to represent the change in the importance of the active text over time. The corresponding forgetting curve value and the active text features obtained in step (1) are then combined. Multiply and sum to obtain the user feature representation u. self '; Subtract the opposing text u after average pooling nay Get the user-initialized representation u self ; (3) User self-driving representation A topological graph G1 of user relationships is established using the relationships between users, and then a two-layer convolutional neural network is used to obtain the user self-driving representation u. effected The initial features of the convolutional neural network user are the user initialization representation u obtained in step (2). self ; (4) User domain representation Based on each user's different domain, the GMM algorithm is used to perform overlapping clustering of the user's domains and construct a domain graph G2. Then, GAT is used to fuse the features, and finally obtain the features of the user's domain, i.e., the user's domain representation u. group ; (5) Integration of group characteristics Based on the outputs of steps (2), (3), and (4), a weighted summation is performed using an attention mechanism to obtain the user representation S in the group. Then, the group features are obtained using LSTM and an attention mechanism. (6) Group opinion prediction Represent the active text features obtained in step (1) and the group features obtained in step (5) The data is then concatenated and input into a classifier composed of multilayer perceptrons for classification, ultimately yielding the prediction result.

2. The content-oriented group opinion prediction method according to claim 1, characterized in that, It also includes the step of system function demonstration, that is, the results obtained in step (6) are visualized and analyzed on the web page, and the accuracy of this method compared with other methods is given.

3. The content-oriented group opinion prediction method according to claim 1, characterized in that, The activity text mentioned in step (1) includes the text title of the activity and a brief textual description of the activity, and the word count is required to be no more than 160 words; In step (1), BERT is used to encode the active text. The feature representation of an active text, i.e. the feature representation of the active text, has a dimension of 1*768. The sentence length processed by BERT is set to 160. In the pre-training process of the active text, four methods are used for training: feature concatenation of the last four layers of BERT, maximum pooling of the last four layers of features, feature of the last layer, and the output of the last layer plus LSTM. When representing the active text supported by the group, for the split words Average pooling is used as its encoding representation, as shown in Formula 1. in, The output of words in the BERT vocabulary, n w This represents the number of words that have corresponding output in this BERT table. This indicates something that is not in the vocabulary.

4. The content-oriented group opinion prediction method according to claim 1, characterized in that, Step (2) consists of the following steps: The forgetting curve f(t) is constructed to represent the change in the importance of the active text over time, as shown in Formula 2. Where f(t) represents the importance of the activity over time, k0, c, and t0 are constants, and t represents time. For an activity proposed by the activity initiator, the activity text features are represented... Sort the data by time, multiply it by the forgetting function designed above, and sum the results to obtain the user feature representation u. self Furthermore, since the opposing text contradicts the user's viewpoint, average pooling was used to obtain the opposing text u. nay and from u self 'Subtract u from the middle nay , obtain the user's initialization representation u self .

5. The content-oriented group opinion prediction method according to claim 1, characterized in that, Step (3) specifically includes: First, we need to establish the topological relationship between users based on the relationship between the initiator and the co-initiator, and obtain the topological relationship graph G1 between users. Based on G1, we obtain the propagation path of influence. In step (2), we obtain the user initialization representation u. self That is, there are initial features. The initial representation for graph convolution computation is the output of the user-initialized representation, as shown in Equation 3. Self-driving ability is represented by Formula 4. in This represents the output after l1+1 convolution operations, taking the output u of the last layer of the network. effected For user-driven representation, U(·) represents the activation function. Where A is the adjacency matrix of topological relation G1, and I is the identity matrix. for The degree matrix, These are the parameters that the model needs to learn.

6. The content-oriented group opinion prediction method according to claim 1, characterized in that... The specific method for calculating the domain representation of the user in step (4) is as follows: A Gaussian Mixture Model (GMM) is used to cluster user domains. GMM assigns a user to multiple domains, and the specific algorithm is shown in Equation 5. Where p(x) represents the distribution of the Gaussian mixture model, k cluster Represents the number of categories. The representative observation data belongs to the i-th cluster The mixing coefficient of each category Let be the probability density function of a random vector x that follows a Gaussian distribution, where The mean vector representing the data cannot be directly solved analytically using the GMM clustering algorithm's optimization formula. Therefore, the EM algorithm is used for iterative optimization to find the solution. in, This represents the GMM optimization objective, where N represents the number of users in the network; Through the above clustering process, users are divided into different interest domains. However, the domain influence has not yet spread between different domains. Therefore, the divided domains are abstracted as nodes in the graph, and a graph G2 with domains as nodes is constructed. The specific construction process is as follows: For users in multiple domains, they are regarded as central nodes, connecting two or more domains. The user's self-driving nature means that under the action of the GMM algorithm, they will cluster into different domains. Then, the domains are abstracted as nodes in the graph, and users who span multiple domains are used as anchor points to connect two or more nodes in G2. Before using a graph attention network, the parameters of the nodes need to be initialized. The initialization method for each node in the constructed graph G2 is as follows: an attention mechanism is used to fuse user features within an interest domain as the representation of the current domain. As shown in Formula 7, in This represents the user in the current domain, and Attention(·) represents the attention mechanism. Subsequently, a graph attention network is used in the constructed G2 to obtain a representation of the group influence on the user. The calculation process of the attention coefficient is shown in Equation 8. in Representing a group For groups Attention coefficient Representatives and groups i g Connected groups, k g Representing the groups within it, Representation and group Connected groups right The resulting impact weight; This weight needs to be obtained by linearly transforming the features of all nodes adjacent to a given node, and then applying the LeakeyReLU activation function, as shown in equations 9 and 10. in This represents the features of a group after a linear transformation. and Let l1 be the matrix of linear transformation, and l2 and l3 be the network layer numbers, respectively. The calculation process for GAT is shown in Formula 11. The influence on the target domain comes from the spread of influence from neighboring domains. Representatives and groups i cluster Connected groups, j cluster Representing the groups within it, Indicate j cluster The influence coefficient of the current domain on the current domain, l4 is the number of layers in the network, and after passing through the activation function σ, the output of layer l4+1 is obtained. The final layer output is the feature after being influenced by the domain. To overlay the obtained domain features onto each user's feature representation, a user-domain relationship map is used. ug To determine a user's domain, for users belonging to only one domain, the obtained domain features are used as part of their overall characteristics. For users belonging to multiple domains (i.e., possessing multiple domain features), average pooling is used to incorporate these multiple domain features as part of the user's domain influence, resulting in the user's domain-specific representation, u. group As shown in Formula 12, Where N cluster This indicates the number of domains a user belongs to.

7. The content-oriented group opinion prediction method according to claim 1, characterized in that, The group feature fusion in step (5) specifically involves using an attention mechanism to fuse the three features obtained in steps (2), (3), and (4)—the user initialization representation, the user self-driven representation, and the user domain representation—to obtain the final representation S of user i. i As shown in Formula 13. in This indicates that user i, based on their own interests, proactively proposed some activity plans, which is the initialization representation; For user i's self-driven characteristic, after the user initiates an activity, they spontaneously seek support from other users, which means the direct impact of actively initiating cooperation with other users; Let i be the domain representation of user i, indicating that the activity proposed by the user belongs to a certain domain. There must be other users in this domain of interest, and the proposal of the activity plan will also be influenced by other users in the domain. Feature fusion was performed using a Long Short-Term Memory (LSTM) network and an attention mechanism to reflect the temporal characteristics of co-initiators joining the group. The output h is obtained by using a Long Short-Term Memory network that satisfies the temporal characteristics. lstm Once the hidden states of the LSTM are output, an attention mechanism can be used to fuse the features, thereby obtaining group features. As shown in formulas 14-16, h u =ReLU(W sg h lstm ) (14) have been i =W gh h u (15) Where h u Represents h lstm The result after linear transformation and ReLU, e i Let α represent the attention coefficient for user i. i Let W represent the attention weight after normalization for user i, τ(i) represent the user set of the group to which user i belongs, and W... sg W gh Indicates the parameters that the model needs to learn; Through α i Hidden state features corresponding to LSTM By performing multiplication, the group features can be obtained, as shown in Formula 17. Its h lstm The features of users in the group are represented by LSTM encoding, and N represents the number of users.

8. The content-oriented group opinion prediction method according to claim 1, characterized in that, The specific method for predicting group opinions in step (6) is as follows: The group characteristics obtained in step (5) and the active text feature representation obtained in step (1) By splicing the data and feeding it into a classifier composed of a multilayer perceptron, the probability p that the group supports initiating the activity can be obtained. Among them W gb The parameter to be learned is the current group's opinion on the corresponding activity text, whether it holds a positive or negative attitude. If the opinion is positive, the model believes that the current group will act as a co-initiator and propose the activity; otherwise, the model believes that the current group will not initiate the activity.

9. The content-oriented group opinion prediction method according to claim 2, characterized in that, The system's functional demonstration includes the display of supplementary data, the display of model performance comparison results, and the display of group opinion prediction results. In the data supplementation display section, the platform uploads group information, including user domain information and information on previously initiated or rejected activities, to supplement the model training dataset. The model performance comparison results display section provides performance analysis results of the proposed model compared with other relevant models. In the group opinion prediction display section, the platform selects a group and uploads a new activity text, and the system can automatically determine whether the group will support the initiation of the activity.

10. A content-oriented group opinion prediction system, used to implement the content-oriented group opinion prediction method as described in claim 1, characterized in that, It includes a data management and storage module, a data preprocessing module, a model training module, and a user interaction module. The data management and storage module is responsible for data supplementation and storage. The data supplementation function uploads new data to the platform to supplement the original dataset. The data storage function is responsible for storing the relevant original data, preprocessed data, related datasets, and the finally trained model. The data preprocessing module preprocesses the raw data in the dataset for subsequent model training. The model training module builds and trains the model, including model parameter initialization, iterative input, and parameter updates. Finally, the user interaction module is responsible for receiving and processing user requests and visualizing the results.