Dynamic social network alignment method based on longitudinal federation and canonical correlation analysis

By using a dynamic spatiotemporal graph autoencoder memory model and longitudinal federated learning, combined with attention mechanisms and LSTM units, the problems of dynamism, sparsity, and privacy leakage in social network alignment are solved, achieving more accurate cross-platform user alignment.

CN119887427BActive Publication Date: 2025-11-07CHONGQING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510069611.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-11-07
Estimated Expiration
2045-01-16

AI Technical Summary

Technical Problem

Existing social network alignment methods are unable to effectively address the dynamic nature of social networks, the sparsity and imbalance of cross-platform network structures, and the differences in cross-domain user data and the risk of privacy leaks, resulting in unstable alignment results and low accuracy.

Method used

We employ a method based on vertical federation and canonical correlation analysis, using a dynamic spatiotemporal graph autoencoder memory model to simulate user spatiotemporal relationships. We construct a user relationship matrix by combining an attention mechanism and LSTM units, and train the model under a vertical federated learning framework. By fusing the user attribute matrix, we achieve cross-platform user alignment.

Benefits of technology

It improves the comprehensiveness and accuracy of user characteristics, effectively avoids the risk of cross-platform data privacy leakage, and achieves more accurate cross-platform user alignment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119887427B_ABST
    Figure CN119887427B_ABST
Patent Text Reader

Abstract

The application belongs to the field of social network analysis, and particularly relates to a dynamic social network alignment method based on longitudinal federation and canonical correlation analysis, comprising the following steps: simulating the spatio-temporal relationship of users through a dynamic spatio-temporal graph self-encoding memory model and an attention mechanism to obtain a user relationship matrix; constructing a user attribute matrix and fusing the user relationship matrix to obtain a user matrix; inputting the user feature matrices of platforms X and Y into a model for training through a training model based on federated learning to obtain a prediction result of cross-domain user alignment; and updating and modeling the dynamic relationship representation in combination with the time sequence characteristics of user relationship and fusing other non-time sequence characteristics for cross-platform user alignment prediction. Through the method, the problems of cross-domain data privacy leakage and social network dynamics can be effectively solved, and finally precise cross-platform network user alignment is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the field of social network analysis, and particularly relates to a dynamic social network alignment method based on longitudinal federation and canonical correlation analysis. BACKGROUND

[0002] Social network alignment is to align the same user on different social networks, also called anchor link prediction. It is widely used in cross-domain recommendation, link prediction and network fusion. In today's information society, social networks have become an important part of people's daily life, and people participate in multiple social media to meet their different spiritual needs. However, with the continuous emergence of various social network platforms, how to effectively integrate the user and relationship information on these platforms to realize cross-platform information sharing and user behavior analysis has become a major challenge in current research. The core of this challenge is social network alignment, that is, matching and associating user information from different social network platforms.

[0003] For social network alignment, existing research methods can be roughly divided into two categories: one is the network alignment method based on graph neural network, which mainly uses the graph structure information of social network and the feature information of users to align users. The second is the heterogeneous social network alignment method based on network representation learning, which uses the structure information in heterogeneous social network and the complex interaction relationship between users to align users. However, the current user alignment still has the following shortcomings:

[0004] 1. Dynamic nature of social networks. The topology of social networks in the real world is constantly changing over time, and this dynamic change makes user behavior and relationship links extremely complex and difficult to accurately capture and predict. Frequent changes in user relationships increase the difficulty of alignment algorithms. Therefore, new technologies and methods need to be explored to effectively deal with network dynamics and improve the accuracy of social network alignment.

[0005] 2. Sparsity and imbalance of cross-platform network structure. Although cross-domain networks may describe the same or similar users and relationships, due to the natural distribution of data, collection methods, and different network construction strategies, the connections between user nodes are relatively sparse, leading to unstable alignment results and low accuracy. At the same time, the imbalance of cross-domain network data increases the risk of privacy leakage of the data disadvantaged party. Therefore, it is an important problem to strengthen data security and alleviate network sparsity under the condition of data imbalance.

[0006] 3. Differences in cross-domain user data leakage and distribution. When aligning cross-platform users, centralized learning has the risk of user privacy data leakage. At the same time, different social platforms have different functions, user groups and network structures, and the distribution of user data may be biased, which may lead to differences in user representation. Therefore, how to ensure that the differences are solved on the basis of data security and precise cross-platform user alignment become the key problems to be solved at present. SUMMARY

[0007] To solve the above technical problems, the application provides a dynamic social network alignment method based on longitudinal federation and canonical correlation analysis, comprising the following steps:

[0008] S1, simulate the user spatio-temporal relationship by a dynamic spatio-temporal graph auto-encoding memory model and an attention mechanism to obtain a user relationship matrix;

[0009] The dynamic spatio-temporal graph auto-encoding memory model is composed of a GAE network and an LSTM unit;

[0010] S2, construct a user attribute matrix and fuse the user relationship matrix to obtain a user matrix;

[0011] S3, input the user feature matrices of platforms X and Y into the model for training by a training model based on federated learning to obtain a prediction result of cross-domain user alignment.

[0012] The application has the following beneficial effects:

[0013] The dynamic spatio-temporal graph auto-encoding memory model can effectively simulate the change of user relationship with time; the application combines network structure enhancement and privacy protection technology to reflect user features from multiple angles, improving the comprehensiveness and accuracy of user features; the application is performed under a longitudinal federated learning framework, combined with canonical correlation analysis technology, effectively avoiding the risk of cross-social platform data privacy leakage, capturing the correlation of user representation between platforms, and more accurately realizing user alignment. BRIEF DESCRIPTION OF DRAWINGS

[0014] Figure 1 The flowchart of the application is shown in the figure;

[0015] Figure 2 The dynamic spatio-temporal graph auto-encoding memory model process is shown in the figure;

[0016] Figure 3 The multi-level attribute embedding diagram is shown in the figure;

[0017] Figure 4 The training model process of federated learning is shown in the figure. DETAILED DESCRIPTION

[0018] With reference to the drawings and embodiments of the present application, the technical solutions in the embodiments of the present application will be described clearly and completely. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the protection scope of the present application.

[0019] A dynamic social network alignment method based on longitudinal federation and canonical correlation analysis, the flow of the present application is as shown in Figure 1 , and specifically includes the following steps:

[0020] S1, simulate the user spatio-temporal relationship through a dynamic spatio-temporal graph auto-encoding memory model and an attention mechanism, to obtain a user relationship matrix;

[0021] The dynamic spatio-temporal graph auto-encoding memory model is composed of LSTM units of a GAE network;

[0022] S2, construct a user attribute matrix, and fuse the user relationship matrix to obtain a user matrix;

[0023] S3, input the user feature matrices of platforms X and Y into a model for training through a training model based on federated learning, to obtain a prediction result of cross-domain user alignment.

[0024] Simulate the user spatio-temporal relationship through a dynamic spatio-temporal graph auto-encoding memory model and an attention mechanism, to obtain a user relationship matrix, as shown in Figure 2 , and specifically includes:

[0025] S11, obtain a user relationship-time series feature binary set G t , which is specifically represented as,

[0026] G t ={(V,E1),(V,E2),...,(V,E T )}={(V,E t )|t∈[0,T]}

[0027] Wherein, the user behavior-time series feature set G t ={(V,E t )|t∈[0,T]} is defined as a set composed of V and E. Wherein:

[0028] V={ν1,ν2,···,ν n}

[0029] E={e ij =(v i ,v j )|v i ,v j ∈V}

[0030] wherein V represents a set of user nodes, E represents a set of n users and a set of undirected edges between users in t time, and each edge e ij represents a binary state between users v i and v j ;

[0031] S12, calculate the input static node embedding Z of the encoding part in the GAE network according to the dynamic user relationship set, specifically represented as:

[0032]

[0033] wherein A represents an adjacency matrix, W0 and W1 are to-be-learned parameters, and D represents a degree matrix of nodes;

[0034] S13, calculate the reconstructed adjacency matrix of the encoding according to the static node embedding Z.

[0035]

[0036] If a good node embedding Z is to be obtained, the reconstructed adjacency matrix should be as similar as possible to the original adjacency matrix A. Therefore, in the training process of GAE, cross-entropy is used as a loss function, specifically represented as:

[0037]

[0038] wherein y represents a value (0 or 1) of an element in the adjacency matrix A, represents a value (between 0 and 1) of a corresponding element in the reconstructed adjacency matrix ;

[0039] S14, for the user relationship-time sequence set G t , a static embedding set {Z t ,t∈[0,T]} is obtained after passing through the GAE network, at time step t=0, the state S0 of the LSTM unit is initialized; from t=1 to t=T, the LSTM unit processes the embedding set Z t of each time step in turn and updates its internal state S t :

[0040] S t = f(S t-1 ,Z t )

[0041] wherein f is a function of the LSTM unit, responsible for updating the state S t-1and the embedding of the current time step Z t updating the current state S t ;

[0042] S141, the output f of the forget gate of the current time step t , specifically expressed as:

[0043] f t = sigmoid(W f · [h t-1 , Z t ] + b f )

[0044] wherein W f represents the weight matrix of the forget gate, h t-1 represents the hidden state of the previous time step, and b f represents the bias term of the forget gate. It is decided by the forget gate which network information will be discarded from the cell state;

[0045] S142, the output i of the input gate of the current time step t and the candidate cell state , specifically expressed as:

[0046] i t = sigmoid(W t · [h t-1 , Z t ] + b t )

[0047]

[0048] wherein W i and W c represent the weight matrix of the input gate and the candidate cell state, and b i and b c represent the bias term of the input gate and the candidate cell state. It is decided by the input gate which new network information in the current time step will be stored in the cell state;

[0049] S143, the cell state C of the current time step t :

[0050]

[0051] wherein C t-1 represents the cell state of the previous time step. Some parts of the old cell state are discarded by the output of the forget gate, and the product of the output of the input gate and the candidate cell state is added to update the cell state;

[0052] S144, the output o of the output gate of the previous time step t , specifically expressed as:

[0053] o t = sigmoid(W O ·[h t-1 , Z t ]+b O )

[0054] wherein W o represents a weight matrix of the output gate, and b o represents a bias term of the output gate. The output gate determines which part of the network information of the next hidden state will be output;

[0055] S145, the LSTM output h t of the current time step, is specifically represented as:

[0056] h t = o t *tanh(C t )

[0057] filtered through a sigmoid function, and then the cell state is passed through tanh and multiplied by the output of the output gate to obtain the required next hidden state h t ;

[0058] S146, the attention weight a t of the t-th time step is calculated according to the LSTM output h t of the t-th time step, and is specifically represented as:

[0059] a t = softmax(tanh(W a ·h t +b a ))

[0060] wherein W a represents a learnable weight matrix, and b a represents a bias term;

[0061] S147, the user relationship embedding Q t of the t-th time step is calculated according to the attention weight a t of the t-th time step and the LSTM output h t of the current time step, and is specifically represented as:

[0062]

[0063] A user attribute matrix is constructed, and a user relationship matrix is fused, as shown in Figure 3 , specifically including the following steps:

[0064] S21, a user attribute set A is constructed:

[0065] A = {A c ,A w ,A t}

[0066] The attribute information of the user includes username, gender, geographical location, educational background and historical generated content, and is divided into three levels: character level A c , word level A w and theme level A t ;

[0067] S22, character level feature matrix:

[0068] S221, for the character level attribute set Use the bag of words (BoW) model for attribute vectorization. The user v i corresponding character level attribute is contains characters, strings and numbers. For construct a vocabulary w = w1, w2, …, w k , …, w m , where k ∈ {1, 2, …, m};

[0069] S222, the vocabulary w vector is expressed as where is the number of w corresponding to w k ; finally, the attribute set A c is vectorized as

[0070]

[0071] The count-weighted matrix X c is reduced using an autoencoder. First, the autoencoder maps the input vector to a latent representation as follows:

[0072]

[0073] where W is the weight matrix and b is the bias vector;

[0074] S223, the obtained latent representation is mapped to a reconstructed vector as follows:

[0075]

[0076] where W * is the weight matrix and b *This is the bias vector. To minimize the average reconstruction loss, the parameters are optimized in this section.

[0077]

[0078] After optimization, a character-level feature matrix is ​​obtained. as follows:

[0079] P c =WX c +b

[0080] In the formula, W and b are the weight matrix and the bias of the autoencoder, respectively;

[0081] S23, Word-level Feature Matrix:

[0082] For word-level attribute sets Use word2vec to capture word-level attribute features of users. User v i The corresponding word-level attribute is Includes gender, geographic location, and educational background.

[0083] S231, Attributes Divide into m word sequences w i =w i1 ,w i2 ,…,w ik ,…,w im In the formula w ik Representing attributes The vocabulary list w for the k-th word;

[0084] S232, User v i Font-level attribute vector It can be represented as w i The word vectors of all words in the text are summed, as shown in formula (23).

[0085]

[0086] S233, Obtain the word-level feature matrix

[0087] S24, Topic-level Feature Matrix:

[0088] For topic-level attribute sets Extracting topic features of user attributes using the latent Dirichlet distribution, user v i Theme-level attributes Composed of paragraphs or articles. The i-th topic-level attribute. It is a topic-level document. i That is, the subject-level attribute set A t Corresponding document set w = w1, w2, ..., wi ,…,w n .

[0089] LDA generates topics from a document by generating a document w according to prior probabilities. i Document w is generated by sampling from the Dirichlet distribution α. i Theme distribution The topic z is generated by sampling from the Dirichlet distribution β. ij Word distribution Multinomial distribution of words and topics Sampled document w i The topic of the j-th word z ij From the multiple distribution of words Mid-sampling generates words w ij .

[0090] Therefore, the joint distribution of all variables in the LDA model is shown below:

[0091]

[0092] Sampling is performed using Gibbs sampling, and the results are updated after each sampling. and θ m,k The value is iterated repeatedly until convergence. After the overall sampling is completed, θ is calculated. m,k Obtain the topic distribution θ of each document in the document set. m As shown below:

[0093]

[0094] Therefore, user v is obtained i Topic-level feature vectors The topic-level attribute matrix for all users is as follows:

[0095] S25, User Attribute Matrix:

[0096] The attribute matrices from the three levels above are merged, and the merged matrix P is processed using a perturbation matrix γ to obtain the user attribute matrix P′. Then, the user relationship matrix Q is merged to obtain the user matrix U. As shown below:

[0097] U=Q+P′=Q+γP=Q+γ(P c +P w +P t ).

[0098] By establishing a federated learning-based training model, user feature matrices from platforms X and Y are input into the model for training, resulting in predictions for cross-domain user alignment. Figure 4As shown, specifically comprising the following steps:

[0099] S31, model training:

[0100] The user matrix U calculated above X and U Y as the input of model training, according to the loss function L CCA , the platform X, Y as the source domain and target domain for federated training. The formula is as follows:

[0101]

[0102] Where, W X represents the projection matrix of X, W Y represents the projection matrix of Y, C XY represents the cross-covariance matrix between X and Y, C XX represents the covariance matrix of X, C YT represents the covariance matrix of Y, λ represents the regularization parameter, Ω(Θ) represents the regularization term, and T represents the matrix transpose.

[0103] The goal of model training is to make the two network parameters W X and W Y such that the two networks have the maximum correlation after CCA calculation. The training steps are as follows:

[0104] S311, initialize the model parameters. Platform X initializes the projection matrix Platform Y initializes the projection matrix

[0105] S312, platform X calculates the typical variable matrix G X = U X W X , and sends G X encrypted to platform Y.

[0106] S313, platform Y calculates the typical variable matrix G Y = U Y W Y . Use the received encrypted G X and local G Y to securely jointly calculate the encrypted cross-covariance matrix [[C XY ]]. Calculate the encrypted residual [[d]] = [[G X -G Y ]]. Where d is the residual, representing the difference between the two typical variable matrices, and [[·]] represents encrypting the data.

[0107] S313, solve the encrypted gradient. Platform X and Y solve the encrypted gradient according to the loss function L CCA, the following encrypted gradients are calculated: The specific formula is:

[0108]

[0109] wherein, denotes the partial derivative, L CCA denotes the loss function L CCA , C XY denotes the cross-covariance matrix between X and Y, C XX denotes the covariance matrix of X, C YY denotes the covariance matrix of Y, and T denotes the matrix transpose.

[0110] S315, upload the solved encrypted gradients to a trusted third-party server for decryption. The decrypted gradients are returned to platforms X and Y respectively. Update the parameters using gradient descent: platforms X and Y update their projection matrices W X and W Y respectively according to the decrypted gradients.

[0111]

[0112] wherein, η denotes the learning rate, t denotes the current training round, and t+1 denotes the next training round.

[0113] Repeat S312 to S315 until the algorithm converges.

[0114] S33, model prediction:

[0115] The third-party server can calculate the maximum correlation according to the following formula, and verify according to the cosine similarity to predict the potential anchor user pair.

[0116]

[0117] wherein, ρ denotes the correlation coefficient, maxcorr() denotes the maximum correlation, U X denotes the user embedding matrix of platform X, U Y denotes the user embedding matrix of platform Y, max denotes the maximum, W X denotes the projection matrix of X, W Y denotes the projection matrix of Y, C XY denotes the cross-covariance matrix between X and Y, C XX denotes the covariance matrix of X, C YY denotes the covariance matrix of Y, and T denotes the matrix transpose.

[0118] While embodiments of the application have been shown and described, it is to be understood that the embodiments described are merely exemplary of the principles and application of the present application. Numerous modifications and adaptions can be effected without departing from the spirit and scope of the present application, which is not limited to the exact construction and arrangement described. It is intended, therefore, to cover all modifications and adaptions that fall within the scope of the claims and their equivalents.

Claims

1. A method for dynamic social network alignment based on longitudinal federated and canonical correlation analysis, characterized in that, Comprise: S1, simulate the user spatio-temporal relationship through a dynamic spatio-temporal graph auto-encoding memory model and an attention mechanism, and obtain a user relationship matrix; The dynamic spatio-temporal graph auto-encoding memory model is composed of a GAE network and an LSTM unit; S2, construct a user attribute matrix, and fuse the user relationship matrix to obtain a user matrix; S3, input the user feature matrices of platforms X and Y into the model for training through a training model based on federated learning, and obtain a prediction result of cross-domain user alignment; Input the user feature matrices of platforms X and Y into the model for training to obtain a prediction result of cross-domain user alignment, comprising: S31, The user matrix U X and U Y As input for model training, according to the loss function L CCA The platform X and Y are used as the source and target domains, respectively, for federated training until the algorithm converges. S32, the third-party server calculates the maximum correlation of the network according to canonical correlation analysis CCA, and then calculates the correlation coefficient between users according to cosine similarity, and predicts the potential anchor user pair, i.e. the aligned user; The user matrix U X and U Y are taken as the input of model training, according to the loss function L CCA , the platform X and Y are taken as the source domain and the target domain for federated training, including: S311, initialize model parameters: platform X initializes the projection matrix Platform Y initializes the projection matrix wherein, denotes the i-th projection vector of the projection matrix W X denotes the i-th projection vector of the projection matrix W Y ​​ S312, Platform X calculates the canonical variable matrix G X =U X W X , will G X After encryption, it is sent to platform Y; where U X This represents the user matrix of platform X; S313, platform Y calculates a typical variable matrix G Y = U Y W Y , uses the received encrypted G X and the local G Y to securely jointly compute an encrypted cross-covariance matrix [[C XY ]]; compute an encrypted residual [[d]] = [[G X -G Y ]]; wherein d represents a residual used to represent the difference between two typical variable matrices, [[ ]] represents encrypting data, U Y represents the user matrix of platform Y; S313, solving the encrypted gradient: the platform X and Y calculate the encrypted gradient according to the loss function L CCA , the encrypted gradient: wherein, denotes a partial derivative, L CCA denotes the loss function L CCA , C XY denotes the cross-covariance matrix between X and Y, C XX denotes the covariance matrix of X, C YY denotes the covariance matrix of Y, T denotes the matrix transpose; S315, upload the solved encrypted gradient to a trusted third-party server for decryption, and the decrypted gradient is returned to platforms X and Y respectively, and platforms X and Y respectively update their projection matrices W according to the decrypted gradient using gradient descent X and W Y : Wherein, η represents the learning rate, t represents the current training round, and t+1 represents the next training round.

2. The method of claim 1, wherein, Simulate the user spatio-temporal relationship through a dynamic spatio-temporal graph auto-encoding memory model and an attention mechanism, comprising: S11. Obtain the user relationship-time series set G t ={(V,E t )|t∈[0,T]};where, V={v1,v2,···,v n } represents a set of nodes, E t Let T represent the set of undirected edges between users, and let T represent the time series. S12. Calculate input static node embeddings of the encoding part in the GAE network according to the user relationship-time sequence set wherein, denotes the symmetric normalized adjacency matrix, ReLU() denotes the activation function, W0denotes the initial weight matrix, and W1denotes the second layer weight matrix. S13. Compute the adjacency matrix of the decoded output reconstruction from the static node embedding Z where sigmoid() denotes the conversion of the scores computed by the decoder into a probability value, and T denotes the matrix transpose. S14, for a user relationship-time sequence set G t , after passing through the GAE network, a static embedding set {Z t , t ∈ [0, T]} is obtained, at time step t = 0, the state S0 of the LSTM unit is initialized; from t = 1 to t = T, the LSTM unit processes the embedding set Z t of each time step in turn and updates its internal state S t : S t = f(S t-1 , Z t ) where f is a function of the LSTM unit responsible for updating the current state S t-1 from the previous time step S t and the embedding Z t of the current time step S15, an attention mechanism is introduced to better capture the importance of different time steps, and the LSTM outputs of all time steps are combined to obtain a user relationship matrix Q: where a t denotes the attention weight at time step t, h t denotes the LSTM output at the current time step.

3. The method of claim 2, wherein, The main process of the S14 step comprises: S141, the first step of LSTM determines which network information will be discarded from the cell state through the forget gate: f t = sigmoid(W f · [h t-1 , Z t ] + b f ) where h t-1 denotes the hidden state of the previous time step, Z t denotes the input of the current time step, f t denotes the output of the forget gate of the current time step, W f denotes the weight matrix of the forget gate, b f denotes the bias term of the forget gate, and [·] denotes the vector concatenation operation. S142, The LSTM decides which new network information at the current time step will be stored in the cell state by the input gate, including: using the sigmoid function to decide which values will be updated and creating a candidate cell state By applying the tanh function: i t = sigmoid(W t · [h t-1 , Z t ] + b t ) where i t denotes the output of the input gate for the current time step, denotes the candidate cell state for the current time step, W t and W c denote the weight matrices for the input gate and the candidate cell state, b t and b c denote the bias terms for the input gate and the candidate cell state; S143, LSTM uses the output of the forget gate to discard some parts of the old cell state, and adds the product of the output of the input gate and the candidate cell state to update the cell state: where C t represents the cell state at the current time step, C t-1 represents the cell state at the previous time step; S144, the LSTM decides through the output gate which part of the network information of the next hidden state will be output, this output will be based on the cell state o t but will be filtered through a sigmoid function and then the cell state o t The desired next hidden state h t is obtained by multiplying the output of the output gate with the output of the tanh function and o t = sigmoid(W O · [h t-1 , Z t ] + b O ) h t = o t tanh(C t ) where W O represents the weight matrix of the output gate, b O represents the bias of the output gate, h t represents the LSTM output at the current time step; S145, update its internal state S t = [C t ,h t ].

4. The method of claim 1, wherein, Construct a user attribute matrix and fuse the user relationship matrix, comprising: S21, construct a user attribute set A = {A c ,A w ,A t}; the user attributes include usernames on two social platforms, gender, geographical location, educational background and user-generated content, and are divided into three levels: character level A c , word level A w and topic level A t ; S22, construct a character-level feature matrix: For the character-level attribute set Using bag-of-words vector model for attribute vectorization, resulting in a count-weighted matrix where, denotes the user v i The corresponding character-level attribute, T denotes the matrix transpose, denotes the bag-of-words vector of the character-level attribute ​ Using an autoencoder on the count-weighted matrix X c Perform dimensionality reduction to get the character-level feature matrix where, denotes the character-level attribute vector of user v i ; S23, construct a word-level feature matrix: For the word-level attribute set Capture the word-level attribute features of the user using word2vec to obtain a word-level feature matrix Wherein, The user v i Corresponding word-level attribute, The word-level attribute vector of the user v i ; S24, construct a topic-level feature matrix: For the topic-level attribute set The topic feature of the user attribute is extracted by using the latent Dirichlet distribution, and the topic-level attribute matrix of all users is obtained as wherein, denotes a topic-level attribute of a user v i , denotes a topic-level feature vector of a user v i . S25, construct a user matrix: The character-level feature matrix, the word-level feature matrix and the topic-level feature matrix are fused to obtain a fusion matrix P, and the fusion matrix P is processed using a perturbation matrix γ to obtain a user attribute matrix P ′ The user relationship matrix Q is fused again to obtain a user matrix U = Q + P ′ = Q + γP = Q + γ(P c + P w + P t ); wherein P c represents the character-level feature matrix, P w represents the word-level feature matrix, and P t represents the topic-level feature matrix.

5. The method of claim 1, wherein, The loss function L CCA , comprises: where W X denotes the projection matrix of X, W Y denotes the projection matrix of Y, C XY denotes the cross-covariance matrix between X and Y, C XX denotes the covariance matrix of X, C YY denotes the covariance matrix of Y, λ denotes a regularization parameter, Ω(Θ) denotes a regularization term, and T denotes matrix transpose.

6. The method of claim 1, wherein, The third-party server calculates the maximum correlation of the network according to canonical correlation analysis CCA, and then calculates the correlation coefficient between users according to cosine similarity, and predicts the potential anchor user pair, comprising: where p denotes the correlation coefficient, maxcorr() denotes maximizing the correlation, U X denotes the user embedding matrix for platform X, U Y denotes the user embedding matrix for platform Y, max denotes maximizing, W X denotes the projection matrix for X, W Y denotes the projection matrix for Y, C XY denotes the cross-covariance matrix between X and Y, C XX denotes the covariance matrix for X, C YY denotes the covariance matrix for Y, T denotes matrix transpose.

Citation Information

Patent Citations

  • Dialogue intention recognition method and training method of model for recognizing dialogue intention

    CN113590798A

  • Signal detection method, device and equipment and computer readable storage medium

    CN118797491A