Data recommendation method based on hypergraph structure entropy pre-training

Through the hypergraph structure entropy pre-training method, hypergraph structure is constructed and optimized, which solves the problem of difficult user-project interaction in the existing recommendation methods, and improves the performance and user experience of the recommendation system.

CN120278159APending Publication Date: 2025-07-08BEIHANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510426729.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

Existing recommendation methods are difficult to effectively capture complex user-project interactions, especially in the absence of user historical behavior data, resulting in information overload and poor recommendation results.

Method used

Using the data recommendation method based on entropy pre-training of hypergraph structure, by constructing a hypergraph structure, using the hypergraph neural network layer to learn user or project node embedding, combining hypergraph structure entropy and pooling operations, the topology of the recommended model is optimized, heterogeneous relationships are captured and hierarchical clustering is performed.

Benefits of technology

It improves the performance of the recommendation system, reduces the pressure of data diffusion, simplifies the social data processing process, and improves information utilization efficiency and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120278159A_ABST
    Figure CN120278159A_ABST
Patent Text Reader

Abstract

The invention discloses a data recommendation method based on hypergraph structure entropy pre-training. The method comprises the following steps: preprocessing collected data; constructing a text graph; aligning nodes and edges in the cross graph; the multi-channel layer is used for mining different dependence characteristics between word pairs; joint training, main task combination and cross graph optimization are carried out; constructing an initial hypergraph, carrying out pre-training, learning items or user node embedding in different hypergraphs by utilizing an optimized hypergraph neural network layer, obtaining priori knowledge from an auxiliary task and a main task, and constructing user or item embedding for a downstream recommendation model; a hypergraph neural network layer is jointly optimized through hypergraph structure entropy and hypergraph pooling, and implementation is achieved from the perspective of hypergraph structure learning and global information capturing. And performing pre-training optimization according to different task types, and outputting a recommendation result. According to the method, information overload is reduced, the community propagation process is enhanced, node embedding is improved, the data diffusion pressure in recommendation is reduced, and the decision and transaction cost of the user is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of platform recommendation, and particularly relates to a data recommendation method based on hypergraph structure entropy pre-training. Background Art

[0002] With the rapid development of information and communication technologies, the problem of information overload has become increasingly serious. Multiple platforms simultaneously disseminate the same content, resulting in individuals being inundated with a large amount of information and making it difficult to effectively screen and utilize relevant information. This phenomenon not only increases the complexity of information processing but may also lead to information fatigue, conflicts, stress, and anxiety, thereby affecting productivity and innovation. Especially in the fields of e-commerce and social media, recommendation systems are widely used to help users filter information and recommend suitable products or content, thereby reducing user decision-making and transaction costs and enhancing the user experience.

[0003] Traditional recommendation methods mainly rely on content-based filtering and collaborative filtering, which recommend by analyzing the similarity between users or items. However, these methods are limited by low-dimensional features and simple classification or sorting algorithms and are difficult to capture complex user-item interaction relationships. Although methods based on matrix factorization and logistic regression have improved the recommendation effect to some extent in sparse data scenarios, these methods still perform poorly in the absence of user historical behavior data. Summary of the Invention

[0004] To solve the above problems, the present invention proposes a data recommendation method based on hypergraph structure entropy pre-training, which reduces information overload, enhances the community propagation process, and improves node embedding, thereby reducing the pressure of data diffusion in the recommendation system and reducing the costs of users in decision-making and transactions.

[0005] To achieve the above object, the technical solution adopted by the present invention is: a data recommendation method based on hypergraph structure entropy pre-training, including the steps of:

[0006] S10, data collection, collecting task data samples; and preprocessing the collected data;

[0007] S20, entity relationship extraction, including the steps of: inputting the preprocessed data into a graph encoder to construct a text graph; aligning nodes and edges in the cross-graph and transferring the matching semantic information from the object to the entity; multi-channel layer, mining different dependency features between word pairs (w i , w j ) to detect the relationship between them; joint training, combining the main task and optimizing the cross-graph;

[0008] S30. Hypergraph structure entropy pre-training. After extraction is completed, construct an initial hypergraph, and then perform pre-training on the hypergraph, including the pre-training process. Use the optimized hypergraph neural network layer to learn the item or user node embeddings in different hypergraphs, obtain prior knowledge from the auxiliary task and the main task, and construct user or item embeddings for the downstream recommendation model; and jointly optimize the hypergraph neural network layer through hypergraph structure entropy and hypergraph pooling, which is achieved from the perspectives of hypergraph structure learning and global information capture.

[0009] S40. Recommendation and optimization. Perform pre-training optimization according to different task types and output the recommendation results.

[0010] Furthermore, the preprocessing of the collected data includes: missing value processing, outlier processing, feature standardization and normalization, text cleaning, word segmentation processing, and stop word removal.

[0011] Furthermore, when constructing the graph text, use the dependency parsing tool to construct the text graph; after parsing, the given sentence is converted into a text graph G T ={V T ,E T}, where V T and E T respectively represent the nodes and edges of the syntactic dependency, and G T is an undirected self-loop graph;

[0012] At the same time, use A T to represent the adjacency mask matrix of the text graph, where represents whether there is an edge between the word pair (w i ,w j );

[0013] In addition, the node V T is input into the text encoder BERT to obtain X T as the text output representation; at the same time, use the edge transition matrix to map the edge type into a trainable vector to obtain the trainable edge matrix Z T ;

[0014] After that, use the attribute attention mechanism to process the data results of the text encoder to update the text graph by integrating the edge type into the keys and values of the self-attention mechanism of the transformer as the attribute transformer.

[0015] Furthermore, use the enhanced edge graph optimal transport method to match the nodes and edges in the cross-graph; use the image-to-text attention mechanism to transfer the matched semantic information from the visual object to the text modality and obtain the improved text representation H T .

[0016] Furthermore, the enhanced edge graph optimization transmission method includes: performing cross-graph matching using two distance metric methods: (1) The Wasserstein distance WD is used for node matching; (2) The Gromov-Wasserstein distance GWD is used for edge matching;

[0017] Obtain the matching nodes H using the Wasserstein distance I to H T of the optimal transport distance D wd (H I , H T );

[0018] Use the Gromov-Wasserstein distance to measure the similarity score D of edges in the cross-graph by calculating the distance between node pairs gwd ;

[0019] Use a unified solver and the Sinkhorn algorithm combined with entropy regularization to iteratively optimize D wd and D gwd , and obtain the objective loss function L for optimizing the cross-graph graph ;

[0020] Then, use the attention mechanism of the text to effectively convert visual semantic information into text representation

[0021] Finally, add the obtained to H T , and perform layer normalization to obtain the final context representation O.

[0022] Furthermore, in the multi-channel layer: After enhanced edge graph alignment, it is divided into three feature matrices including a part-of-speech matrix, a morphological distance matrix, and a word co-occurrence matrix; Each matrix is modeled by a weighted graph convolutional network to obtain the representation of each channel; Each matrix first passes through an embedding layer to obtain a trainable representation R l ;

[0023] Combine the obtained representations and send them to the multi-layer perceptron MLP layer to obtain the final word representation S;

[0024] Connect the enhanced representations of the two word representations S i , w j ) of the word pair (w i and S j to obtain the connection result r i,j ;

[0025] Send r i,j to the linear prediction layer to obtain the probability distribution p i,j ;

[0026] Use cross - entropy error to measure the difference between the baseline true distribution and the predicted label distribution L main ;

[0027] The ultimate goal is the combination of the main task and the optimized cross - graph:

[0028] L = L main + λL graph ;

[0029] where λ is a trade - off hyperparameter used to control the contribution of the optimized cross - graph.

[0030] Furthermore, during the pre - training process, it includes steps:

[0031] Hyper - graph construction, including auxiliary task hyper - graph and main task hyper - graph;

[0032] After the hyper - graph construction is completed, the hyper - graph needs to be learned and trained; all auxiliary task hyper - graphs and main task hyper - graphs are fed into the optimized HGCN layer for training to generate corresponding representations for item and user nodes.

[0033] Furthermore, for the hyper - graph structure entropy, convert the hyper - edge information in the hyper - sub - graph matrix into a node adjacency matrix to transmit hyper - edge information and design the hyper - graph structure entropy; subsequently, by constructing an optimal coding tree, dividing node communities, and combining the principle of minimizing the structure entropy, the best high - dimensional community division is finally achieved; Figure 2 After establishing the main task hyper - graph and the auxiliary task hyper - graph, use the structural information theory to classify closely related users or items into the same community; the structural information theory evaluates the structural uncertainty based on the random walk of nodes through connected edges in the graph.

[0034] Furthermore, learn the topological structure of the hyper - graph through the optimal coding tree, divide the nodes into two - level communities; add pooling and non - cooling layers to enhance the propagation process of node information; the pooling operation is used to aggregate the embedding information of low - level nodes to form the representation of high - level communities; on the other hand, the non - cooling operation propagates the obtained high - level community representation back to the low - level nodes;

[0035]

[0036] ​In the Hypergraph Pooling Encoder (HP), two upsampling modules are adopted, and each module consists of an HGCN layer and a pooling layer; the Hypergraph Uncooling Decoder (HPU) utilizes two downfeedback modules, and each module includes an uncooling layer and an HGCN layer; between the encoder and the decoder, an additional HGCN layer is inserted to facilitate the propagation of the highest-level community information; in order to capture both the global community representation and the direct neighbor dependencies simultaneously, the HGCN layer in the decoder does not directly use the output of the uncooling layer. Instead, it combines the representations of the output of the HGCN layer and the output of the uncooling layer in the corresponding encoder by weighting; this encoder-decoder skip connection operation allows the upper-level community information to be embedded into the embeddings of the lower-level nodes, avoiding the situation where slightly different feedbacks are received by different underlying nodes.

[0037] Furthermore, for the main task, i.e., the recommendation task, the inner product calculation is used to calculate the ranking score Y of the user-item pair. rec :

[0038] Y rec =E u ·E i T ;

[0039] Among them, E u and E i respectively represent the user and item embedding matrices output from the main task by the optimized HGCN layer.

[0040] The alignment loss is used to optimize the recommendation task:

[0041]

[0042] Among them, represents the set of user-item interaction pairs in the training set, and e u and e i respectively represent the embedding vectors of user u and item i; Θ represents all trainable parameters, including the initial user and item embeddings, denoted as Θ = {E u ∪E i}, and λ Θ is used for regularization.

[0043] For the auxiliary task, since the hyperedges represent relationships and attributes, the inner product between the embedding representing this relationship or attribute hyperedge and the embedding of the corresponding item or user node is used as the prediction score:

[0044] Among them, and E i respectively represent the embedding of the specific hyperedge of the auxiliary task and the embedding of the task-related nodes.

[0045] Subsequently, the Bayesian Personalized Ranking (BPR) loss is used to optimize the auxiliary tasks:

[0046]

[0047] Among them, represents the set of node-hyperedge interaction pairs in the training set of items with the same cate for the auxiliary task. Each node i is connected to a positive example attribute hyperedge e and a negative example attribute hyperedge e'; y cate (i,e) and y cate (i,e') respectively represent the predicted scores with the positive example attribute and negative example attribute of the node; σ is the Sigmoid function;

[0048] Finally, the pre-training is jointly optimized using the recommendation main task and the auxiliary tasks:

[0049]

[0050] Among them, λ rec is used as the coefficient to balance the losses between the main task and the auxiliary tasks; represents the additional auxiliary losses of all auxiliary tasks; during the fine-tuning process, only the main task of the pre-training is retained.

[0051] Beneficial effects of adopting this technical solution:

[0052] The present invention is based on the recommendation pre-training of hypergraph structure entropy and encodes the topological structure of the recommendation system. This method first uses two forms of pre-training tasks to capture the heterogeneous relationships between users or items, constructs multiple auxiliary task hypergraphs, compensates for the sparse interactions between users and items, and reveals potential information. Then, the method of optimizing the hypergraph structure entropy is used to convert the hyper-edge information in the hypergraph into a high-dimensional coding tree. The hypergraph structure entropy helps to decode the basic structure of the recommendation bipartite graph and achieve hierarchical clustering of users or items. Using the hypergraph pooling training method, this method combines the pooling layer into the hypergraph convolutional network to aggregate high-order information. By passing the high-level community insights to the main users or items, the social propagation process is enhanced, thereby improving the quality of node embeddings. The present invention improves the performance of the recommendation system by encoding the topological structure of the recommendation system, capturing the heterogeneous relationships between users and items, and optimizing the hypergraph structure entropy.

[0053] The present invention constructs a hypergraph structure by referring to entity relationship extraction and obtains the required data through graph encoding. Such methods can help identify the entities in the text and the relationships between them, and then construct a graph structure, where the entities are nodes and the relationships are edges. Through the graph encoding ability, using methods such as attention mechanism and distance measurement, the semantic associations between entities can be better captured, thereby constructing a richer hypergraph structure.

[0054] The present invention aims to address the complexity and diversity in the process of social data processing, simplify the processing flow, and improve efficiency. Through a series of steps, including data preprocessing, entity relationship extraction, and hypergraph construction, the processing of social data becomes more automated and efficient. The key advantage of this technology lies in its ability to comprehensively analyze and mine data from multiple perspectives, helping users quickly obtain and understand important information in social data. Through more accurate recommendations and analyses, users can better utilize social data to achieve more effective decision-making and actions. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 is a schematic flowchart of a data recommendation method based on hypergraph structure entropy pre-training according to the present invention;

[0056] Figure 2 is a framework diagram of entity relationship extraction in an embodiment of the present invention;

[0057] Figure 3 is a hypergraph pre-training framework diagram in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0058] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described below with reference to the accompanying drawings.

[0059] In this embodiment, referring to Figure 1 as shown, the present invention proposes a data recommendation method based on hypergraph structure entropy pre-training, including the steps of:

[0060] S10, data collection, collecting task data samples; through social media platforms, online forum blogs, news media and other channels, monitoring and analyzing content such as posts, comments, tags, and topics, obtaining a large amount of user data, analyzing it, and ensuring the quality and availability of the data to prepare for subsequent recommendation tasks. And preprocessing the collected data; the preprocessing process converts the original data into a format suitable for model training, while removing noise, handling missing values, normalizing the data, etc., to improve the quality of the data and the generalization ability of the model.

[0061] S20, entity relationship extraction, including the steps of: inputting the preprocessed data into a graph encoder to construct a text graph; aligning nodes and edges in the cross-graph and transferring the matching semantic information from the object to the entity; multi-channel layer, mining different dependency features between word pairs (w i , w j ) to detect the relationship between them; joint training, combining the main task and optimizing the cross-graph;

[0062] S30. Hypergraph structure entropy pre-training. After extraction is completed, an initial hypergraph is constructed, and then pre-training of the hypergraph is performed, including the pre-training process. The optimized hypergraph neural network layer is used to learn item or user node embeddings in different hypergraphs, obtain prior knowledge from auxiliary tasks and main tasks, and construct user or item embeddings for the downstream recommendation model. The hypergraph neural network layer is jointly optimized by hypergraph structure entropy and hypergraph pooling, which is achieved from the perspectives of learning from hypergraph structure and capturing global information.

[0063] S40. Recommendation and optimization. Perform pre-training optimization according to different task types and output recommendation results.

[0064] As an optimized solution of the above embodiment, the preprocessing of the collected data includes: missing value processing, outlier processing, feature standardization and normalization, text cleaning, word segmentation processing, and stop word removal.

[0065] Missing value processing: For the case where the proportion of missing values is small, it can be considered to directly delete the samples containing missing values; for the missing values of numerical features, statistical quantities such as mean, median, and mode can be used for filling; for categorical features, the mode can be used for filling.

[0066] Outlier processing: Identify outliers through methods such as box plots and Z-scores, and then choose to delete the outliers or replace them with appropriate methods to process the outliers, such as replacing the outliers with upper and lower limits.

[0067] Feature standardization and normalization: Standardize the features according to the Gaussian distribution so that their mean is 0 and variance is 1; scale the features to a range for normalization, such as [0, 1] or [-1, 1].

[0068] Text cleaning: Remove interference information such as special characters and punctuation marks in the text; unify the letters in the text to lowercase to avoid the same word in different cases being regarded as different words.

[0069] Word segmentation processing: Split the text into meaningful units, and space word segmentation or more complex word segmentation techniques can be used.

[0070] Stop word removal: Remove commonly used but meaningless words, such as "de", "shi", "zai", etc., which do not contribute much to the text feature representation.

[0071] As an optimized solution of the above embodiment, when constructing the graph text, a dependency syntax analysis tool should be used to construct the text graph; after parsing, the given sentence is converted into a text graph G T ={V T ,E T}, where V T and E Trespectively represent the nodes and edges of syntactic dependencies, G T is an undirected self-loop graph;

[0072] Meanwhile, use A T to represent the adjacency mask matrix of the text graph, where represents whether there is an edge between word pairs (w i , w j );

[0073] In addition, the node V T is input into the text encoder BERT to obtain X T as the text output representation; meanwhile, use the edge transition matrix to map the edge types into trainable vectors to obtain the trainable edge matrix Z T ;

[0074] After that, use the attribute attention mechanism to process the data results of the text encoder to update the text graph, by integrating the edge types into the keys and values in the self-attention mechanism of the transformer as the attribute transformer.

[0075] Use the attribute transformer to effectively update the node states in the cross-graph and effectively integrate the relational edges;

[0076] The attribute transformer, in the text modality, the hidden representation of the i-th token is defined as:

[0077]

[0078] where, is the adjacency mask set of the i-th node, and are the matrices corresponding to the queries, keys, and values of the i-th word in the text respectively, and their definitions are as follows:

[0079]

[0080] Other operations of the attribute transformer are the same as those of ordinary transformers: add to X T and use the feed-forward neural network FFN and layer normalization LayerNorm to obtain the text representation H T .

[0081] Use the enhanced edge graph optimal transport method to match the nodes and edges in the cross-graph; use the image-to-text attention mechanism to transfer the matched semantic information from visual objects to the text modality and obtain the improved text representation H T .

[0082] The enhanced edge graph optimization transmission method includes: to explicitly encourage the simultaneous alignment of nodes and edges across graphs, the optimal transport method originally proposed in transfer learning is adopted; different from the original optimal transport method that treats text as a fully connected graph, we only consider the nodes and edges with adjacency relationships across graphs. Two distance metrics are used for cross-graph matching: (1) The Wasserstein distance WD is used for node matching; (2) The Gromov-Wasserstein distance GWD is used for edge matching;

[0083] Use the Wasserstein distance to obtain the matching node H I to H T of the optimal transport distance D wd (H I ,H T );

[0084] Defined as:

[0085]

[0086] Where, represents to the cosine distance between, defined as T i,j represents the amount of cost transferred from node to .

[0087] Use the Gromov-Wasserstein distance to measure the similarity score D of the edges across graphs by calculating the distance between node pairs gwd ;

[0088] Specifically, its definition is as follows:

[0089]

[0090] Where, and are respectively and the sets of adjacent nodes in the text graph, L(·) is regarded as the distance cost of the cross-graph edge to , that is The learned matrix T^ now represents the transport plan that helps to align the edges across graphs.

[0091] Use a unified solver and use the Sinkhorn algorithm combined with entropy regularization to iteratively optimize D wd and D gwd , to obtain the objective loss function L for optimizing the cross-graphgraph ;

[0092]

[0093] where α is a hyperparameter that balances the importance of cost.

[0094] Then, the attention mechanism of the text is used to effectively transform the visual semantic information into a text representation

[0095] The text representation is:

[0096] where ATT cross represents the multi-head attention across modalities.

[0097] Finally, the obtained is added to H T and layer normalization is performed to obtain the final context representation O.

[0098] In the multi-channel layer:

[0099] (a) The part-of-speech Pos can provide lexical information for word pairs; for example, the part-of-speech of most entities belongs to nouns NOUN and proper nouns PEROPN, such as NBA, Curry, and Thompson; (b) Encoding the syntactic distance Sd between word pairs can improve the model's ability to capture long-distance syntactic information; (c) The word co-occurrence matrix Co can provide corpus-level information between word pairs. Based on these three considerations.

[0100] After enhanced edge graph alignment, it is divided into three feature matrices including the part-of-speech matrix, the morphological distance matrix, and the word co-occurrence matrix; each matrix is modeled by a weighted graph convolutional network to obtain the representation of each channel; each matrix first passes through an embedding layer to obtain a trainable representation R l ;

[0101] For the i-th word in the l-th matrix, the calculation of the W-GCN process is:

[0102] where is the i-th word in the l-th language matrix, and are shared weights used to perform a linear layer to learn language features and representation capabilities.

[0103] The obtained representations are combined and sent to the multi-layer perceptron MLP layer to obtain the final word representation S;

[0104]

[0105] where S i is the representation of the i-th word. Thus, the output representation of the multi-channel layer is S = [S1, S2, S3, …, S n .

[0106] Concatenate the enhanced representations of the two word representations S i , w j ) of the word pair (w i and S j to obtain the concatenation result r i,j ;

[0107] Send r i,j to the linear prediction layer to obtain the probability distribution p i,j ;

[0108] p i,j = Softmax(W p r i,j + b p );

[0109] where W p and are trainable parameters, and d y is the number of labels.

[0110] Use cross-entropy error to measure the difference L main between the reference true distribution and the predicted label distribution;

[0111]

[0112] where S and θ represent the number of training samples and all trainable parameters respectively.

[0113] The ultimate goal is the combination of the main task and the optimization across graphs:

[0114] L = L main + λL graph

[0115] where λ is a trade-off hyperparameter used to control the contribution of the optimization across graphs.

[0116] As an optimization solution for the above embodiments, during the pre-training process, it includes the steps of:

[0117] Hypergraph construction, including an auxiliary task hypergraph and a main task hypergraph;

[0118] First, a hypergraph is constructed for each auxiliary task to uniformly handle the processing of various auxiliary tasks; these tasks may involve specific attributes related to users and projects, such as age, occupation, price, category, etc., and may also include relationships such as substitution and supplementation. Taking the age attribute of users as an example, users can be represented as nodes in the hypergraph, and user groups are connected to user groups within the same age range that super-match. This method can also construct other hypergraphs for auxiliary tasks, such as "items with the same category", items with the same rate, "items purchased together", "items purchased together", "items compared with", and "users using the same job".

[0119] After the hypergraph construction is completed, the hypergraph needs to be learned and trained; all auxiliary task hypergraphs and the main task hypergraph are fed into the optimized HGCN layer for training to generate corresponding representations for item and user nodes.

[0120] For example, E i represents the item node embedding in the main task hypergraph "items purchased by the same user", and E u represents the user node embedding in the main task hypergraph "users purchasing the same item". represents the item node embedding in the auxiliary hypergraph of "items with the same category", represents the user node embedding in the auxiliary hypergraph of "users of the same age". The optimized HGCN combines the learning of hypergraph topology and captures global relationships to complete the enhancement of the initial HGCN.

[0121] The above representations are used as optimization objectives during pre-training, and separate losses are calculated for each representation. The losses of the auxiliary tasks are aggregated to form the auxiliary loss while the losses from the main task are combined to obtain the recommendation loss The final pre-training loss is obtained by applying a weighted combination of these two components. During the recommendation stage, recommendations are generated by comparing the two main task representations E i and E u to generate a recommendation sequence.

[0122] For the hypergraph structure entropy, the hyper-edge information in the hyper-subgraph matrix is converted into a node adjacency matrix to transfer hyper-edge information and design the hypergraph structure entropy; subsequently, by constructing an optimal coding tree, dividing node communities, and combining the principle of minimizing the structure entropy, the best high-dimensional community division is finally achieved; Figure 2 The structure information theory shows stronger adaptability in learning graph structures and clustering. Although it is currently limited to simple homogeneous graphs and multi-relational graphs, it overcomes the challenges of constructing coding trees and dividing communities on hypergraphs.

[0123] The structure information theory shows stronger adaptability in learning graph structures and clustering. Although it is currently limited to simple homogeneous graphs and multi-relational graphs, it overcomes the challenges of constructing coding trees and dividing communities on hypergraphs.

[0124] After establishing the main task hypergraph and the auxiliary task hypergraph, the structural information theory is used to classify closely related users or items into the same community; the structural information theory evaluates structural uncertainty based on the random walk of nodes through connected edges in the graph.

[0125] Since the hyperedges in the hypergraph can connect multiple nodes, these hyperedges represent the relationships between groups, rather than just one-to-one relationships; this is the most significant difference between the hypergraph and the homogeneous graph / multi-relational graph. Therefore, the key to generalizing the concept of structural entropy to the hypergraph lies in converting the group relationships (i.e., hyperedges) in the hypergraph into abstract one-to-one relationships (i.e., node adjacency matrix) between nodes while retaining the characteristics of the hypergraph. For the bipartite graph matrix A of the hypergraph G Perform the following operations to obtain the node adjacency matrix

[0126]

[0127] where D e represents the hyperedge degree matrix. represents the degree of the m-th hyperedge. a i,m represents whether there is a connection relationship between node v i and hyperedge e m or not.

[0128] By calculating the weighted sum of the hyperedges shared between two nodes, assign values to the adjacency relationship of the nodes. (D e ) -1 serves to amplify the influence of hyperedges with fewer connected nodes. After obtaining the adjacency matrix of the hypergraph nodes, the hypergraph structural entropy can be calculated through the following formula:

[0129]

[0130] where is the sum of the degrees of all vertices in graph . vol(α) is the volume of T α , which is the sum of the degrees of all vertices in the vertex subset; g α is the sum of the weights of all edges from the vertex subset to the vertex subset , understood as the total weight of the edges from the vertices outside the vertex subset; represents the probability of randomly entering T α ; the structural entropy of graph is the minimum

[0131] Let represent a tree with a coding height not greater than k, and then the k-structural entropy of is defined as follows:

[0132]

[0133] In addition, the one-dimensional structure entropy has its particularity because there are only root nodes and leaf nodes in a single-layer coding tree. In the graph , all vertices are grouped into the same community λ under the one-dimensional coding tree, which makes this community unique in terms of community division.

[0134] Therefore, the one-dimensional structure entropy can be directly expressed as:

[0135]

[0136] where d i is the sum of the weights of all edges connected to vertex v in the graph i , called the degree of vertex v i ; the one-dimensional structure entropy measures the uncertainty of the graph without stratification.

[0137] Specifically, the degrees of all nodes are represented by the corresponding row sums of the adjacency matrix:

[0138]

[0139] where diag(·) represents the diagonal elements of the matrix, is the number of nodes in the hypergraph , and is a column vector of length , i.e., (1, 1, …, 1) T .

[0140] Next, an optimal three-dimensional coding tree is formed by constructing an optimal two-dimensional coding tree in two rounds; once the optimal two-dimensional coding tree is obtained, these community nodes will be used as hypergraph nodes to build a two-dimensional coding tree again.

[0141] The construction process of the two-dimensional coding tree is as follows: The initial one-dimensional coding tree represents the simplest two-level structure, where the leaf nodes in the graph are directly connected to the root node; through the merging operation, nodes are grouped together, and the best coding tree is constructed using the greedy search strategy; the merging operation combines two subtrees under the same parent node to form a single subtree, i.e., the merger of two small communities; in the optimal coding tree, the graph structure exhibits the minimum uncertainty, the nodes reach a balanced and stable state, and the best node division is achieved.

[0142] After the merging operation , record as the coding tree, and the difference in the structure entropy of the graphs and determined by the two coding trees is:

[0143]

[0144] Among them, is the structural information of subtree α, is the structural information of subtree β. is the structural information of subtree δ of α and β; is the structural information of subtree α after merging subtree β into α; similarly, is the structural information of subtree δ after merging subtree β into α. If then the merging operation runs successfully, denoted as

[0145] The initial one-dimensional encoding tree means that each node v i forms an independent community. Next, calculate the difference in structural entropy before and after performing the merging operation on any two nodes, and select no more than the maximum number of pairs from the node pairs with the largest reduction in structural entropy, denoted as Q, to perform the merging operation. Continue this process until no node pairs that allow the merging operation to be successfully executed can be found. The Q value for each iteration is determined as follows:

[0146] Q = ceil((n cur - 1) × p)

[0147] where n cur represents the current number of communities / nodes, and p is a hyperparameter between 0 and 1 used to control the parallel operation speed of the merging operator. The operator ceil is a mathematical function that rounds the given number to the nearest integer. The final community partition can be represented mathematically as S1. Therefore, by treating the community set as a set of new nodes, the bipartite graph matrix A' of the community hypergraph can be obtained G :

[0148] A' G = S1 T · A G

[0149]

[0150] where C1 is the initial community set, and s ij represents whether node v i belongs to the community A G is the bipartite graph matrix of the initial hypergraph. The obtained in this way represents the connection relationship of the hyperedges between communities and there will be self-loops. Through iteration, the final community set C2 with a higher-level community partition S2 can be obtained. At this time, the construction of the three-dimensional optimal encoding tree and the corresponding community partition have been completed.

[0151] Learn the topological structure of the hypergraph through the optimal coding tree, and divide the nodes into two-level communities; add pooling and uncooling layers to enhance the propagation process of node information; the pooling operation is used to aggregate the embedding information of low-level nodes to form the representation of high-level communities; on the other hand, the uncooling operation propagates the obtained high-level community representation back to the lower-level nodes;

[0152] These operations can be expressed mathematically as: E c = S T ·E v ; E v = S·E c

[0153] The above two are the pooling operation and the uncooling operation respectively, where S is the community partition matrix. E c and E v are the embedding representations of the community set and the node set respectively.

[0154] In the hypergraph pooling encoder HP, two upsampling modules are adopted, and each module consists of an HGCN layer and a pooling layer; the hypergraph uncooling decoder HPU uses two down feedback modules, and each module includes an uncooling layer and an HGCN layer; between the encoder and the decoder, an additional HGCN layer is inserted to promote the propagation of the highest-level community information; in order to capture the global community representation and the direct neighbor dependencies simultaneously, the HGCN layer in the decoder does not directly use the output of the uncooling layer. Instead, it combines the representations of the output of the HGCN layer and the output of the uncooling layer in the corresponding encoder through weighting; this encoder-decoder skip connection operation allows the upper-level community information to be embedded into the embeddings of the lower-level nodes, avoiding the situation where slightly different feedbacks are received by different lower-level nodes.

[0155] Specifically, for the initial hypergraph First, randomly initialize the node embedding matrix with size After one round of HGCN layer information propagation, an embedding matrix of the same size is obtained Based on the community partition matrix S1 on the coding tree, perform embedding pooling for upsampling to obtain the initial community embedding matrix of size |C1|×d The updated hypergraph is the community hypergraph After another round of HGCN information propagation, an embedding matrix of size |C1|×d is obtained Then, based on the community partition matrix S2, perform initial community pooling and upsampling again to obtain the final community embedding matrix of size |C2|×d This represents the completion of the encoder. The updated hypergraph is the final community hypergraph After one round of HGCN information propagation, an embedding matrix of size |C2|×d is obtained. The decoder starts to work. First, it performs the final community annealing operation based on the community partition matrix S2, providing an initial community embedding matrix of size |C1|×d. Then, the initial community embedding matrix at the corresponding level in the encoder is used. For weighting, it is input into the HGCN layer. After one round of propagation, an embedding matrix of size |C1|×d is obtained. Next, based on the community partition S1 matrix, the initial community annealing operation is performed, providing a node embedding matrix of size . Finally, the node embedding matrix is used for weighting and passed to the HGCN layer to obtain a final node embedding matrix of size .

[0156] Essentially, the above pooling and annealing processes aim to obtain the upper-level community embedding of nodes. For the original hypergraph , this is equivalent to performing two rounds of hypergraph encoding:

[0157] 1) Initial node embedding After one round of HGCN propagation, new node embeddings with direct neighbor dependencies are generated.

[0158] 2) The new node embeddings are weighted with the corresponding upper-level community embeddings . After the second round of HGCN propagation, the final node embeddings are obtained with both direct neighbor dependencies and global collective representations.

[0159] As an optimized solution for the above embodiment, for the main task, i.e., the recommendation task, the inner product calculation is used to calculate the ranking score Y of the user-item pair. rec :

[0160] Y rec = E u · E i T ;

[0161] Where, E u and E i respectively represent the user and item embedding matrices output from the main task by the optimized HGCN layer;

[0162] The alignment loss is used to optimize the recommendation task:

[0163]

[0164] Among them, represents the set of user-item interaction pairs in the training set, and e u and e i represent the embedding vectors of user u and item i respectively; Θ represents all trainable parameters, including the initial user and item embeddings, denoted as Θ = {E u ∪E i}, and λ Θ is used for regularization;

[0165] For the auxiliary task, since the hyperedge represents the relationship and attributes, the inner product between the embedding of the hyperedge representing this relationship or attribute and the embedding of the corresponding item or user node is used as the prediction score:

[0166] Among them, and E i represent the embedding of the specific hyperedge of the auxiliary task and the embedding of the task-related node respectively;

[0167] Subsequently, the Bayesian Personalized Ranking BPR loss is used to optimize the auxiliary task:

[0168]

[0169] Among them, represents the set of node-hyperedge interaction pairs in the training set of items with the same cate in the auxiliary task. Each node i is connected to a positive example attribute hyperedge e and a negative example attribute hyperedge e'; y cate (i,e) and y cate (i,e') represent the prediction scores with the positive example attribute and negative example attribute of the node respectively; σ is the Sigmoid function;

[0170] Finally, the pre-training is jointly optimized using the recommendation main task and the auxiliary task:

[0171]

[0172] Among them, λ rec is used as the coefficient to balance the losses between the main task and the auxiliary task; represents the additional auxiliary losses of all auxiliary tasks; during the fine-tuning process, only the main task of the pre-training is retained.

[0173] Specifically, during the optimization, only the information of the optimized hypergraph encoder regarding the user-item hypergraph and the item-user hypergraph is aggregated. The loss calculation part is the same as that before training the same.

[0174] In the process of entity relationship extraction of the present invention: Entity relationship extraction is one of the key technologies in the field of natural language processing. It aims to identify the relationships between entities from text, helping computers understand the meaning in the text. With the rapid growth of information volume, entity relationship extraction has become increasingly important and has extensive applications in many fields, such as search engines, question answering systems, information extraction, etc. The goal of entity relationship extraction is to extract the relationships between entities from text, and these entities can be people, places, organizations, times, events, etc. By identifying the relationships between entities, computers can better understand the text and obtain useful information from it. In entity relationship extraction, there are some common methods and technologies, including rule-based methods, machine learning-based methods, and deep learning-based methods. In search engines, entity relationship extraction can help search engines understand the query intent of users and provide more accurate search results. In question answering systems, entity relationship extraction can help the system better answer users' questions. In the field of information extraction, entity relationship extraction can help extract useful information from massive text and provide support for decision-making. The goal of entity relationship extraction is to extract accurate and comprehensive entities and their relationships so that the subsequent graph construction process can build a graph structure with rich semantics based on this information.

[0175] The enhanced pre-training method of hypergraph structure entropy proposed by the present invention: This method uses a pre-training framework integrating auxiliary tasks to effectively capture the heterogeneous relationships between users (items) and solve the sparsity of the interaction between users and items. In addition, by integrating the structural information theory into hypergraph learning, hierarchical user (item) community division is achieved when constructing a high-dimensional coding tree, deeply exploring the topological structure in the hypergraph. At the same time, the hypergraph pooling architecture is extended, and global dependencies are allowed to be extracted, enhancing the simulation of the community diffusion process and improving node embedding.

[0176] The present invention adopts the structural information theory: The structural information theory was initially proposed to measure the structural information contained in a graph. Specifically, this theory aims to calculate the structural entropy of a homogeneous graph, reflecting its uncertainty during hierarchical division.

[0177] In the present invention when performing a recommendation task: The recommendation task aims to determine the best ranking of items for each user based on the given user-item interaction information. This information includes two separate sets of nodes (user set U and item set I) and user-item interaction edges. Therefore, the graph representing user-item interaction can be... In the personalized recommendation task of users, the goal is to predict the list of items that users have not interacted with in graph G. The higher the ranking, the greater the possibility that users will interact with these items.

[0178] The foregoing has shown and described the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments, and what is described in the above embodiments and the specification is only to illustrate the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements fall within the scope of the present invention claimed. The scope of the present invention claimed is defined by the appended claims and their equivalents.

Claims

1. A data recommendation method based on pre-training of hypergraph structure entropy, characterized in that Including the steps: S10, Data collection, collecting task data samples; and preprocessing the collected data; S20, Entity relationship extraction, including the steps of: inputting the preprocessed data into a graph encoder to construct a text graph; aligning nodes and edges in the cross-graph and transferring the matching semantic information from the object to the entity; multi-channel layer, mining different dependency features between word pairs (w i , w j ) to detect the relationship between them; joint training, combining the main task and optimizing the cross-graph; S30, Hypergraph structure entropy pre-training. After extraction is completed, construct an initial hypergraph, and then pre-train the hypergraph, including the pre-training process. Use the optimized hypergraph neural network layer to learn item or user node embeddings in different hypergraphs, obtain prior knowledge from auxiliary tasks and the main task, and construct user or item embeddings for the downstream recommendation model; and jointly optimize the hypergraph neural network layer through hypergraph structure entropy and hypergraph pooling, which is achieved from the perspectives of learning from hypergraph structure and capturing global information; S40, Recommendation and optimization, perform pre-training optimization according to different task types and output recommendation results.

2. The data recommendation method based on hypergraph structure entropy pre-training according to claim 1, characterized in that Preprocessing the collected data includes: missing value processing, outlier processing, feature standardization and normalization, text cleaning, word segmentation processing, and stop word removal.

3. The data recommendation method based on hypergraph structure entropy pre-training according to claim 1, characterized in that When constructing the graph text, a dependency parsing tool is used to construct the text graph; after parsing, the given sentence is converted into a text graph G T ={V T ,E T}, where V T and E T represent the nodes and edges of syntactic dependencies respectively, and G T is an undirected self-loop graph; Meanwhile, use A T to represent the adjacency mask matrix of the text graph, where represents whether there is an edge between word pairs (w i , w j ); In addition, node V T is input into the text encoder BERT to obtain X T as the text output representation; meanwhile, the edge types are mapped into trainable vectors using the edge transition matrix, obtaining the trainable edge matrix Z ; T ; After that, use the attribute attention mechanism to process the data results of the text encoder to update the text graph by incorporating edge types into the keys and values of the self-attention mechanism of the transformer as an attribute transformer.

4. A data recommendation method based on hypergraph structure entropy pre-training according to claim 3, characterized in that, Use the enhanced edge graph optimal transport method to match nodes and edges across graphs; use the image-to-text attention mechanism to transfer the matched semantic information from visual objects to the text modality and obtain the improved text representation H T .

5. A data recommendation method based on pre-training of hypergraph structure entropy according to claim 4, characterized in that The enhanced edge graph optimization transmission method includes: using two distance metric methods for cross-graph matching: (1) Wasserstein distance WD for node matching; (2) Gromov-Wasserstein distance GWD for edge matching; Obtain the matching node H using the Wasserstein distance I to H T The optimal transport distance D wd (H I , H T ) Measure the similarity score D of edges across graphs by calculating the distance between node pairs using the Gromov-Wasserstein distance gwd ; Use a unified solver and use the Sinkhorn algorithm combined with entropy regularization to iteratively optimize D wd and D gwd to obtain the optimized cross-image objective loss function L graph ; Then, the attention mechanism of the text is used to effectively convert the visual semantic information into a text representation Finally, add the obtained to H T and perform layer normalization to obtain the final context representation O.

6. The data recommendation method based on hypergraph structure entropy pre-training according to claim 5, characterized in that In the multi-channel layer: After enhanced edge map alignment, it is divided into three feature matrices including a part-of-speech matrix, a morphological distance matrix, and a word co-occurrence matrix; a weighted graph convolutional network is used to model each matrix to obtain the representation of each channel; each matrix first passes through an embedding layer to obtain a trainable representation R l ; Combine the obtained representations and send them to the multi-layer perceptron MLP layer to obtain the final word representation S; Concatenate the two-word representations S i and S j of the word pair (w i , w j ) to obtain the concatenation result r i,j ; Send r i,j to the linear prediction layer to obtain the probability distribution p i,j ; Use cross-entropy error to measure the difference between the ground-truth distribution and the predicted token distribution L main ; The ultimate goal is the combination of the main task and optimizing the cross-graph: L = L main + λL graph ; where λ is a trade-off hyperparameter used to control the contribution of optimizing the cross-graph.

7. A data recommendation method based on hypergraph structure entropy pre-training according to claim 1, characterized in that, During the pre-training process, it includes the steps: Hypergraph construction, including the auxiliary task hypergraph and the main task hypergraph; After the hypergraph construction is completed, the hypergraph needs to be learned and trained; all auxiliary task hypergraphs and the main task hypergraph are fed into the optimized HGCN layer for training to generate corresponding representations for item and user nodes.

8. A data recommendation method based on pre-training of hypergraph structure entropy according to claim 1, characterized in that, For the hypergraph structure entropy, convert the hyper-edge information in the hypergraph bipartite matrix into a node adjacency matrix to transmit hyper-edge information and design the hypergraph structure entropy; Subsequently, by constructing an optimal coding tree, dividing node communities, and combining the principle of minimizing structural entropy, the best high-dimensional community division is finally achieved; After establishing the main task hypergraph and the auxiliary task hypergraph, use the structural information theory to classify closely related users or items into the same community; The structural information theory evaluates structural uncertainty based on the random walk of nodes in the graph through connected edges.

9. A data recommendation method based on hypergraph structure entropy pre-training according to claim 8, characterized in that, Learn the topological structure of the hypergraph through the optimal coding tree, divide the nodes into two levels of communities; add pooling and non-cooling layers to enhance the propagation process of node information; the pooling operation is used to aggregate the embedding information of low-level nodes to form the representation of high-level communities; on the other hand, the non-cooling operation propagates the obtained high-level community representation back to the low-level nodes; In the Hypergraph Pooling Encoder (HP), two upsampling modules are adopted, and each module consists of an HGCN layer and a pooling layer; the Hypergraph Uncooling Decoder (HPU) utilizes two down feedback modules, and each module includes an uncooling layer and an HGCN layer; between the encoder and the decoder, an additional HGCN layer is inserted to facilitate the propagation of the highest-level community information; in order to capture both the global community representation and the direct neighbor dependencies simultaneously, the HGCN layer in the decoder does not directly use the output of the uncooling layer. Instead, it combines the representations of the output of the HGCN layer and the output of the uncooling layer in the corresponding encoder through weighting. This encoder-decoder skip connection operation allows the upper-level community information to be embedded into the embeddings of the lower-level nodes, avoiding the situation where slightly different feedback is received by different underlying nodes.

10. A data recommendation method based on hypergraph structure entropy pre-training according to claim 1, characterized in that, For the main task, i.e., the recommendation task, the inner product calculation is used to calculate the ranking score Y of the user-item pair rec : Y rec = E u · E i T ; Among them, E u and E i respectively represent the user and item embedding matrices output from the main task by the optimized HGCN layer; The alignment loss is used to optimize the recommendation task: Among them, represents the set of user-item interaction pairs in the training set, and e u and e i represent the embedding vectors of user u and item i respectively; Θ represents all trainable parameters, including the initial user and item embeddings, denoted as Θ = {E u ∪ E i}, and λ Θ is used for regularization; For the auxiliary task, since the hyperedges represent relationships and attributes, the inner product between the embedding of the hyperedge representing this relationship or attribute and the embedding of the corresponding item or user node is used as the prediction score: Among them, and E i respectively represent the embedding of specific hyperlinks of the auxiliary task and the embedding of task-related nodes; Subsequently, the Bayesian Personalized Ranking (BPR) loss is adopted to optimize the auxiliary task: Among them, represents the set of node hyperedge interaction pairs in the training set of items with the same cate for the auxiliary task, where each node i is connected to a positive example attribute hyperedge e and a negative example attribute hyperedge e′; y cate (i, e) and y cate (i, e′) represent the predicted scores with the positive example attribute and negative example attribute of the node respectively; σ is the Sigmoid function; Finally, the recommendation main task and the auxiliary task are jointly used to optimize the pre-training: Among them, λ rec is used as a coefficient to balance the losses between the main task and the auxiliary tasks; represents the auxiliary losses added by all auxiliary tasks; during the fine-tuning process, only the pre-trained main task is retained.