Hypergraph convolution session recommendation method based on reinforcement learning

By constructing a hypergraph convolutional conversational recommendation method based on reinforcement learning, the problem of insufficient flexibility of existing models when facing sudden changes in user interests is solved, a more accurate recommendation strategy that is more in line with users' long-term experience is achieved, and the effect of conversational recommendations is improved.

CN120744237APending Publication Date: 2025-10-03LANZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510865329.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing conversational recommendation models lack flexibility and are difficult to adapt to complex situations such as sudden changes in user interests. They also focus too much on short-term accuracy and ignore the user's long-term experience.

Method used

A hypergraph convolutional conversation recommendation method based on reinforcement learning is constructed, including a conversation hypergraph construction module, a conversation hypergraph convolutional network encoding module, a self-supervised learning enhancement module and a reinforcement learning module. The conversation representation is embedded through the hypergraph convolutional layer, and the recommendation strategy is optimized by combining self-supervision and reinforcement learning to capture the complex high-order conversion relationship of user interactions.

Benefits of technology

It improves the accuracy and user satisfaction of conversational recommendations, can adapt to changes in user interests more quickly, provides more comprehensive recommendation strategies, and enhances the long-term user experience of the recommendation system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120744237A_ABST
    Figure CN120744237A_ABST
Patent Text Reader

Abstract

The invention provides a hypergraph convolution session recommendation method based on reinforcement learning, and relates to the technical field of session recommendation, and the method comprises the steps: constructing a hypergraph convolution session recommendation model based on reinforcement learning, the model comprises a session hypergraph construction module, a session hypergraph convolutional network coding module, a self-supervised learning enhancement session recommendation module, a reinforcement learning module and a model prediction module, and session recommendation is performed based on the model. By introducing the reinforcement learning method and combining with the hypergraph convolutional neural network, hypergraph convolutional session recommendation based on reinforcement learning is realized, the problem of focusing on short-term sessions in a traditional recommendation system is effectively relieved, and a more comprehensive framework and method are provided for sequence recommendation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of conversation recommendation, and in particular to a hypergraph convolutional conversation recommendation method based on reinforcement learning. Background Art

[0002] With the rapid development of internet technology, global data is growing rapidly. E-commerce platforms' shopping records, video viewing history, social media likes and comments, and other interactive data related to consumer behavior contain a wealth of information on consumer interests, preferences, and behavioral patterns, providing greater opportunities and a larger platform for the development of recommendation systems. In the recommendation field, conversational recommendations, which closely align with users' real-time interaction scenarios, have become a cutting-edge research trend in recommendation technology.

[0003] Conversational recommendation is based on historical user conversation data, which includes user behavior sequences and contextual information such as time, location, and device. It can more accurately capture short-term changes in user preferences and better enhance the user experience. However, some existing conversational recommendation models suffer from data sparsity and bias, making them difficult to adapt to new situations in real time. They lack flexibility in dealing with complex situations such as sudden shifts in user interests, and their excessive focus on short-term accuracy can neglect the user's long-term experience. Therefore, it is necessary to design a hypergraph convolutional conversational recommendation method based on reinforcement learning. Summary of the Invention

[0004] In order to overcome the shortcomings of the prior art, the purpose of the present invention is to provide a hypergraph convolutional conversation recommendation method based on reinforcement learning.

[0005] To achieve the above object, the present invention provides the following solutions:

[0006] The present invention also provides a hypergraph convolutional session recommendation method based on reinforcement learning, comprising:

[0007] Step 1: Construct a hypergraph convolutional conversation recommendation model based on reinforcement learning, which includes a conversation hypergraph construction module, a conversation hypergraph convolutional network encoding module, a self-supervised learning enhanced conversation recommendation module, a reinforcement learning module, and a model prediction module;

[0008] Step 2: Perform conversational recommendations based on the model.

[0009] Preferably, in step 1, a hypergraph convolutional conversation recommendation model based on reinforcement learning is constructed, which specifically includes the following steps:

[0010] Step 101: Construct a conversation hypergraph and generate a conversation hypergraph construction module;

[0011] Step 102: Perform session hypergraph convolutional network coding to generate a session hypergraph convolutional network coding module;

[0012] Step 103: Perform self-supervised enhanced conversation recommendation and generate a self-supervised learning enhanced conversation recommendation module;

[0013] Step 104: Integrate the reinforcement learning algorithm into the model to generate a reinforcement learning module;

[0014] Step 105: Perform model prediction and generate a model prediction module.

[0015] Preferably, in step 101, a conversation hypergraph is constructed and a conversation hypergraph construction module is generated, specifically:

[0016] The session data is defined as a time series set of user interaction items, where I is the item set and a user session is represented as s = {i s,1 ,i s,2 ,...,i s,n}, where i s,n Represents the click item of the item, maps each interactive item to the d-dimensional latent feature space, and uses the vector represents the embedding representation of item i in the network layer l, and the set represents the complete item embedding matrix, where each session s corresponds to a session-level representation vector This vector captures the overall semantic information of the session and the user's intention;

[0017] Let G = {V, E} represent a hypergraph structure, where V is the vertex set, E is the hyperedge set, each hyperedge can connect any number of vertices, and each hyperedge is assigned a weight W ∈∈ , these weights form a diagonal matrix Hypergraph structure through incidence matrix Represents that its elements are defined as: if the vertex v i is included by the hyperedge, then H i∈ =1, otherwise H i∈ = 0, the degree of vertex and hyperedge is defined as: vertex v i degree represents the weight and degree of all hyperedges associated with the vertex, and the degree of hyperedge ∈ Indicates the number of vertices connected by the hyperedge;

[0018] For a hypergraph G, the line graph of the hypergraph is defined as L(G), each node of L(G) is a hyperedge in G, and if there is at least one common node between two nodes of L(G) on the corresponding hyperedge in G, then it is defined as: L(G) = (V L ,E L ), where V L={v e :v e ∈E}, E L ={(v ep ,v eq ):e p ,e q ∈E,|e p ∩e q ≥1|}, where e p ,e q It refers to a subset of hyperedges, and each edge of the line graph is assigned a weight W p,q , defined as W p,q =|e p ∩e q | / |e p ∩e q |;

[0019] Each user's session is presented in the form of a hyperedge, that is, each hyperedge is represented as {i s,1 ,i s,2 ,i s,3 ,...,i s,n}∈E, where i s,n ∈V represents the click item of each item. The original conversation sequence will be constructed into a linear sequence, in which the two item click items i s,n-1 ,i s,n , only when you click on item i s,n Before and item click item i s,n-1 Only when there is interaction will there be corresponding connection relationships.

[0020] Preferably, in step 102, session hypergraph convolutional network coding is performed to generate a session hypergraph convolutional network coding module, specifically:

[0021] A hypergraph convolutional network is used as a feature learning framework to capture the complex high-order transformation relationships between items. In terms of the computational process, an information propagation mechanism is established based on the hypergraph Laplacian operator. The standard graph convolution is extended to the hypergraph structure so that it can handle one-to-many complex connection patterns. The information propagation items are learned and deep encoding of item semantic information is achieved through parameterized matrix mapping. Based on the spectral hypergraph convolution, the hypergraph convolution is defined as:

[0022]

[0023] Among them, the model does not use nonlinear activation functions and convolution filter parameter matrices. For W, each hyperedge is assigned the same weight 1, and the row normalization of the normalized hypergraph convolution is:

[0024]

[0025] The hypergraph convolution operation follows the node-hyperedge-node message passing mode. The node passes its feature representation to the hyperedge to which it belongs through the association matrix. Each hyperedge aggregates the multi-node information received to generate a hyperedge-level representation. The processed information is then passed back to the original node through the transpose of the association matrix, completing a complete feature propagation cycle. This is the information aggregation calculation, and then multiply it with H to aggregate the information in the second part. Then, the item embedding obtained at each layer is calculated as follows:

[0026]

[0027] Define a learnable position matrix, denoted as P r =[p1,p2,p3,...,p m ], integrating the learned item click items and reverse position embeddings, where m is the length of the current session, and the tth session s={i s,1 ,i s,2 ,...,i s,m} is embedded as:

[0028]

[0029] Where, is a learnable parameter;

[0030] The session embedding is obtained by aggregating the items in the session. According to the method used in the SR-GNN model, the session s = {i s,1 ,i s,2 ,...,i s,m},for:

[0031]

[0032] Where θ h represents the user’s overall interest embedding in this session, are all learning project weights α t The attention parameter, represents the embedding of session s, which is represented by the average value of the embeddings of the items contained in session s. That is, the embedding of the t-th item in session s is expressed as:

[0033]

[0034] For a given session s, it is necessary to calculate the current scores of all candidate click items i∈I, which is:

[0035]

[0036] With the help of the softmax function, we can calculate the probability that an item in the session will become the next clicked item:

[0037]

[0038] The cross entropy loss function is used as the learning optimization objective, which is:

[0039]

[0040] Where N represents the number of samples, y i is the true predicted probability of the i-th sample, It is the model prediction probability, which ranges from 0 to 1. The cross entropy loss value of the entire model is obtained by summing the calculation results of each sample and taking the negative.

[0041] Preferably, in step 103, self-supervised enhanced conversation recommendation is performed to generate a self-supervised learning enhanced conversation recommendation module, specifically:

[0042] Initialize embedding learning for the nodes in the line graph by calculating the mean embedding of the items contained in each session. The correlation matrix of the line graph is defined as Where M represents the number of nodes in the line graph, that is, the number of hyperedges in the original hypergraph, which is:

[0043]

[0044] Define line graph convolution as:

[0045]

[0046] During the calculation process of each convolutional layer, the conversation node aggregates information from its neighborhood through the graph structured message passing mechanism. After completing the l-layer graph convolution propagation, the model calculates the weighted average of the representations at different levels to obtain the final conversation embedding representation, which is:

[0047]

[0048] A dual-channel architecture is designed to obtain two complementary conversation representations. The first channel uses a hypergraph to learn the connections between items within a conversation, revealing the interaction information within the conversation. The line graph encoded in the second channel focuses on depicting structural information at the inter-session level, namely the potential connections between different conversations, revealing the relationship between sessions.

[0049] We consider the two channels as complementary perspectives for conversational learning, each capturing different conversational characteristics. We further introduce a contrastive learning framework to systematically compare and align the two sets of learned conversational embedding representations. A noise contrastive objective is used as the learning objective for the model. The loss generated by this objective is a standard binary cross-entropy loss. This loss is derived from two sets of samples: one is the real-world sample, i.e., the positive sample, and the other is the negative sample, which is defined as:

[0050]

[0051] Where, or Yes or Negative samples after sorting in column and row form, where f D (·): It is a discriminant function that takes two vectors as input and can capture the consistency between them.

[0052] Preferably, in step 104, the reinforcement learning algorithm is integrated into the model to generate a reinforcement learning module, specifically:

[0053] A near-end policy optimization algorithm is integrated into the model. During the reinforcement learning process, various interaction data are collected in the environment through the policy network and the target network, including state, behavior, reward, and possible next state. The state is represented by the score output by the session hypergraph convolutional network encoding. The behavior is to select a specific item from the recommended candidate items. The agent selects an action based on the current state to maximize the cumulative reward. The possible next state is represented by the output score of the model, which reflects the new situation of the recommendation system after making a recommendation action. The reward is the negative of the model loss value.

[0054] Compute advantage estimates and adjust the parameters of the policy network;

[0055] Optimize the objective function and iteratively update the parameters θ of the policy network to continuously optimize the agent's strategy;

[0056] Calculate target loss;

[0057] Update the model strategy parameters and repeat the above steps until a certain number of iterations is reached or the model performance no longer improves.

[0058] Preferably, in step 105, model prediction is performed to generate a model prediction module, specifically:

[0059] During prediction training, based on the session embedding encoded by the hypergraph convolutional network and the reward signal fed back by reinforcement learning, the model ensures better accuracy in recommending the next item. The model uses the recommendation task generated by the hypergraph convolutional network channel as the main learning objective and the self-supervised learning task generated by the line graph convolutional channel as the auxiliary learning objective. The model also integrates the objective function of reinforcement learning for joint learning, and defines the loss function as:

[0060]

[0061] Where β and γ are parameters used to balance the loss.

[0062] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:

[0063] The present invention provides a hypergraph convolutional conversation recommendation method based on reinforcement learning, which includes constructing a hypergraph convolutional conversation recommendation model based on reinforcement learning. The model includes a conversation hypergraph construction module, a conversation hypergraph convolutional network coding module, a self-supervised learning enhanced conversation recommendation module, a reinforcement learning module and a model prediction module, and performs conversation recommendation based on the model. The present invention constructs a hypergraph and a line graph to model the complex high-order conversion relationship of user interaction, intuitively represents the conversation relationship of user interaction, and then embeds the items through the hypergraph convolution layer to generate the corresponding conversation representation, which can ensure the integrity and accuracy of the conversation information. Secondly, the proximal policy optimization (PPO) algorithm of reinforcement learning is integrated into the model, in which the optimization objective function and advantage estimation are used to maximize the cumulative reward, and the strategy is updated in the trust domain to avoid the problem of poor performance caused by the excessively large traditional gradient step size. The PPO algorithm uses the relationship between the new and old strategies to constrain the objective function to ensure the stability and efficiency of the strategy update, and introduces reinforcement learning in the conversation recommendation scenario. The intelligent agent selects recommended products based on the current state of the conversation, provides feedback (such as clicks, purchases, etc.) through the environment, and continuously learns and optimizes the recommendation strategy. During the learning process, the model converges faster and learns better strategies, thereby improving the recommendation accuracy and user satisfaction. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0065] Figure 1 A schematic diagram of the process flow of a hypergraph convolutional conversation recommendation method based on reinforcement learning provided by an embodiment of the present invention;

[0066] Figure 2 Schematic diagram of the hypergraph convolutional conversation recommendation model structure based on reinforcement learning;

[0067] Figure 3 Schematic diagram of the impact of different RL algorithms on the model on different indicators;

[0068] Figure 4 A schematic diagram of the impact of different components on the model on different indicators;

[0069] Figure 5 Schematic diagram of the impact of trimming parameters on the model on different indicators;

[0070] Figure 6 This is a schematic diagram of the indicator results of the Diginetica dataset on different layers of neural networks;

[0071] Figure 7 This is a schematic diagram of the indicator results of the Tmall dataset on different layers of neural networks. DETAILED DESCRIPTION

[0072] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0073] The purpose of the present invention is to provide a hypergraph convolution conversation recommendation method based on reinforcement learning, which constructs a hypergraph and line graph to model the complex high-order conversion relationship of user interaction, intuitively represents the conversation relationship of user interaction, and then embeds the items through the hypergraph convolution layer to generate the corresponding conversation representation, which can ensure the integrity and accuracy of the conversation information. Secondly, the proximal policy optimization (PPO) algorithm of reinforcement learning is integrated into the model, in which the optimization objective function and advantage estimation are used to maximize the cumulative reward, and the strategy is updated in the trust domain to avoid the problem of poor performance caused by the excessively large traditional gradient step size. The PPO algorithm uses the relationship between the new and old strategies to constrain the objective function to ensure the stability and efficiency of the strategy update. Reinforcement learning is introduced in the conversation recommendation scenario. The intelligent agent selects recommended products according to the current state of the conversation, provides feedback (such as clicks, purchases, etc.) through the environment, and continuously learns and optimizes the recommendation strategy. During the learning process, the model converges faster and learns better strategies, thereby improving the recommendation accuracy and user satisfaction.

[0074] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0075] Figure 1 A flow chart of a hypergraph convolutional session recommendation method based on reinforcement learning provided by an embodiment of the present invention is shown in FIG. Figure 1 As shown, the present invention provides a hypergraph convolutional session recommendation method based on reinforcement learning, comprising:

[0076] Step 1: Construct a hypergraph convolutional conversation recommendation model based on reinforcement learning, which includes a conversation hypergraph construction module, a conversation hypergraph convolutional network encoding module, a self-supervised learning enhanced conversation recommendation module, a reinforcement learning module, and a model prediction module;

[0077] Step 2: Perform conversational recommendations based on the model.

[0078] This paper introduces reinforcement learning method and combines it with hypergraph convolutional neural network to propose a hypergraph convolutional conversation recommendation model based on PPO (HCPPO), which includes a conversation hypergraph construction module, a conversation hypergraph convolutional network encoding module, a self-supervised learning enhanced conversation recommendation module, a reinforcement learning module and a model prediction module. Its structural diagram is shown in the figure below. Figure 2 As shown in Figure 1, the model first constructs two different types of conversation graphs based on the given conversation data: a hypergraph and a line graph. In the model, each user conversation is treated as a hyperedge, and all item nodes within the same hyperedge are considered fully connected. Different hyperedges are interconnected through shared item nodes, thus constructing a structured representation that captures high-order item associations. Based on this hypergraph representation, the model further constructs its line graph, mapping each hyperedge in the original hypergraph to a node in the line graph and establishing connections based on the intersection between hyperedges. This transformation enables the model to explicitly model conversation-level relationships, namely, inter-conversation information. Subsequently, the model learns item-level and conversation-level embedding representations separately by designing dedicated convolutional neural network layers. Building on this, the model introduces a self-supervised learning module that further enhances representational capabilities by maximizing the information objective function of inter-conversation interactions, thereby improving overall recommendation performance. A key innovation of this invention is that within the reinforcement learning module, using the PPO algorithm, the entire model collects interaction data, calculates advantage estimates, optimizes the objective function, calculates the loss, and updates model parameters in both the policy and target networks until a certain number of iterations is reached or policy performance stops improving. Ultimately, the model achieves joint training of recommendation tasks and strategy optimization tasks through a multi-objective optimization framework, effectively alleviating the problem of traditional recommendation systems focusing on short-term sessions, and providing a more comprehensive framework and method for sequential recommendations.

[0079] Next, we will introduce the model in detail. The construction of the hypergraph convolutional conversation recommendation model based on reinforcement learning includes the following steps:

[0080] Step 101: Construct a conversation hypergraph and generate a conversation hypergraph construction module;

[0081] Step 102: Perform session hypergraph convolutional network coding to generate a session hypergraph convolutional network coding module;

[0082] Step 103: Perform self-supervised enhanced conversation recommendation and generate a self-supervised learning enhanced conversation recommendation module;

[0083] Step 104: Integrate the reinforcement learning algorithm into the model to generate a reinforcement learning module;

[0084] Step 105: Perform model prediction and generate a model prediction module.

[0085] In step 101, a conversation hypergraph is constructed and a conversation hypergraph construction module is generated, specifically:

[0086] Conversational recommendation aims to predict the next item based on the user's real-time behavior. However, due to the lack of long-term user profiles, it is crucial to explore the user's intention to convert items. Although simple graph relationships and graph neural networks have achieved certain results, they still ignore complex high-order information relationships. Therefore, this paper proposes to construct a hypergraph to capture the relationship between items beyond pairwise relationships. Because its hyperedge characteristics are suitable for representing non-strict sequential dependencies and correlations between items, it also introduces line graphs on this basis to enhance the modeling capabilities of the hypergraph and improve the recommendation effect.

[0087] In this model, session data is defined as a time series set of user interaction items, where I is the item set and a user session is represented as s = {i s,1 ,i s,2 ,...,i s,n}, where i s,n Represents the click item of the item, maps each interactive item to the d-dimensional latent feature space, and uses the vector represents the embedding representation of item i in the network layer l, and the set represents the complete item embedding matrix, where each session s corresponds to a session-level representation vector This vector captures the overall semantic information of the conversation and the user's intent. From the perspective of computational objectives, the core task of a conversational recommendation system is to accurately predict the user's next potential interaction item in a given conversation based on the historical interaction sequence of the conversation, thereby more accurately capturing the user's immediate interests.

[0088] In the representation of hypergraph, let G = {V, E} represent a hypergraph structure, where V is the vertex set, E is the hyperedge set, each hyperedge can connect any number of vertices, and each hyperedge is assigned a weight W ∈∈ , these weights form a diagonal matrix Hypergraph structure through incidence matrix Represents that its elements are defined as: if the vertex vi is included by the hyperedge, then H i∈ =1, otherwise H i∈ = 0, the degree of vertex and hyperedge is defined as: vertex v i degree represents the weight and degree of all hyperedges associated with the vertex, and the degree of hyperedge ∈ Indicates the number of vertices connected by the hyperedge;

[0089] For a hypergraph G, the line graph of the hypergraph is defined as L(G), each node of L(G) is a hyperedge in G, and if there is at least one common node between two nodes of L(G) on the corresponding hyperedge in G, then it is defined as: L(G) = (V L ,E L ), where V L ={v e :v e ∈E}, E L ={(v ep ,v eq ):e p ,e q ∈E,|e p ∩e q ≥1|}, where e p ,e q It refers to a subset of hyperedges, and each edge of the line graph is assigned a weight W p,q , defined as W p,q =|e p ∩e q | / |e p ∪e q |;

[0090] In order to obtain the complex high-order pairwise transformation relationship in the conversation in the recommendation system, the present invention constructs a hypergraph G = {V, E}, and presents each user's conversation in the form of a hyperedge, that is, each hyperedge is represented as {i s,1 ,i s,2 ,i s,3 ,...,i s,n}∈E, where i s,n ∈V represents the click item of each item. The original conversation sequence will be constructed into a linear sequence, in which the two item click items i s,n-1 ,i s,n , only when you click on item i s,n Before and item click item i s,n-1 Only when there is interaction will there be corresponding connection relationship;

[0091] In the process of mapping session data to a hypergraph structure, the system converts the set of co-occurring items in each session into a fully connected substructure. This conversion process requires special attention to the correspondence between the temporal characteristics of the session sequence and the undirected graph representation. Specifically, although the original session data has an inherent temporal sequence, when constructing a hypergraph-based representation, this sequence data must be converted into an undirected associative structure to accurately reflect the co-occurrence pattern between items rather than a strict sequence. This conversion mechanism ensures that the hypergraph structure can capture the inherent correlation of the item set in the user session, providing reasonable input for subsequent graph neural network processing, and effectively integrating the two complementary data characteristics of sequence information and associative information.

[0092] In addition, the model constructs a line graph structure. This two-layer graph structure design enables the model to simultaneously model micro-item associations and macro-session associations, thereby capturing user behavior patterns at different granularity levels. In this line graph, each session instance is abstracted as a node, and connections are established between different session nodes through shared elements, which is different from the hypergraph that depicts item-level association patterns, such as Figure 2 The line graph shown focuses on capturing session-level relationships. In the formal representation of the line graph, weights are assigned to the edges connecting session nodes. The weights characterize the similarity between two session instances, providing higher accuracy for the session recommendation system.

[0093] In step 102, session hypergraph convolutional network coding is performed to generate a session hypergraph convolutional network coding module, specifically:

[0094] A hypergraph convolutional network is used as a feature learning framework to capture the complex high-order transformation relationships between items. In terms of the computational process, an information propagation mechanism is established based on the hypergraph Laplacian operator. The standard graph convolution is extended to the hypergraph structure so that it can handle one-to-many complex connection patterns. The information propagation items are learned and deep encoding of item semantic information is achieved through parameterized matrix mapping. Based on the spectral hypergraph convolution, the hypergraph convolution is defined as:

[0095]

[0096] Among them, the model does not use nonlinear activation functions and convolution filter parameter matrices. For W, each hyperedge is assigned the same weight 1, and the row normalization of the normalized hypergraph convolution is:

[0097]

[0098] The hypergraph convolution operation can be understood from the perspective of information transmission as a structured multi-stage feature transformation process. Specifically, the calculation follows the "node-hyperedge-node" message passing mode. First, the node transmits its feature representation to the hyperedge to which it belongs through the association matrix. Second, each hyperedge aggregates the received multi-node information to generate a hyperedge-level representation. Then, the processed information is transmitted back to the original node again through the transpose of the association matrix, completing a complete feature propagation cycle. In formula (2), This is the information aggregation calculation, and then multiply it with H to aggregate the information in the second part. Then, the item embedding obtained at each layer is calculated as follows:

[0099]

[0100] In order to make the recommendation effect more accurate, location information cannot be ignored. Location information is a concept introduced in Transformer, which can record the location of items. In this paper, a learnable location matrix is ​​defined, denoted as P r =[p1,p2,p3,...,p m ], integrating the learned item click items and reverse position embeddings, where m is the length of the current session, and the tth session s={i s,1 ,i s,2 ,...,i s,m} is embedded as:

[0101]

[0102] Where, is a learnable parameter;

[0103] The session embedding is obtained by aggregating the items in the session. According to the method used in the SR-GNN model, the session s = {i s,1 ,i s,2 ,...,i s,m},for:

[0104]

[0105] Where θ h represents the user’s overall interest embedding in this session, are all learning project weights α t The attention parameter, represents the embedding of session s, which is represented by the average value of the embeddings of the items contained in session s. That is, the embedding of the t-th item in session s is expressed as:

[0106]

[0107] In common models, GRU units have a significant advantage in processing time dependencies when processing sequence data. However, the computational process is too complex, resulting in more time required for model training and a significant resource consumption. While the self-attention mechanism can effectively capture the connections between elements in a sequence, creating an efficient and lightweight model does not require these complexities. Therefore, this invention uses position embedding to reflect the time factor. Position embedding can simply and effectively reflect the position information of elements in a sequence, thus reflecting the order of the conversation process to a certain extent. In this way, the model avoids the burden of complex sequence modeling techniques. It can not only rationally utilize the temporal features in the conversation data, but also greatly improve efficiency and reduce the complexity of the model, truly achieving both efficiency and lightweightness, providing a strong guarantee for the deployment and operation of the model in practical application scenarios.

[0108] For a given session s, it is necessary to calculate the current scores of all candidate click items i∈I, which is:

[0109]

[0110] With the help of the softmax function, we can calculate the probability that an item in the session will become the next clicked item:

[0111]

[0112] This model uses the cross-entropy loss function as the learning optimization objective. This loss function can effectively quantify the difference between the predicted distribution and the true distribution, providing a rigorous and efficient optimization method for the recommendation task. Specifically, it is:

[0113]

[0114] Where N represents the number of samples, y i is the true predicted probability of the i-th sample, It is the model prediction probability, which ranges from 0 to 1. The cross entropy loss value of the entire model is obtained by summing the calculation results of each sample and taking the negative.

[0115] In step 103, self-supervised enhanced conversation recommendation is performed to generate a self-supervised learning enhanced conversation recommendation module, specifically:

[0116] Hypergraph construction and hypergraph convolutional encoding can significantly improve model performance. However, the small amount of session data cannot be ignored. This makes hypergraph modeling difficult and reduces recommendation performance. To overcome this difficulty, a self-supervised learning module is added to the model network to improve the effectiveness of hypergraph modeling.

[0117] Introduce them respectively:

[0118] (1) Linear channel and convolution

[0119] Figure 2 The process of constructing the line graph L(G) by the model is shown. The line graph can accurately depict the connection relationship between hyperedges and provide a basis for capturing session-level relationships. Since the nodes in the line graph structure correspond to the hyperedges in the original hypergraph, the model first needs to initialize the embedding learning for the nodes in the line graph. This initialization process is achieved by calculating the mean embedding of the items contained in each session. The association matrix of the line graph is defined as Where M represents the number of nodes in the line graph, that is, the number of hyperedges in the original hypergraph, and is:

[0120]

[0121] Define line graph convolution as:

[0122]

[0123] During the computation of each convolutional layer, the session node aggregates information from its neighborhood through a graph-structured message passing mechanism. This cross-session approach effectively captures the transition patterns and semantic associations between sessions. After completing the l-layer graph convolution propagation, the model calculates the weighted average of representations at different levels to obtain the final session embedding representation, which is:

[0124]

[0125] (2) Generate self-supervisory signals

[0126] This model uses a dual-channel architecture to obtain two complementary conversational representations. The first channel learns the connections between items within a conversation based on a hypergraph, revealing the interactive information within the conversation. The line graph encoded in the second channel focuses on depicting structural information at the inter-session level, that is, the potential connections between different conversations, revealing the relationship between sessions. Because the two channels learn different levels of information, the mutual understanding between the two sets of embeddings is relatively limited. However, they are not completely unrelated, but rather complementary. From the perspective of information completeness, the information missing from one set of embeddings may be provided by the other set. The combination of the two sets of embeddings can more comprehensively and completely represent the structural characteristics of the conversation-induced hypergraph.

[0127] During model training, a small batch of data is used each time, which contains n sessions. In this small batch of data, there is a special relationship between the two sets of session embeddings, called a bijective mapping relationship. That is, for each session, unique corresponding data can be found in both sets of embeddings. From the perspective of self-supervised learning, these two sets of embeddings can serve as each other's standard reference, that is, the true label. This one-to-one mapping is equivalent to increasing the number of labels and can provide more information. The specific judgment rules are as follows: If two session embeddings learned from two different channel views represent the same session, then this pair of session embeddings is marked as a positive sample. Conversely, if they do not represent the same session, then they are marked as a negative sample. In this way, it is clear for self-supervised learning how to label samples, and the model can better learn the relationship between sessions and the overall characteristics of the session.

[0128] (3) Contrastive Learning

[0129] We consider the two channels as complementary perspectives for conversational learning, each capturing different conversational characteristics. We further introduce a contrastive learning framework to systematically compare and align the two sets of learned conversational embedding representations. A noise contrastive objective is used as the learning objective for the model. The loss generated by this objective is a standard binary cross-entropy loss. This loss is derived from two sets of samples: one is the real-world sample, i.e., the positive sample, and the other is the negative sample, which is defined as:

[0130]

[0131] Where, or Yes or Negative samples after sorting in column and row form, where f D (·): It is a discriminant function that takes two vectors as input and can capture the consistency between them.

[0132] In step 104, the reinforcement learning algorithm is integrated into the model to generate a reinforcement learning module, specifically:

[0133] When the model encounters a complex and changing environment, its decision-making ability needs to be further improved. Reinforcement learning allows the intelligent agent to refer to its current state in the process of interacting with the environment, and then combine past experience to flexibly choose the optimal action strategy. The goal is to maximize long-term benefits. In addition, the environment will feedback reward signals to the intelligent agent, and the intelligent agent will learn based on this signal. The goal of the intelligent agent is to obtain as many rewards as possible by adjusting its own behavioral strategies. This feedback-based learning mechanism allows the intelligent agent to adapt to various complex environmental changes. It is similar to the process of learning and making decisions in real life. It is through interacting with the surrounding environment and constantly adjusting itself based on the feedback received. In reinforcement learning, the Proximal Policy Optimization (PPO) algorithm performs well in optimizing dynamic decision-making and maximizing long-term cumulative rewards. The PPO algorithm can learn effective strategies based on known data samples. It can limit the amplitude of policy updates to avoid excessive changes in the policy during the update process and ensure the stability of the learning process. Compared with other policy gradient-based algorithms, PPO does not require complex and time-consuming precise searches to determine the optimal step size. Instead, it uses a simpler and more effective approximate method to control policy updates. Therefore, incorporating the PPO algorithm into the model and leveraging its advantages provides new ideas and methods for solving complex data processing and decision-making problems.

[0134] A detailed introduction to it:

[0135] (1) Collect data

[0136] In the process of reinforcement learning, various interaction data are collected in the environment through the policy network and the target network, including state, action, reward, and the possible next state (next_state);

[0137] State represents the current state of the agent's environment. The state in this model refers to the current state information of the recommendation system, which is represented by the score output by the session hypergraph convolutional network encoding. This score can reflect the possibility of each item being recommended. The agent decides which items to recommend based on this score. Action is the specific action taken by the agent in a specific state. In this model, action can be understood as selecting specific items from the recommended candidate items. The agent selects an action based on the current state in order to maximize the cumulative reward. The possible next state next_state is the next state to which the environment transfers after the agent takes an action. In this model, the next state is also represented by the output score of the model, reflecting the new situation of the recommendation system after a certain recommendation action is taken. Reward is the feedback of the environment to the agent's action. In this invention, the reward is the negative of the model loss value, which means that the smaller the model loss, the greater the reward. Because the loss value reflects the gap between the model prediction result and the true target, the negative loss value is used to measure the recommendation effect.

[0138] (2) Calculate the advantage estimate

[0139] Advantage estimation plays a core role in guiding the agent to learn the optimal strategy. The essence of advantage estimation is to measure the pros and cons of taking a certain action compared to the average action in a specific state, which is specifically reflected in the expected state action value Q(s t ,a t ) and the current state action value V(s t ), the expected state action value Q(s t ,a t ) reflects the agent’s t Next take action a t After that, the total cumulative reward that can be obtained in the future is expected. This value takes into account the rewards brought by the current action and a series of subsequent possible actions, and the current state action value V(s t ) represents the agent in state s t Under this circumstance, the average value that can be obtained by executing actions based on the current strategy is obtained by calculating the difference between the two, which is:

[0140]

[0141] Under the current training conditions, it is necessary to consider the advantage of each action compared to the average situation. If the advantage value of an action is positive, it means that this action is more valuable than the average action in the current state, and it is likely to bring more cumulative rewards to the agent. If the advantage value is negative, it means that the performance of this action is not as good as the average level. In this case, the agent should try to avoid choosing this action. During the training process of PPO, the advantage estimate is used to adjust the parameters of the policy network. It can help the policy network determine which actions need to have a higher probability of selection and which actions need to have a lower probability of selection. In this way, the agent is guided step by step to learn to select the optimal strategy that can obtain the maximum cumulative reward in various states.

[0142] (3) Optimize the objective function

[0143] According to the target clipping function L in PPO CLIP (θ), which is used to guide the update of the policy network's parameters θ, and the formula is as follows:

[0144]

[0145] Among them, r t (θ) is the probability ratio, which is the new policy π θ With the old strategy π θold In state s t Next take action a t This ratio is used to measure the importance of samples generated using the new strategy relative to samples generated using the old strategy, and plays a key role in strategy updating. t (θ) is expressed as follows:

[0146]

[0147] In formula (15), clip(r t (θ),1-ε,1+ε) is the truncation operation, r t (θ) is limited to the range of [1-ε, 1+ε], where ε is a pre-set hyperparameter. Through this operation and the setting of hyperparameters, the PPO algorithm can prevent the policy update from being too large, which leads to unstable training, and ensure that the new policy does not deviate too far from the old policy;

[0148] In calculating the objective function L CLIP When (θ), we first take the minimum value of the product of the ratio and the advantage function. The purpose of this is to take advantage of the improvements brought by the new strategy while ensuring the stability of the strategy update when the strategy is updated. Ultimately, by minimizing this objective function, we iteratively update the parameters θ of the policy network, so that the strategy of the agent is continuously optimized.

[0149] (4) Target loss

[0150] The above parts are combined to guide the training of the policy network. The target loss function is an expression based on expectation, E t It represents taking the expectation at time step α, that is, averaging the samples of multiple time steps. This total target loss function optimizes the policy network parameters θ by balancing different components, so that the agent can learn a better strategy. The calculation formula is as follows:

[0151]

[0152] in, is an objective function based on the importance sampling ratio and advantage function, It is used to measure the gap between the predicted value of the value network and the true value. c1 is a hyperparameter used to control the weight of the value function loss term in the total target loss. By minimizing this term, the value network can estimate the state value more accurately, provide more reliable reference information for the learning of the policy network, and reduce the variance of the policy gradient estimation. S[π θ ](s t ) is the entropy of the policy, which measures the policy in state s t The uncertainty of the action distribution under the policy, c2, is used to adjust the proportion of the policy entropy term in the total target loss. A higher entropy means that the policy is more random and diverse when choosing actions, which helps the agent to better explore the environment in the early stages of training and avoid falling into a local optimal solution too early. By incorporating the policy entropy term into the total target loss, the agent is encouraged to maintain its exploration ability while utilizing the current experience. The calculation formula is as follows:

[0153] S[π θ ](s t )=-∫π θ (a t ∣s t )log(π θ (a t ∣s t ))da t ; (18)

[0154] (5) Update the model strategy parameters and repeat the above steps until a certain number of iterations is reached or the model performance no longer improves.

[0155] In step 105, model prediction is performed to generate a model prediction module, specifically:

[0156] During prediction training, the model uses the session embedding encoded by the hypergraph convolutional network and the reward signal from reinforcement learning feedback to ensure better accuracy in recommending the next item. The model uses the recommendation task generated by the hypergraph convolutional network channel as the primary learning objective and the self-supervised learning task generated by the line graph convolution channel as the auxiliary learning objective. This combination of primary and secondary tasks enables the model to maintain core recommendation performance while introducing additional constraints through self-supervisory signals, effectively alleviating data sparsity and enhancing the model's generalization ability. The model also integrates the objective function of reinforcement learning for joint learning, and defines the loss function as:

[0157]

[0158] Where β and γ are parameters used to balance the loss.

[0159] This paper provides experimental and analytical examples based on the method of the present invention. Relevant experiments are conducted on three public datasets commonly used in the recommendation field. The performance of the model is evaluated based on two indicators: recall rate and average reciprocal rank. The model is also compared with other models in the field.

[0160] In order to better evaluate and analyze the performance of the proposed HCPPO model in conversational recommendation, we selected three real-world benchmark datasets for evaluation: Diginetica, Nowplaying, and Tmall.

[0161] The Diginetica dataset, released at the CIKM2016 conference, covers a wealth of e-commerce-related information. When using this dataset, we selected only a subset of purchase data and subjected it to rigorous filtering and partitioning based on the rules proposed by Li et al. The filtering criteria removed item clicks with a session length of one and items that appeared fewer than five times. On the one hand, these sessions of length one contain too little information to extract valuable insights. On the other hand, item clicks that appeared fewer than five times in the dataset were filtered out because these low-frequency clicks may be outliers or noise that has little impact on the overall model. After processing, the Diginetica dataset contained 719,470 training sessions and 60,858 testing sessions. This sufficient number of training sessions helps the model learn more complex and accurate user purchasing behavior patterns. The dataset contains 43,097 clicks with an average session length of 5.12, meaning that each user clicked on an item approximately 5.12 times in a single session. This relatively short average session length may reflect some characteristics of user purchasing behavior in this dataset.

[0162] The Nowplaying dataset was collected from the social media platform Twitter. Similar to the processing of the Diginetica dataset, the Nowplaying dataset was also processed accordingly. First, short conversation sequences containing only a single interactive item were removed. Such conversations lack sufficient sequence information to support pattern mining. Second, items with a frequency of less than a threshold of 5 in the overall data corpus were removed to alleviate data sparsity. After this processing process, the dataset contains 825,304 training sessions and 89,824 test sessions. The large number of training sessions allows the model to fully learn users' music listening preferences and habits. There are a total of 60,417 click items in the dataset, with an average session length of 7.42. This means that users perform an average of 7.42 click operations in a music listening session. These operations may include play, pause, and switch songs.

[0163] The Tmall dataset, a public benchmark dataset for the IJCAI-15 international competition, has broad application value in e-commerce research. This dataset is collected from real user interaction records on Tmall, a well-known Chinese B2C e-commerce platform, and includes anonymized user shopping behavior trajectory data. To ensure that the data is more representative of the actual situation and more effectively represents the required data information, a series of similar filtering operations were performed on this dataset. After filtering, the dataset was divided into a training set containing 351,268 sessions and a test set containing 25,898 sessions. The dataset contains 40,727 item clicks, with an average session length of 6.69.

[0164] To further improve the model's training effectiveness and generalization capabilities, the model uses sequence splitting to expand the dataset and complete labeling. The last item click in each sequence is set as the label for that sequence. This approach can provide the model with more training samples, helping it better learn the patterns and regularities in the sequence, thereby improving its performance in practical applications. The statistical data of the dataset are shown in Table 1.

[0165] Table 1 Dataset information table

[0166]

[0167] In order to better complete the experimental content, the specific hardware and software configuration of the experiment is shown in Table 2. This model runs in this performance environment;

[0168] Table 2 Experimental software and hardware environment configuration table

[0169]

[0170] The parameter settings are as follows: First, the Adam optimization algorithm is used for the optimization of the HCPPO model. Secondly, in model training, epoch (round), batch_size (batch size), embedding vector dimension and learning rate are all very critical parameters. After practice, the optimal epoch value of the model of the present invention is 30. If the epoch scale is too large, it will bring many problems, such as a significant increase in training time, excessive memory usage, and easy to cause overfitting. In order to balance training time and memory overhead, during training, the data of an epoch is generally split into multiple batches. The batch_size here refers to the number of training samples contained in each batch. A larger batch tch_size can speed up the training process. Although a smaller batch_size can reduce the risk of overfitting, the training time will be greatly extended. Experimental results show that the model reaches the optimal performance balance point when the batch size is set to 128. The dimension parameter emb_size of the item embedding vector is randomly initialized using a standard Gaussian distribution, and the initialized value is normalized to 100. The learning rate determines the step size in the direction of the loss function gradient each time the parameters are updated. When the initial learning rate is set to 0.001, the performance of this model is most stable. L2 regularization is one of the most effective regularization methods. The regularization value is set to 1e-5 in the model. The optimal parameter values ​​of the HCPPO model after multiple experiments are shown in Table 3.

[0171] Table 3 Experimental parameters

[0172]

[0173] For the HCPPO proposed in this invention, in order to evaluate the recommendation accuracy of the model, the Recall index and Mrr are mainly used, as shown in Table 4;

[0174] Table 4 Basic concepts required for evaluation indicators

[0175]

[0176] 1. Recall

[0177] The recall rate, also known as the recall rate, calculates the proportion of successful hits among the top K recommended items of all items that the user is actually interested in. It reflects the accuracy of the recommendation system in identifying items of potential interest to the user and is:

[0178]

[0179] 2. Average reciprocal ranking (Mrr)

[0180] For each user session, find the rank of the first correct recommended item in the recommendation list, and then take its reciprocal. The average of the reciprocal ranks of all user sessions is Mrr. For example, in one session, the item the user is interested in is ranked 3rd in the recommendation list, and its reciprocal rank is 1 / 3. If it is ranked 6th in another session, the reciprocal rank is 1 / 6. If the rank is 20 higher, the reciprocal rank is set to zero. Mrr is a normalized score with a value range of [0,1]. An increase in the value indicates that most recommended items or projects will be higher in the ranking order of the recommendation list, which indicates that the performance of the recommendation system is better. i For the i-th user, the ranking position of the first item in the recommendation sequence result in the recommendation list is:

[0181]

[0182] The comparative experiment was carried out as follows:

[0183] 1. Baseline comparison experiment

[0184] This paper compares the HCPPO method with classic conversational recommendation method models at home and abroad, mainly comparing the following ten baseline models:

[0185] Item-KNN: The model recommends items based on the user's preferences in the current session and how similar these preferences are to those in other sessions.

[0186] FPMC: This model uses a combination of matrix factorization and first-order Markov chain methods to capture sequence characteristics and user preferences. Furthermore, referring to previous research, the model does not consider users' latent characteristics when calculating recommendation scores.

[0187] GRU4Rec: This model is based on RNN and uses a gated recurrent unit (GRU) to model user conversation sequences.

[0188] NARM: This model improves GRU4Rec by introducing an attention mechanism into the RNN used for conversational recommendations.

[0189] STAMP: Unlike previous studies, this model does not use an RNN encoder, but instead incorporates an attention layer. Furthermore, it relies entirely on retrieving the last item in the current session to capture the user's short-term attention.

[0190] SR-GNN: The model uses a gated GNN layer to capture item embeddings, then calculates the self-attention of the last item, and finally learns a representation of the session-level embedding;

[0191] CSRM: The model uses a memory network to query the latest session to better predict the user's interest preferences in the current session;

[0192] FGNN: In the process of learning item embedding, the model designs a weighted attention layer and graph-level feature extractor, which are used to learn the conversation for recommending the next item;

[0193] GCE-GNN: GCE-GNN creates local and global conversation graphs, and then integrates elements from both levels through GNN learning. A soft attention mechanism is used to aggregate the features of the learned elements to enhance the final conversation recommendation effect.

[0194] DHCN: DHCN first uses hypergraph convolution and line graph convolution to learn information within and between sessions. Then, through self-supervised learning, it fully exploits the session interaction information in these two graphs to optimize the session representation.

[0195] The experimental comparison results are shown in Table 5;

[0196] Table 5 Experimental comparison results

[0197]

[0198]

[0199] Table 5 presents the experimental comparison results of different methods on the Diginetica, Nowplaying, and Tmall datasets based on the Recall@K and Mrr@K metrics (K values ​​are 10 and 20). The following is a detailed analysis from the perspective of the overall method and the key methods:

[0200] As shown in Table 5, the performance of traditional machine learning-based recommendation models (Item_KNN and FPMC) is significantly poor. The fundamental reason is that these models treat conversations as unordered sets of items or use simple statistical features, failing to capture the rich contextual information and transition patterns inherent in item interaction sequences. This lack of information prevents the models from understanding the dynamics of evolving user interests and identifying sequential dependencies between items, thus limiting the predictive accuracy and explanatory power of recommendation systems in sequential decision-making scenarios. However, GRU4Rec, NARM, and STAMP outperform traditional methods, demonstrating the critical role of sequence effects in conversational recommendations. In comparison, GRU4Rec performs slightly worse than NARM and STAMP. This is because GRU4Rec only considers sequential behavior, while NARM and STAMP both utilize advanced neural network architectures and incorporate attention allocation mechanisms to weight item interaction sequences within a conversation. This further demonstrates that using only RNNs to learn user representations can lead to biased recommendations, and that introducing attention mechanisms can effectively mitigate this problem. Graph neural network recommendation models that emerged in 2019, such as SR-GNN, CSRN, FGNN, GCE-GNN, and DHCN, generally outperform previous models. This is attributed to the unique learning ability of graph neural networks. By modeling user conversations as graph representations, GNNs can systematically perform message passing and feature aggregation to generate more accurate item embeddings.

[0201] The experimental results of the proposed HCPPO model demonstrate superior overall performance. It particularly stands out on the Diginetica dataset, achieving a 3.59% improvement in Recall and an 8.64% improvement in Mrr compared to the DHCN model. This is because the encoding and embedding of session information within the hypergraph convolutional neural network better captures user information and complex high-order pairwise relationships between data. Furthermore, self-supervised learning exploits self-supervisory signals from different perspectives, maximizing the interaction information between sessions. A further contributing factor is the PPO algorithm's feedback and reward mechanisms, which guide the recommendation system to more effectively process diverse data types, exploring potential connections between user behavior and items, thereby achieving more accurate recommendations within an appropriate range. To verify PPO's superior learning capabilities, experiments were conducted using the GCE-GNN model. The experimental results shown in the table above demonstrate improvements across various datasets, with a particularly significant 9.89% improvement in Mrr on the Tmall dataset. This demonstrates that the PPO algorithm effectively optimizes the model's policy learning process, adapting it to the characteristics of diverse datasets and enhancing its generalization and recommendation effectiveness.

[0202] The present invention also conducts comparative experiments in the reinforcement learning method part;

[0203] In reinforcement learning, commonly used algorithms include DQN, DDQN, and PPO algorithms. In order to verify the performance of the PPO algorithm in the conversational recommendation system, different reinforcement algorithms are also integrated into this model to compare the experimental results. The experiment mainly uses the two indicators of Recall@K and Mrr@K to evaluate HCDQN (a model that integrates the DQN algorithm), HCDDQN (a model that integrates the DDQN algorithm), and HCPPO (a model that integrates the PPO algorithm) under different K values ​​(5, 10, and 20). Figure 3 As shown in the figure, experimental evaluation on the Diginetica dataset shows that the PPO algorithm has obvious performance advantages over the DQN and DDQN algorithms in conversational recommendation systems, regardless of which indicator is used.

[0204] First, because the PPO algorithm limits the difference between the old and new policies, it ensures that the policy update amplitude remains within a reasonable range. This proximal policy optimization approach avoids the training instability caused by large policy updates, as seen in DQN and DDQN. In conversational recommendation systems, user needs and conversation scenarios are constantly changing. Only with stable policy updates can the model better adapt to these changes, accurately grasp user preferences, and improve recommendation performance. Second, the PPO algorithm has advantages in handling long-term dependencies in sequential data. In conversational recommendation, user interests and needs are not static, resulting in long-term dependencies. The PPO algorithm effectively captures these changes, comprehensively considering the user's behavior throughout the entire conversation, rather than focusing solely on a single moment's behavior. The PPO algorithm connects these previous and subsequent behaviors to recommend items that better align with the user's long-term interests. Furthermore, the PPO algorithm is highly sample-efficient. During training, it fully utilizes the collected sample data, learning more effective policies from a small number of examples. In contrast, when faced with complex environments, DQN and DDQN find it difficult to accurately learn the optimal strategy due to some limitations of the model itself.

[0205] The present invention also conducted ablation experiments, specifically including:

[0206] 1. Impact of different components on model performance

[0207] In order to study the performance of each component of the PPO algorithm in the HCPPO model, this paper mainly designs two variants: HCPPO-NoA and HCPPO-NoC. HCPPO-NoA means that there is no advantage estimation function in the reinforcement learning module, and HCPPO-NoC means that there is no optimization objective function, that is, no clipping strategy. The experiments are mainly conducted on the Diginetica dataset, and the results are as follows. Figure 4 shown.

[0208] As can be seen from the figure, all metrics show an upward trend as the K value increases from 5 to 20. This indicates that as the length of the recommendation list increases, items relevant to the user's actual needs are more likely to be ranked higher in the recommended list, resulting in a better average reciprocal ranking. Furthermore, longer recommendation lists are more likely to recommend items relevant to the user's needs. In the PPO algorithm, the advantage estimation function measures the relative merit of taking an action compared to the average action under a specific state. The advantage estimation function can better understand user behavior patterns and capture user interests and preferences. Therefore, if the advantage estimation function is removed (i.e., HCPPO-NoA), the model cannot accurately determine the advantage of each action, resulting in biased recommendations and difficulty prioritizing items that truly meet user needs, ultimately leading to decreased model performance. The pruning strategy in the optimization objective function acts like a reasonable setting for the model's learning pace. This limits the magnitude of policy updates, allowing the model to better learn user preferences and avoid instability caused by excessive policy updates. Without a pruning strategy (i.e., HCPPO-NoC), the model may experience significant fluctuations during learning or over-adjust the strategy. Consequently, the model will struggle to capture user needs, inaccurately assess user preferences, and cause issues with the ranking of recommended items. Therefore, the full HCPPO model performs best in this scenario.

[0209] 2. Impact of cropping parameters on model performance

[0210] When using the PPO algorithm in the HCPPO model, the clipping parameter ε in the objective function is a key hyperparameter that directly controls the magnitude of policy updates and model performance. If the value is too small, the model will learn more slowly, requiring more training time and samples to converge to a better policy, and may become trapped in a local optimum. Conversely, a larger clipping parameter value allows the new policy to deviate more from the old policy, resulting in a larger policy update. Figure 5 The effects of different values ​​of clipping parameters on experimental performance on different indicators are studied.

[0211] After an in-depth analysis of the two indicators, the experiment found that the model performed best when the clipping parameter was set to 0.1. This is because this value found a particularly suitable balance point in controlling the amplitude of the policy update. If the clipping parameter is too small, the model learning speed will slow down, and it will not be able to timely and fully extract useful information from the data, and it will be difficult to accurately grasp user preferences. If the clipping parameter is too large, the model update strategy will be too large, causing the model to become unstable. However, when the clipping parameter is set to 0.1, the model can be trained at a reasonable pace, showing a higher accuracy rate and a more reasonable ranking of recommended items. Performance has been significantly improved in both the key indicators of Recall and Mrr.

[0212] 3. The impact of different neural network layers on model performance

[0213] In order to study the impact of different numbers of neural network layers on the model, an experiment was conducted for the HCPPO model to investigate the impact of the depth of the hypergraph convolutional network on model performance. Figure 6 and Figure 7 The results of network depth impact analysis based on two benchmark datasets, Diginetica and Tmall, are presented. The experiment systematically examines the impact of the number of model layers on various evaluation indicators of recommendation performance.

[0214] As shown in the figure, as the number of layers in the hypergraph convolutional network increases, the metric values ​​also improve. This is because with more layers, the model can capture richer high-order relationships in the data, more clearly analyzing user behavior patterns, which naturally improves the metric values. However, more layers are not necessarily better. Once the number of layers is too large, the network degrades, and the additional layers are unable to effectively transmit information, especially during backpropagation. This makes the model difficult to optimize, reduces its learning ability, and ultimately reduces performance indicators.

[0215] In the case of the Diginetica dataset, if the number of layers exceeds three, further additions will prevent the model from fully utilizing the newly added layers to extract valuable information and may instead introduce noise, resulting in performance degradation. On the Tmall dataset, adding layers beyond two may cause the model to overfit to the noise and specific patterns in the training data, making it difficult to accurately capture general user behavior patterns, resulting in reduced performance.

[0216] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0217] The present invention uses specific examples to illustrate the principles and implementation methods of the present invention. The above examples are only intended to help understand the method and core concept of the present invention. At the same time, those skilled in the art will find that the specific implementation methods and application scopes may vary based on the concept of the present invention. In summary, the contents of this specification should not be construed as limiting the present invention.

Claims

1. A hypergraph convolutional conversation recommendation method based on reinforcement learning, characterized in that: include: Step 1: Construct a hypergraph convolutional conversation recommendation model based on reinforcement learning, which includes a conversation hypergraph construction module, a conversation hypergraph convolutional network encoding module, a self-supervised learning enhanced conversation recommendation module, a reinforcement learning module, and a model prediction module; Step 2: Perform conversational recommendations based on the model.

2. The method according to claim 1, characterized in that In step 1, a hypergraph convolutional conversation recommendation model based on reinforcement learning is constructed, which specifically includes the following steps: Step 101: Construct a conversation hypergraph and generate a conversation hypergraph construction module; Step 102: Perform session hypergraph convolutional network coding to generate a session hypergraph convolutional network coding module; Step 103: Perform self-supervised enhanced conversation recommendation and generate a self-supervised learning enhanced conversation recommendation module; Step 104: Integrate the reinforcement learning algorithm into the model to generate a reinforcement learning module; Step 105: Perform model prediction and generate a model prediction module.

3. The method according to claim 2, characterized in that In step 101, a conversation hypergraph is constructed and a conversation hypergraph construction module is generated, specifically: The session data is defined as a time series set of user interaction items, where I is the item set and a user session is represented as s = {i s,1 ,i s,2 ,...,i s,n }, where i s,n Represents the click item of the item, maps each interactive item to the d-dimensional latent feature space, and uses the vector represents the embedding representation of item i in the network layer l, and the set represents the complete item embedding matrix, where each session s corresponds to a session-level representation vector This vector captures the overall semantic information of the session and the user's intention; Let G = {V, E} represent a hypergraph structure, where V is the vertex set, E is the hyperedge set, each hyperedge can connect any number of vertices, and each hyperedge is assigned a weight W ∈∈ , these weights form a diagonal matrix Hypergraph structure through incidence matrix Represents that its elements are defined as: if the vertex v i is included by the hyperedge, then H i∈ =1, otherwise H i∈ = 0, the degree of vertex and hyperedge is defined as: vertex v i degree represents the weight and degree of all hyperedges associated with the vertex, and the degree of hyperedge ∈ Indicates the number of vertices connected by the hyperedge; For a hypergraph G, the line graph of the hypergraph is defined as L(G), each node of L(G) is a hyperedge in G, and if there is at least one common node between two nodes of L(G) on the corresponding hyperedge in G, then it is defined as: L(G) = (V L ,E L ), where V L ={v e :v e ∈E}, E L ={(v ep ,v eq ):e p ,e q ∈E,|e p ∩e q ≥1|}, where e p ,e q It refers to a subset of hyperedges, and each edge of the line graph is assigned a weight W p,q , defined as W p,q =|e p ∪e q | / |e p ∪e q |; Each user's session is presented in the form of a hyperedge, that is, each hyperedge is represented as {i s,1 ,i s,2 ,i s,3 ,...,i s,n }∈E, where i s,n ∈V represents the click item of each item. The original conversation sequence will be constructed into a linear sequence, in which the two item click items i s,n-1 ,i s,n , only when you click on item i s,n Before and item click item i s,n-1 Only when there is interaction will there be corresponding connection relationships.

4. The method according to claim 3, characterized in that In step 102, session hypergraph convolutional network coding is performed to generate a session hypergraph convolutional network coding module, specifically: A hypergraph convolutional network is used as a feature learning framework to capture the complex high-order transformation relationships between items. In terms of the computational process, an information propagation mechanism is established based on the hypergraph Laplacian operator. The standard graph convolution is extended to the hypergraph structure so that it can handle one-to-many complex connection patterns. The information propagation items are learned and deep encoding of item semantic information is achieved through parameterized matrix mapping. Based on the spectral hypergraph convolution, the hypergraph convolution is defined as: Among them, the model does not use nonlinear activation functions and convolution filter parameter matrices. For W, each hyperedge is assigned the same weight 1, and the row normalization of the normalized hypergraph convolution is: The hypergraph convolution operation follows the node-hyperedge-node message passing mode. The node passes its feature representation to the hyperedge to which it belongs through the association matrix. Each hyperedge aggregates the multi-node information received to generate a hyperedge-level representation. The processed information is then passed back to the original node through the transpose of the association matrix, completing a complete feature propagation cycle. This is the information aggregation calculation, and then multiply it with H to aggregate the information in the second part. Then, the item embedding obtained at each layer is calculated as follows: Define a learnable position matrix, denoted as P r =[p1,p2,p3,...,p m ], integrating the learned item click items and reverse position embeddings, where m is the length of the current session, and the tth session s={i s,1 ,i s,2 ,...,i s,m } is embedded as: Where, is a learnable parameter; The session embedding is obtained by aggregating the items in the session. According to the method used in the SR-GNN model, the session s = {i s,1 ,i s,2 ,...,i s,m },for: Where θ h represents the user’s overall interest embedding in this session, are all learning project weights α t The attention parameter, represents the embedding of session s, which is represented by the average value of the embeddings of the items contained in session s. That is, the embedding of the t-th item in session s is expressed as: For a given session s, it is necessary to calculate the current scores of all candidate click items i∈I, which is: With the help of the softmax function, we can calculate the probability that an item in the session will become the next clicked item: The cross entropy loss function is used as the learning optimization objective, which is: Where N represents the number of samples, y i is the true predicted probability of the i-th sample, It is the model prediction probability, which ranges from 0 to 1. The cross entropy loss value of the entire model is obtained by summing the calculation results of each sample and taking the negative.

5. The method according to claim 4, characterized in that In step 103, self-supervised enhanced conversation recommendation is performed to generate a self-supervised learning enhanced conversation recommendation module, specifically: Initialize embedding learning for the nodes in the line graph by calculating the mean embedding of the items contained in each session. The correlation matrix of the line graph is defined as Where M represents the number of nodes in the line graph, that is, the number of hyperedges in the original hypergraph, which is: Define line graph convolution as: During the calculation process of each convolutional layer, the conversation node aggregates information from its neighborhood through the graph structured message passing mechanism. After completing the l-layer graph convolution propagation, the model calculates the weighted average of the representations at different levels to obtain the final conversation embedding representation, which is: A dual-channel architecture is designed to obtain two complementary conversation representations. The first channel uses a hypergraph to learn the connections between items within a conversation, revealing the interaction information within the conversation. The line graph encoded in the second channel focuses on depicting structural information at the inter-session level, namely the potential connections between different conversations, revealing the relationship between sessions. We consider the two channels as complementary perspectives for conversational learning, each capturing different conversational characteristics. We further introduce a contrastive learning framework to systematically compare and align the two sets of learned conversational embedding representations. A noise contrastive objective is used as the learning objective for the model. The loss generated by this objective is a standard binary cross-entropy loss. This loss is derived from two sets of samples: one is the real-world sample, i.e., the positive sample, and the other is the negative sample, which is defined as: Where, or Yes or Negative samples after sorting in column and row form, where It is a discriminant function that takes two vectors as input and can capture the consistency between them.

6. The method according to claim 5, characterized in that In step 104, the reinforcement learning algorithm is integrated into the model to generate a reinforcement learning module, specifically: A near-end policy optimization algorithm is integrated into the model. During the reinforcement learning process, various interaction data are collected in the environment through the policy network and the target network, including state, behavior, reward, and possible next state. The state is represented by the score output by the session hypergraph convolutional network encoding. The behavior is to select a specific item from the recommended candidate items. The agent selects an action based on the current state to maximize the cumulative reward. The possible next state is represented by the output score of the model, which reflects the new situation of the recommendation system after making a recommendation action. The reward is the negative of the model loss value. Compute advantage estimates and adjust the parameters of the policy network; Optimize the objective function and iteratively update the parameters θ of the policy network to continuously optimize the agent's strategy; Calculate target loss; Update the model strategy parameters and repeat the above steps until a certain number of iterations is reached or the model performance no longer improves.

7. The method according to claim 6, characterized in that In step 105, model prediction is performed to generate a model prediction module, specifically: During prediction training, based on the session embedding encoded by the hypergraph convolutional network and the reward signal fed back by reinforcement learning, the model ensures better accuracy in recommending the next item. The model uses the recommendation task generated by the hypergraph convolutional network channel as the main learning objective and the self-supervised learning task generated by the line graph convolutional channel as the auxiliary learning objective. The model also integrates the objective function of reinforcement learning for joint learning, and defines the loss function as: Where β and γ are parameters used to balance the loss.