A cross-domain combination recommendation system and method for digital rights
By constructing a hierarchical temporal graph federated architecture and a causal reinforcement learning policy network, combined with a large language model to generate natural language expressions, the accuracy and user experience issues in cross-domain equity portfolio recommendation are solved, achieving efficient utilization of equity assets and improved user satisfaction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JIANGXI MALI DIGITAL TECHNOLOGY CO LTD
- Filing Date
- 2026-03-31
- Publication Date
- 2026-07-31
AI Technical Summary
In existing technologies, rights and benefits recommendations are limited to a single platform or a single type, lacking the ability to discover cross-domain and cross-type rights and benefits combinations. This makes it impossible to effectively utilize users' rights and benefits assets on different platforms, resulting in a large number of idle rights and benefits. The accuracy of recommendations and user experience are poor, and there is a lack of a quantitative evaluation mechanism for the overall value of rights and benefits combinations, making it difficult to dynamically balance platform revenue and user rights and benefits value.
By constructing a hierarchical temporal graph federated architecture to mine cross-domain interest associations, employing a causal reinforcement learning policy network for dynamic interest combination pattern recognition, and combining a large language model to generate natural language expressions and interactive policy generation, cross-domain interest combination recommendation is realized.
It enables cross-domain rights association mining under privacy protection, improves the accuracy of recommendations and user experience, enhances the utilization efficiency of rights assets, increases user satisfaction and system reliability, and reduces system maintenance costs.
Smart Images

Figure CN122492307A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of digital rights recommendation technology, and relates to a cross-domain combination recommendation system and method for digital rights. Background Technology
[0002] With the rapid development of the digital economy, various digital benefits (such as memberships, coupons, points, data packages, and audio-visual content subscriptions) have become important means for businesses to enhance user stickiness and promote consumption conversion. Different business platforms (such as e-commerce, video, music, and travel) have all built independent benefit systems, and users have accumulated a large number of heterogeneous and scattered digital benefits across these platforms.
[0003] However, existing rights recommendation schemes typically suffer from the following problems:
[0004] The recommendation of rights and benefits is limited to a single platform or a single type, lacking the ability to discover cross-domain and cross-type rights and benefits combinations. This makes it impossible to effectively utilize users' rights and benefits assets on different platforms, resulting in a large number of idle rights and benefits and the failure to fully release the value of rights and benefits.
[0005] The recommendation strategy is mainly based on simple associations of users' historical behavior, and fails to deeply analyze the complex relationships between benefits such as complementarity, mutual exclusion, and usage order. The recommended combinations often do not conform to the user's actual usage scenario or consumption path, resulting in poor recommendation accuracy and user experience.
[0006] The lack of a quantitative assessment mechanism for the overall value of the rights and benefits portfolio makes it impossible to dynamically balance platform revenue and user rights and benefits value, and it is also difficult to adaptively optimize the recommendation strategy based on real-time feedback, resulting in the difficulty in continuously improving the long-term revenue and user satisfaction of the recommendation system. Summary of the Invention
[0007] In view of the problems existing in the prior art, the present invention provides a cross-domain combination recommendation system and method for digital rights, which is used to solve the above-mentioned technical problems.
[0008] To achieve the above and other objectives, the technical solution adopted by the present invention is as follows:
[0009] This invention provides a cross-domain combination recommendation method for digital rights, which includes the following steps:
[0010] Step S1: Asynchronously collect user rights data and rights metadata from multiple heterogeneous rights platforms through data interfaces. Clean, normalize, and align the collected data with heterogeneous data to obtain standardized rights data. Extract rights features from the standardized rights data, including rights type, usage scenario, validity period, and value attributes to obtain rights feature data.
[0011] Step S2: Based on the rights and interests feature data, cross-domain rights and interests associations are mined. By calculating the co-occurrence frequency and association strength of different rights and interests in time series, usage scenarios and user groups, a hierarchical time series graph federated architecture is constructed. Each rights and interests platform maintains a local time series graph network locally. The global coordinator constructs a global time series graph network through a time series-aware federated aggregation algorithm to generate dynamic rights and interests association network data. Based on the dynamic rights and interests association network data, rights and interests combination patterns are identified. A graph clustering algorithm is used to divide the rights and interests nodes into combination patterns to obtain rights and interests combination pattern data.
[0012] Step S3: Train a joint decision-making model based on the equity combination pattern data. Use the output of the dynamic equity association network as the state space of the reinforcement learning policy network to construct a causal reinforcement learning policy network. Adapt to graph structure changes through a meta-learning mechanism to obtain the joint decision-making model. Use the joint decision-making model to make policy decisions and score rankings for the equity combinations of the target user, and output candidate equity combinations and their recommended scores.
[0013] Step S4: Generate interactive strategies based on candidate benefit combinations and their recommendation scores. Input the basic actions output by the reinforcement learning strategy network into the large language model. The large language model generates natural language expression variants as the interaction interface. At the same time, the user's feedback on the natural language expression is fed back to the reinforcement learning strategy network as an implicit reward signal to obtain adaptive interaction strategy data. Execute recommendation operations based on the adaptive interaction strategy data. Call the distribution or redemption interfaces of various benefit platforms through APIs, and track the user's adoption and usage behavior of the recommended combinations. Optimize model parameters through feedback loops.
[0014] Preferably, the specific steps for constructing the hierarchical temporal graph federation architecture in step S2 include: Step S211: Constructing a local temporal graph network on each rights platform, with rights within the platform as nodes and user behavior sequences as edges, and learning the evolution and embedding of rights nodes in the time dimension through a temporal graph convolutional network to obtain local rights embedding data.
[0015] Step S212: Each platform uploads the parameter statistics (including mean vector and covariance matrix) of the local rights embedding to the global coordinator, but does not upload the original user interaction data;
[0016] Step S213: The global coordinator performs time-aware aggregation on the parameter statistics uploaded by each platform, uses a dynamic time warping algorithm to align the time window offsets of different platforms, calculates the aggregation parameters of global equity embedding, and obtains global equity embedding data.
[0017] Step S214: The global coordinator distributes the aggregated global stake embedding parameters to each platform. Each platform integrates the global information in its local model for the next round of iterative training until the model converges, thus obtaining dynamic stake association network data.
[0018] Preferably, the specific steps for constructing the causal reinforcement learning policy network in step S3 include:
[0019] Step S311: Construct a causal graph of equity association on historical data, and use a structural causal model to identify causal effects and spurious associations between equity nodes to obtain causal constraint data;
[0020] Step S312: Use the output graph snapshots of the dynamic stake association network at different timestamps as the state space of reinforcement learning, and define the state as St=[Ut,Gt,Ht], where Ut is the user's current embedding, Gt is the stake association graph snapshot at time t, and Ht is the historical state sequence;
[0021] Step S313: Use a meta-reinforcement learning algorithm to pre-train the policy network on graph snapshots of multiple time segments, so that it has the ability to generalize to adapt to changes in graph structure, and obtain the initial policy network;
[0022] Step S314: Use causal constraint data as the action filtering layer of the policy network. For candidate actions output by the policy network, if the associated equity combination is marked as a false association in the causal graph, reduce the selection probability of the action or require additional user confirmation to obtain causal-enhanced policy output.
[0023] Step S315: Using the causal reinforcement policy output and user feedback signal as rewards, optimize the policy network parameters through the policy gradient method to obtain a joint decision model.
[0024] Preferably, the specific steps for generating adaptive interaction strategy data in step S4 include: Step S411: Inputting the basic actions (including equity combination identifier and recommendation score) output by the joint decision model into the large language model;
[0025] Step S412: The large language model generates multiple natural language expression variants based on the basic actions and user context information, including recommendation reasons, usage scenario descriptions, and combined value descriptions, to obtain natural language expression data;
[0026] Step S413: Output the natural language expression data to the user terminal as an interactive interface, collect the user's feedback behavior on the natural language expression, including clicks, follow-up questions, clarifications and direct acceptance, and obtain user interaction feedback data;
[0027] Step S414: The user interaction feedback data is used as an implicit reward signal and fed back to the reinforcement learning policy network through the reward function enhancement mechanism, so that the policy network can simultaneously optimize the "recommended content" and "recommended expression" to obtain adaptive interaction policy data.
[0028] Preferably, the method further includes step S5: offline verification and online constraints based on graph causal reasoning, specifically including:
[0029] Step S51: In the offline stage, construct a causal graph of stake association based on historical data, identify causal effects and false associations between stake nodes, and inject the causal graph as prior knowledge into the initialization of the time sequence graph network;
[0030] Step S52: In the online phase, the causal graph is used as the action constraint layer. For candidate actions output by the joint decision model, if the associated portfolio of interests is marked as a false association in the causal graph, the recommendation priority of the action is reduced or a counterfactual explanation is required.
[0031] Step S53: When a recommendation is rejected by the user, the large language model generates counterfactual explanations based on the causal graph, including alternative combination suggestions and adjusted usage paths, thus obtaining counterfactual explanation data.
[0032] Preferably, the method further includes step S6: inter-module knowledge distillation based on self-supervised learning, specifically including:
[0033] Step S61: Design self-supervised pre-training tasks, including "predicting the equity portfolio for the next period", "reconstructing time series snapshots" and "inferring causal relationships", and perform multi-task joint pre-training on unlabeled data;
[0034] Step S62: The temporal graph network, reinforcement learning policy network and large language model share the same latent space. Knowledge distillation enables knowledge transfer between modules, achieving co-evolution among modules.
[0035] Step S63: During joint training, a gradient coordination mechanism is adopted to dynamically adjust the loss weights of each module to avoid training instability caused by gradient conflicts.
[0036] This invention also provides a cross-domain combination recommendation system for digital rights, the system comprising:
[0037] The data collection and standardization module is used to asynchronously collect user rights data and rights metadata from multiple heterogeneous rights platforms through data interfaces, and to clean, normalize and align the collected data with heterogeneous data to obtain standardized rights data.
[0038] The rights feature extraction module is used to extract rights features from standardized rights data, including rights type, usage scenario, validity period and value attributes, to obtain rights feature data.
[0039] The cross-domain association mining module is used to mine cross-domain equity association relationships based on equity feature data. By constructing a hierarchical time-series graph federated architecture, dynamic equity association network data is generated, and equity combination pattern recognition is performed based on the dynamic equity association network data to obtain equity combination pattern data.
[0040] The joint decision-making module is used to train a joint decision-making model based on equity combination pattern data. The output of the dynamic equity association network is used as the state space of the reinforcement learning policy network to construct a causal reinforcement learning policy network, obtain the joint decision-making model, and use the joint decision-making model to make policy decisions and score rankings for the equity combinations of the target user, and output candidate equity combinations and their recommended scores.
[0041] The interaction strategy generation module is used to generate interactive strategies based on candidate benefit combinations and their recommended scores. It inputs the basic actions output by the reinforcement learning strategy network into the large language model, which generates natural language expression variants as interaction interfaces. At the same time, it feeds back the user's feedback on the natural language expression as an implicit reward signal to the reinforcement learning strategy network to obtain adaptive interaction strategy data.
[0042] The recommendation execution and feedback module is used to perform recommendation operations based on adaptive interaction strategy data. It calls the distribution or redemption interfaces of various benefit platforms through APIs, tracks users' adoption and usage behavior of recommended combinations, and optimizes model parameters through feedback loops.
[0043] As described above, the cross-domain combination recommendation system and method for digital rights provided by the present invention has at least the following beneficial effects:
[0044] This invention provides a cross-domain combination recommendation system and method for digital rights. By constructing a hierarchical temporal graph federated architecture, it integrates global temporal patterns without exposing the original data, achieving privacy-preserving cross-domain rights association mining and solving the data silo problem. By constructing a causal-enhanced reinforcement learning policy network, the output of the dynamic rights association network is used as the state space. A meta-learning mechanism adapts to changes in the graph structure, achieving robust policy learning in dynamic environments. By deeply integrating the reinforcement learning policy network with a large language model, the decision network simultaneously optimizes both "recommended content" and "recommended expression," achieving synergistic optimization of decision accuracy and interaction richness. By introducing graph causal reasoning as a unified constraint layer, it filters false associations and generates counterfactual explanations, improving the system's interpretability and user trust.
[0045] The aforementioned deeply integrated architecture enables each technical module to be mutually dependent, mutually reinforcing, and jointly optimized. On one hand, this method, through a hierarchical temporal graph federated architecture, significantly improves the accuracy of cross-platform equity association mining, solving the problem of cross-domain data incompatibility in current equity recommendation systems. On the other hand, through causal reinforcement learning and interactive policy generation, it significantly improves the long-term cumulative returns of recommendations and user satisfaction, while reducing system maintenance costs. This intelligent and collaborative recommendation mechanism not only effectively improves the utilization efficiency of equity assets but also enhances the user experience in cross-platform scenarios, ensuring the system's economic efficiency and security, and greatly enhancing business continuity and reliability in complex environments. Attached Figure Description
[0046] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0047] Figure 1 This is a schematic diagram showing the connections between the steps of the method of the present invention. Detailed Implementation
[0048] The following description, in conjunction with the implementation of this invention, is merely an example and illustration of the concept of this invention. Those skilled in the art can make various modifications or additions to the specific embodiments described, or use similar methods to replace them, as long as they do not deviate from the inventive concept or exceed the scope defined in these claims, all of which should fall within the protection scope of this invention.
[0049] Example 1
[0050] Please see Figure 1 As shown, a cross-domain combination recommendation method for digital rights includes the following steps:
[0051] Step S1: Asynchronously collect user rights data and rights metadata from multiple heterogeneous rights platforms through data interfaces. Clean, normalize, and align the collected data with heterogeneous data to obtain standardized rights data. Extract rights features from the standardized rights data, including rights type, usage scenario, validity period, and value attributes to obtain rights feature data.
[0052] For example, step S1 includes:
[0053] Step S11: Asynchronously obtain user rights data and rights metadata from multiple heterogeneous rights platforms through data interfaces, and record the data returned by each platform with timestamp and platform source identifier to obtain the original user rights dataset and the original rights metadata dataset.
[0054] Step S12: Perform format validation and field mapping analysis on the original user rights dataset and the original rights metadata dataset, perform unified field mapping on heterogeneous data models of different platforms, complete missing fields, and remove or correct abnormal data to obtain formatted user rights data and formatted rights metadata.
[0055] Step S13: Based on the formatted user rights data and formatted rights metadata, merge and associate the data according to the unified rights identifier field, and perform time-series normalization processing according to the rights issuance time and validity period to obtain standardized rights data;
[0056] Step S14: Extract static features of standardized rights data, including the type, face value, platform, and applicable category of rights, to obtain static feature data of rights.
[0057] Step S15: Extract dynamic features of rights from standardized rights data, including the distribution of rights usage periods, historical redemption rates, and user preference attributes, to obtain dynamic feature data of rights.
[0058] Step S16: Perform feature fusion on the static feature data and dynamic feature data of rights and interests, calculate the comprehensive value score and scenario label vector of each right and interest, and obtain the feature data of rights and interests.
[0059] It should be added that step S14 includes:
[0060] Step S141: Based on standardized rights data, extract the type identifiers of rights, including discount coupons, full reduction coupons, memberships, points, and physical redemption categories. Encode and convert the type identifiers to obtain rights type feature data.
[0061] Step S142: Based on standardized rights data, extract the face value attributes of the rights, including discount ratio, reduction amount, value points, and market reference price. Perform numerical normalization processing on the face value attributes to obtain the face value feature data of the rights.
[0062] Step S143: Based on standardized rights data, extract the platform identifier and applicable category tree of the rights, perform one-hot encoding on the platform identifier, and perform hierarchical feature extraction on the applicable category to obtain rights source feature data;
[0063] Step S144: Based on the equity type feature data, equity face value feature data, and equity source feature data, integrate the static features of equity to form an equity static feature vector and output the equity static feature data.
[0064] Step S2: Based on the rights and interests feature data, cross-domain rights and interests associations are mined. By calculating the co-occurrence frequency and association strength of different rights and interests in time series, usage scenarios and user groups, a hierarchical time series graph federation architecture is constructed. Each rights and interests platform maintains a local time series graph network locally. The global coordinator constructs a global time series graph network through a time series-aware federation aggregation algorithm to generate dynamic rights and interests association network data. Based on the dynamic rights and interests association network data, rights and interests combination patterns are identified. A graph clustering algorithm is used to divide the rights and interests nodes into combination patterns to obtain rights and interests combination pattern data.
[0065] For example, step S2 includes:
[0066] Step S21: Based on the rights and benefits feature data, serialize the user's historical rights and benefits usage records, and construct the rights and benefits usage behavior of the same user on different platforms and at different times in chronological order as rights and benefits usage sequence data;
[0067] Step S22: Perform co-occurrence analysis on the rights and interests use sequence data, calculate the number of times and intervals in which any two rights and interests co-occur in the same user sequence, and obtain rights and interests co-occurrence frequency data;
[0068] Step S23: Based on the co-occurrence frequency data of rights and interests, and combined with the similarity of the scenario tags and the complementarity of the value of rights and interests, construct a hierarchical time-series graph federation architecture to generate dynamic rights and interests association network data;
[0069] Step S24: Based on the dynamic equity association network data, a graph convolutional network is used to vectorize the equity nodes, extract the structural features and neighborhood features of the equity nodes, and obtain the embedded vector data of the equity nodes.
[0070] Step S25: Based on the embedded vector data of equity nodes, a clustering algorithm is used to divide the equity nodes into combination patterns. Equity nodes with similar association patterns are clustered into the same combination type to obtain equity combination pattern data.
[0071] It should be added that the specific steps for constructing the hierarchical sequence diagram federated architecture in step S23 include:
[0072] Step S231: Construct a local temporal graph network on each rights platform. With rights within the platform as nodes and user behavior sequences as edges, learn the evolution and embedding of rights nodes in the time dimension through a temporal graph convolutional network to obtain local rights embedding data.
[0073] Step S232: Each platform uploads the parameter statistics (including mean vector and covariance matrix) of the local rights embedding to the global coordinator, but does not upload the original user interaction data;
[0074] Step S233: The global coordinator performs time-aware aggregation on the parameter statistics uploaded by each platform, uses a dynamic time warping algorithm to align the time window offsets of different platforms, calculates the aggregation parameters of global equity embedding, and obtains global equity embedding data.
[0075] Step S234: The global coordinator distributes the aggregated global stake embedding parameters to each platform. Each platform integrates the global information in its local model for the next round of iterative training until the model converges, thus obtaining dynamic stake association network data.
[0076] Step S3: Train a joint decision-making model based on the equity combination pattern data. Use the output of the dynamic equity association network as the state space of the reinforcement learning policy network to construct a causal reinforcement learning policy network. Adapt to graph structure changes through a meta-learning mechanism to obtain the joint decision-making model. Use the joint decision-making model to make policy decisions and score rankings for the equity combinations of the target user, and output candidate equity combinations and their recommended scores.
[0077] For example, step S3 includes:
[0078] Based on the data of benefit combination patterns, samples are constructed and positive and negative samples are divided into user historical interaction data. The benefit combinations adopted by users are used as positive samples, and the combinations that users exposed but did not adopt are used as negative samples. Training sample set and validation sample set are established to obtain sample stratification data.
[0079] Feature engineering is performed on the stratified sample data to extract user-benefit interaction features, including users' historical click-through rate, redemption rate and usage preferences for each benefit in the combination. At the same time, benefit-benefit association features are extracted, including the association strength, scenario similarity and value complementarity between benefits in the combination. The features are then normalized and encoded to obtain the combined feature vector data.
[0080] Based on combined feature vector data, a causal-enhanced reinforcement learning policy network is constructed. The output of the dynamic stake association network is used as the state space. The joint decision model is obtained by adapting to graph structure changes through a meta-learning mechanism.
[0081] The joint decision-making model is validated by calculating its prediction accuracy, recall, and normalized depreciation cumulative gain. The model output is compared with the historical actual adoption results to obtain model validation data.
[0082] Based on model validation data, the parameters of the joint decision-making model are adjusted and optimized to determine a stable model version that can be used for real-time recommendation, and the final joint decision-making model is output. The final joint decision-making model is used to make combined recommendations to target users, and the candidate benefit combinations are ranked and scored according to strategy. The top K candidate benefit combinations and their corresponding recommendation scores are output, and the candidate benefit combination data is obtained.
[0083] It should be added that the specific steps in step S3 for constructing the causal reinforcement learning policy network include:
[0084] Step S311: Construct a causal graph of equity association on historical data, and use a structural causal model to identify causal effects and spurious associations between equity nodes to obtain causal constraint data;
[0085] Step S312: Use the output graph snapshots of the dynamic stake association network at different timestamps as the state space of reinforcement learning, and define the state as St=[Ut,Gt,Ht], where Ut is the user's current embedding, Gt is the stake association graph snapshot at time t, and Ht is the historical state sequence;
[0086] Step S313: Use a meta-reinforcement learning algorithm to pre-train the policy network on graph snapshots of multiple time segments, so that it has the ability to generalize to adapt to changes in graph structure, and obtain the initial policy network;
[0087] Step S314: Use causal constraint data as the action filtering layer of the policy network. For candidate actions output by the policy network, if the associated equity combination is marked as a false association in the causal graph, reduce the selection probability of the action or require additional user confirmation to obtain causal-enhanced policy output.
[0088] Step S315: Use the causal reinforcement policy output and user feedback signal as rewards, optimize the policy network parameters through the policy gradient method, and obtain the joint decision model;
[0089] In a preferred embodiment of the present invention, the construction of the three-dimensional state representation in step S312 is specifically implemented through the following sub-steps:
[0090] Step S3121: Construct the user's real-time state feature vector
[0091] Based on the standardized rights data obtained in step S1 and the user-rights interaction features extracted in step S3, a real-time user status feature vector is constructed. Specifically:
[0092] First, collect the target user's real-time context information at the current moment, including:
[0093] User static profile features: age range, gender, membership level, and historical spending power score, which are processed through one-hot encoding and numerical normalization to form a static feature vector;
[0094] User short-term behavior characteristics: The user's rights browsing, clicking, claiming, and redemption behavior sequence in the past 24 hours is statistically analyzed, and the behavior sequence is encoded using a gated recurrent unit network to obtain short-term behavior embedding vectors;
[0095] Long-term user preference characteristics: The redemption frequency distribution of users on different types of benefits (discount coupons, memberships, points, physical redemptions, etc.) in the past 90 days is statistically analyzed, and feature mapping is performed through a multilayer perceptron to obtain the long-term preference embedding vector;
[0096] Real-time scene features: current time period (morning / noon / evening / late night), user geographical location (based on location service identification of business districts / residential / office areas), device type (mobile / PC), which are combined to form a scene feature vector.
[0097] The four feature vectors are concatenated and fused using a two-layer fully connected network to output a real-time user embedding vector. This embedding vector is updated every 15 minutes, or in real time when the user generates a new interaction, and is used to represent the user's personalized state at the moment of decision-making.
[0098] Step S3122: Obtain a snapshot of the rights and interests association topology.
[0099] Based on the dynamic stake association network constructed in step S2, a snapshot of the stake association topology at the current moment is extracted. This snapshot is organized in a graph structure and includes the following elements:
[0100] Rights Node Set: Contains all active digital rights in the system. Each node is associated with the rights feature vector extracted in step S1, including attributes such as rights type, face value, platform, and applicable category.
[0101] Edge set: Represents the relationship between equity nodes. An edge connection is established when the overall correlation strength between two equity nodes exceeds a preset threshold.
[0102] Edge weight matrix: Stores the correlation strength values between each equity pair, with values normalized to the range of 0 to 1, reflecting the tightness of the correlation between equity pairs.
[0103] The specific method for generating topology snapshots is as follows:
[0104] (1) Obtain the embedded representation of each stake node at the current time from the dynamic stake association network output in step S23;
[0105] (2) Calculate the cosine similarity between any two interest nodes as the basic association metric;
[0106] (3) Combining the historical co-occurrence frequency, scene similarity and value complementarity calculated in step S23, the comprehensive correlation strength is calculated by weighted fusion;
[0107] (4) Normalize the association strength and set a threshold to filter weak associations, and finally obtain a snapshot of the equity association topology stored in the form of a sparse graph.
[0108] In the actual system, the topology snapshot is updated every 15 minutes, or an incremental update is triggered when there are new rights or significant changes in rights usage behavior, to ensure that the dynamic evolution of rights relationships can be reflected in real time.
[0109] Step S3123: Constructing a historical evolutionary trajectory sequence
[0110] The historical evolution trajectory sequence is defined as a concatenation of state records from multiple previous moments, used to capture the temporal evolution patterns of user states and stake association topologies. Specifically, a sliding window mechanism is employed for implementation.
[0111] (1) Maintain a state cache queue to store state records of several past moments in chronological order. Each state record contains the user's real-time state feature vector and a snapshot of the rights and interests association topology at the corresponding moment.
[0112] (2) To avoid dimensionality explosion, the topological structure snapshot of the rights and interests at each historical moment is compressed using a graph convolutional encoder to extract its core structural features and form a fixed-dimensional topological structure embedding vector.
[0113] (3) The user state feature vector at a historical moment is concatenated with the topological structure embedding vector to form a compressed state representation at that moment;
[0114] (4) The compressed state representations of all historical moments are spliced together in chronological order to form a complete historical evolution trajectory sequence vector.
[0115] In this embodiment, the historical window length is set to 6, corresponding to the evolutionary trajectory over the past 90 minutes (one sampling point every 15 minutes). By introducing the historical trajectory sequence, the reinforcement learning policy network can identify the periodic changes and trend evolution of user behavior patterns and interest relationships.
[0116] Step S3124: Fusion and normalization of three-dimensional state representations
[0117] The user real-time state feature vector obtained in step S3121, the graph embedding representation of the interest association topology snapshot obtained in step S3122, and the historical evolution trajectory sequence vector obtained in step S3123 are spliced and fused to form a complete reinforcement learning three-dimensional state representation.
[0118] The dimension of this three-dimensional state representation is the sum of three dimensions, where:
[0119] The user's real-time status feature vector has 256 dimensions;
[0120] The embedded dimension of the equity association topology is 64 dimensions;
[0121] The historical evolution trajectory sequence has a dimension of 6 times 320 (i.e., 1920 dimensions).
[0122] The total dimensions after merging the three are 2240. To improve the stability and convergence speed of model training, each dimension of the three-dimensional state representation is normalized to have a mean of 0 and a standard deviation of 1, resulting in the final state vector input to the reinforcement learning policy network.
[0123] Step S3125: Dynamic update mechanism for three-dimensional state representation
[0124] To ensure the real-time performance and accuracy of state representation, the system employs a multi-level update mechanism:
[0125] Real-time updates: When a user generates a new interaction (click, claim, redeem), the user's real-time state feature vector is immediately updated, and the updated three-dimensional state representation is pushed into the policy network for immediate decision-making.
[0126] Periodic updates: Every 15 minutes, an update of the equity association topology snapshot is triggered, and the sliding window of the historical evolution trajectory sequence is updated at the same time, and the three-dimensional state representation is recalculated.
[0127] Event-triggered update: When a significant change in the equity pool is detected (new equity launched, equity expires, equity attributes are adjusted), an incremental update of the equity association topology is immediately triggered, and the three-dimensional state representation is recalculated.
[0128] Through the above mechanism, reinforcement learning policy networks can perceive changes in user personalized status, equity-related environmental topology, and historical evolution trajectory in real time, providing complete state observation information for robust policy learning in dynamic environments.
[0129] Step S4: Generate interactive strategies based on candidate benefit combinations and their recommendation scores. Input the basic actions output by the reinforcement learning strategy network into the large language model. The large language model generates natural language expression variants as the interaction interface. At the same time, the user's feedback on the natural language expression is fed back to the reinforcement learning strategy network as an implicit reward signal to obtain adaptive interaction strategy data. Execute recommendation operations based on the adaptive interaction strategy data. Call the distribution or redemption interfaces of various benefit platforms through APIs, and track the user's adoption and usage behavior of the recommended combinations. Optimize model parameters through feedback loops.
[0130] For example, step S4 includes:
[0131] Step S41: Based on the candidate benefit combination data, analyze the user's current real-time context information, including the user's geographical location, time period, device type, and current active scenario, to obtain user context feature data;
[0132] Step S42: Based on user context feature data and combined value assessment data, calculate the real-time adaptability of candidate benefit combinations, evaluate the value prominence and user operation convenience of each combination in the current scenario, and obtain the real-time adaptability data of the combinations.
[0133] Step S43: Input the basic actions (including equity portfolio identifier and recommendation score) output by the joint decision-making model into the large language model to generate adaptive interaction strategy data;
[0134] Step S44: Analyze the resource consumption and platform interface availability of the adaptive interaction strategy data. Combine the current issuance amount, interface response time and system load status of each rights platform to filter and dynamically adjust the interaction strategy to obtain the final interaction strategy data.
[0135] Step S45: Based on the final interaction strategy data, execute the recommendation operation by calling the rights distribution or redemption interface of each rights platform through API, monitor the interface return status and user reach in real time, and obtain recommendation execution data;
[0136] Step S46: Track user behavior on the recommended execution data, collect user clicks, claims, redemptions and sharing behaviors of recommended combinations, construct user feedback behavior sequences, and obtain user feedback data;
[0137] Step S47: Based on user feedback data, evaluate the adoption rate, redemption rate, and user satisfaction of the recommended combinations, calculate the comprehensive benefit index of the recommendation strategy, and use the feedback data as incremental samples to input into the joint decision-making model for parameter updates, outputting the final recommendation effect data.
[0138] It should be added that the specific steps for generating adaptive interaction strategy data in step S43 include:
[0139] Step S431: Input the basic actions (including equity portfolio identifier and recommendation score) output by the joint decision-making model into the large language model;
[0140] Step S432: The large language model generates multiple natural language expression variants based on the basic actions and user context information, including recommendation reasons, usage scenario descriptions, and combined value descriptions, to obtain natural language expression data;
[0141] Step S433: Output the natural language expression data to the user terminal as an interactive interface, collect the user's feedback behavior on the natural language expression, including clicks, follow-up questions, clarifications and direct acceptance, and obtain user interaction feedback data;
[0142] Step S434: The user interaction feedback data is used as an implicit reward signal and fed back to the reinforcement learning policy network through the reward function enhancement mechanism, so that the policy network can simultaneously optimize the "recommended content" and "recommended expression" to obtain adaptive interaction policy data.
[0143] Preferably, the method further includes step S5: offline verification and online constraints based on graph causal reasoning, specifically including:
[0144] Step S51: In the offline stage, construct a causal graph of stake association based on historical data, identify causal effects and false associations between stake nodes, and inject the causal graph as prior knowledge into the initialization of the time sequence graph network;
[0145] Step S52: In the online phase, the causal graph is used as the action constraint layer. For candidate actions output by the joint decision model, if the associated portfolio of interests is marked as a false association in the causal graph, the recommendation priority of the action is reduced or a counterfactual explanation is required.
[0146] Step S53: When a recommendation is rejected by the user, the large language model generates counterfactual explanations based on the causal graph, including alternative combination suggestions and adjusted usage paths, thus obtaining counterfactual explanation data.
[0147] Preferably, the method further includes step S6: inter-module knowledge distillation based on self-supervised learning, specifically including:
[0148] Step S61: Design self-supervised pre-training tasks, including "predicting the equity portfolio for the next period", "reconstructing time series snapshots" and "inferring causal relationships", and perform multi-task joint pre-training on unlabeled data;
[0149] Step S62: The temporal graph network, reinforcement learning policy network and large language model share the same latent space. Knowledge distillation enables knowledge transfer between modules, achieving co-evolution among modules.
[0150] Step S63: During joint training, a gradient coordination mechanism is adopted to dynamically adjust the loss weights of each module to avoid training instability caused by gradient conflicts.
[0151] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0152] It should be understood that determining B based on A does not mean determining B solely based on A; it also means determining B based on A and / or other information.
[0153] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0154] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for cross-domain combination recommendation of digital rights, characterized in that, The method includes: Step S1: Asynchronously collect user rights data and rights metadata from multiple heterogeneous rights platforms through data interfaces. Clean, normalize, and align the collected data with heterogeneous data to obtain standardized rights data. Extract rights features from the standardized rights data, including rights type, usage scenario, validity period, and value attributes to obtain rights feature data. Step S2: Based on the rights and interests feature data, cross-domain rights and interests associations are mined. By calculating the co-occurrence frequency and association strength of different rights and interests in time series, usage scenarios and user groups, a hierarchical time series graph federated architecture is constructed. Each rights and interests platform maintains a local time series graph network locally. The global coordinator constructs a global time series graph network through a time series-aware federated aggregation algorithm to generate dynamic rights and interests association network data. Based on the dynamic rights and interests association network data, rights and interests combination patterns are identified. A graph clustering algorithm is used to divide the rights and interests nodes into combination patterns to obtain rights and interests combination pattern data. Step S3: Train a joint decision-making model based on the equity combination pattern data. Use the output of the dynamic equity association network as the state space of the reinforcement learning policy network to construct a causal reinforcement learning policy network. Adapt to graph structure changes through a meta-learning mechanism to obtain the joint decision-making model. Use the joint decision-making model to make policy decisions and score rankings for the equity combinations of the target user, and output candidate equity combinations and their recommended scores. Step S4: Generate interactive strategies based on candidate benefit combinations and their recommendation scores. Input the basic actions output by the reinforcement learning strategy network into the large language model. The large language model generates natural language expression variants as the interaction interface. At the same time, the user's feedback on the natural language expression is fed back to the reinforcement learning strategy network as an implicit reward signal to obtain adaptive interaction strategy data. Execute recommendation operations based on the adaptive interaction strategy data. Call the distribution or redemption interfaces of various benefit platforms through APIs, and track the user's adoption and usage behavior of the recommended combinations. Optimize model parameters through feedback loops.
2. The cross-domain combination recommendation method for digital rights according to claim 1, characterized in that, The specific steps in step S2 to construct the hierarchical sequence diagram federated architecture include: Step S211: Construct a local temporal graph network on each rights platform. With rights within the platform as nodes and user behavior sequences as edges, learn the evolution and embedding of rights nodes in the time dimension through a temporal graph convolutional network to obtain local rights embedding data. Step S212: Each platform uploads the parameter statistics of the embedded local rights to the global coordinator, but does not upload the original user interaction data; Step S213: The global coordinator performs time-aware aggregation on the parameter statistics uploaded by each platform, uses a dynamic time warping algorithm to align the time window offsets of different platforms, calculates the aggregation parameters of global equity embedding, and obtains global equity embedding data. Step S214: The global coordinator distributes the aggregated global stake embedding parameters to each platform. Each platform integrates the global information in its local model for the next round of iterative training until the model converges, thus obtaining dynamic stake association network data.
3. The cross-domain combination recommendation method for digital rights according to claim 1, characterized in that, The specific steps in step S3 for constructing the causal reinforcement learning policy network include: Step S311: Construct a causal graph of equity association on historical data, and use a structural causal model to identify causal effects and spurious associations between equity nodes to obtain causal constraint data; Step S312: Use the output graph snapshots of the dynamic stake association network at different timestamps as the state space of reinforcement learning, and define the state as St=[Ut,Gt,Ht], where Ut is the user's current embedding, Gt is the stake association graph snapshot at time t, and Ht is the historical state sequence; Step S313: Use a meta-reinforcement learning algorithm to pre-train the policy network on graph snapshots of multiple time segments, so that it has the ability to generalize to adapt to changes in graph structure, and obtain the initial policy network; Step S314: Use causal constraint data as the action filtering layer of the policy network. For candidate actions output by the policy network, if the associated equity combination is marked as a false association in the causal graph, reduce the selection probability of the action or require additional user confirmation to obtain causal-enhanced policy output. Step S315: Using the causal reinforcement policy output and user feedback signal as rewards, optimize the policy network parameters through the policy gradient method to obtain a joint decision model.
4. The cross-domain combination recommendation method for digital rights according to claim 1, characterized in that, The specific steps in step S4 for generating adaptive interaction strategy data include: Step S411: Input the basic actions output by the joint decision-making model into the large language model; Step S412: The large language model generates multiple natural language expression variants based on the basic actions and user context information, including recommendation reasons, usage scenario descriptions, and combined value descriptions, to obtain natural language expression data; Step S413: Output the natural language expression data to the user terminal as an interactive interface, collect the user's feedback behavior on the natural language expression, including clicks, follow-up questions, clarifications and direct acceptance, and obtain user interaction feedback data; Step S414: The user interaction feedback data is used as an implicit reward signal and fed back to the reinforcement learning policy network through the reward function enhancement mechanism, so that the policy network can simultaneously optimize the "recommended content" and "recommended expression" to obtain adaptive interaction policy data.
5. The cross-domain combination recommendation method for digital rights according to claim 1, characterized in that, It also includes step S5: offline verification and online constraints based on graph causal reasoning, specifically including: Step S51: In the offline stage, construct a causal graph of stake association based on historical data, identify causal effects and false associations between stake nodes, and inject the causal graph as prior knowledge into the initialization of the time sequence graph network; Step S52: In the online phase, the causal graph is used as the action constraint layer. For candidate actions output by the joint decision model, if the associated portfolio of interests is marked as a false association in the causal graph, the recommendation priority of the action is reduced or a counterfactual explanation is required. Step S53: When a recommendation is rejected by the user, the large language model generates counterfactual explanations based on the causal graph, including alternative combination suggestions and adjusted usage paths, thus obtaining counterfactual explanation data.
6. The cross-domain combination recommendation method for digital rights according to claim 1, characterized in that, It also includes step S6: inter-module knowledge distillation based on self-supervised learning, specifically including: Step S61: Design self-supervised pre-training tasks, including "predicting the equity portfolio for the next period", "reconstructing time series snapshots" and "inferring causal relationships", and perform multi-task joint pre-training on unlabeled data; Step S62: The temporal graph network, reinforcement learning policy network and large language model share the same latent space. Knowledge distillation enables knowledge transfer between modules, achieving co-evolution among modules. Step S63: During joint training, a gradient coordination mechanism is adopted to dynamically adjust the loss weights of each module to avoid training instability caused by gradient conflicts.
7. The cross-domain combination recommendation method for digital rights according to claim 2, characterized in that, The parameter statistics mentioned in step S212 include the mean vector and the covariance matrix; the time-aware aggregation mentioned in step S213 uses a dynamic time warping algorithm to align the time window offsets of different platforms.
8. The cross-domain combination recommendation method for digital rights according to claim 3, characterized in that, The historical state sequence Ht mentioned in step S312 includes the preceding graph snapshot sequence and the corresponding user feedback sequence.
9. The cross-domain combination recommendation method for digital rights according to claim 4, characterized in that, The natural language expression variants described in step S412 are generated using low-rank adaptation fine-tuning techniques, enabling large language models to quickly adapt to changes in reinforcement learning strategies.
10. A cross-domain combination recommendation system for digital rights, characterized in that, The system includes: The data collection and standardization module is used to asynchronously collect user rights data and rights metadata from multiple heterogeneous rights platforms through data interfaces, and to clean, normalize and align the collected data with heterogeneous data to obtain standardized rights data. The rights feature extraction module is used to extract rights features from standardized rights data, including rights type, usage scenario, validity period and value attributes, to obtain rights feature data. The cross-domain association mining module is used to mine cross-domain equity association relationships based on equity feature data. By constructing a hierarchical time-series graph federated architecture, dynamic equity association network data is generated, and equity combination pattern recognition is performed based on the dynamic equity association network data to obtain equity combination pattern data. The joint decision-making module is used to train a joint decision-making model based on equity combination pattern data. The output of the dynamic equity association network is used as the state space of the reinforcement learning policy network to construct a causal reinforcement learning policy network, obtain the joint decision-making model, and use the joint decision-making model to make policy decisions and score rankings for the equity combinations of the target user, and output candidate equity combinations and their recommended scores. The interaction strategy generation module is used to generate interactive strategies based on candidate benefit combinations and their recommended scores. It inputs the basic actions output by the reinforcement learning strategy network into the large language model, which generates natural language expression variants as interaction interfaces. At the same time, it feeds back the user's feedback on the natural language expression as an implicit reward signal to the reinforcement learning strategy network to obtain adaptive interaction strategy data. The recommendation execution and feedback module is used to perform recommendation operations based on adaptive interaction strategy data, call the distribution or redemption interfaces of various benefit platforms through API, track users' adoption and usage behavior of recommended combinations, and optimize model parameters through feedback loops. The cross-domain association mining module includes a hierarchical time-series graph federated architecture unit, which builds a local time-series graph network on each rights platform, and the global coordinator builds a global time-series graph network through a time-aware federated aggregation algorithm; The joint decision-making module includes a causal enhancement unit, which constructs a causal graph of interest association based on historical data and uses causal constraints as an action filtering layer of the policy network. The interaction strategy generation module includes a large language model interaction unit, which converts the basic actions output by the reinforcement learning policy network into variations of natural language expressions and feeds back user feedback as an implicit reward signal.