Cross-platform user behavior data intelligent aggregation and analysis processing method and system
By building a decentralized federated reinforcement learning framework and multi-layer causal relationship diagram, combined with graph attention network and counterfactual reasoning, the privacy protection and prediction accuracy problems in cross-platform user behavior analysis are solved, and cross-platform data collaborative training and personalized services are realized.
Patent Information
- Application Number
- CN202510898495.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-01
- Publication Date
- 2025-08-01
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing cross-platform user behavior analysis technology has difficulty in protecting data privacy, lack of interpretability in predicting results, and difficulty in capturing the causal relationships and long-term dependencies behind user behavior, resulting in insufficient accuracy in user intention identification and behavior prediction.
Build a decentralized federated reinforcement learning framework, use behavioral transfer probability graph training graph attention network and comparison learning network, combine multi-layer causal graph and counterfactual reasoning methods to realize cross-platform collaborative training and behavior prediction, and protect user privacy through differential privacy mechanisms.
On the premise of protecting user privacy, the accuracy of user intention recognition and the accuracy of behavior prediction are improved, the interpretability of prediction results and the self-optimization of the model are enhanced, and personalized service support is provided.
Smart Images

Figure CN120408101A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to data processing technologies, and particularly to an intelligent aggregation and analysis processing method and system for cross-platform user behavior data. Background Art
[0002] With the rapid development of the Internet and mobile Internet, users generate a large amount of behavior data on different platforms, and this data contains rich user intention and preference information. Cross-platform user behavior data analysis has become a key technology for understanding user needs, enhancing user experience, and achieving precision marketing. Traditional user behavior analysis methods mainly focus on data mining and pattern recognition within a single platform, and it is difficult to comprehensively capture the complex behavior patterns of users in a multi-platform environment. In recent years, with the progress of artificial intelligence technologies, user behavior modeling methods based on deep learning, reinforcement learning, and graph neural networks have gradually emerged, providing new technical paths for cross-platform user intention recognition and behavior prediction.
[0003] However, the existing cross-platform user behavior analysis technologies still have the following defects and deficiencies: First, the existing technologies usually adopt a centralized data processing architecture, which requires centralized analysis of user data from each platform. This not only faces strict data privacy protection regulations but also increases the risk of data leakage, making it difficult to achieve effective cross-platform data sharing and collaborative analysis while protecting user privacy. Second, most of the existing user behavior prediction models are based on statistical correlation analysis, lacking in-depth understanding of the causal relationships behind user behaviors, resulting in the lack of interpretability of prediction results and relatively low prediction accuracy when facing sparse or new user behavior patterns. In addition, when dealing with time-series behavior data, the existing technologies often ignore the long-term dependencies and context information in the behavior sequence, making it difficult to accurately capture the dynamic change process of user intentions, thereby affecting the accuracy of user intention recognition and the forward-looking nature of behavior prediction. Summary of the Invention
[0004] Embodiments of the present invention provide an intelligent aggregation and analysis processing method and system for cross-platform user behavior data, which can solve the problems in the existing technologies. [[ID=##]] [[ID=##]]
[0005] In the first aspect of the embodiments of the present invention, an intelligent aggregation and analysis processing method for cross-platform user behavior data is provided, including: Obtaining user historical behavior data from multiple different platforms, where the user historical behavior data includes behavior timestamps, and performing time-series sorting on the user historical behavior data through the behavior timestamps; Performing session segmentation on the time-series sorted user historical behavior data to obtain a behavior session sequence, and constructing a behavior transition probability graph according to the transition relationship between adjacent behaviors in the behavior session sequence; Construct a decentralized federated reinforcement learning framework, train a graph attention network using the behavior transition probability graph to implement an intention recognition agent, train a contrastive learning network based on the behavior session sequence to implement a behavior prediction agent, and deploy the intention recognition agent and the behavior prediction agent locally on each platform. Cross-platform collaborative training is achieved by encrypting and transmitting model gradient information through a differential privacy mechanism; Construct a reward signal generator based on the predicted behavior sequence generated by the behavior prediction agent, calculate the reward value of the predicted behavior sequence based on the behavior session sequence, and use the reward value to drive the model optimization of the decentralized federated reinforcement learning framework; Construct a multi-layer causal relationship graph based on the behavior transition probability graph, use graph neural networks and counterfactual reasoning methods to analyze the rationality of the predicted behavior sequence, generate a correction plan through causal intervention and select the optimal correction result using causal effect evaluation. Generate a user intention portrait based on the intention recognition result output by the intention recognition agent and the optimal correction result.
[0006] Perform session segmentation on the user's historical behavior data after time series sorting to obtain a behavior session sequence, and construct a behavior transition probability graph based on the transition relationship between adjacent behaviors in the behavior session sequence, including: Calculate the average behavior interval time and the standard deviation of the behavior interval time in the user's historical behavior data, and construct an adaptive time threshold function based on the standard deviation; Compare the time interval between adjacent behaviors with the calculation result of the adaptive time threshold function. When the time interval is greater than the calculation result of the adaptive time threshold function, determine the session boundary. Segment the user's historical behavior data based on the session boundary to obtain a behavior session sequence, and construct an intra-session behavior association graph to depict the temporal dependence relationship of behaviors within the same session. Use the temporal dependence relationship to optimize the division of the session boundary; Extract multi-modal features for each behavior in the behavior session sequence. The multi-modal feature extraction includes extracting behavior type features, target object features, and environmental parameter features, and performing weighted fusion on the behavior type features, the target object features, and the environmental parameter features through a feature weight matrix to obtain a behavior vector representation. Based on the behavior vector representation, construct an attention enhancement layer to capture the dynamic correlation between different feature dimensions; Construct a behavior transition probability graph based on the transition relationship between adjacent behaviors in the behavior vector representation, integrate the temporal dependence relationship in the intra-session behavior association graph into the behavior transition probability graph, model long-term behavior dependencies through a recurrent neural network, and use a gating mechanism to control the forgetting and updating of historical information to achieve in-depth modeling of the user's behavior evolution law.
[0007] Construct a decentralized federated reinforcement learning framework, use the behavior transfer probability map to train a graph attention network to implement an intention recognition agent, and train a contrastive learning network based on the behavior session sequence to implement a behavior prediction agent, including: Construct a decentralized federated reinforcement learning framework based on the theory of brain-like synaptic plasticity, model the dynamic synaptic connection strength of knowledge transfer between federated learning nodes, and construct a knowledge transfer path between the federated learning nodes based on the dynamic synaptic connection strength; Map the transfer probability map to a continuous-discrete hybrid state space, perform stability constraints on the dynamic synaptic connection strength based on the continuous-discrete hybrid state space, and construct a Lyapunov stability criterion through the stability constraints; Construct a geometric structure representing the knowledge transfer of the gauge field on the knowledge transfer path, use the Lyapunov stability criterion to constrain the dynamic evolution of the gauge field, and establish a coupling equation between the knowledge field and the gauge field based on Yang-Mills gauge theory to achieve knowledge distillation between the federated learning nodes; Introduce a neurotransmitter regulation mechanism based on the coupling equation between the knowledge field and the gauge field, construct a contrastive learning network through the neurotransmitter regulation mechanism and train the behavior session sequence, and the neurotransmitter regulation mechanism adaptively adjusts the temperature parameter of the contrastive learning using the nonlinear combination of multiple types of neurotransmitter concentrations; Map the multiple types of neurotransmitter concentrations to an Ising model to construct a critical state prediction framework, and the critical state prediction framework calculates the correlation length through the second moment of the spatio-temporal correlation function, and evaluates the spatio-temporal correlation characteristics of the conditional transition probability between behavior states based on the correlation length to achieve a behavior prediction agent.
[0008] Map the transfer probability map to a continuous-discrete hybrid state space, perform stability constraints on the dynamic synaptic connection strength based on the continuous-discrete hybrid state space, and construct a Lyapunov stability criterion including: Map the behavior transfer probability map to a continuous-discrete hybrid state space, the continuous-discrete hybrid state space includes continuous state variables and discrete state variables, and construct a system state representation based on the continuous state variables and the discrete state variables; Construct an entropy production rate evaluation model based on the system state representation, the entropy production rate evaluation model calculates the time derivative of the system entropy through the product sum of the generalized flow and the generalized force, designs a nonlinear fluctuation compensator using the time derivative of the system entropy, and the nonlinear fluctuation compensator calculates the compensation amount based on the entropy-dependent compensation gain matrix and the gradient of the potential function; Apply the compensation amount to the stability constraint of the dynamic synaptic connection strength. The stability constraint controls the influence degree of the entropy gradient term through the dissipation coefficient, and constructs a bifurcation parameter by using the time derivative and spatial derivative of the system entropy to achieve dynamic regulation; Construct a dynamic symbolic encoder to map the evolution trajectory of the dynamic synaptic connection strength into a symbol sequence, and calculate the topological entropy based on the symbol sequence. The topological entropy is obtained through the logarithmic limit of the number of allowed sequences of different lengths in the symbol sequence; Integrate and weight-combine the topological entropy and the time derivative of the system entropy output by the entropy production rate evaluation model to construct a Lyapunov stability criterion, and optimize the dynamic synaptic connection strength by minimizing the Lyapunov stability criterion.
[0009] Construct a reward signal generator according to the predicted behavior sequence generated by the behavior prediction agent, calculate the reward value of the predicted behavior sequence based on the behavior session sequence, and use the reward value to drive the model optimization of the decentralized federated reinforcement learning framework, including: Generate a predicted behavior sequence based on the behavior prediction agent, and construct a reward signal generator for the predicted behavior sequence. The reward signal generator constructs a temporal difference evaluation module by weighting multiple reward components of the system state and the predicted behavior. The temporal difference evaluation module calculates the temporal difference value based on the immediate reward value, the discount factor, and the state value function, and constructs a prediction accuracy evaluator by using the temporal difference value. The prediction accuracy evaluator obtains the accuracy evaluation value by calculating the Gaussian similarity between the predicted behavior feature and the real behavior feature; Calculate the information entropy and conditional entropy based on the system state and the predicted behavior, construct an information gain metric through the difference between the information entropy and the conditional entropy, and weight-combine the information gain metric and the accuracy evaluation value to obtain a comprehensive reward value; Construct a policy gradient update rule by using the comprehensive reward value. The policy gradient update rule obtains the gradient of the model parameters by calculating the expected value of the product of the logarithmic derivative of the parameterized policy function and the comprehensive reward value; Construct an asynchronous update mechanism for the decentralized federated reinforcement learning framework based on the gradient of the model parameters. The asynchronous update mechanism includes a parameter difference term and a local gradient term between nodes. Control the influence degree of the parameter difference term through the synchronization rate, and control the influence degree of the local gradient term through the learning rate to achieve the model optimization of the decentralized federated reinforcement learning framework.
[0010] Construct a multi-layer causal relationship graph based on the behavior transition probability graph, analyze the rationality of the predicted behavior sequence using graph neural networks and counterfactual reasoning methods, generate a correction plan through causal intervention, and select the optimal correction result using causal effect evaluation, including: Construct a multi-layer causal relationship graph based on the behavior transition probability graph, describe the evolution process of the multi-layer causal relationship graph using the Fokker-Planck equation, which includes a drift term and a diffusion term, describe the time evolution of the system state distribution through the drift term and the diffusion term, and construct a state transition matrix based on the characteristics of the time evolution; Construct a counterfactual reasoning module based on the multi-layer causal relationship graph. The counterfactual reasoning module generates a counterfactual trajectory by minimizing the Lagrangian action, and evaluates the rationality of the counterfactual trajectory using the state transition matrix; Conduct a stability analysis of the counterfactual trajectory, calculate the steady-state distribution based on the Fokker-Planck equation, construct a causal intervention generator according to the steady-state distribution, and the causal intervention generator generates a correction plan based on the system state distribution. The correction plan realizes system optimization by minimizing the perturbation of the state transition matrix; Evaluate the system state distribution of the correction plan, analyze the evolution characteristics of the correction plan through the Fokker-Planck equation, calculate the convergence and stability of the system state distribution based on the evolution characteristics, and select the optimal correction plan according to the convergence and the stability. The optimal correction plan is obtained by minimizing the time evolution deviation of the system state distribution.
[0011] Construct a counterfactual reasoning module based on the multi-layer causal relationship graph. The counterfactual reasoning module generates a counterfactual trajectory by minimizing the Lagrangian action, and evaluates the rationality of the counterfactual trajectory using the state transition matrix, including: Construct a counterfactual reasoning module based on the multi-layer causal relationship graph. The counterfactual reasoning module includes a system state vector, node features, and a time variable, and defines a Lagrangian in the counterfactual reasoning module. The Lagrangian establishes a system dynamics model through a combination of a system kinetic energy term, a potential energy term, and a constraint term. The system dynamics model describes the evolution law of the system state vector under the time variable; Construct a minimum action functional based on the system dynamics model. The minimum action functional performs an integral operation on the Lagrangian over a time interval, solves the minimum action functional through the variational principle to obtain the Euler-Lagrange equation, and calculates the evolution trajectory of the system state vector using the Euler-Lagrange equation to obtain a counterfactual trajectory; Calculate the state transition probability based on the counterfactual trajectory. The state transition probability maps the calculation result of the minimum action functional to the probability space through the Boltzmann distribution, and constructs a state transition matrix using the state transition probability to evaluate the rationality of the counterfactual trajectory.
[0012] In the second aspect of the embodiments of the present invention, a cross-platform intelligent aggregation and analysis processing system for user behavior data is provided, including: A first unit for obtaining user historical behavior data of multiple different platforms, where the user historical behavior data includes behavior timestamps, and sorting the user historical behavior data in time series through the behavior timestamps; A second unit for performing session segmentation on the user historical behavior data after time series sorting to obtain a behavior session sequence, and constructing a behavior migration probability graph according to the transition relationship between adjacent behaviors in the behavior session sequence; A third unit for constructing a decentralized federated reinforcement learning framework, using the behavior migration probability graph to train a graph attention network to implement an intention recognition agent, training a contrastive learning network to implement a behavior prediction agent based on the behavior session sequence, and deploying the intention recognition agent and the behavior prediction agent locally on each platform, and realizing cross-platform collaborative training by encrypting and transmitting model gradient information through a differential privacy mechanism; A fourth unit for constructing a reward signal generator according to the predicted behavior sequence generated by the behavior prediction agent, calculating the reward value of the predicted behavior sequence based on the behavior session sequence, and using the reward value to drive the model optimization of the decentralized federated reinforcement learning framework; A fifth unit for constructing a multi-layer causal relationship graph based on the behavior migration probability graph, using a graph neural network and a counterfactual reasoning method to analyze the rationality of the predicted behavior sequence, generating a correction plan through causal intervention and selecting the optimal correction result using causal effect evaluation, and generating a user intention portrait according to the intention recognition result output by the intention recognition agent and the optimal correction result.
[0013] In the third aspect of the embodiments of the present invention, an electronic device is provided, including: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.
[0014] In the fourth aspect of the embodiments of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0015] The beneficial effects of this application are as follows: Through the intelligent aggregation and analysis processing method of multi-platform user behavior data, the present invention realizes the cross-platform integration and in-depth analysis of user behavior data, improves the accuracy of user intention recognition and the precision of behavior prediction. By constructing a decentralized federated reinforcement learning framework, collaborative training of cross-platform data is achieved on the premise of protecting user privacy, effectively solving the data silo problem and reducing the risk of data leakage.
[0016] The present invention combines an intention recognition agent based on a graph attention network and a behavior prediction agent based on contrastive learning, which can comprehensively capture user behavior patterns and intention transformation rules, systematically analyze user behavior characteristics on different platforms, so as to provide more personalized and accurate service recommendations for users. The construction of the behavior migration probability graph and the multi-layer causal relationship analysis method effectively improve the interpretability of the prediction results.
[0017] The reward signal generation mechanism and the prediction correction scheme based on counterfactual reasoning of the present invention significantly improve the model self-optimization ability and adaptability, and can dynamically adjust to adapt to changes in user behavior patterns. The introduction of the multi-layer causal relationship graph enables the system to deeply understand the causal logic behind user behavior, further enhancing the accuracy and practical value of the user intention portrait, and providing strong support for cross-platform personalized services and precision marketing. Brief Description of the Drawings
[0018] Figure 1 It is a schematic flow chart of the intelligent aggregation and analysis processing method of cross-platform user behavior data in the embodiment of the present invention; Figure 2 It is a bar chart comparing the performance of the user behavior time series analysis model in the embodiment of the present invention; Figure 3 It is an attention graph of the influence of the correlation length on the behavior prediction performance in the Ising model critical state prediction framework in the embodiment of the present invention; Figure 4 It is a flow chart of the reward signal generation and the model optimization of the decentralized federated reinforcement learning framework in the embodiment of the present invention; Figure 5 It is a bar chart comparing the performance of the multi-layer causal relationship graph and the counterfactual reasoning method in the embodiment of the present invention. Detailed Embodiments
[0019] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0020] The technical solution of the present invention will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0021] Figure 1 It is a schematic flowchart of the cross-platform user behavior data intelligent aggregation and analysis processing method according to an embodiment of the present invention, as Figure 1 shown, the method includes: Obtain user historical behavior data of multiple different platforms, where the user historical behavior data includes behavior timestamps, and perform time series sorting on the user historical behavior data through the behavior timestamps; Perform session segmentation on the user historical behavior data after time series sorting to obtain a behavior session sequence, and construct a behavior migration probability graph according to the transition relationship between adjacent behaviors in the behavior session sequence; Construct a decentralized federated reinforcement learning framework, use the behavior migration probability graph to train a graph attention network to implement an intent recognition agent, train a contrastive learning network based on the behavior session sequence to implement a behavior prediction agent, and deploy the intent recognition agent and the behavior prediction agent locally on each platform, and realize cross-platform collaborative training by encrypting and transmitting model gradient information through a differential privacy mechanism; Construct a reward signal generator according to the predicted behavior sequence generated by the behavior prediction agent, calculate the reward value of the predicted behavior sequence based on the behavior session sequence, and use the reward value to drive the model optimization of the decentralized federated reinforcement learning framework; Construct a multi-layer causal relationship graph based on the behavior migration probability graph, analyze the rationality of the predicted behavior sequence by using a graph neural network and a counterfactual reasoning method, generate a correction plan through causal intervention and select the optimal correction result by using causal effect evaluation, and generate a user intent portrait according to the intent recognition result output by the intent recognition agent and the optimal correction result.
[0022] In an alternative embodiment, performing session segmentation on the user historical behavior data after time series sorting to obtain a behavior session sequence, and constructing a behavior migration probability graph according to the transition relationship between adjacent behaviors in the behavior session sequence includes: Calculate the average behavior interval time and the standard deviation of the behavior interval time in the user historical behavior data, and construct an adaptive time threshold function based on the standard deviation; Compare the time interval between adjacent behaviors with the calculation result of the adaptive time threshold function. When the time interval is greater than the calculation result of the adaptive time threshold function, determine the session boundary. Based on the session boundary, segment the user historical behavior data to obtain a sequence of behavior sessions, and construct an intra-session behavior association graph to depict the temporal dependence relationship of behaviors within the same session. Use the temporal dependence relationship to optimize the division of session boundaries; Extract multi-modal features for each behavior in the sequence of behavior sessions. The multi-modal feature extraction includes extracting behavior type features, target object features, and environmental parameter features, and weighted fusion of the behavior type features, the target object features, and the environmental parameter features through a feature weight matrix to obtain a behavior vector representation. Based on the behavior vector representation, construct an attention enhancement layer to capture the dynamic correlation between different feature dimensions; Construct a behavior transition probability graph based on the transition relationship between adjacent behaviors in the behavior vector representation, integrate the temporal dependence relationship in the intra-session behavior association graph into the behavior transition probability graph, model long-term behavior dependence through a recurrent neural network, and use a gating mechanism to control the forgetting and updating of historical information to achieve in-depth modeling of the evolution law of user behaviors.
[0023] Calculate the average behavior interval time and the standard deviation of the behavior interval time in the user historical behavior data. For example, assume that a user's behavior sequence in a day is: browse product A (08:00), view product details (08:02), add to cart (08:05), browse product B (12:30), view reviews (12:32), place an order (12:40), browse product C (20:00). By calculating the time interval between adjacent behaviors, the interval sequence is obtained as: 2 minutes, 3 minutes, 265 minutes, 2 minutes, 8 minutes, 440 minutes. The calculated average behavior interval time is 120 minutes, and the standard deviation is 175 minutes.
[0024] Construct an adaptive time threshold function based on the standard deviation. The threshold function can be set as the average behavior interval time plus a multiple of the standard deviation. In this example, if the multiple is set to 1.5, the adaptive time threshold is 120 + 175×1.5 = 382.5 minutes. According to this threshold, the system compares the time interval between adjacent behaviors with the threshold. If the time interval is greater than the threshold, it is determined as the session boundary. In this example, 265 minutes is less than the threshold of 382.5 minutes, while 440 minutes is greater than the threshold. Therefore, no session boundary is divided between browsing product B and placing an order, and a session boundary is divided between placing an order and browsing product C. Finally, two behavior sessions are obtained: Session 1 (browse product A, view product details, add to cart, browse product B, view reviews, place an order) and Session 2 (browse product C).
[0025] Based on session segmentation, the system constructs an in - session behavior association graph to depict the temporal dependence relationship of behaviors within the same session. Taking Session 1 as an example, by analyzing the temporal patterns in the behavior sequence, it can be found that the two subsequences of "browsing Product A → viewing product details → adding to cart" and "browsing Product B → viewing reviews → placing an order" show strong temporal dependence relationships. The system records these temporal dependence relationships in the in - session behavior association graph for subsequent optimization of session boundary division.
[0026] For each behavior in the behavior session sequence, the system performs multi - modal feature extraction, including behavior type features, target object features, and environmental parameter features. Taking "browsing Product A" as an example, the behavior type feature can be the vector representation of "browsing", the target object feature can be the attribute vector of "Product A" (such as category, price range, brand, etc.), and the environmental parameter features can include the time when the behavior occurs (8 am) and the type of device the user is using (such as mobile or PC).
[0027] These features are weighted and fused through a feature weight matrix to obtain a behavior vector representation. Suppose for the behavior of "browsing Product A", the extracted behavior type feature vector is [0.8, 0.2, 0.1], the target object feature vector is [0.7, 0.4, 0.6], the environmental parameter feature vector is [0.3, 0.9, 0.2], and the feature weight matrix is [0.4, 0.4, 0.2] respectively. Then the behavior vector representation after weighted fusion is [0.8×0.4 + 0.7×0.4 + 0.3×0.2, 0.2×0.4 + 0.4×0.4 + 0.9×0.2, 0.1×0.4 + 0.6×0.4 + 0.2×0.2] = [0.66, 0.38, 0.32].
[0028] Based on the behavior vector representation, the system constructs an attention - enhanced layer to capture the dynamic correlation between different feature dimensions. By calculating the importance scores of each dimension in the feature vector and re - weighting the features accordingly, the system can adaptively focus on the most relevant feature dimensions. For example, for the behavior of "adding to cart", the system pays more attention to the price of the product and the user's historical purchase behavior, while for the "browsing" behavior, the system pays more attention to the category and display position of the product.
[0029] Based on the transition relationship between adjacent behaviors in the behavior vector representation, the system constructs a behavior transition probability graph. The nodes in this graph represent different behaviors, and the edges represent the transition probabilities between behaviors. For example, from historical data, it can be found that after "viewing product details", there is a 70% probability that the user will "add to cart", a 20% probability that the user will "view reviews", and a 10% probability that the user will "return to browse other products". These transition probabilities are recorded in the behavior transition probability graph.
[0030] The system integrates the temporal dependencies in the in - session behavior association graph into the behavior transition probability graph and models long - term behavior dependencies through a recurrent neural network. For example, the system can identify a complete purchase path such as "browse → view details → add to cart → place an order", and controls the forgetting and updating of historical information through a gating mechanism. Specifically, when detecting that the user is performing a sequence of behaviors with a clear intention (such as the purchase path), the system retains more historical information related to this intention; while when the user's behavior pattern changes significantly (such as switching from shopping to watching videos), the system updates the current state representation and reduces the dependence on old behavior information.
[0031] Through the above - mentioned method, the system can deeply model the evolution law of user behavior, providing a basis for subsequent user behavior prediction and personalized recommendation.
[0032] Figure 2 This is the bar chart for the performance comparison of the user behavior temporal analysis model in the embodiments of the present invention: This figure shows the comparison results of three different models on the key performance indicators of the dialogue system. In terms of the accuracy of session boundary recognition, the basic temporal model reaches 58.6%, the adaptive threshold model is improved to 72.4%, and the multi - modal feature fusion model is further increased to 83.7%, showing a gradually increasing trend. In terms of the capture rate of behavior dependency relationship, the basic temporal model is 50.2%, the adaptive threshold model is significantly improved to 73.8%, and the multi - modal feature fusion model reaches 93.2%, reflecting the superiority of the fusion model. In the evaluation of user intention prediction accuracy, the basic temporal model achieves an accuracy of 60.5%, the adaptive threshold model is improved to 80.3%, and the multi - modal feature fusion model finally reaches a high accuracy of 94.1%. Through the comparison of the three key indicators, it can be seen that the multi - modal feature fusion model significantly outperforms the other two models in terms of various performance manifestations, demonstrating the comprehensive advantages of this model in the dialogue system.
[0033] In an optional implementation, a decentralized federated reinforcement learning framework is constructed. Using the behavior transition probability graph to train a graph attention network to implement an intention recognition agent, and training a contrastive learning network based on the behavior session sequence to implement a behavior prediction agent includes: Construct a decentralized federated reinforcement learning framework based on the theory of brain - like synaptic plasticity, model the dynamic synaptic connection strength of knowledge transfer between federated learning nodes, and construct a knowledge transfer path between the federated learning nodes based on the dynamic synaptic connection strength; Map the transition probability graph to a continuous - discrete hybrid state space, perform stability constraints on the dynamic synaptic connection strength based on the continuous - discrete hybrid state space, and construct a Lyapunov stability criterion through the stability constraints. Construct a geometric structure that represents the knowledge transfer of the gauge field on the knowledge transfer path, use the Lyapunov stability criterion to constrain the dynamic evolution of the gauge field, and establish a coupling equation between the knowledge field and the gauge field based on Yang-Mills gauge theory to achieve knowledge distillation between the federated learning nodes; Introduce a neurotransmitter regulation mechanism based on the coupling equation between the knowledge field and the gauge field, construct a contrastive learning network through the neurotransmitter regulation mechanism and train the behavior session sequence. The neurotransmitter regulation mechanism adaptively adjusts the temperature parameter of contrastive learning using the non-linear combination of multiple types of neurotransmitter concentrations; Map the concentrations of multiple types of neurotransmitters to the Ising model to construct a critical state prediction framework. The critical state prediction framework calculates the correlation length through the second moment of the spatio-temporal correlation function, and evaluates the spatio-temporal correlation characteristics of the conditional transition probability between behavior states based on the correlation length to achieve a behavior prediction agent.
[0034] Model the dynamic synaptic connection strength of knowledge transfer between federated learning nodes. The connection strength value ranges from 0 to 1, representing the efficiency of knowledge transfer between nodes. Specifically, use the modified Hebbian learning rule to calculate the synaptic connection strength. When the knowledge representations of two nodes are highly similar, the synaptic connection strength between them is enhanced; otherwise, it is weakened. Based on the calculated dynamic synaptic connection strength matrix, construct a knowledge transfer path between federated learning nodes. This path can be represented as a directed weighted graph, where the nodes represent federated learning participants and the edge weights are the corresponding synaptic connection strength values.
[0035] After constructing the knowledge transfer path, map the behavior transition probability graph to a continuous-discrete hybrid state space. This hybrid state space contains continuous variables (such as user movement trajectories, operation durations, etc.) and discrete variables (such as click behaviors, page jumps, etc.). Based on this hybrid state space, perform stability constraints on the dynamic synaptic connection strength. The stability constraint condition is defined as: when the state in the hybrid state space deviates from the equilibrium point, the change rate of the synaptic connection strength should prompt the system to return to the equilibrium point. In actual implementation, set the threshold parameter α = 0.05, and trigger the adjustment mechanism when the connection strength change rate exceeds this threshold. Construct the Lyapunov stability criterion through this stability constraint to ensure the stability of knowledge transfer during the federated learning process.
[0036] In the constructed knowledge transfer pathway, a gauge field is used to characterize the geometric structure of knowledge transfer. The gauge field consists of two parts: the gauge potential and the field strength. The gauge potential describes the direction and strength of knowledge transfer between nodes, and the field strength describes the spatial distribution characteristics of knowledge transfer. The dynamic evolution of the gauge field is constrained by the aforementioned Lyapunov stability criterion. When the system deviates from the stable state, the evolution direction of the gauge field should reduce the system energy. Based on the Yang-Mills gauge theory, a coupling equation between the knowledge field and the gauge field is established. This coupling equation describes the propagation law of the knowledge field under the action of the gauge field, thus realizing knowledge distillation between federated learning nodes. In practical applications, the knowledge distillation process is achieved through soft label transfer, with the temperature parameter set to 2.5 and the distillation weight coefficient to 0.4.
[0037] Based on the coupling equation between the knowledge field and the gauge field, a neurotransmitter regulation mechanism is introduced. This mechanism mimics the way in which multiple neurotransmitters in the human brain cooperate to regulate the cognitive process. Four main neurotransmitters are considered in the implementation: dopamine, serotonin, acetylcholine, and norepinephrine. The concentration range of each neurotransmitter is between 0 and 1. A contrastive learning network is constructed through the neurotransmitter regulation mechanism and trained on the behavioral session sequence. This network adopts a two-tower structure, with the left tower encoding the current behavior and the right tower encoding the target behavior. The neurotransmitter regulation mechanism adaptively adjusts the temperature parameter of contrastive learning using the non-linear combination of the concentrations of multiple neurotransmitters. The adjustment formula is based on the weighted sum of the concentrations of various neurotransmitters. The initial temperature is set to 0.07, and the maximum value is limited to 0.2. Experimental data show that this adaptive temperature parameter improves the performance of contrastive learning by approximately 12.3% compared to the fixed temperature parameter.
[0038] The concentrations of multiple neurotransmitters are mapped to the Ising model to construct a critical state prediction framework. In this mapping process, the neurotransmitter concentration values are converted into the spin states and interaction strengths in the Ising model. The dopamine concentration affects the probability of spin-up, and serotonin affects the interaction strength between spins. This critical state prediction framework obtains the correlation length by calculating the second moment of the spatio-temporal correlation function. The correlation length is defined as the distance corresponding to the decay of the correlation function to 1 / e of the initial value. Based on the calculated correlation length, the spatio-temporal correlation characteristics of the conditional transition probability between behavioral states are evaluated. When the correlation length exceeds the preset threshold of 3.5, the system is considered to be in a critical state, and the accuracy of behavior prediction is significantly improved at this time. Experimental verification shows that near the critical state, the accuracy of behavior prediction is improved by 18.7% compared to the non-critical state.
[0039] In the actual application of this behavior prediction agent, taking the user behavior sequence on an e-commerce platform as an example, the input is a session sequence containing behaviors such as browsing, adding to cart, favoriting, and placing an order, with a length ranging from 20 to 50. The behavior prediction agent first identifies the user's intention (such as price comparison, impulse consumption, planned purchase, etc.) through a graph attention network, and then predicts the next behavior based on a contrastive learning network. On a test set containing 1 million user sessions, the behavior prediction accuracy of this method reaches 78.3%, which is 9.2 percentage points higher than that of traditional sequence prediction models. Moreover, in a federated learning environment with 50 nodes, the communication cost is reduced by 43.5%.
[0040] Figure 3 Attention graph for the influence of the correlation length on the behavior prediction performance in the Ising model critical state prediction framework of the embodiments of the present invention: This figure shows the comparison of the prediction accuracies of three different prediction methods as the correlation length changes. The performance curve of this technical solution shows a trend of first rising and then falling. Starting from 70.5% at ξ = 1.0, passing through 76.6% at 2.0 and 88.3% at ξ = 3.0, it reaches the optimal performance of 98.5% at the critical point ξ = 3.5. After that, the performance gradually decreases and drops to 73.3% at 6.0. The performance curve of the standard Ising model also shows a similar trend, but the overall performance is lower than that of this technical solution, rising from the initial 66.7% to the peak of 85.0% at ξ = 3.5, and then dropping to the final 66.7%. The prediction method based on the Markov chain is the most stable, with a small fluctuation range, and its performance always fluctuates between 63.3% and 67.5%, but the overall performance is significantly lower than the other two methods. The comparison of the three methods clearly shows that this technical solution has a significant advantage at the critical point, and its prediction accuracy of 98.5% far exceeds the other two comparison methods.
[0041] In an optional implementation manner, mapping the migration probability graph to a continuous-discrete hybrid state space, performing stability constraints on the dynamic synaptic connection strength based on the continuous-discrete hybrid state space, and constructing a Lyapunov stability criterion through the stability constraints includes: Mapping the behavior migration probability graph to a continuous-discrete hybrid state space, where the continuous-discrete hybrid state space includes continuous state variables and discrete state variables, and constructing a system state representation based on the continuous state variables and the discrete state variables; Constructing an entropy production rate evaluation model based on the system state representation. The entropy production rate evaluation model calculates the time derivative of the system entropy through the product sum of the generalized flux and the generalized force, designs a non-linear fluctuation compensator using the time derivative of the system entropy, and the non-linear fluctuation compensator calculates the compensation amount based on the entropy-dependent compensation gain matrix and the gradient of the potential function; Apply the compensation amount to the stability constraint of the dynamic synaptic connection strength. The stability constraint controls the influence degree of the entropy gradient term through the dissipation coefficient, and constructs a bifurcation parameter using the temporal derivative and spatial derivative of the system entropy to achieve dynamic regulation; Construct a dynamic symbolic encoder to map the evolution trajectory of the dynamic synaptic connection strength into a symbol sequence, and calculate the topological entropy based on the symbol sequence. The topological entropy is obtained through the logarithmic limit of the number of allowed sequences of different lengths in the symbol sequence; Perform an integration operation and weighted combination on the topological entropy and the temporal derivative of the system entropy output by the entropy production rate evaluation model, construct a Lyapunov stability criterion, and optimize the dynamic synaptic connection strength by minimizing the Lyapunov stability criterion.
[0042] This embodiment provides a method for mapping a behavior transition probability map to a continuous-discrete hybrid state space and constructing a Lyapunov stability criterion. This method first obtains a behavior transition probability map, which represents the connection relationships and their probability intensities between different neuron nodes. In a specific example, a network containing 10 neuron nodes can be considered, where the transition probabilities between nodes are represented by a 10×10 matrix, and the matrix element values are between 0 and 1, representing the probability of transferring from one node to another.
[0043] When mapping the above behavior transition probability map to a continuous-discrete hybrid state space, continuous state variables and discrete state variables are introduced. The continuous state variable can be represented as an n-dimensional vector x, where each element represents the membrane potential or activation level of a neuron, and the value range is usually in the interval [-1,1]; the discrete state variable can be represented as an m-dimensional vector q, where each element represents the activation state of a neuron (such as 0 for resting and 1 for activated), and the value is 0 or 1. In the above example, n = 10 can be set to represent the continuous membrane potential states of 10 neurons, and m = 10 can be set to represent the discrete activation states of these neurons. Based on these two sets of variables, construct a system state representation S=(x,q) as the basis for subsequent stability analysis.
[0044] Based on the above system state representation, construct an entropy production rate evaluation model. This model estimates the temporal derivative of the system entropy by calculating the product sum of the generalized flux and the generalized force. In practical applications, the generalized flux can be represented as a vector J of the change rates of state variables, and the generalized force can be represented as a vector F of the potential energy gradient inside the system. The temporal derivative of the system entropy is calculated as the inner product sum of the two, and can be expressed as d S / d t=∑JᵢFᵢ. For example, when n = 10, there can be 10 product terms of generalized fluxes and generalized forces. In a specific case, if at a certain moment, J = [0.02, 0.05, -0.03, 0.01, 0.04, -0.02, 0.03, -0.01, 0.02, 0.01] and F = [0.1, 0.2, -0.1, 0.3, 0.2, -0.1, 0.1, -0.2, 0.1, 0.3] are measured, then the calculated entropy change rate is 0.041.
[0045] Design a non - linear fluctuation compensator using the time derivative of the system entropy. This compensator calculates the compensation amount u = K (S) ·∇V (S) through the entropy - dependent compensation gain matrix K (S) and the gradient of the potential function ∇V (S) . In actual implementation, K (S) can be designed as an n×n diagonal matrix, and its diagonal element values are dynamically adjusted according to the current system entropy state; while the potential function V (S) can be designed as a quadratic function of the state variables. In the above case, if K (S) is selected as a diagonal matrix with all diagonal elements being 0.05 and the gradient of the potential function ∇V (S) = [0.1, 0.15, -0.12, 0.08, 0.14, -0.09, 0.11, -0.07, 0.09, 0.13], then the calculated compensation amount is u = [0.005, 0.0075, -0.006, 0.004, 0.007, -0.0045, 0.0055, -0.0035, 0.0045, 0.0065].
[0046] Apply the above - mentioned compensation amount to the stability constraint of the dynamic synaptic connection strength. This constraint controls the influence degree of the entropy gradient term by introducing a dissipation coefficient α, and the stability constraint is expressed as d w / d t =f (w) -α·u, where w represents the synaptic connection strength and f (w) represents the original dynamics. At the same time, construct a bifurcation parameter β = g(d S / d t , ∇ S ) using the time derivative and spatial derivative of the system entropy to dynamically adjust the strength of the stability constraint. In specific implementation, the initial value of α can be set to 0.1 and adjusted dynamically according to the system entropy change; the calculation of the bifurcation parameter β can adopt a weighted combination of d S / d t and ∇ S , such as β = 0.7·d S / d t +0.3·|∇ S|. Based on the above case, if it is calculated that β = 0.048, the updated α value can be adjusted to 0.1×(1 + 0.048) = 0.1048.
[0047] Construct a dynamic symbolic encoder to map the evolution trajectory of the dynamic synaptic connection strength into a symbol sequence. In actual implementation, the continuous synaptic strength state space can be divided into a finite number of regions, and each region corresponds to a symbol. For example, the synaptic strength range [-1, 1] is evenly divided into 4 regions, corresponding to the symbols A, B, C, and D respectively. By recording the transition sequence of the system state between these regions, a symbol sequence is generated. Taking a specific evolution trajectory as an example, if the system states fall in the regions B, B, C, A, D, C, B, A, A, D in sequence within 10 time steps, the generated symbol sequence is "BBCADCBAAD".
[0048] Calculate the topological entropy H based on the above symbol sequence. This entropy is obtained through the logarithmic limit of the number of allowed sequences of different lengths in the symbol sequence, and the calculation formula is H = lim(n→∞)log(N (n) ) / n, where N (n) represents the number of all subsequences of length n. In actual calculation, a finite length n (such as n = 4) can be selected for approximation. For the above symbol sequence, there are 4 different subsequences of length 1, namely {A, B, C, D}, and 8 different subsequences of length 2, namely {BB, BC, CA, AD, DC, CB, BA, AA}. The approximate value of the calculated topological entropy is log(8) / 2 ≈ 1.04.
[0049] Integrate and weight - combine the topological entropy with the time - derivative of the system entropy output by the entropy production rate evaluation model to construct the Lyapunov stability criterion L = γ·H+(1 - γ)·∫(d S / d t )d t , where γ is the weight coefficient. Optimize the dynamic synaptic connection strength by minimizing this stability criterion. In a specific application, γ = 0.4 can be set. If the calculated topological entropy H = 1.04 and the integral value of the time - derivative of the system entropy is 0.85, then the value of the Lyapunov stability criterion is L = 0.4×1.04+(1 - 0.4)×0.85 = 0.926. When this value is less than a preset threshold (such as 0.95), it is considered that the system reaches a stable state; otherwise, the synaptic connection strength needs to be continuously adjusted until the stability condition is met.
[0050] In an alternative implementation, constructing a reward signal generator based on the predicted behavior sequence generated by the behavior prediction agent, calculating the reward value of the predicted behavior sequence based on the behavior session sequence, and using the reward value to drive the model optimization of the decentralized federated reinforcement learning framework includes: Generate a predicted behavior sequence based on the behavior prediction agent, construct a reward signal generator for the predicted behavior sequence, the reward signal generator constructs a temporal difference evaluation module by weighting multiple reward components of the system state and the predicted behavior, the temporal difference evaluation module calculates a temporal difference value based on the immediate reward value, the discount factor, and the state value function, and constructs a prediction accuracy evaluator using the temporal difference value, the prediction accuracy evaluator obtains an accuracy evaluation value by calculating the Gaussian similarity between the predicted behavior features and the true behavior features; Calculate the information entropy and the conditional entropy based on the system state and the predicted behavior, construct an information gain metric by the difference between the information entropy and the conditional entropy, and perform a weighted combination of the information gain metric and the accuracy evaluation value to obtain a comprehensive reward value; Construct a policy gradient update rule using the comprehensive reward value, the policy gradient update rule obtains the gradient of the model parameters by calculating the expected value of the product of the logarithmic derivative of the parameterized policy function and the comprehensive reward value; Construct an asynchronous update mechanism for the decentralized federated reinforcement learning framework based on the gradient of the model parameters, the asynchronous update mechanism includes a parameter difference term and a local gradient term between nodes, controls the influence degree of the parameter difference term through a synchronization rate, and controls the influence degree of the local gradient term through a learning rate, to achieve model optimization of the decentralized federated reinforcement learning framework.
[0051] As Figure 4 shown, the method further includes: In the decentralized federated reinforcement learning framework, effectively optimize the model by constructing a reward signal generator according to the predicted behavior sequence generated by the behavior prediction agent.
[0052] In practical applications, the behavior prediction agent first generates a predicted behavior sequence based on the user's historical behavior data. For example, for an e-commerce platform, it can predict behaviors such as browsing, searching, adding to cart, and purchasing by the user within the next 24 hours. Taking user A as an example, his historical behaviors include "browsing mobile phones - searching for performance parameters - viewing reviews - adding to cart", and the predicted behavior sequence generated by the prediction agent is "viewing prices - comparing different models - purchasing".
[0053] Construct a reward signal generator for the above-mentioned predicted behavior sequence, which weights multiple reward components of the system state and the predicted behavior. The system state includes information such as the current scenario, time period, and device type of the user, and the predicted behavior includes features such as behavior type, target object, and duration. Taking the example of user A continuing the operation, the system state can be described as "weekend evening - mobile device - home scenario", and the predicted behavior is "comparing different models". The reward components can include behavior relevance (weight 0.3), time accuracy (weight 0.2), and behavior coherence (weight 0.5).
[0054] The temporal difference evaluation module is constructed based on the above weighted rewards. Calculating the temporal difference value by this module involves three key factors: the immediate reward value, the discount factor, and the state value function. In actual implementation, the immediate reward value can be set as the matching degree between the predicted behavior and the actual observed behavior. For user A, if the prediction "comparing different models" is consistent with the actual behavior, an immediate reward of 9 points (out of 10) can be given. The discount factor is set to 0.85, indicating the degree of emphasis on future rewards. The state value function estimates the long-term value of a specific system state. For example, the state value of "weekend evening - mobile device - home scenario" is 7.5.
[0055] Through the above parameters, the temporal difference value is calculated to be 3.2. This value is used to construct a prediction accuracy estimator, which calculates the Gaussian similarity between the predicted behavior features and the true behavior features. For the "comparing different models" behavior of user A, the extracted predicted behavior feature vector can be represented as [0.8, 0.6, 0.9], and the true behavior feature vector is [0.75, 0.65, 0.85]. The calculated Gaussian similarity is 0.92, indicating a high degree of matching between the predicted behavior and the true behavior.
[0056] To further improve the effectiveness of the reward signal, the system calculates the information entropy and conditional entropy based on the system state and the predicted behavior. For the case of user A, the information entropy of the system state is calculated to be 1.8, indicating the uncertainty of the system state; the conditional entropy is calculated to be 0.7, indicating the uncertainty of the system state given the predicted behavior. The information gain metric is the difference between the information entropy and the conditional entropy, that is, 1.1, indicating the amount of information provided by the predicted behavior. The information gain metric and the accuracy evaluation value are weighted and combined with weights of 0.4 and 0.6 respectively, and the comprehensive reward value is 0.98.
[0057] Using this comprehensive reward value to construct a policy gradient update rule, which calculates the expected value of the product of the logarithmic derivative of the parameterized policy function and the comprehensive reward value, thereby obtaining the gradient of the model parameters. In actual implementation, the policy function can be represented by a neural network. For the example of User A, the probability that the policy function outputs the behavior of "comparing different models" is 0.78, the logarithmic derivative is calculated to be 0.12, and after multiplying by the comprehensive reward value of 0.98, 0.1176 is obtained as a sample of the expected value. Through the accumulation of multiple samples, the gradient value vector [0.08, 0.12, -0.05, 0.09] is finally obtained.
[0058] Based on the gradient of the above model parameters, an asynchronous update mechanism for the decentralized federated reinforcement learning framework is constructed. This mechanism consists of two key components: the parameter difference term between nodes and the local gradient term. In a federated learning network composed of 10 nodes, taking Node 2 as an example, its current model parameters are [1.2, 0.8, 1.5, 0.6], and the parameters of neighboring Node 1 and Node 3 are [1.3, 0.75, 1.55, 0.58] and [1.15, 0.85, 1.48, 0.63] respectively. The parameter difference term is calculated to be [-0.025, 0.05, -0.035, 0.025].
[0059] The synchronization rate is set to 0.3 to control the influence degree of the parameter difference term, and the adjusted parameter difference contribution is [-0.0075, 0.015, -0.0105, 0.0075]. The local gradient term is the previously calculated gradient value vector [0.08, 0.12, -0.05, 0.09], the learning rate is set to 0.05, and the adjusted local gradient contribution is [0.004, 0.006, -0.0025, 0.0045]. Combining these two parts gives the parameter update amount [-0.0035, 0.021, -0.013, 0.012], and the finally updated model parameters are [1.1965, 0.821, 1.487, 0.612].
[0060] Through the above asynchronous update mechanism, the decentralized federated reinforcement learning framework can effectively integrate the model updates of each node, realize the optimization of the global model while maintaining data privacy, and improve the accuracy of prediction behavior and system performance.
[0061] In an alternative implementation manner, a multi-layer causal relationship graph is constructed based on the behavior migration probability graph, the rationality of the predicted behavior sequence is analyzed by using a graph neural network and a counterfactual reasoning method, a correction scheme is generated through causal intervention, and the optimal correction result is selected by using causal effect evaluation, including: Construct a multi-layer causal relationship graph based on the behavior transition probability graph, use the Fokker-Planck equation to describe the evolution process of the multi-layer causal relationship graph. The Fokker-Planck equation includes a drift term and a diffusion term, and describes the time evolution of the system state distribution through the drift term and the diffusion term, and construct a state transition matrix based on the characteristics of the time evolution; Construct a counterfactual reasoning module based on the multi-layer causal relationship graph. The counterfactual reasoning module generates counterfactual trajectories by minimizing the Lagrangian action, and uses the state transition matrix to evaluate the rationality of the counterfactual trajectories; Conduct a stability analysis on the counterfactual trajectories, calculate the steady-state distribution based on the Fokker-Planck equation, construct a causal intervention generator according to the steady-state distribution. The causal intervention generator generates a correction scheme based on the system state distribution, and the correction scheme realizes system optimization by minimizing the perturbation of the state transition matrix; Evaluate the system state distribution of the correction scheme, analyze the evolution characteristics of the correction scheme through the Fokker-Planck equation, calculate the convergence and stability of the system state distribution based on the evolution characteristics, and select the optimal correction scheme according to the convergence and the stability. The optimal correction scheme is obtained by minimizing the time evolution deviation of the system state distribution.
[0062] Construct a behavior transition probability graph from the collected behavior data, and then construct a multi-layer causal relationship graph. Then use graph neural networks and counterfactual reasoning to analyze and predict the rationality of the behavior sequence, and generate a correction scheme through causal intervention.
[0063] In the stage of constructing the multi-layer causal relationship graph, the system extracts the causal dependence relationship between nodes based on the behavior transition probability graph. In this process, the Fokker-Planck equation is used to describe the evolution process of the multi-layer causal relationship graph. The drift term reflects the deterministic change trend of the system state, and the diffusion term characterizes the influence intensity of random fluctuations. For example, for the user shopping behavior sequence, the drift term represents the deterministic transfer trend of the user from browsing products to adding products to the shopping cart, while the diffusion term reflects the random behavior changes of the user affected by promotional activities. By discretizing the time and state space, the state transition matrix can be calculated, and each element represents the transition probability from state i to state j. In practical applications, if the probability of a user transferring from the product browsing state A to the adding-to-cart state B is 0.35, then the corresponding element value in the state transition matrix is 0.35.
[0064] During the construction phase of the counterfactual reasoning module, the system generates counterfactual scenarios based on a multi-layer causal relationship graph. To minimize the Lagrangian action, the system searches for the optimal path from the initial state to the target state, minimizing the overall energy consumption of the system. For example, when it detects that the user abnormally skips the product comparison step and directly makes a purchase, the system compares the value of the Lagrangian action in the normal shopping path (e.g., 23.7 units) with that of the current abnormal path (e.g., 42.1 units), determining that the current path does not conform to the user's habits. Subsequently, the system uses the previously constructed state transition matrix to evaluate the rationality of the counterfactual trajectory. If the state transition probability in the trajectory is lower than a threshold (e.g., 0.15), the trajectory is considered irrational.
[0065] During the analysis of the stability of the counterfactual trajectory, the system calculates the steady-state distribution of the system based on the Fokker-Planck equation. Through numerical iteration methods, when the change in the system state distribution is less than a preset threshold (e.g., 0.001), the system is considered to have reached a steady state. For example, after 500 iterations, the change in the system state distribution drops to 0.0008. The obtained steady-state distribution shows the long-term residence probabilities of the user in various behavioral states. Based on this steady-state distribution, a causal intervention generator is constructed. By making small perturbations to the state transition matrix (with the perturbation intensity controlled within 0.05), multiple correction schemes are generated. In an application scenario, if it is found that the user abnormally skips the product evaluation step, the system generates three correction schemes: increasing the probability of product comparison recommendations (perturbation intensity 0.03), providing historical evaluation information (perturbation intensity 0.04), or showing purchase records of similar products (perturbation intensity 0.02).
[0066] During the evaluation phase of the correction schemes, the system once again uses the Fokker-Planck equation to analyze the evolution characteristics of the system state distribution under each correction scheme. By calculating the convergence speed and stability index of the system state distribution from the initial state to the steady state, the effects of each correction scheme are evaluated. The convergence speed is determined by calculating the Euclidean distance between consecutive iterations of the state distribution. When this distance is less than a threshold (e.g., 0.005) and remains stable, the system is determined to have reached a stable state. In an actual case, the first correction scheme requires 125 iterations to reach stability, the second requires 92 iterations, and the third requires 108 iterations; the stability indices are 0.87, 0.93, and 0.81 (out of a full score of 1.0), respectively. Finally, the system selects the second scheme as the optimal correction scheme because it has a faster convergence speed and a higher stability index, and this scheme results in the smallest time evolution deviation of the system state distribution (deviation value 0.023, lower than 0.038 and 0.042 of the other schemes).
[0067] Through this method for correcting behavior sequences based on multi-layer causal relationship graphs and counterfactual reasoning, the system can effectively identify abnormal behavior sequences, generate optimal correction schemes, ensure the rationality and reliability of behavior prediction results, and improve user experience and system performance. Practical applications show that this method can increase the behavior prediction accuracy rate from 85.7% of the benchmark method to 93.2%, while reducing the user behavior anomaly rate by 41.3%.
[0068] Figure 5 This is a bar chart comparing the performance of the multi-layer causal relationship graph and counterfactual reasoning method in the embodiments of the present invention: This figure compares the performance of three counterfactual reasoning methods. Among them, the traditional causal analysis method represents the causal inference algorithm based on Bayesian networks, the standard counterfactual reasoning represents the counterfactual reasoning framework based on Pearl's do-operator, and the technical solution of the present invention represents the counterfactual reasoning method based on Lagrangian dynamics and non-equilibrium statistical mechanics. In terms of the convergence speed of the steady-state distribution, the traditional method based on Bayesian networks reaches 54.5%, the standard method based on the do-operator is improved to 66.4%, while the technical solution of the present invention based on Lagrangian dynamics is significantly improved to 90.7%. In the evaluation of the rationality of counterfactual trajectories, the Bayesian network method is 56.3%, the do-operator method is improved to 76.2%, and the Lagrangian dynamics method reaches a high accuracy rate of 94.5%. In the evaluation of the optimality of the correction scheme, the Bayesian network method only reaches 50.9%, the do-operator method is improved to 72.1%, and the Lagrangian dynamics method reaches the optimal performance of 96.6%.
[0069] In an optional implementation manner, a counterfactual reasoning module is constructed based on the multi-layer causal relationship graph. The counterfactual reasoning module generates a counterfactual trajectory by minimizing the Lagrangian action, and evaluates the rationality of the counterfactual trajectory by using the state transition matrix, including: A counterfactual reasoning module is constructed based on the multi-layer causal relationship graph. The counterfactual reasoning module includes a system state vector, node features, and a time variable, and a Lagrangian is defined in the counterfactual reasoning module. The Lagrangian establishes a system dynamics model through a combination of a system kinetic energy term, a potential energy term, and a constraint term. The system dynamics model describes the evolution law of the system state vector under the time variable; A minimum action functional is constructed based on the system dynamics model. The minimum action functional performs an integral operation on the Lagrangian over a time interval, solves the minimum action functional through the variational principle to obtain the Euler-Lagrange equation, and calculates the evolution trajectory of the system state vector by using the Euler-Lagrange equation to obtain a counterfactual trajectory; A state transition probability is calculated based on the counterfactual trajectory, the state transition probability maps the calculation result of the minimum action functional to a probability space through Boltzmann distribution, and a state transition matrix is constructed using the state transition probability. The rationality of the counterfactual trajectory is evaluated through the state transition matrix.
[0070] Define the system state vector, which is represented as an n-dimensional vector S=(s1,s2...s n ), where each component s i Represents the state value of a specific node in the system. The node feature is represented by a vector F=(f1,f2...f m ) means that each feature f i Describes the properties or behavioral characteristics of a node. The time variable t is used to track the evolution of the system state over time. In the counterfactual reasoning module, the Lagrangian L(S,d S / d t ,F,t) is defined as the combination of the system kinetic energy term T, potential energy term V and constraint term C: L=T-V+C.
[0071] The kinetic energy term T represents the rate of change of the system state and is calculated as the sum of the squares of the time derivatives of the state vector multiplied by the system parameter α: T = α·∑(d si / d t ) 2 Where α is an adjustable parameter, which is determined to be 0.75 by training data. The potential energy term V represents the interaction between nodes within the system, based on the connection weights w between nodes in the multi-layer causal relationship graph. ij Calculation: V = β·∑w ij ·(s i -s j ) 2 , where β is the weight coefficient, and a value of 1.2 has been found to be optimal. The constraint term C ensures that the state transition complies with the constraints of the causal graph: C = γ·∑g(S,F), where γ is the constraint strength parameter, set to 0.8, and g(S,F) is the constraint function defined based on the node feature F and the multi-layer causal graph structure.
[0072] The system dynamics model uses the Lagrangian L to describe the evolution of the state vector S over time t. For implementation, a multi-layer causal relationship graph with five nodes is used as an example, where the nodes represent different event states. The initial state vector is set to S0 = (0.2, 0.5, 0.3, 0.1, 0.4), and the node eigenvectors are F = (0.6, 0.8, 0.3, 0.7, 0.5). The causal relationship graph consists of two layers: the first layer represents direct causal relationships, and the second layer represents indirect causal relationships. The weight matrices are W1 and W2, respectively.
[0073] Construct the least action functional \(A\) based on the system dynamics model. This functional is obtained by integrating the Lagrangian \(L\) over the time interval \([t_0, t_1]\): \(A=\int L(S, d S / d t , F, t)d t . In actual calculations, the time interval is discretized into \(N\) time points, and the integral is approximated by numerical methods. For example, for the time interval \([0, 10]\), it is divided into 100 time points, and each time step is 0.1.
[0074] Solve the least action functional through the variational principle to obtain the Euler - Lagrange equation: \(d / d t (\partial L / \partial(d S / d t ))-\partial L / \partial S = 0. This equation describes the optimal evolution path of the state vector \(S\). In actual implementation, numerical iteration methods are used to solve this equation. Set the initial state \(S_0\) and the assumed intervention conditions \(I=(i_1, i_2...i n )), where some components are artificially set to specific values, representing the intervention on the system. By iteratively calculating the state vector values at each time step, the state evolution trajectory of the system under the intervention conditions, that is, the counterfactual trajectory \(S(t)\), is obtained.
[0075] Take a specific case as an example. Suppose an intervention is made on node 2, changing its state value from 0.5 to 0.8, forming the intervention condition \(I=(0.2, 0.8, 0.3, 0.1, 0.4)\). By solving the Euler - Lagrange equation, the counterfactual trajectory of the system under this intervention is calculated. After 10 time steps of evolution, the final state vector becomes \(S 10 =(0.25, 0.78, 0.42, 0.18, 0.55)\), indicating that the intervention has different degrees of influence on the states of each node.
[0076] Calculate the state transition probability based on the counterfactual trajectory, and map the calculation result of the least action functional to the probability space. The state transition probability \(P(S'|S)\) is calculated through the Boltzmann distribution: \(P(S'|S)=\exp(-A(S\rightarrow S') / \tau) / Z\), where \(A(S\rightarrow S')\) is the action from state \(S\) to state \(S'\), \(\tau\) is the temperature parameter set to 0.3, and \(Z\) is the normalization factor to ensure that the sum of probabilities is 1. By calculating the transition probabilities between different states, construct the state transition matrix \(M\), where each element \(M ij represents the probability of transitioning from state \(i\) to state \(j\).
[0077] For the above case, a 5×5 state transition matrix is constructed to record the probability of the system transitioning from the current state to the next state at each time step. For example, the transition probability from the initial state S0 to the state S1 at the first time step is 0.85, indicating a reasonable transition with a high probability. By calculating the average of the transition probabilities between adjacent states on the entire counterfactual trajectory, the overall rationality score of the trajectory is obtained as 0.78, which is higher than the preset threshold of 0.7. Therefore, it is determined that the counterfactual trajectory is reasonable.
[0078] To evaluate the rationality of the counterfactual trajectory, the stability and consistency of the trajectory can also be calculated through the state transition matrix. The stability is determined by eigenvalue analysis, and the consistency is evaluated by comparing the similarity between the counterfactual trajectory and historical data. In the above case, the trajectory stability score is 0.82, the consistency score is 0.75, and the comprehensive score is 0.78, indicating that the generated counterfactual trajectory conforms to the system dynamics laws and causal constraint conditions and has a high degree of rationality.
[0079] In the second aspect of the embodiments of the present invention, a cross-platform intelligent aggregation and analysis processing system for user behavior data is provided, including: A first unit for obtaining user historical behavior data of multiple different platforms, where the user historical behavior data includes behavior timestamps, and sorting the user historical behavior data in time series through the behavior timestamps; A second unit for performing session segmentation on the user historical behavior data sorted in time series to obtain a behavior session sequence, and constructing a behavior migration probability graph according to the transfer relationship between adjacent behaviors in the behavior session sequence; A third unit for constructing a decentralized federated reinforcement learning framework, using the behavior migration probability graph to train a graph attention network to implement an intention recognition agent, training a contrastive learning network to implement a behavior prediction agent based on the behavior session sequence, and deploying the intention recognition agent and the behavior prediction agent locally on each platform, and realizing cross-platform collaborative training by encrypting and transmitting model gradient information through a differential privacy mechanism; A fourth unit for constructing a reward signal generator according to the predicted behavior sequence generated by the behavior prediction agent, calculating the reward value of the predicted behavior sequence based on the behavior session sequence, and using the reward value to drive the model optimization of the decentralized federated reinforcement learning framework; A fifth unit for constructing a multi-layer causal relationship graph based on the behavior migration probability graph, analyzing the rationality of the predicted behavior sequence by using a graph neural network and a counterfactual reasoning method, generating a correction plan through causal intervention and selecting the optimal correction result by using causal effect evaluation, and generating a user intention portrait according to the intention recognition result output by the intention recognition agent and the optimal correction result.
[0080] In a third aspect of the embodiments of the present invention, there is provided an electronic device, comprising: a processor; a memory for storing instructions executable by the processor; wherein the processor is configured to call the instructions stored in the memory to execute the method described above.
[0081] In a fourth aspect of the embodiments of the present invention, there is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0082] The present invention may be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for performing various aspects of the present invention.
[0083] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. Cross-platform intelligent aggregation and analysis method for user behavior data, characterized in that, Including: Obtain user historical behavior data from multiple different platforms. The user historical behavior data includes behavior timestamps, and the user historical behavior data is sorted in time series according to the behavior timestamps; Based on the user historical behavior data sorted in time series, perform session segmentation to obtain a behavior session sequence, and construct a behavior transition probability graph according to the transition relationship between adjacent behaviors in the behavior session sequence; Construct a decentralized federated reinforcement learning framework. Use the behavior transition probability graph to train a graph attention network to implement an intent recognition agent, train a contrastive learning network based on the behavior session sequence to implement a behavior prediction agent, and deploy the intent recognition agent and the behavior prediction agent locally on each platform. Cross-platform collaborative training is achieved by encrypting and transmitting model gradient information through a differential privacy mechanism; Construct a reward signal generator according to the predicted behavior sequence generated by the behavior prediction agent, calculate the reward value of the predicted behavior sequence based on the behavior session sequence, and use the reward value to drive the model optimization of the decentralized federated reinforcement learning framework; Construct a multi-layer causal relationship graph based on the behavior transition probability graph, use a graph neural network and a counterfactual reasoning method to analyze the rationality of the predicted behavior sequence, generate a correction plan through causal intervention and select the optimal correction result using causal effect evaluation. Generate a user intent portrait according to the intent recognition result output by the intent recognition agent and the optimal correction result; 2. The method according to claim 1, wherein Based on the user historical behavior data sorted in time series, perform session segmentation to obtain a behavior session sequence, and constructing a behavior transition probability graph according to the transition relationship between adjacent behaviors in the behavior session sequence includes: Calculate the average behavior interval time and the standard deviation of the behavior interval time in the user historical behavior data, and construct an adaptive time threshold function based on the standard deviation; Compare the time interval between adjacent behaviors with the calculation result of the adaptive time threshold function. When the time interval is greater than the calculation result of the adaptive time threshold function, determine the session boundary. Based on the session boundary, segment the user historical behavior data to obtain a behavior session sequence, and construct an intra-session behavior association graph to depict the temporal dependence relationship of behaviors within the same session. Use the temporal dependence relationship to optimize the division of the session boundary; Extract multi-modal features for each behavior in the behavior session sequence. The multi-modal feature extraction includes extracting behavior type features, target object features, and environmental parameter features, and weighted fusion of the behavior type features, the target object features, and the environmental parameter features is performed through a feature weight matrix to obtain a behavior vector representation. On the basis of the behavior vector representation, construct an attention enhancement layer to capture the dynamic correlation between different feature dimensions; Construct a behavior transition probability graph based on the transition relationship between adjacent behaviors in the behavior vector representation, integrate the temporal dependence relationship in the intra-session behavior association graph into the behavior transition probability graph, model long-term behavior dependencies through a recurrent neural network, and use a gating mechanism to control the forgetting and updating of historical information to achieve in-depth modeling of the evolution law of user behavior.
3. The method according to claim 1, characterized in that, Construct a decentralized federated reinforcement learning framework, use the behavior transfer probability map to train a graph attention network to implement an intent recognition agent, and train a contrastive learning network based on the behavior session sequence to implement a behavior prediction agent, including: Construct a decentralized federated reinforcement learning framework based on the theory of brain-like synaptic plasticity, model the dynamic synaptic connection strength of knowledge transfer between federated learning nodes, and construct a knowledge transfer path between the federated learning nodes based on the dynamic synaptic connection strength; Map the transfer probability map to a continuous-discrete hybrid state space, perform stability constraints on the dynamic synaptic connection strength based on the continuous-discrete hybrid state space, and construct a Lyapunov stability criterion through the stability constraints; Construct a gauge field on the knowledge transfer path to represent the geometric structure of knowledge transfer, use the Lyapunov stability criterion to constrain the dynamic evolution of the gauge field, and establish a coupling equation between the knowledge field and the gauge field based on Yang-Mills gauge theory to achieve knowledge distillation between the federated learning nodes; Introduce a neurotransmitter regulation mechanism based on the coupling equation between the knowledge field and the gauge field, construct a contrastive learning network through the neurotransmitter regulation mechanism and train the behavior session sequence, and the neurotransmitter regulation mechanism adaptively adjusts the temperature parameter of contrastive learning using the nonlinear combination of multiple types of neurotransmitter concentrations; Map the multiple types of neurotransmitter concentrations to an Ising model to construct a critical state prediction framework, and the critical state prediction framework calculates the correlation length through the second moment of the spatio-temporal correlation function, and evaluates the spatio-temporal correlation characteristics of the conditional transition probability between behavior states based on the correlation length to implement a behavior prediction agent.
4. The method according to claim 3, characterized in that Map the transfer probability map to a continuous-discrete hybrid state space, perform stability constraints on the dynamic synaptic connection strength based on the continuous-discrete hybrid state space, and construct a Lyapunov stability criterion including: Map the behavior transfer probability map to a continuous-discrete hybrid state space, where the continuous-discrete hybrid state space includes continuous state variables and discrete state variables, and construct a system state representation based on the continuous state variables and the discrete state variables; Construct an entropy production rate evaluation model based on the system state representation, and the entropy production rate evaluation model calculates the time derivative of the system entropy through the product sum of the generalized flow and the generalized force, design a nonlinear fluctuation compensator using the time derivative of the system entropy, and the nonlinear fluctuation compensator calculates the compensation amount based on the entropy-dependent compensation gain matrix and the gradient of the potential function; Apply the compensation amount to the stability constraint of the dynamic synaptic connection strength, and the stability constraint controls the influence degree of the entropy gradient term through the dissipation coefficient, and constructs a bifurcation parameter using the time derivative and the spatial derivative of the system entropy for dynamic regulation; Construct a dynamic symbol encoder to map the evolution trajectory of the dynamic synaptic connection strength into a symbol sequence, calculate the topological entropy based on the symbol sequence, and the topological entropy is obtained through the logarithmic limit of the number of allowed sequences of different lengths in the symbol sequence; Integrate the topological entropy with the time derivative of the system entropy output by the entropy production rate evaluation model and perform weighted combination to construct a Lyapunov stability criterion, and optimize the dynamic synaptic connection strength by minimizing the Lyapunov stability criterion.
5. The method according to claim 1, wherein Construct a reward signal generator based on the predicted behavior sequence generated by the behavior prediction agent, calculate the reward value of the predicted behavior sequence based on the behavior session sequence, and use the reward value to drive the model optimization of the decentralized federated reinforcement learning framework, including: Generate a predicted behavior sequence based on the behavior prediction agent, construct a reward signal generator for the predicted behavior sequence, the reward signal generator constructs a temporal difference evaluation module by weighting multiple reward components of the system state and the predicted behavior, the temporal difference evaluation module calculates the temporal difference value based on the immediate reward value, the discount factor and the state value function, and uses the temporal difference value to construct a prediction accuracy evaluator, and the prediction accuracy evaluator obtains an accuracy evaluation value by calculating the Gaussian similarity between the predicted behavior feature and the real behavior feature; Calculate the information entropy and the conditional entropy based on the system state and the predicted behavior, construct an information gain metric through the difference between the information entropy and the conditional entropy, and perform weighted combination of the information gain metric and the accuracy evaluation value to obtain a comprehensive reward value; Construct a policy gradient update rule using the comprehensive reward value, and the policy gradient update rule obtains the gradient of the model parameters by calculating the expected value of the product of the logarithmic derivative of the parameterized policy function and the comprehensive reward value; Construct an asynchronous update mechanism for the decentralized federated reinforcement learning framework based on the gradient of the model parameters. The asynchronous update mechanism includes a parameter difference term and a local gradient term between nodes. Control the influence degree of the parameter difference term through the synchronization rate, and control the influence degree of the local gradient term through the learning rate to achieve the model optimization of the decentralized federated reinforcement learning framework.
6. The method according to claim 1, wherein Construct a multi-layer causal relationship graph based on the behavior transition probability graph, use graph neural network and counterfactual reasoning method to analyze the rationality of the predicted behavior sequence, generate a correction plan through causal intervention and select the optimal correction result using causal effect evaluation, including: Construct a multi-layer causal relationship graph based on the behavior transition probability graph, use the Fokker-Planck equation to describe the evolution process of the multi-layer causal relationship graph. The Fokker-Planck equation includes a drift term and a diffusion term, describe the time evolution of the system state distribution through the drift term and the diffusion term, and construct a state transition matrix based on the characteristics of the time evolution; Construct a counterfactual reasoning module based on the multi-layer causal relationship graph. The counterfactual reasoning module generates a counterfactual trajectory by minimizing the Lagrangian action, and evaluates the rationality of the counterfactual trajectory using the state transition matrix; Perform stability analysis on the counterfactual trajectory, calculate the steady-state distribution based on the Fokker-Planck equation, construct a causal intervention generator according to the steady-state distribution, the causal intervention generator generates a correction scheme based on the system state distribution, and the correction scheme realizes system optimization by minimizing the perturbation of the state transition matrix; Evaluate the system state distribution of the correction scheme, analyze the evolution characteristics of the correction scheme through the Fokker-Planck equation, calculate the convergence and stability of the system state distribution based on the evolution characteristics, and select the optimal correction scheme according to the convergence and the stability. The optimal correction scheme is obtained by minimizing the time evolution deviation of the system state distribution.
7. The method according to claim 6, wherein Construct a counterfactual reasoning module based on the multi-layer causal relationship graph. The counterfactual reasoning module generates a counterfactual trajectory by minimizing the Lagrangian action, and evaluates the rationality of the counterfactual trajectory using the state transition matrix, including: Construct a counterfactual reasoning module based on the multi-layer causal relationship graph. The counterfactual reasoning module includes a system state vector, node features, and a time variable, and defines a Lagrangian in the counterfactual reasoning module. The Lagrangian establishes a system dynamics model through a combination of a system kinetic energy term, a potential energy term, and a constraint term. The system dynamics model describes the evolution law of the system state vector under the time variable; Construct a minimum action functional based on the system dynamics model. The minimum action functional performs an integral operation on the Lagrangian over a time interval, solves the minimum action functional through the variational principle to obtain the Euler-Lagrange equation, and calculates the evolution trajectory of the system state vector using the Euler-Lagrange equation to obtain a counterfactual trajectory; Calculate the state transition probability based on the counterfactual trajectory. The state transition probability maps the calculation result of the minimum action functional to the probability space through the Boltzmann distribution, constructs a state transition matrix using the state transition probability, and evaluates the rationality of the counterfactual trajectory through the state transition matrix.
8. Cross-platform intelligent aggregation and analysis processing system for user behavior data, which is used to implement the method described in any one of the foregoing claims 1-7, characterized in that Including: The first unit is used to obtain user historical behavior data of multiple different platforms. The user historical behavior data includes behavior timestamps, and the user historical behavior data is sorted in time sequence through the behavior timestamps; The second unit is used to perform session segmentation on the user historical behavior data after time sequence sorting to obtain a behavior session sequence, and construct a behavior migration probability graph according to the transition relationship between adjacent behaviors in the behavior session sequence; The third unit is used to construct a decentralized federated reinforcement learning framework, use the behavior migration probability graph to train a graph attention network to implement an intent recognition agent, train a contrastive learning network based on the behavior session sequence to implement a behavior prediction agent, and deploy the intent recognition agent and the behavior prediction agent locally on each platform, and realize cross-platform collaborative training by encrypting and transmitting model gradient information through the differential privacy mechanism; A fourth unit, configured to construct a reward signal generator according to the predicted behavior sequence generated by the behavior prediction agent, calculate the reward value of the predicted behavior sequence based on the behavior session sequence, and drive the model optimization of the decentralized federated reinforcement learning framework by using the reward value; A fifth unit, configured to construct a multi-layer causal relationship graph based on the behavior transition probability graph, analyze the rationality of the predicted behavior sequence by using a graph neural network and a counterfactual reasoning method, generate a correction plan through causal intervention and select an optimal correction result by using causal effect evaluation, and generate a user intention portrait according to the intention recognition result output by the intention recognition agent and the optimal correction result.
9. An electronic device, characterized in that, Comprising: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Fault diagnosis model construction method based on distributed causal discovery and federated learning
CN120670792A
Cross-platform user behavior analysis method and system based on transfer learning
CN120811792A
A Cross-Platform User Behavior Analysis Method and System Based on Transfer Learning
CN120811792B
Personalized information accurate pushing system and method based on artificial intelligence
CN120823019A
E-commerce operation platform user portrait generation method based on big data
CN121032556A