A debiased cognitive recommendation method and system based on counterfactual learning

Through a multi-task debiased cognitive optimization recommendation model based on counterfactual learning, the problem of poor robustness of existing recommendation methods in dynamic environments is solved, an in-depth understanding of user behavioral intentions and intention evolution is achieved, and the accuracy and personalization of recommendations are improved.

CN117972197BActive Publication Date: 2025-09-26SUN YAT SEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410080165.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-19
Publication Date
2025-09-26
Estimated Expiration
2044-01-19

AI Technical Summary

Technical Problem

Existing recommendation methods have poor robustness in dynamic environments and are easily affected by data sparsity and bias problems. They find it difficult to accurately understand user behavioral intentions and their evolution, resulting in inaccurate recommendation results and poor personalization effects.

Method used

A multi-task debiased cognitive optimization recommendation model based on counterfactual learning is adopted. By constructing a user behavior intention evolution model and a memory network, combined with a dual robust debiased recommendation prediction loss estimator, user behavior intentions and their dynamic evolution are mined to predict user-item preference ratings.

Benefits of technology

It improves the robustness and accuracy of the recommendation method, can better understand the cognitive level of user behavior, provide personalized and accurate recommendation results, and effectively combat the impact of data bias.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117972197B_ABST
    Figure CN117972197B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of recommendation systems, and in particular to a debiased cognitive recommendation method and system based on counterfactual learning, comprising inputting an acquired sequence of user behavior representations into a user behavior intention evolution model to dynamically learn the evolution of user behavior intentions, obtaining a user behavior intention evolution state representation, and constructing a user behavior state representation in combination with a user preference representation; utilizing a multi-task debiased cognitive optimization recommendation model based on counterfactual learning to predict user behavior state representations and item representations to obtain a user-item preference optimization score; and sorting items according to the user-item preference optimization score to obtain an item recommendation list. The present invention utilizes a multi-task debiased cognitive optimization recommendation model based on counterfactual learning, employs a multi-task learning optimization mechanism to elevate the recommendation method from the perceptual intelligence stage to the cognitive intelligence stage, utilizes counterfactual learning to combat the influence of data bias, and improves the accuracy and robustness of the recommendation method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of recommendation systems, and in particular to a debiased cognitive recommendation method and system based on counterfactual learning. Background Art

[0002] Existing recommendation methods mainly include recommendation methods based on deep learning and knowledge graphs. Among them, for recommendation methods based on deep learning, relevant studies have proposed a collaborative deep filtering method. This method processes user review text information through a stacked denoising autoencoder to learn the semantic representation of users and items, and uses a collaborative filtering method to predict user ratings to improve recommendation accuracy. At the same time, some studies have also proposed a collaborative attention multi-task learning framework. Based on the traditional encoder-decoder framework, this framework introduces an encoder-selector-decoder structure. By designing a hierarchical collaborative attention selector, it promotes the cross-semantic knowledge transfer between the two tasks of recommendation prediction and interpretation generation, thereby improving the recommendation accuracy. In addition, some researchers have optimized the graph collaborative filtering model and proposed a collaborative filtering model based on enhanced graph learning. This model constructs an enhanced graph based on the similarity relationship between users and items, and uses a graph convolutional network to learn the feature representation of users and items, and designs a local-global optimization function to achieve model optimization.

[0003] However, deep learning models usually contain deeply nested nonlinear structures, which makes it difficult to determine the specific factors affecting recommendation decisions, resulting in a significant decrease in their robustness in dynamically changing environments, adversarial interference and false information. Moreover, most deep learning-based recommendation methods rely on data-driven methods. Therefore, data sparsity and data bias problems also seriously affect the effectiveness of recommendation systems.

[0004] On the other hand, for recommendation methods based on knowledge graphs, some researchers have proposed feature representation learning methods to process structured entity association information, unstructured text and visual content information introduced by external knowledge bases, and realized item recommendation by designing a joint learning framework of collaborative filtering and knowledge embedding representation; at the same time, other researchers have proposed an end-to-end knowledge graph-aware recommendation method, which automatically discovers propagation paths from the association paths of the knowledge graph and uses the information propagation mechanism to discover the user's hierarchical potential interests; in addition, some researchers have combined the knowledge graph enhancement feature representation learning at the word level and the entity level, used the mutual information maximization mechanism to align the semantic space of the word level and the entity level, and designed a knowledge-enhanced recommendation algorithm to realize item recommendation.

[0005] However, current recommendation methods based on knowledge graphs are prone to introducing redundant noise data and have difficulty processing unrelated entities, which affects the efficiency and results of recommendations. In addition, current recommendation methods based on knowledge graphs still use knowledge-enhanced learning to represent user behavior preferences from the perception level, ignoring fine-grained user behavior intentions and their evolution, and cannot achieve cognitive understanding of user behavior, which limits the depth and accuracy of recommendation methods.

[0006] In summary, although existing recommendation methods have achieved improvements in recommendation effects, due to the sparsity of user behavior data and data bias problems, it is still difficult to accurately learn user behavior preference representations. At the same time, the current user behavior understanding modeling still remains at the perception level, making it difficult to explore the potential intentions and intention evolution patterns of user behavior, and it is difficult to generate high-quality and robust recommendation results. Summary of the Invention

[0007] The purpose of the present invention is to provide a debiased cognitive recommendation method and system based on counterfactual learning, which combats the influence of data bias by using a multi-task debiased cognitive optimization recommendation model based on counterfactual learning, and mines user behavioral intentions and their dynamic evolution mechanism to improve the accuracy and robustness of recommendations.

[0008] To solve the above technical problems, the present invention provides a debiased cognitive recommendation method and system based on counterfactual learning.

[0009] In a first aspect, the present invention provides a debiased cognitive recommendation method based on counterfactual learning, the method comprising the following steps:

[0010] Constructing a user behavior representation sequence, and inputting the user behavior representation sequence into a pre-constructed user behavior intention evolution model to dynamically learn the user behavior intention evolution, thereby obtaining a user behavior intention evolution state representation;

[0011] Constructing a user behavior state representation based on the user behavior intention evolution state representation and the user preference representation;

[0012] A multi-task debiased cognitive optimization recommendation model based on counterfactual learning is used to predict user behavior state representations and item representations to obtain user-item preference optimization scores; the multi-task debiased cognitive optimization recommendation model based on counterfactual learning includes a user-item interaction recommendation prediction model and a dual robust debiased recommendation prediction loss estimator based on counterfactual learning;

[0013] Items are sorted according to the user-item preference optimization scores to obtain an item recommendation list.

[0014] In a further embodiment, the user behavior intention evolution model includes a recurrent neural network and a user behavior intention memory network model connected in series; the step of inputting the user behavior representation sequence into a pre-built user behavior intention evolution model to dynamically learn the user behavior intention evolution and obtain the user behavior intention evolution state representation includes:

[0015] Inputting the user behavior representation sequence into the recurrent neural network to learn the semantic dependency information of the user behavior intention and generate the user behavior semantic intention state representation;

[0016] Inputting the user behavior semantic intention state representation into the user behavior intention memory network model to dynamically learn the user behavior intention evolution, thereby obtaining the user behavior intention memory state representation;

[0017] The user behavior semantic intention state representation and the user behavior intention memory state are merged to generate a user behavior intention evolution state representation.

[0018] In a further embodiment, the user behavior intention memory network model includes a memory matrix and a memory controller, wherein the memory matrix includes a plurality of memory slots;

[0019] The memory controller is used to perform content addressing based on the user behavior semantic intention state representation, read the user behavior intention memory state representation from the memory slot of the memory matrix using read and write operations, and update the memory matrix. The update expression of the memory matrix is:

[0020]

[0021] in,

[0022]

[0023] Where, is the user behavior intention state representation stored in the kth memory slot in the memory matrix; g is a gate vector, which is used to determine the retention ratio of the user behavior intention state representation used to update the memory matrix; is the semantic intention state representation of user behavior; σ(·) is the activation function.

[0024] In a further embodiment, the dual robust debiased recommendation prediction loss estimator based on counterfactual learning includes a propensity score estimator and an interpolation error estimator; the step of using the multi-task debiased cognitive optimization recommendation model based on counterfactual learning to predict user behavior state representation and item representation to obtain a user-item preference optimization score includes:

[0025] A dual robust debiased recommendation prediction loss estimator based on counterfactual learning is used to estimate user behavior state representation and item representation to obtain propensity score estimation and interpolation error assessment value; the item representation includes item representation and user-item interaction relationship representation;

[0026] The user-item interaction recommendation prediction model is used to perform fusion learning on the user behavior state representation and the item representation to obtain a predicted recommendation score, and the predicted recommendation score is debiased according to the propensity score estimate and the interpolation error assessment value to predict the user-item preference optimization score.

[0027] In a further embodiment, the user-item interaction recommendation prediction model adopts a hierarchical neural collaborative filtering recommendation prediction model, and the hierarchical neural collaborative filtering recommendation prediction model includes two neural network layers;

[0028] The first neural network layer is used to aggregate the user behavior state representations at different times using the attention mechanism to obtain the global user state representation;

[0029] The second neural network layer is used to use a multi-layer perceptron to perform fusion learning on the global user state representation and item representation, and to perform debiasing processing in combination with the propensity score estimate and the interpolation error assessment value to predict the user-item preference optimization score.

[0030] In a further embodiment, the multi-task debiased cognitive optimization recommendation model based on counterfactual learning is trained and optimized using a multi-task loss function consisting of recommendation prediction loss, propensity score loss and data interpolation loss.

[0031] In a further embodiment, the mathematical expression of the multi-task loss function is:

[0032]

[0033] Where L(·) is the multi-task loss function; p u Represents the user's behavior status; v For items; r uv represents the user-item interaction relationship; θ represents all parameters in the user-item interaction recommendation prediction model; are all parameters of the propensity score estimator; are all parameters of the imputation error estimator; B is the user-item combination pair used; o u,i is the observed rating of user u; r u,i The real rating of user u; For prediction recommendation score; The recommendation prediction loss between the true rating and the predicted recommendation score; Estimated values ​​for interpolation errors; is the propensity score estimate predicted by the propensity score estimator; λ is the parameter of the regularization term.

[0034] In a second aspect, the present invention provides a debiased cognitive recommendation system based on counterfactual learning, the system comprising:

[0035] A behavior intention evolution module is used to construct a user behavior representation sequence and input the user behavior representation sequence into a pre-built user behavior intention evolution model to dynamically learn the user behavior intention evolution and obtain the user behavior intention evolution state representation;

[0036] A behavior state generation module, configured to construct a user behavior state representation based on the user behavior intention evolution state representation and the user preference representation;

[0037] A preference score prediction module is used to predict user behavior state representations and item representations using a multi-task debiased cognitive optimization recommendation model based on counterfactual learning to obtain user-item preference optimization scores; the multi-task debiased cognitive optimization recommendation model based on counterfactual learning includes a user-item interaction recommendation prediction model and a dual robust debiased recommendation prediction loss estimator based on counterfactual learning;

[0038] The recommendation list generation module is used to sort the items according to the user-item preference optimization score to obtain an item recommendation list.

[0039] At the same time, in a third aspect, the present invention also provides a computer device, comprising a processor and a memory, wherein the processor is connected to the memory, the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory, so that the computer device performs the steps of implementing the above method.

[0040] In a fourth aspect, the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0041] The present invention provides a debiased cognitive recommendation method and system based on counterfactual learning. The method inputs a constructed sequence of user behavior representations into a user behavior intention evolution model to dynamically learn the evolution of user behavior intentions, obtaining a representation of the user behavior intention evolution state; constructs a user behavior state representation based on the user behavior intention evolution state representation and the user preference representation; utilizes a multi-task debiased cognitive optimization recommendation model based on counterfactual learning to predict the user behavior state representation and item representation to obtain a user-item preference optimization score; and sorts items based on the user-item preference optimization score to obtain an item recommendation list. Compared with existing technologies, this method fully exploits user behavior intentions and their dynamic evolution to achieve a comprehensive and fine-grained cognitive understanding of user behavior, breaking through the technical bottlenecks of existing technologies in user behavior cognitive understanding modeling. Furthermore, the multi-task debiased cognitive optimization recommendation model based on counterfactual learning improves the robustness and accuracy of the recommendation method from a cognitive intelligence perspective, effectively combating the impact of data bias on the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 1 is a flow chart of a method for debiased cognitive recommendation based on counterfactual learning provided by an embodiment of the present invention;

[0043] Figure 2 This is a block diagram of a user behavior intention evolution model provided by an embodiment of the present invention;

[0044] Figure 3 1 is a schematic diagram of the overall process of multi-task debiasing cognitive optimization recommendation based on counterfactual learning provided by an embodiment of the present invention;

[0045] Figure 4 is a block diagram of a debiased cognitive recommendation system based on counterfactual learning provided by an embodiment of the present invention;

[0046] Figure 5 It is a structural diagram of a computer device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0047] The following describes the embodiments of the present invention in detail with reference to the accompanying drawings. The embodiments are provided for illustrative purposes only and are not to be construed as limiting the present invention. The accompanying drawings are provided for reference and illustration only and do not constitute a limitation on the scope of protection of the present invention. Many changes may be made to the present invention without departing from the spirit and scope of the present invention.

[0048] Current recommendation systems are still in the perceptual intelligence stage, relying primarily on machine learning and deep learning technologies to mine user behavior patterns and predict user behavior preferences in a data-driven manner. However, in dynamically changing environments and in the presence of adversarial interference and false information, the accuracy of such recommendation methods decreases significantly. Therefore, how to model cognitive understanding of user behavior and establish corresponding cognitive recommendation methods has become a major challenge facing current recommendation system research. In addition, the shortcomings of existing recommendation methods in user behavior analysis modeling and solving data sparsity and data bias problems are the main challenges facing the current recommendation system in its transition from the perceptual intelligence stage to the cognitive intelligence stage, mainly including:

[0049] (1) User behavior data in recommendation systems suffer from sparsity and data bias, which seriously affects the accuracy and effectiveness of recommendation results. Existing recommendation methods mainly focus on designing complex models to mine users' potential behavior patterns from sparse user behavior data. However, this method is difficult to accurately learn user behavior preference representations, resulting in inaccurate recommendation results. In addition, data bias is also an important challenge. Due to the influence of factors such as user selection bias, item popularity, and item exposure, the credibility and fairness of recommendation results are seriously affected, which leads to deviations in personalized recommendation results and fails to meet user needs and expectations.

[0050] (2) Existing recommendation methods mainly mine user behavior and contextual information from the perceptual level to predict user preferences. However, this method ignores the potential intentions of user behavior and the evolution of user behavior preferences, and cannot clearly represent the user's real feedback and emotional polarity. In addition, existing recommendation methods also rarely consider the dynamic characteristics of user behavior, ignore the evolution process of user behavior intentions, and cannot deeply mine the user's potential implicit preferences. This makes it difficult for the recommendation system to achieve cognitive understanding modeling of user behavior from the cognitive level, and cannot provide users with more personalized and accurate recommendations.

[0051] To solve these problems, refer to Figure 1 , the embodiment of the present invention provides a debiased cognitive recommendation method based on counterfactual learning, such as Figure 1 As shown, the method includes the following steps:

[0052] S1. Construct a user behavior representation sequence, and input the user behavior representation sequence into a pre-constructed user behavior intention evolution model to dynamically learn the user behavior intention evolution, and obtain the user behavior intention evolution state representation.

[0053] In this embodiment, the user behavior intention evolution model includes a serially connected recurrent neural network and a user behavior intention memory network model; the step of inputting the user behavior representation sequence into the pre-built user behavior intention evolution model to dynamically learn the user behavior intention evolution and obtain the user behavior intention evolution state representation includes:

[0054] Inputting the user behavior representation sequence into the recurrent neural network to learn the semantic dependency information of the user behavior intention and generate the user behavior semantic intention state representation;

[0055] Inputting the user behavior semantic intention state representation into the user behavior intention memory network model to dynamically learn the user behavior intention evolution, thereby obtaining the user behavior intention memory state representation;

[0056] The user behavior semantic intention state representation and the user behavior intention memory state are merged to generate a user behavior intention evolution state representation.

[0057] In a specific embodiment, in order to track the evolution of user behavior intentions, this embodiment constructs a user behavior representation sequence based on the time sequence of user-item interactions. in, Represents the user behavior representation of user u at the current time t, such as Figure 2 As shown, this embodiment first uses a recurrent neural network to learn the semantic dependency information of user behavior intention from the user behavior representation sequence to obtain the user behavior semantic intention state representation Among them, the recurrent neural network adopted in this embodiment preferably selects the gated recurrent unit (GRU). Those skilled in the art can select other recurrent neural network architectures according to the specific implementation situation, which is not limited to the embodiments of the present invention.

[0058] Then, this embodiment inputs the user behavior semantic intention state representation into the user behavior intention memory network model, uses the user behavior intention memory network model to simulate the user behavior state, dynamically learns the evolution of user behavior intention, and obtains the user behavior intention memory state representation. In this embodiment, the user behavior intention memory network model includes a memory matrix and a memory controller, wherein the memory matrix is ​​used to represent and store user state, and the memory controller is used to extract the user behavior intention memory state representation from the memory matrix using read and write operations, and update the memory matrix. In this embodiment, the main function of the controller is to extract the current state and update the state through read and write operations, that is, the controller can obtain the current state information through read operations, and can update the state through write operations to keep it consistent with the latest data. This combination of read and write operations enables the memory controller to effectively manage and control the state of the system.

[0059] The memory matrix includes several memory slots, which are used to store the latest several user behavior intention state representations of the user behavior sequence. At the current moment t, according to the user behavior semantic intention state representation Perform content addressing and read the user behavior intention memory state representation from the memory matrix The read operation is defined as follows:

[0060]

[0061] in,

[0062]

[0063]

[0064] Where, is the user behavior intention memory state representation of user u at the current time t; z tk Scale the intermediate representation variables for user states; M u is the memory matrix of user u, d m is the vector dimension of the memory slot; is the user behavior intention state representation stored in the kth memory slot in the memory matrix; ω is the scaling parameter; s tk is the user state at time t, which is the intermediate representation variable of the user state in the user behavior intention evolution model; is the semantic intention state representation of user behavior of user u at the current time t.

[0065] When updating the memory matrix using a write operation, this embodiment draws on the idea of ​​the Neural Turing model and sets a gate vector g to determine how much information is retained and updated to the memory matrix. The update expression of the memory matrix is:

[0066]

[0067] in,

[0068]

[0069] Where, is the user behavior intention state representation stored in the kth memory slot in the memory matrix; g is a gate vector, which is used to determine the retention ratio of the user behavior intention state representation used to update the memory matrix; is the semantic intention state representation of user behavior; σ(·) is the activation function.

[0070] Finally, this embodiment concatenates the user behavior semantic intention state representation and the user behavior intention memory state Output the evolution state representation of the user behavior intention at the current time t

[0071] S2. Construct a user behavior state representation based on the user behavior intention evolution state representation and the user preference representation.

[0072] S3. A multi-task debiased cognitive optimization recommendation model based on counterfactual learning is used to predict user behavior state representations and item representations to obtain user-item preference optimization scores; the multi-task debiased cognitive optimization recommendation model based on counterfactual learning includes a user-item interaction recommendation prediction model and a dual robust debiased recommendation prediction loss estimator based on counterfactual learning.

[0073] like Figure 3 As shown, in this embodiment, the user behavior intention evolution state representation and the user preference representation are combined to obtain the user behavior state representation p u Afterwards, the user behavior state is represented by the shared feature representation layer u and object representation v ,r uv Input is a multi-task debiased cognitive optimization recommendation model based on counterfactual learning, which uses a multi-task learning optimization mechanism to simultaneously learn recommendation prediction, propensity score estimator and interpolation error estimator, wherein the item representation includes item representation e v and user-item interaction relationship representation r uv The multi-task debiased cognitive optimization recommendation model based on counterfactual learning includes a user-item interactive recommendation prediction model and a dual robust debiased recommendation prediction loss estimator based on counterfactual learning; the dual robust debiased recommendation prediction loss estimator based on counterfactual learning includes a propensity score estimator and an interpolation error estimator. In this embodiment, the multi-task debiased cognitive optimization recommendation model based on counterfactual learning is used to predict user behavior state representation and item representation to obtain a user-item preference optimization score, including the following steps:

[0074] A dual robust debiased recommendation prediction loss estimator based on counterfactual learning is used to estimate user behavior state representation and item representation to obtain propensity score estimation and interpolation error assessment value;

[0075] The user-item interaction recommendation prediction model is used to perform fusion learning on the user behavior state representation and the item representation to obtain a predicted recommendation score, and the predicted recommendation score is debiased according to the propensity score estimate and the interpolation error assessment value to predict the user-item preference optimization score.

[0076] Among them, the user-item interaction recommendation prediction model adopts a hierarchical neural collaborative filtering recommendation prediction model, which includes two layers of neural network layers. This embodiment designs a user-item interaction recommendation prediction algorithm with a two-layer neural network structure. The first layer of the neural network layer is used to use the attention mechanism to aggregate the user behavior state representation at different times to obtain the global user state representation. Among them, W I is a learning parameter, f1 is an activation function, and in this embodiment, the ReLU function can be selected as the activation function; the second neural network layer is used to use a multi-layer perceptron MLP to perform fusion learning on the global user state representation and item representation, and to perform debiasing processing in combination with the propensity score estimate and the interpolation error assessment value to predict the user-item preference optimization score. User-item preference optimization score The mathematical expression is:

[0077]

[0078] Where W z is the learning parameter; s u is the global user state representation; z v is the item representation; σ(·) is the sigmoid function.

[0079] The debiased cognitive recommendation method based on counterfactual learning proposed in this embodiment adopts a multi-task learning optimization mechanism. It not only constructs a user-item interaction recommendation prediction algorithm based on the evolution of user behavior intention through the user behavior intention evolution state representation output by the user behavior intention evolution model, realizes cognitive understanding modeling of user behavior, and breaks through the technical bottleneck that existing methods are difficult to achieve cognitive understanding modeling of user behavior; at the same time, on the basis of the user-item interaction recommendation prediction algorithm based on user behavior intention evolution reasoning, from the perspective of causal reasoning, based on counterfactual learning combined with inverse propensity score weighting (IPS) and data imputation (Data Imputation) method, a dual robust debiased recommendation prediction loss estimator is proposed to combat interference such as user data bias in actual scenarios and achieve robust recommendation.

[0080] like Figure 3 As shown in FIG, in the multi-task learning optimization mechanism, the multi-task debiased cognitive optimization recommendation model based on counterfactual learning is trained and optimized using a multi-task loss function consisting of recommendation prediction loss, propensity score loss, and data interpolation loss. In the hierarchical user-item interaction recommendation prediction algorithm based on recommendation tasks, the mathematical expression of the recommendation prediction loss is:

[0081]

[0082] Where, E R is the recommendation prediction loss; R is the true rating set; is the set of predicted recommendation scores of the user-item interaction recommendation prediction model; B is the user-item combination pair used; r u,i The real rating of user u; The predicted recommendation score of user u output by the user-item interaction recommendation prediction model; In order to calculate the loss index between the true score and the predicted recommendation score, the mean square error (MSE) and root mean square error (RMSE) can be used. These indicators can objectively reflect the accuracy and reliability of the prediction results and provide an important basis for optimizing the recommendation system.

[0083] In this embodiment, the loss ξ of the user-item interaction recommendation prediction model in the observed data R The estimate is defined as:

[0084]

[0085] Where O is the observed user rating set. If there is a true rating r u,i , then o u,i =1, otherwise for missing data, o u,i =0;o u,i is the observed rating of user u; R o is the set of observed true ratings.

[0086] However, if we only rely on the recommendation prediction loss, there will be a large deviation between the prediction loss and the expected loss of the recommendation prediction due to data bias problems (such as user selection bias), which will affect the accuracy of the recommendation. It should be noted that since the missing data is not random, this deviation will cause a systematic deviation in the recommendation results. In order to remove the data bias, this embodiment combines the inverse propensity score weighting and data interpolation method to design a dual robust debiased recommendation estimator, in which the loss estimator ξ based on the inverse propensity score weighting is used. IPS Defined as:

[0087]

[0088] Where, The propensity score estimated value predicted by the propensity score estimator can be predicted by an independent logistic regression model in this embodiment.

[0089] Meanwhile, the data interpolation method is to design an interpolation error estimator The interpolation error estimator is Figure 3 As shown in the right box, the data interpolation error estimator ξ Imp The loss is defined as:

[0090]

[0091] This embodiment uses an interpolation error estimator to estimate the loss of missing data and interpolate it with the loss error of observed data. Therefore, this embodiment combines the loss estimator of the inverse propensity score and the interpolation error estimator to design a dual robust debiased recommendation prediction loss estimator ξ DR It can be defined as:

[0092]

[0093] In summary, this embodiment adopts a multi-task learning optimization mechanism, through the shared feature representation layer p u ,e v ,r uv , while learning recommendation prediction, propensity score estimator and imputation error estimator, the multi-task loss function adopted by the multi-task learning optimization mechanism can be defined as:

[0094]

[0095] Where L(·) is the multi-task loss function; p u Represents the user's behavior status; v For items; r uv represents the user-item interaction relationship; θ represents all parameters in the user-item interaction recommendation prediction model; are all parameters of the score estimator; are all parameters of the imputation error estimator; B is the user-item combination pair used; o u,i is the observed rating of user u; r u,i The real rating of user u; For prediction recommendation score; The recommendation prediction loss between the true rating and the predicted recommendation score; Estimated values ​​for interpolation errors; is the propensity score estimate predicted by the propensity score estimator; λ is the parameter of the regularization term.

[0096] S4. Sort the items according to the user-item preference optimization score to obtain an item recommendation list.

[0097] After iteratively training the multi-task debiased cognitive optimization recommendation model using the multi-task loss function, this embodiment optimizes the user-item preference score predicted by the trained multi-task debiased cognitive optimization recommendation model. Sort the items and output a top-N item recommendation list to user u.

[0098] In view of the shortcomings of existing recommendation methods, such as their inability to deeply explore and recognize the potential intentions and intention evolution patterns of user behavior, difficulty in processing sparse and biased data, and poor robustness against interference and false information, this embodiment proposes a de-biased cognitive recommendation framework based on counterfactual learning, which improves the robustness and accuracy of existing recommendation methods from the perspective of cognitive intelligence. This framework not only takes the lead in designing a dual-robust de-biased recommendation prediction loss estimation framework, using counterfactual learning to effectively combat the impact of data bias on the model, but also innovatively proposes a user-item interaction recommendation prediction model based on user behavior intention evolution reasoning, explores user behavior intentions and their dynamic evolution mechanism, and constructs a comprehensive and fine-grained user behavior cognitive understanding model, breaking through the technical bottleneck that existing methods have difficulty in achieving user behavior cognitive understanding modeling, and achieving more accurate recommendation results.

[0099] An embodiment of the present invention provides a debiased cognitive recommendation method based on counterfactual learning. The method dynamically learns the evolution of user behavior intention by inputting an acquired user behavior representation sequence into a user behavior intention evolution model to obtain a user behavior intention evolution state representation; constructs a user behavior state representation based on the user behavior intention evolution state representation and the user preference representation; uses a multi-task debiased cognitive optimization recommendation model based on counterfactual learning to predict the user behavior state representation and item representation to obtain a user-item preference optimization score; and sorts items according to the user-item preference optimization score to obtain an item recommendation list. Compared with existing recommendation methods, this embodiment elevates the recommendation method from the perceptual intelligence stage to the cognitive intelligence stage. By deeply exploring the potential intentions and preference evolution of user behavior and paying attention to the dynamic characteristics of user behavior, it better understands the evolution process of user behavior, constructs a comprehensive and fine-grained user behavior cognitive understanding model, and breaks through the technical bottleneck that existing methods have difficulty in achieving user behavior cognitive understanding modeling. At the same time, from the perspective of causal reasoning, a dual robust de-biased recommendation estimator is proposed based on counterfactual learning combined with inverse propensity score weighting and data interpolation methods to combat interference factors such as user data bias in actual scenarios, implement a robust recommendation method, and enhance users' trust and acceptance of recommendation results.

[0100] It should be noted that the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of this application.

[0101] In one embodiment, Figure 4 As shown, an embodiment of the present invention provides a debiased cognitive recommendation system based on counterfactual learning, the system comprising:

[0102] A behavior intention evolution module is used to construct a user behavior representation sequence and input the user behavior representation sequence into a pre-built user behavior intention evolution model to dynamically learn the user behavior intention evolution and obtain the user behavior intention evolution state representation;

[0103] A behavior state generation module, configured to construct a user behavior state representation based on the user behavior intention evolution state representation and the user preference representation;

[0104] A preference score prediction module is used to predict user behavior state representations and item representations using a multi-task debiased cognitive optimization recommendation model based on counterfactual learning to obtain user-item preference optimization scores; the multi-task debiased cognitive optimization recommendation model based on counterfactual learning includes a user-item interaction recommendation prediction model and a dual robust debiased recommendation prediction loss estimator based on counterfactual learning;

[0105] The recommendation list generation module is used to sort the items according to the user-item preference optimization score to obtain an item recommendation list.

[0106] For the specific definition of a de-biased cognitive recommendation system based on counterfactual learning, please refer to the above-mentioned definition of a de-biased cognitive recommendation method based on counterfactual learning, which will not be repeated here. Those of ordinary skill in the art will appreciate that the various modules and steps described in conjunction with the embodiments disclosed in this application can be implemented in hardware, software, or a combination of both. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0107] Embodiments of the present invention provide a debiased cognitive recommendation system based on counterfactual learning. The system's behavior intention evolution module dynamically learns the evolution of user behavior intention by inputting a sequence of user behavior representations into a user behavior intention evolution model, obtaining a representation of the user behavior intention evolution state. The behavior state generation module constructs a user behavior state representation based on the user behavior intention evolution state representation and user preference representation. The preference score prediction module predicts the user behavior state representation and item representation using a multi-task debiased cognitive optimization recommendation model based on counterfactual learning, obtaining an optimized user-item preference score. The recommendation list generation module ranks items based on the optimized user-item preference score to obtain an item recommendation list. Compared to existing technologies, this application uses a user-item interaction recommendation prediction algorithm based on user behavior intention evolution reasoning to deeply explore the potential intentions and preference evolution of user behavior, focusing on the dynamic characteristics of user behavior and gaining a deeper understanding of the cognitive aspects of user behavior. Furthermore, to better address data sparsity and data bias, a dual robust debiased recommendation prediction loss estimator is designed. Counterfactual learning is used to combat the effects of data bias, thereby improving the robustness and accuracy of the recommendation method and enhancing the quality and credibility of personalized recommendations.

[0108] Figure 5 A computer device provided in an embodiment of the present invention includes a memory, a processor and a transceiver, which are connected via a bus; the memory is used to store a set of computer program instructions and data, and can transmit the stored data to the processor, and the processor can execute the program instructions stored in the memory to perform the steps of the above method.

[0109] The memory may include volatile memory or non-volatile memory, or may include both volatile and non-volatile memory; the processor may be a central processing unit, a microprocessor, an application-specific integrated circuit, a programmable logic device, or a combination thereof. By way of example and not limitation, the programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic, or any combination thereof.

[0110] Additionally, the memory may be a physically separate unit or integrated with the processor.

[0111] It can be understood by those skilled in the art that Figure 5 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have the same component arrangement.

[0112] In one embodiment, the present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the above method are implemented.

[0113] The embodiment of the present invention provides a de-biased cognitive recommendation method and system based on counterfactual learning. The de-biased cognitive recommendation method based on counterfactual learning addresses the limitations of prior art in terms of insufficient cognitive understanding of user behavior, poor robustness against interference and false information, and data sparsity and bias. It utilizes knowledge reasoning to mine user behavior intentions and their dynamic evolution to achieve cognitive understanding of user behavior, and constructs a comprehensive and fine-grained user behavior cognitive understanding model, breaking through the technical bottleneck of prior recommendation methods that are difficult to achieve user behavior cognitive understanding modeling. At the same time, a dual de-biased cognitive recommendation framework is designed based on user behavior cognitive understanding modeling, thereby improving the robustness of recommendation results.

[0114] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., an SSD).

[0115] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing related hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods.

[0116] The above-described embodiments merely represent several preferred implementations of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent. It should be noted that a person skilled in the art could make several improvements and substitutions without departing from the technical principles of the present invention, and these improvements and substitutions should also be considered within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be based on the scope of protection of the claims.

Claims

1. A debiased cognitive recommendation method based on counterfactual learning, characterized in that: The following steps are involved: Constructing a user behavior representation sequence, and inputting the user behavior representation sequence into a pre-constructed user behavior intention evolution model to dynamically learn the user behavior intention evolution, thereby obtaining a user behavior intention evolution state representation; Constructing a user behavior state representation based on the user behavior intention evolution state representation and the user preference representation; A multi-task debiased cognitive optimization recommendation model based on counterfactual learning is used to predict user behavior state representation and item representation to obtain user-item preference optimization scores; The multi-task debiased cognitive optimization recommendation model based on counterfactual learning includes a user-item interaction recommendation prediction model and a dual robust debiased recommendation prediction loss estimator based on counterfactual learning; Sort the items according to the user-item preference optimization score to obtain an item recommendation list; The dual robust debiased recommendation prediction loss estimator based on counterfactual learning includes a propensity score estimator and an interpolation error estimator. The step of using the multi-task debiased cognitive optimization recommendation model based on counterfactual learning to predict user behavior state representation and item representation to obtain user-item preference optimization scores includes: A dual robust debiased recommendation prediction loss estimator based on counterfactual learning is used to estimate user behavior state representation and item representation to obtain propensity score estimation and interpolation error assessment value; the item representation includes item representation and user-item interaction relationship representation; Using a user-item interactive recommendation prediction model to perform fusion learning on the user behavior state representation and the item representation to obtain a predicted recommendation score, and debiasing the predicted recommendation score based on the propensity score estimate and the interpolation error assessment value to predict a user-item preference optimization score; The multi-task debiased cognitive optimization recommendation model based on counterfactual learning is trained and optimized using a multi-task loss function consisting of recommendation prediction loss, propensity score loss, and data interpolation loss. The mathematical expression of the multi-task loss function is: Where, is the multi-task loss function; Represent user behavior status; For items; Represents the user-item interaction relationship; Recommend all parameters in the prediction model for user-item interactions; are all parameters of the propensity score estimator; are all parameters of the imputation error estimator; B is the user-item combination pair used; is the observed rating of user u; The real rating of user u; For prediction recommendation score; The recommendation prediction loss between the true rating and the predicted recommendation score; Estimated values ​​for interpolation errors; The propensity score estimate predicted by the propensity score estimator; is the parameter of the regularization term.

2. The method for debiasing cognitive recommendation based on counterfactual learning according to claim 1, wherein: The user behavior intention evolution model includes a serially connected recurrent neural network and a user behavior intention memory network model; the step of inputting the user behavior representation sequence into the pre-built user behavior intention evolution model to dynamically learn the user behavior intention evolution and obtain the user behavior intention evolution state representation includes: Inputting the user behavior representation sequence into the recurrent neural network to learn the semantic dependency information of the user behavior intention and generate the user behavior semantic intention state representation; Inputting the user behavior semantic intention state representation into the user behavior intention memory network model to dynamically learn the user behavior intention evolution, thereby obtaining the user behavior intention memory state representation; The user behavior semantic intention state representation and the user behavior intention memory state are merged to generate a user behavior intention evolution state representation.

3. The method for debiasing cognitive recommendation based on counterfactual learning according to claim 2, characterized in that: The user behavior intention memory network model includes a memory matrix and a memory controller, wherein the memory matrix includes a plurality of memory slots; The memory controller is used to perform content addressing based on the user behavior semantic intention state representation, read the user behavior intention memory state representation from the memory slot of the memory matrix using read and write operations, and update the memory matrix. The update expression of the memory matrix is: in, Where, It represents the user behavior intention state stored in the kth memory slot in the memory matrix; is a gate vector, the gate vector being used to determine a retention ratio of the user behavior intention state representation for updating the memory matrix; Represent the semantic intention state of user behavior; is the activation function.

4. The method for debiasing cognitive recommendation based on counterfactual learning according to claim 1, characterized in that: The user-item interaction recommendation prediction model adopts a hierarchical neural collaborative filtering recommendation prediction model, which includes two neural network layers; The first neural network layer is used to aggregate the user behavior state representations at different times using the attention mechanism to obtain the global user state representation; The second neural network layer is used to use a multi-layer perceptron to perform fusion learning on the global user state representation and item representation, and to perform debiasing processing in combination with the propensity score estimate and the interpolation error assessment value to predict the user-item preference optimization score.

5. A debiased cognitive recommendation system based on counterfactual learning, characterized in that: The system comprises: A behavior intention evolution module is used to construct a user behavior representation sequence and input the user behavior representation sequence into a pre-built user behavior intention evolution model to dynamically learn the user behavior intention evolution and obtain the user behavior intention evolution state representation; A behavior state generation module, configured to construct a user behavior state representation based on the user behavior intention evolution state representation and the user preference representation; A preference score prediction module is used to predict user behavior state representations and item representations using a multi-task debiased cognitive optimization recommendation model based on counterfactual learning to obtain user-item preference optimization scores; the multi-task debiased cognitive optimization recommendation model based on counterfactual learning includes a user-item interaction recommendation prediction model and a dual robust debiased recommendation prediction loss estimator based on counterfactual learning; A recommendation list generation module is used to sort items according to the user-item preference optimization score to obtain an item recommendation list; The dual robust debiased recommendation prediction loss estimator based on counterfactual learning includes a propensity score estimator and an interpolation error estimator. The multi-task debiased cognitive optimization recommendation model based on counterfactual learning is used to predict user behavior state representation and item representation to obtain a user-item preference optimization score, specifically including: A dual robust debiased recommendation prediction loss estimator based on counterfactual learning is used to estimate user behavior state representation and item representation to obtain propensity score estimation and interpolation error assessment value; the item representation includes item representation and user-item interaction relationship representation; Using a user-item interactive recommendation prediction model to perform fusion learning on the user behavior state representation and the item representation to obtain a predicted recommendation score, and debiasing the predicted recommendation score based on the propensity score estimate and the interpolation error assessment value to predict a user-item preference optimization score; The multi-task debiased cognitive optimization recommendation model based on counterfactual learning is trained and optimized using a multi-task loss function consisting of recommendation prediction loss, propensity score loss, and data interpolation loss. The mathematical expression of the multi-task loss function is: Where, is the multi-task loss function; Represent user behavior status; For items; Represents the user-item interaction relationship; Recommend all parameters in the prediction model for user-item interactions; are all parameters of the propensity score estimator; are all parameters of the imputation error estimator; B is the user-item combination pair used; is the observed rating of user u; The real rating of user u; For the prediction recommendation score; The recommendation prediction loss between the true rating and the predicted recommendation score; Estimated values ​​for interpolation errors; The propensity score estimate predicted by the propensity score estimator; is the parameter of the regularization term.

6. A computer device, characterized in that: The computer device comprises a processor and a memory, wherein the processor is connected to the memory, the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory, so that the computer device performs the method according to any one of claims 1 to 4.

7. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • Model agnostic anti-fact interpretation method based on multi-behavior recommendation model

    CN116071119A

  • Optimized smith-waterman search

    US20080250016A1