Sample selection and recommendation model construction method based on causal inference debiasing
Patent Information
- Application Number
- CN202611094468.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-22
- Publication Date
- 2026-09-18
AI Technical Summary
可以理解为,系统在历史运行过程中,会根据已有推荐模型、项目热度及曝光规则向用户展示部分项目,即推荐给该用户的数据,一些是系统经过历史曝光筛选后的结果数据,即有偏数据,这类基于有偏数据构建的推荐机制,并不能提供准确且适配用户个体的适应性推荐
本发明通过从目标用户历史交互行为日志中提取每次交互的行为类型、交互场景和时间信息,在按照组合规则整合为交互的行为状态后,进一步构建基于行为状态转移关系的异构节点拓扑图,并结合结构偏置、时间偏置和行为惯性偏置对状态关联权重进行调整,实现对目标用户历史交互行为的去偏处理,降低由项目关联、时间变化以及重复行为导致的数据偏差。
Smart Images

Figure CN122778052A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, specifically to a method for sample selection and recommendation model construction based on causal inference and bias removal. Background Technology
[0002] In recommendation mechanisms within software systems, user preferences are typically learned from historical interaction data between the user and recommended items. However, these historical interactions are not entirely determined by the user's active interests; a portion is influenced by the system's existing exposure strategies. This can be understood as the system displaying some items to the user based on existing recommendation models, item popularity, and exposure rules—this is the data recommended to the user. Some of this data is the result of filtering based on historical exposures, i.e., biased data. Recommendation mechanisms built on this biased data cannot provide accurate and adaptive recommendations tailored to individual users.
[0003] In summary, the interaction data used to train the recommendation model in the system suffers from biased data selection due to the historical exposure mechanism. This makes the model prone to learning non-user interest factors caused by exposure opportunities, item popularity, and repeated interaction behaviors, thereby reducing the accuracy of user interest representation and affecting the personalized recommendation effect for target users. Summary of the Invention
[0004] (a) Technical problems to be solved This invention provides a sample selection and recommendation model construction method based on causal inference to reduce the introduction of biased data in the training samples of the recommendation model.
[0005] (II) Technical Solution To achieve the above objectives, the present invention provides the following technical solution: a sample selection and recommendation model construction method based on causal inference bias removal, comprising the following execution steps: Extract interaction items, behavior types, interaction scenarios, and time information from the target user's historical interaction behavior logs. Generate behavior states based on the combination relationship between the behavior types and interaction scenarios. Count the number of transitions between adjacent behavior states and generate a transition probability matrix of behavior states after normalization. Construct a heterogeneous node topology graph that includes project nodes and behavior state nodes; wherein, establish association edges between the project nodes and behavior state nodes, and establish state transition edges between adjacent behavior state nodes; Based on the association between the behavioral state nodes at both ends of the state transition edge and the same project node, the structural bias feature generated by the project node is calculated; the time bias feature is calculated based on the time interval between the corresponding behavioral states of the state transition edge; and the behavioral inertia bias feature is calculated based on the repetition of the corresponding behavioral states of the continuous state transition edge. The initial weights of each state transition edge are determined according to the transition probability matrix. The initial weights are corrected after calculating the weight adjustment coefficients according to the bias features of each item. The weights of the state association edges after debiasing are obtained. The association strength between adjacent behavioral states in the behavioral state sequence is adjusted to generate a debiased behavioral state sequence. The debiased behavioral state sequence is used as input and the corresponding interaction item is used as the prediction target to construct training samples. The biased behavior state sequence is input into the long short-term memory network recommendation model. The state is updated at each time step through the gating unit to generate the user behavior representation. The interactive items in the training samples are vector encoded, the matching score between the user behavior representation and the encoded vector is calculated, and the model parameters are optimized according to the matching score to obtain the trained recommendation model.
[0006] In some feasible embodiments, the historical interaction behavior logs of the target user are obtained, each interaction behavior log is parsed, and the interaction item identifier, behavior type, interaction scenario, and behavior occurrence timestamp in the record are identified. The parsed fields are uniformly formatted to generate interaction records corresponding to each interaction behavior event. Based on the behavior occurrence timestamp corresponding to each interaction record, the historical interaction record sequence of the target user is generated after sorting. Each interaction record corresponds to one interaction behavior event between the user and the interaction item.
[0007] In some feasible embodiments, based on the behavior type and interaction scenario in each interaction record, the behavior type and interaction scenario are combined and encoded into a corresponding behavior state according to a predefined behavior state mapping rule; the generated behavior states are then organized sequentially according to the order of the interaction records in the historical interaction record sequence to form a behavior state sequence corresponding to the target user. Traverse adjacent combinations of behavioral states in the sequence of behavioral states, count the number of transitions between each behavioral state, and obtain a state transition count matrix; Based on the total number of transitions corresponding to each initial behavioral state, the state transition number matrix is normalized to generate a behavioral state transition probability matrix.
[0008] In some feasible embodiments, the historical interaction records are traversed, and the interaction items of the same type are mapped to the same item node. At the same time, the behavior state sequence is traversed, and the behavior states of the same type are mapped to the same behavior state node, thereby generating a set of item nodes and a set of behavior state nodes respectively. Based on the interactive items and corresponding behavior states in each interactive record, the corresponding item nodes and behavior state nodes are located respectively, and an association edge is established between the two types of nodes, so that each interactive record is mapped to the association relationship between item nodes and behavior state nodes. According to the order of the behavior state sequence, read two adjacent behavior states in sequence and locate the corresponding behavior state node; establish a state transition edge between the previous behavior state node and the next behavior state node; the state transition edge is a directed edge.
[0009] In some feasible embodiments, when calculating the structural bias characteristics, for any of the state transition edges... And the behavior state node of the state transition edge. and behavior state nodes Traversing the heterogeneous node topology graph yields a set of commonly associated project nodes: ; in, Represents the behavior state node Establish a set of project nodes with associated edges. Represents the behavior state node Establish a set of project nodes with associated edges. This represents a set of items that are connected to two state nodes simultaneously. Calculate associated edges Structural bias characteristics: ; in, Represents the behavior state node With project nodes The weight of the associated edges between them Represents the behavior state node With project nodes The weights of the associated edges between them.
[0010] In some feasible embodiments, when calculating the time bias characteristics, for any of the aforementioned state transition edges Obtain its corresponding behavior state node Corresponding behavior timestamp , with behavior state nodes Corresponding behavior timestamp Calculate the time interval between adjacent behavioral states. The time offset characteristics are calculated using a time decay function based on the time interval: ; in, This represents the time decay coefficient, used to control the effect of time intervals on the degree of bias; When calculating the behavioral inertia bias characteristics, the number of repetitions of the same target behavioral state during continuous state transitions is counted. : ; in, This represents the number of behavioral states in a continuous state transition path. This represents a conditional function that takes the value 1 if the condition is true and 0 otherwise. Calculate the behavioral inertia bias characteristics based on the number of repetitions and the number of consecutive states: ; in, This represents the behavioral inertia bias feature corresponding to continuous state transition edges.
[0011] In some feasible embodiments, the state transition probability between two behavioral states is used as the initial weight of the state transition edge of the corresponding node according to the transition probability matrix; for each state transition edge, the weight adjustment coefficient of the corresponding state transition edge is calculated by weighted fusion according to the structural bias feature, time bias feature and behavioral inertia bias feature, and the initial weight of the state transition edge is corrected according to the weight adjustment coefficient.
[0012] In some feasible embodiments, positive sample items are determined based on the actual interaction items corresponding to the debiased behavior state sequence, and the credibility of the behavior path is calculated based on the weight of the debiased state transition edge. Interaction items corresponding to low-credibility behavior paths are filtered out, and interaction items that meet the credibility conditions are retained as positive samples. Obtain a set of historical exposures of target users that did not generate interactions, and then filter out interactive items that did not appear in the positive sample set as negative samples.
[0013] In some feasible embodiments, in the constructed recommendation model based on long short-term memory network, based on the gating unit in the recommendation model architecture, each input vector is processed for time-step state update. In each time step, the hidden state corresponding to the current time step is generated by combining the current input vector and the hidden state and memory state corresponding to the previous time step, and the hidden state vector sequence is output. The hidden state vector corresponding to the final time step is used as the user's behavior representation.
[0014] In some feasible embodiments, during the training of the recommendation model, the user behavior representation output by the recommendation model is combined with the interaction items corresponding to the positive and negative samples to form positive and negative item matching pairs for the same user; the matching score between the user behavior representation and the behavior representation output by the positive sample interaction item, and the matching score between the user behavior representation and the behavior representation output by the negative sample interaction item are calculated respectively, and ranking constraints are constructed based on the relative size of the matching scores of the positive and negative samples; wherein the matching score corresponding to the positive sample is higher than the matching score corresponding to the negative sample. The training loss of the recommendation model is calculated based on the ranking constraints. The parameters of the recommendation model are adjusted through backpropagation so that the matching score of the positive sample items is gradually higher than the matching score of the negative sample items. The model training is repeated iteratively until the training loss meets the convergence condition, and the trained recommendation model is obtained.
[0015] (III) Beneficial Effects: Compared with the prior art, this invention has the following beneficial effects: This invention extracts the behavior type, interaction scenario, and time information of each interaction from the historical interaction behavior logs of target users. After integrating these into the interaction behavior states according to combination rules, it further constructs a heterogeneous node topology graph based on the behavior state transition relationship. By combining structural bias, time bias, and behavior inertia bias to adjust the state association weights, it achieves de-biasing of the target user's historical interaction behavior and reduces data deviations caused by project association, time changes, and repetitive behaviors.
[0016] Building recommendation model training samples based on the debiased behavior state sequence can preserve the time dependence and long-term behavior trends in user behavior, enabling the recommendation model to be trained on data that better matches the evolution of user behavior, thereby improving the accuracy of user behavior representation and the matching degree of recommendation results. Attached Figure Description
[0017] Figure 1 This is a schematic diagram of the sample selection and recommendation model construction method based on causal inference bias removal provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the sample selection process for causal inference-based bias removal in the first stage of the sample selection and recommendation model construction method provided in the embodiments of the present invention. Figure 3 This is a schematic diagram of the process of building a recommendation model and performing sample update training in the second stage of the sample selection and recommendation model construction method based on causal inference and bias removal provided in the embodiments of the present invention. Figure 4This is a schematic diagram of the heterogeneous node topology graph constructed by the sample selection and recommendation model construction method based on causal inference and bias removal provided in the embodiments of the present invention. Detailed Implementation
[0018] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0019] It should be noted that, where there is no conflict, the features in the embodiments of the present invention can be combined with each other.
[0020] refer to Figure 1 and Figure 2 In embodiments of the present invention, the overall execution method is divided into two stages.
[0021] In the first phase (sample selection), the historical behavior logs of individual target users recorded on the platform are initially constructed into a topology graph of heterogeneous nodes (two types of nodes: behavior state nodes and item nodes) representing individual user behaviors. Then, edge weight correction is performed. The reason only edge weights are adjusted, and not the nodes themselves, is that the two types of behavior states (defined by combinations of interactive behaviors and interactive scenarios) and interactive items are objectively real; only the strength of the connections between nodes can be adjusted. For details on behavior state nodes, please refer to Table 1 below.
[0022] Table 1 Rules for Combining Behavioral States Regarding the heterogeneous node topology in the first phase, it is important to understand that if only a single type of node is used (e.g., only behavior and state nodes), a problem arises: the data information related to each type of behavior cannot be separated.
[0023] Taking a typical purchase action as an example, it generally follows this process: view (search page) → click → buy, or view → click (details page) → buy. However, if there are only action state nodes, it's impossible to distinguish which item caused the click, whether the click was triggered by the context, or whether the view and buy belong to the same item. refer to Figure 2First, when acquiring the historical interaction behavior log data of the target user, the system first filters out all the user's interaction records from the log table by user ID. Each record is essentially a piece of event data. For example, a record can be understood as "the user performs a certain behavior on a certain product at a certain time, accompanied by a certain scene".
[0024] Therefore, each piece of raw data contains at least four fields: The project identifier indicates who the target audience is; Behavior type indicates what action was performed; Time information, indicating when it happened; Interaction scenario information indicates the environment in which the event occurred.
[0025] From an engineering execution perspective, this step corresponds to retrieving multiple log records scattered across the database by user dimension and forming an event list aggregated by user. For example, the system will centralize all clicks, views, and purchase records of the same user, instead of storing them in different tables or different time slices.
[0026] After aggregation, the operating system performing the operation sorts these events according to the time field, for example, placing the earliest behavior first and the latest behavior last, thus forming a strict time chain. The essence of this time chain is that each record can be regarded as a state point on a time index, that is, what behavior occurred at the first time point, what behavior occurred at the second time point, and so on.
[0027] This process can be understood as transforming a log stack into a timeline sequence, for example: t1: Browsing product A (mobile phone, homepage scenario); t2: Click on the product details page for product A (mobile device, search scenario); t3: Add to cart (mobile, mobile device, product page); t4: Browse product B (recommendation stream).
[0028] Each record in this sequence has a fixed time position; they are no longer isolated events, but rather a continuous trajectory of behavior with a sequential relationship.
[0029] After obtaining the behavioral sequence, the system further transforms the sequence structure into a heterogeneous node topology graph. Based on the above analysis, it can be seen that the core of this step is not adding data, but rather re-expressing the relationships.
[0030] In some embodiments of the present invention, the information in the sequence is first categorized, and the item identifiers are uniformly mapped to a set of item nodes. For example, product A and product B become nodes in the graph, and the behavior type is mapped to a set of behavior state nodes, such as browsing, clicking, and adding to cart, which become different types of nodes. The interaction scenario information is mapped to a set of interaction scenario nodes, and environmental factors such as mobile devices, homepage, and recommendation stream also exist as nodes.
[0031] After generating the various types of nodes described above, the edge relationships are constructed. The heterogeneous node topology graph generated in this step can be referenced. Figure 4 The edges are not constructed by random connections, but are generated strictly according to the three relationships in the original sequence.
[0032] The first type of relationship is the behavior association relationship, which comes from the same record. That is, a certain behavior occurs on a certain item. Therefore, it is represented in the diagram as the connection between the behavior status node and the item node. For example, "click" is connected to "product A", and "add to cart" is connected to "product A".
[0033] The second type of relationship is state transition, which comes from continuous actions at adjacent time points, such as continuous changes from t1 to t2, t2 to t3. Therefore, in the graph, it is represented by directed connections between action state nodes, such as "browse" pointing to "click", and "click" pointing to "add to cart", indicating the advancement of action state over time.
[0034] After completing the construction of the heterogeneous node topology graph, we began to analyze the changing patterns of user behavior over time.
[0035] Since the previous graph structure only connected the relationships between user behaviors, but did not yet know which transitions between behaviors were stable and which were just accidental, it is necessary to further calculate the transition probabilities between behavior states and analyze whether there are causal bias behaviors in the graph structure.
[0036] Before constructing the topology graph, we need to look at the calculation process of the state transition probability function. Referring to Markov chains, the formula for calculating the state transition probability function is as follows: ; in, The number of state transitions. This represents the total number of occurrences of the preceding state.
[0037] Specifically, taking Table 1 above as an example, let's assume the user behavior is: browsing the search page → clicking the details page → purchasing the recommendation page; this can be transformed into: S1→S2→S3; and two state transition relationships can be constructed: S1→S2 and S2→S3.
[0038] For the sequence of behavioral states during transformation, iterate through adjacent states. Statistical states. , transferred to Number of times Then, normalization is performed using each state as a starting point:
[0039] ; in, For state Transferred to The number of times; For state Transferred to The probability is used to obtain the state transition matrix.
[0040] Regarding the heterogeneous node topology, project nodes are constructed based on the interaction item identifiers in the historical interaction records, and behavior state nodes are constructed based on the different state types in the behavior state sequence, forming a heterogeneous node topology that includes project nodes and behavior state nodes.
[0041] Next, based on the mapping relationship between behavioral states and corresponding interactive items, state-item association edges (hereinafter referred to as association edges) are established to represent the association relationship between behavioral states and interactive items. Based on the transition relationship between adjacent behavioral states in the behavioral state sequence, state transition edges are established to represent the temporal evolution relationship between user behavioral states. In the subsequent bias feature calculation, state-item association edges (association edges) correspond to structural bias, and state transition edges correspond to temporal bias and behavioral inertia bias; specific calculations are described below.
[0042] Regarding the state transition edge : → Based on the behavioral state transition probability matrix generated in the previous step, where... For state Transferred to The probability of transition is used as the initial weight of the transition edge for that state. This weight represents the probability of the user's historical behavior. → The intensity of the original occurrence of this behavioral evolutionary relationship.
[0043] When calculating structural bias characteristics, for any state transition edge And the initial behavior of the state transition edge is the state node. and target behavior state node Traverse the set of item nodes that have established association edges with the two behavior state nodes respectively, calculate the intersection relationship of the corresponding associated items of the two behavior state nodes, and obtain the set of commonly associated items: ; in, Represents the behavior state node Establish a set of project nodes with associated edges. Represents the behavior state node Establish a set of project nodes with associated edges. This represents a set of items that are connected to two state nodes simultaneously.
[0044] The reason for calculating structural bias first is that many behaviors in a recommendation system do not necessarily stem from genuine interests, but may originate from structural concentration. For example, certain popular products naturally receive a lot of exposure, which is the biased data mentioned above. Even if users do not have genuine interest in this type of data, they will click on it frequently. Therefore, the system needs to determine whether the connections in the graph structure are too concentrated.
[0045] Based on the edge weights corresponding to each item node in the shared set of items, calculate the structural bias features corresponding to the state transition edges: ; in, Represents the behavior state node With project nodes The weight of the associated edges between them Represents the behavior state node With project nodes The weights of the associated edges between them.
[0046] This process statistically analyzes the connection distribution between project nodes and behavioral state nodes. If the number of behavioral state nodes connected to certain project nodes is significantly higher than that of other nodes, then the project node is identified as a node with high connection concentration.
[0047] Similarly, if the project nodes or interactive scene nodes connected to certain behavioral state nodes are too concentrated in the overall distribution, then the behavioral state node is judged as a highly concentrated node. This type of statistics essentially corresponds to the non-uniformity of the connection relationship in the graph structure. Therefore, the value of the structural bias feature is directly calculated from the degree of concentration of the node connection distribution.
[0048] When calculating the time bias feature, for the state transition edge Obtain its corresponding behavior state node Corresponding behavior timestamp , with behavior state nodes Corresponding behavior timestamp Calculate the time interval between adjacent behavioral states. The time offset characteristics are calculated using a time decay function based on the time interval: ; in, This represents the time decay coefficient, used to control the effect of time intervals on the degree of bias.
[0049] If a large number of transfers occur within a very short time interval, the user behavior exhibits a time-dense bias; if the transfers are mainly distributed over a longer time interval, it exhibits a time-sparse bias. This is because the time bias characteristic is composed of the proportion of transfers within different time intervals.
[0050] Finally, the behavioral inertia bias feature is calculated, which statistically analyzes whether repeated paths occur during the transition of behavioral state nodes.
[0051] If behavioral states frequently repeat between adjacent time steps, a self-transition structure is formed. If certain behavioral paths (e.g., specific combinations of state sequences) repeatedly appear in the graph, a high-frequency repetitive path structure is formed.
[0052] Therefore, the behavioral inertia bias characteristic consists of two parts; One part is the proportion of self-transition of behavioral states, that is, the proportion of the same state before and after the state transition; The other part is the percentage of repeated paths, which is the frequency of the same state appearing in the global transition set.
[0053] Specifically, when calculating the behavioral inertia bias characteristics, the number of repetitions of the same target behavioral state during continuous state transitions is counted. : ; in, This represents the number of behavioral states in a continuous state transition path. This represents a judgment function, which takes a value of 1 when the condition is true and a value of 0 otherwise; the behavioral inertia bias characteristic is calculated based on the number of repetitions and the number of consecutive states. ; After completing the calculation of the state transition probability function and the three types of bias features, the system begins to recalibrate the edge weights in the heterogeneous node topology graph. This step essentially addresses the question of which behavioral connections in the graph are driven by genuine interest, and which are merely biased strong connections caused by biased outputs such as popular exposure, short-term repetitive operations, or behavioral inertia.
[0054] The heterogeneous node topology graph constructed earlier simply records all behavioral relationships as they are, but there will be many natural deviations in real user behavior.
[0055] For example, a popular product might generate a large number of browsing edges due to its high exposure; some users might repeatedly refresh the same page within a short period; and certain behavioral paths might be repeatedly executed due to habitual actions. All of these factors can cause the weight of certain edges to be abnormally amplified, thus requiring a recalculation of the importance of each edge in the graph.
[0056] First, the state transition probability function is used as the base value for edge weight updates because the state transition probability already describes the stability of a certain behavioral state transitioning to the next behavioral state. Then, three types of bias corrections are added, as detailed in Table 2 below.
[0057] Table 2 Behavioral Topology Graph Update Parameters and Corresponding Update Methods After the above three types of bias corrections, the edge weight distribution in the entire heterogeneous node topology graph is updated again, and the behavioral path structure is reorganized.
[0058] This weakens strong edges that were originally formed due to popular exposure, short-term repetitive behavior, or habitual cycles; while paths with stable behavior conversion capabilities are preserved or even strengthened. The resulting behavioral topology is no longer just a raw behavior record diagram, but a behavioral relationship structure that has been causally debiased.
[0059] After completing the construction of the behavior topology graph after causal bias correction for a single user, now refer to Figure 2 Then, the corresponding recommendation model construction process is executed in the second stage.
[0060] First, path traversal is performed along the state transition edges between behavioral state nodes, concatenating behavioral state nodes that can form stable connections between consecutive time steps into a set of behavioral paths. During path generation, each path corresponds to a set of continuous behavioral state transition structures, and a corresponding project node is associated at the end of the path. For each state transition edge in the path, the system reads the edge weight result updated in the previous stage and accumulates the edge weights in the entire path to form the overall causal weight of the path.
[0061] If the behavioral state transitions in a certain path maintain a high transition probability over a long period of time, and the project nodes corresponding to the path do not exhibit obvious structural concentration, and the behavioral interval distribution in the path is stable with a low proportion of repeated cycles, then the path will still maintain a high causal weight after the accumulation of edge weights.
[0062] The reason why items corresponding to the endpoints of high causal weight paths are considered positive samples is that if a behavioral path still maintains a high weight after bias removal, it means that the path simultaneously meets several conditions: Behavioral state transitions are stable; The concentration of structures is not due to popular nodes; This does not constitute abnormally repetitive behavior over a short period of time; It does not belong to a high-frequency inertial loop path.
[0063] In other words, behavioral changes in such paths more closely resemble the actual evolution of interests, and the endpoint of the path best represents the final point of interest formed by that path. Conversely, some negative samples come from the endpoint items corresponding to paths with low causal weight.
[0064] If some paths, despite having a large number of behavioral records, exhibit behavior states that remain within repetitive browsing, self-looping transitions, or short-term high-frequency repetitive structures for extended periods, then these paths have already undergone behavioral inertia bias compression and structural bias weakening in the previous stage, resulting in a decrease in the overall causal weight of the path. The system will mark the corresponding project nodes at the end of such low-causal-weight paths as candidate negative sample projects.
[0065] Based on the candidate negative samples, and further combined with the platform's exposure statistics, projects that have been exposed multiple times but have never generated behavioral conversions are screened.
[0066] Since these projects have already achieved sufficient exposure, but users have not yet clicked, added to cart, or taken any further action, the system merges them with projects corresponding to low causal weight paths into a negative sample set.
[0067] Regarding the calculation of path credibility mentioned above, in some embodiments, the behavioral path credibility is calculated based on the weights of each state transition edge in the path, using cumulative path credibility, such as geometric mean or multiplication. It is considered that if the credibility of any state transition segment in a path is very low, then the overall path credibility should also decrease.
[0068] Finally, the system constructs positive sample items corresponding to the endpoints of high causal weight paths, constructs negative sample items corresponding to the endpoints of low causal weight paths and high exposure low behavioral feedback items, and establishes user-item training sample pairs with the target users for subsequent training of the long short-term memory network recommendation model.
[0069] After constructing the positive and negative samples, the system begins to build a Long Short-Term Memory (LSTM) network recommendation model. The reason for using a LSTM network is that user behavior itself has obvious temporal continuity.
[0070] After performing the bias removal process on the user's historical interaction behavior as described above, the bias removal behavior state sequence corresponding to the user is obtained as follows: (Click on the homepage) → (Add to Favorites) → (Click to search again); The weights of the state transition edges between each state have been corrected based on structural bias, temporal bias, and behavioral inertia bias, resulting in corresponding weights. These weights represent the reliability of the relationship between the corresponding states; the larger the weight, the better the state transition reflects the evolution of the user's true interests.
[0071] Since Long Short-Term Memory (LSTM) networks cannot directly process discrete state representations, it is first necessary to convert the behavioral states into continuous vector representations. Let the behavioral state mapping function be... The sequence of behavioral states is mapped to the sequence of input vectors:
[0072] ; The data is then input into the Long Short-Term Memory (LSTM) network in chronological order.
[0073] For the At each time step, the network reads the current input vector. Hidden state in the previous time step and memory state The information is updated through the input gate, forget gate, and output gate, respectively. The calculation for the input gate is as follows:
[0074] ; in, For input weights, In hidden state, For bias terms, This is the Sigmoid function. The input gate here is mainly used to determine the current behavioral state and how much of it needs to be written to memory.
[0075] The calculation principle for the forget gate is the same as above: ; The forget gate is used to determine the proportion of historical memory retained. When a certain historical behavior state contributes little to the current interest, the corresponding forget gate output is smaller, thereby reducing its impact.
[0076] The following calculations are then performed on the candidate memory states: ; Then, the input gate and forget gate are used together to update the new memory state: ; Among them This represents element-wise multiplication.
[0077] Final output: ; This is the current hidden state.
[0078] Based on the examples above, when the input... generate This is used to indicate the user's interest status after clicking on the homepage; Next, enter (corresponding to the above) )generate The act of collecting further enhances user interest, making... Integrate click and save behaviors; Finally, enter (corresponding to the above) Generate the final hidden state Since historical information from the entire debiasing behavior state sequence has been accumulated and integrated, therefore... As a representation of current user behavior . Used to represent the current user's true interest characteristics.
[0079] During model training, after constructing positive and negative samples, the user behavior representation is matched with both positive and negative sample items. Regarding the matching score between the user behavior representation (vector) and the corresponding interaction items for each type of sample, in some embodiments, the vector dot product is used for calculation.
[0080] ; Among them This is the item vector of the interactive item corresponding to the sample.
[0081] Furthermore, in some embodiments, a ranking loss function is used for model training, such as BPR loss: ; Among them For items corresponding to positive samples For items corresponding to negative samples This is the Sigmoid function. This loss function aims to maximize the difference in matching scores between positive and negative samples.
[0082] For example, if negative samples actually get higher matching values than positive samples in a certain training session, it means that the current parameters of the model are still unable to correctly distinguish between real interests and noisy behavior. In this case, the system will update the gating parameters and item encoding parameters in the Long Short-Term Memory network in reverse.
[0083] In summary, by jointly training the matching relationship between positive and negative samples, the recommendation model can gradually learn the stable matching relationship between the user's real behavior state and the items, thereby improving the adaptability of the recommendation results to changes in the user's actual interests and enhancing the accuracy and stability of the recommendation results.
[0084] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. The scope of protection of the present invention is defined by the claims. Similarly, any equivalent structural changes made based on the description and drawings of the present invention should also be included within the scope of protection of the present invention.
Claims
1. A method for sample selection and recommendation model construction based on causal inference and bias removal, characterized in that, The execution steps include the following: Extract interaction items, behavior types, interaction scenarios, and time information from the target user's historical interaction behavior logs. Generate behavior states based on the combination relationship between the behavior types and interaction scenarios. Count the number of transitions between adjacent behavior states and generate a transition probability matrix of behavior states after normalization. Construct a heterogeneous node topology graph that includes project nodes and behavior state nodes; wherein, establish association edges between the project nodes and behavior state nodes, and establish state transition edges between adjacent behavior state nodes; Based on the association between the behavioral state nodes at both ends of the state transition edge and the same project node, the structural bias feature generated by the project node is calculated; the time bias feature is calculated based on the time interval between the corresponding behavioral states of the state transition edge; and the behavioral inertia bias feature is calculated based on the repetition of the corresponding behavioral states of the continuous state transition edge. The initial weights of each state transition edge are determined according to the transition probability matrix. The initial weights are corrected after calculating the weight adjustment coefficients according to the bias features of each item. The weights of the state association edges after debiasing are obtained. The association strength between adjacent behavioral states in the behavioral state sequence is adjusted to generate a debiased behavioral state sequence. The debiased behavioral state sequence is used as input and the corresponding interaction item is used as the prediction target to construct training samples. The biased behavior state sequence is input into the long short-term memory network recommendation model. The state is updated at each time step through the gating unit to generate the user behavior representation. The interactive items in the training samples are vector encoded, the matching score between the user behavior representation and the encoded vector is calculated, and the model parameters are optimized according to the matching score to obtain the trained recommendation model.
2. The sample selection and recommendation model construction method based on causal inference bias removal according to claim 1, characterized in that, The system acquires the target user's historical interaction behavior logs, parses each log entry to identify the interaction item identifier, behavior type, interaction scenario, and behavior timestamp, and formats the parsed fields to generate interaction records corresponding to each interaction event. Based on the behavior timestamp in each interaction record, the system sorts the logs to generate a sequence of the target user's historical interaction records. Each interaction record corresponds to one interaction event between the user and the interaction item.
3. The sample selection and recommendation model construction method based on causal inference bias removal according to claim 2, characterized in that, Based on the behavior type and interaction scenario in each interaction record, the behavior type and interaction scenario are combined and encoded into the corresponding behavior state according to the predefined behavior state mapping rules. The generated behavioral states are organized sequentially according to the order of the interaction records in the historical interaction record sequence to form a behavioral state sequence corresponding to the target user. Traverse adjacent combinations of behavioral states in the sequence of behavioral states, count the number of transitions between each behavioral state, and obtain a state transition count matrix; Based on the total number of transitions corresponding to each initial behavioral state, the state transition number matrix is normalized to generate a behavioral state transition probability matrix.
4. The sample selection and recommendation model construction method based on causal inference bias removal according to claim 1, characterized in that, Traverse the historical interaction records, and map the interaction items of the same type to the same item node. At the same time, traverse the behavior state sequence, and map the behavior state of the same type to the same behavior state node, and generate item node set and behavior state node set respectively. Based on the interactive items and corresponding behavior states in each interactive record, the corresponding item nodes and behavior state nodes are located respectively, and an association edge is established between the two types of nodes, so that each interactive record is mapped to the association relationship between item nodes and behavior state nodes. According to the order of the behavior state sequence, read two adjacent behavior states in sequence and locate the corresponding behavior state node; establish a state transition edge between the previous behavior state node and the next behavior state node; the state transition edge is a directed edge.
5. The sample selection and recommendation model construction method based on causal inference bias removal according to claim 1, characterized in that, When calculating the structural bias features, for any of the state transition edges And the behavior state node of the state transition edge. and behavior state nodes Traversing the heterogeneous node topology graph yields a set of commonly associated project nodes: ; in, Represents the behavior state node Establish a set of project nodes with associated edges. Represents the behavior state node Establish a set of project nodes with associated edges. This represents a set of items that are connected to two state nodes simultaneously. Calculate associated edges Structural bias characteristics: ; in, Represents the behavior state node With project nodes The weight of the associated edges between them Represents the behavior state node With project nodes The weights of the associated edges between them.
6. The sample selection and recommendation model construction method based on causal inference bias removal according to claim 5, characterized in that, When calculating the time bias characteristics, for any of the aforementioned state transition edges Obtain its corresponding behavior state node Corresponding behavior timestamp , with behavior state nodes Corresponding behavior timestamp Calculate the time interval between adjacent behavioral states. The time offset characteristics are calculated using a time decay function based on the time interval: ; in, This represents the time decay coefficient, used to control the effect of time intervals on the degree of bias; When calculating the behavioral inertia bias characteristics, the number of repetitions of the same target behavioral state during continuous state transitions is counted. : ; in, This represents the number of behavioral states in a continuous state transition path. This represents a conditional function that takes the value 1 if the condition is true and 0 otherwise. Calculate the behavioral inertia bias characteristics based on the number of repetitions and the number of consecutive states: ; in, This represents the behavioral inertia bias feature corresponding to continuous state transition edges.
7. The sample selection and recommendation model construction method based on causal inference bias removal according to claim 6, characterized in that, Based on the transition probability matrix, the state transition probability between the two behavioral states is used as the initial weight of the state transition edge of the corresponding node; For each state transition edge, a weighted fusion is performed based on the structural bias features, temporal bias features, and behavioral inertia bias features to calculate the weight adjustment coefficient of the corresponding state transition edge, and the initial weight of the state transition edge is corrected based on the weight adjustment coefficient.
8. The sample selection and recommendation model construction method based on causal inference bias removal according to claim 1, characterized in that, Positive sample items are determined based on the actual interaction items corresponding to the debiased behavior state sequence, and the credibility of the behavior path is calculated based on the weight of the debiased state transition edge. Interaction items corresponding to low-credibility behavior paths are filtered out, and interaction items that meet the credibility conditions are retained as positive samples. Obtain a set of historical exposures of target users that did not generate interactions, and then filter out interactive items that did not appear in the positive sample set as negative samples.
9. The sample selection and recommendation model construction method based on causal inference bias removal according to claim 1, characterized in that, In the constructed recommendation model based on long short-term memory network, based on the gating unit in the recommendation model architecture, each input vector is processed for time-step state update. In each time step, the hidden state corresponding to the current time step is generated by combining the current input vector and the hidden state and memory state corresponding to the previous time step, and the hidden state vector sequence is output. The hidden state vector corresponding to the final time step is used as the user's behavior representation.
10. The sample selection and recommendation model construction method based on causal inference bias removal according to claim 8, characterized in that, In the training of the recommendation model, the user behavior representation output by the recommendation model is combined with the corresponding interaction items of the positive and negative samples to form positive and negative item matching pairs for the same user. The matching score between the user behavior representation and the behavior representation output by the positive sample interaction item, and the matching score between the user behavior representation and the behavior representation output by the negative sample interaction item are calculated respectively. Ranking constraints are constructed based on the relative size of the matching scores of the positive and negative samples. The matching score of the positive sample is higher than the matching score of the negative sample. The training loss of the recommendation model is calculated based on the ranking constraints. The parameters of the recommendation model are adjusted through backpropagation so that the matching score of the positive sample items is gradually higher than the matching score of the negative sample items. The model training is repeated iteratively until the training loss meets the convergence condition, and the trained recommendation model is obtained.