User Portrait Label Generation Method and System Based on a Homogeneous Transition Model

By building a state transfer probability matrix of user behavior based on the time-horizontal transfer model, a dynamically updated user portrait tag is generated, which solves the problem of overly one-sided user portrait generation in the prior art, and achieves a more accurate and reliable user behavior feature description.

CN119989032BActive Publication Date: 2025-06-24CHINA RONGXIN CLOUD TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510466876.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-06-24
Estimated Expiration
2045-04-15

AI Technical Summary

Technical Problem

The prior art fails to fully explore the complex relationship between user behavior operation types when generating user portraits, and fails to consider the dynamic evolution characteristics of user behavior over time, resulting in the generated portrait labels being too one-sided and cannot accurately reflect the real patterns and dynamic changes of user behavior.

Method used

Using a method based on the time-horizontal transfer model, by obtaining the user behavior trajectory data of the target user in the preset time period, constructing a state transfer probability matrix, determining the absorbed state behavior operation type set, generating a picture tag set, and matching it with the preset picture tag knowledge base, and outputting a dynamically updated picture tag sequence.

Benefits of technology

This method can comprehensively capture the timing correlation between user behavior, accurately describe the user's stable characteristics in the specified behavior dimension, avoid misjudgments caused by short-term behavior fluctuations in users, improve the scientificity and reliability of the image tag generation process, and enable the generated image tag to more accurately reflect the user's real behavior characteristics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119989032B_ABST
    Figure CN119989032B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for generating user portrait tags based on a stationary transition model. First, behavioral trajectory data including behavioral operation types and context attributes within a preset time period of a target user is obtained. Then, a state transition probability matrix is constructed based on the temporal relevance of the behavioral operation types to reflect the transition probability between behaviors. Next, according to the steady-state convergence of the matrix state nodes, a set of absorbing state behavioral operation types that meet specific conditions is determined. Based on this set of absorbing state behavioral operation types and the corresponding context attributes, a set of portrait tags is generated to describe the stable characteristics of the target user in the specified behavioral dimension. Finally, the set of portrait tags is matched and verified with a preset knowledge base, and a dynamically updated portrait tag sequence is output, which can comprehensively and accurately generate portrait tags reflecting the stable characteristics of the user and improve the quality of the user portrait.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of digital services, and in particular, to a method and system for generating user portrait tags based on a homogeneous transition model. Background Art

[0002] In today's digital age, user portrait technology plays a crucial role in enabling enterprises to deeply understand user needs, provide personalized services, and formulate precise marketing strategies. By constructing user portraits, enterprises can abstract and quantify complex and diverse user behaviors and characteristics, thereby better grasping user behavior patterns and preferences.

[0003] However, many existing methods simply collect and statistically analyze user behavior data, failing to fully explore the potential complex relationships between different types of user behavior operations. For example, only focusing on the occurrence frequency of a single behavior while ignoring the temporal correlation between different behaviors, resulting in overly one-sided user portraits that cannot accurately reflect the true patterns and dynamic changes of user behavior, leading to an insufficiently in-depth and comprehensive understanding of user behavior.

[0004] In addition, most existing technologies do not take into account the dynamic evolution characteristics of user behavior over time. User behavior patterns are not static but change with factors such as time and environment. However, traditional methods often generate portrait tags in a static manner and cannot timely capture the dynamic changes in user behavior, causing the generated portrait tags to lose their accurate description of the user's true behavior after a certain period of time, reducing the effectiveness and value of user portraits in practical applications. Summary of the Invention

[0005] In view of the above-mentioned problems, in combination with the first aspect of the present invention, embodiments of the present invention provide a method for generating user portrait tags based on a homogeneous transition model, the method comprising:

[0006] Obtaining user behavior trajectory data of a target user within a preset time period, the user behavior trajectory data including types of behavior operations and corresponding context attributes at multiple discrete time points;

[0007] Based on the temporal correlation of each type of behavior operation in the user behavior trajectory data, constructing a state transition probability matrix corresponding to the target user, where each element in the state transition probability matrix represents the transition probability of the target user migrating from a first type of behavior operation to a second type of behavior operation;

[0008] Determine the set of absorbing state behavior operation types corresponding to the target user according to the steady-state convergence of each state node in the state transition probability matrix, and each behavior operation type in the set of absorbing state behavior operation types satisfies the state transition probability threshold condition within a continuous time period;

[0009] Generate a set of portrait tags for the target user based on the behavior operation types and corresponding context attributes in the set of absorbing state behavior operation types, and each portrait tag in the set of portrait tags is used to describe the stable characteristics of the target user in a specified behavior dimension;

[0010] After matching and verifying the set of portrait tags with a preset portrait tag knowledge base, output the dynamically updated portrait tag sequence corresponding to the target user.

[0011] On the other hand, an embodiment of the present invention further provides a user portrait tag generation system based on a time-homogeneous transition model, including a processor and a machine-readable storage medium. The machine-readable storage medium is connected to the processor. The machine-readable storage medium is used to store programs, instructions, or codes, and the processor is used to execute the programs, instructions, or codes in the machine-readable storage medium to implement the above method.

[0012] Based on the above aspects, the embodiments of the present application can comprehensively capture the temporal correlation between user behaviors by obtaining the detailed behavior trajectory data of the target user within a preset time period, including behavior operation types and their context attributes, and constructing a state transition probability matrix based on the time-homogeneous transition model. On this basis, determining the set of absorbing state behavior operation types according to the state transition probability matrix, and then generating a set of portrait tags, can accurately describe the stable characteristics of the user in a specified behavior dimension, effectively avoiding misjudgment caused by short-term behavior fluctuations of the user. Next, using the time-homogeneous transition model, converting user behavior data into a state transition probability matrix, and determining the absorbing state behavior operation type by analyzing the steady-state convergence of the state nodes in the state transition probability matrix greatly improves the scientificity and reliability of the portrait tag generation process, making the generated portrait tags more accurately reflect the true behavior characteristics of the user. Finally, after generating the set of portrait tags, matching and verifying with a preset portrait tag knowledge base, and then outputting the dynamically updated portrait tag sequence, ensures that the portrait tags can not only reflect the latest behavior characteristics of the user at present, but also fit in with the existing knowledge base standards and knowledge systems. On the one hand, the dynamic update mechanism enables the portrait tags to be adjusted in time with the change of user behaviors, always maintaining real-time tracking and accurate portrayal of user behaviors; on the other hand, the matching and verification process ensures the consistency and accuracy of the generated portrait tags within the entire system knowledge framework, improving the generality and usability of the portrait tags in different application scenarios. Description of the Drawings

[0013] Figure 1 It is a schematic diagram of the execution process of the user portrait label generation method based on the homogeneous transition model provided by an embodiment of the present invention.

[0014] Figure 2 It is a schematic diagram of exemplary hardware and software components of the user portrait label generation system based on the homogeneous transition model provided by an embodiment of the present invention. Detailed implementation manners

[0015] The present invention will be specifically described below in conjunction with the accompanying drawings of the specification. Figure 1 It is a schematic diagram of the process of the user portrait label generation method based on the homogeneous transition model provided by an embodiment of the present invention. The user portrait label generation method based on the homogeneous transition model will be introduced in detail below.

[0016] Step S110: Obtain the user behavior trajectory data of the target user within a preset time period. The user behavior trajectory data includes the types of behavior operations and corresponding context attributes at multiple discrete time points.

[0017] For example, taking a financial service platform as an example, financial institutions usually record various operation behaviors of users. Assume that the preset time period is the past year and the target user is a credit card user.

[0018] First, the original behavior event sequence can be extracted from the credit card interaction log of this user. For example, the event trigger timestamp may include the specific date and time when the user uses the credit card for each transaction, such as 10:30:00 on March 15, 2022; the event operation type identifier can be the transaction type, such as consumption, repayment, cash withdrawal, etc.; the event context parameters can include the transaction amount, transaction location (such as the name or address code of a certain shopping mall), the industry to which the transaction merchant belongs (such as catering, retail, etc.).

[0019] Then, perform data cleaning on the original behavioral event sequence. Event records with timestamps earlier than the system recording start time (assuming the system starts recording from January 1, 2020) or later than the current system time (assuming the current time is July 2023) are detected and marked as abnormal timestamp events. For example, there may be incorrect time records due to system failures. Event records with the event operation type identifier being empty (possibly due to data transmission errors resulting in some transaction records having no clear operation type) or not conforming to the preset coding specifications (such as the coding being tampered with or wrongly entered) are marked as invalid operation type events. Traverse every two consecutive event records. If the timestamp interval between these two consecutive event records is less than the preset minimum operation interval (for example, for credit card transactions, normally the interval between two transactions is not less than 1 minute, but if it is less than 1 minute, it may be a duplicate record), then it is marked as a repeated trigger event. Remove these abnormal timestamp events, invalid operation type events, and repeated trigger events from the original behavioral event sequence to generate an intermediate cleaned event sequence. Then, perform integrity verification on the remaining event records in the intermediate cleaned event sequence. For example, if a certain transaction record is found to be missing the context parameter of the transaction amount, complete its default value (which can be completed according to the average transaction amount of this user or the common amount of similar transactions) to generate the cleaned behavioral event sequence.

[0020] Segment and cut the cleaned behavioral event sequence according to the preset time window granularity. Assume that according to the distribution density of the event trigger timestamps in the cleaned behavioral event sequence, during the peak period of credit card usage (such as holidays or promotional seasons), if the number of event triggers exceeds the density threshold (assuming 50 transactions per day), then reduce the time window granularity to the first sub-granularity (for example, change the window division from weekly to daily); during the off-peak period of credit card usage (such as the non-peak hours on weekdays), if the number of event triggers is below the sparsity threshold (assuming 5 transactions per day), then expand the time window granularity to the second sub-granularity (for example, change the window division from daily to bi-weekly). Based on the adjusted time window granularity, divide the cleaned behavioral event sequence into multiple consecutive or partially overlapping time window intervals. Count the frequencies of the event operation type identifiers within each time window interval, and take the operation type with the highest frequency as the main behavioral operation type. For example, if the frequency of consumption is the highest within a certain time window, then consumption is the main behavioral operation type of this time window.

[0021] Perform clustering analysis on the event operation type identifiers in each behavior subsequence. Convert the behavior subsequence corresponding to each time window into a sequence vector of event operation type identifiers. For example, [consumption, repayment, consumption, cash withdrawal, repayment] is converted into the corresponding vector representation. Use a preset clustering model (such as the K-means clustering algorithm) to perform unsupervised clustering on the sequence vectors to obtain multiple candidate behavior clusters. Calculate the similarity between the center vector of each candidate behavior cluster and the preset benchmark behavior pattern vector (such as the normal credit card usage behavior pattern vector, including the proportional relationship of consumption, repayment, cash withdrawal, etc.). Determine the candidate behavior cluster with the highest similarity as the main behavior cluster, and use the operation type corresponding to the main behavior cluster as the main behavior operation type. Add the types with the occurrence frequency of operation types in other candidate behavior clusters greater than the auxiliary frequency threshold (assumed to be 3 times) to the set of auxiliary behavior operation types. For example, in addition to consumption as the main behavior operation type, if the cash withdrawal appears more than 3 times in a certain candidate behavior cluster, the cash withdrawal is added to the set of auxiliary behavior operation types. Finally, based on the event context parameters, the main behavior operation type, and the corresponding set of auxiliary behavior operation types, generate the behavior operation type and context attributes corresponding to each discrete time point in the user behavior trajectory data. For example, the behavior operation type at a certain discrete time point is consumption, and the context attributes are the transaction amount of 500 yuan, the transaction location is a large shopping mall, and the industry of the transaction merchant is retail.

[0022] Step S120, based on the temporal correlation of each behavior operation type in the user behavior trajectory data, construct a state transition probability matrix corresponding to the target user. Each element in the state transition probability matrix represents the transition probability of the target user migrating from the first behavior operation type to the second behavior operation type.

[0023] Specifically, according to the unique identifier set of all behavior operation types in the user behavior trajectory data, such as consumption (identified as C), repayment (identified as R), cash withdrawal (identified as W), etc., define the state space dimension of the state transition probability matrix. Here, the state transition probability matrix is a 3×3 matrix (assuming only these three behavior operation types), and the rows and columns respectively correspond to these three behavior operation types.

[0024] Align the discrete time points in the user behavior trajectory data in chronological order to generate the behavior state transition chain of the target user. For example, the sequence of behavior operation types recorded in chronological order is [C, R, C, W, R], which constitutes the behavior state transition chain. Count the number of transitions between every two adjacent behavior operation types in the behavior state transition chain to generate the initial state transition count matrix. Assume that in the entire behavior state transition chain, the number of transitions from consumption (C) to repayment (R) is 5 times, the number of transitions from consumption (C) to cash withdrawal (W) is 3 times, the number of transitions from repayment (R) to consumption (C) is 4 times, etc. Construct the initial state transition count matrix based on these data.

[0025] Perform normalization processing on the initial state transition count matrix. Traverse each row element in the initial state transition count matrix and calculate the sum value of all elements in that row. Assume that the sum value of the elements in a certain row (corresponding to consumption C) is 8 (the number of times to repayment 3 + the number of times to cash withdrawal 5). Compare the sum value with the preset minimum row sum threshold (assume it is 5). If the sum value is less than the preset minimum row sum threshold, add a pseudo-count compensation value (assume the compensation value is 2) to this row to obtain the compensated row elements. Perform probability normalization on the compensated row elements so that the sum of all elements in that row is 1 to obtain the normalized probability values. Round the normalized probability values to the preset decimal precision (assume it is two decimal places) to generate the standardized transition probability matrix. Replace the elements in the standardized transition probability matrix that are lower than the preset probability lower limit value (assume it is 0.05) with the lower limit value to generate the final state transition probability matrix. For example, in this 3×3 matrix, the transition probability corresponding to the element (C, R) is 0.38, and the transition probability corresponding to the element (C, W) is 0.62, etc., indicating the transition probabilities from consumption (C) to repayment (R) and cash withdrawal (W). Replace the elements with zero transition probability values in the state transition probability matrix with the preset minimum probability value (assume it is 0.01) to generate the smoothed target state transition probability matrix to avoid calculation problems caused by zero probability in subsequent calculations.

[0026] Step S130, determine the set of absorbing state behavior operation types corresponding to the target user according to the steady-state convergence of each state node in the state transition probability matrix, and each behavior operation type in the set of absorbing state behavior operation types satisfies the state transition probability threshold condition within a continuous time period.

[0027] Specifically, power iteration calculation can be performed on the state transition probability matrix. Assume the initial state transition probability matrix is:

[0028]

[0029] Perform power iteration calculation on the above initial state transition probability matrix until the maximum difference of the elements in each row of the state transition probability matrix is less than the preset convergence threshold (assumed to be 0.001) to obtain the steady-state distribution vector. After multiple iterations, the steady-state distribution vector is [0.33, 0.33, 0.34].

[0030] Select the state nodes with probability values greater than the preset absorbing state threshold (assumed to be 0.3) from the steady-state distribution vector to generate the candidate set of absorbing state behavior operation types. Here, the probability values corresponding to consumption (C) and cash withdrawal (W) meet the conditions and enter the candidate set of absorbing state behavior operation types.

[0031] For each behavior operation type in the candidate set of absorbing state behavior operation types, verify whether the state self-transition probability of this behavior operation type in a continuous preset number of time periods (assumed to be 3 consecutive months) is continuously greater than the preset stability threshold (assumed to be 0.6). Assume that for consumption (C), the self-transition probabilities in each of these 3 months are 0.7, 0.65, and 0.62 respectively, which meet the conditions; for cash withdrawal (W), the self-transition probabilities are 0.5, 0.55, and 0.58 respectively, which do not meet the conditions. Merge the behavior operation types that are continuously greater than the preset stability threshold (here only consumption C) into the set of absorbing state behavior operation types, and delete the behavior operation types that do not pass the verification in the candidate set of absorbing state behavior operation types (here it is cash withdrawal W).

[0032] Increase the self-transition probability of the state nodes corresponding to each behavior operation type in the set of absorbing state behavior operation types (here it is consumption C) by the preset reinforcement coefficient (assumed to be 1.2), reduce the transition probability value from the absorbing state behavior operation type to the non-absorbing state behavior operation type, and evenly distribute the difference to the migration paths of other absorbing state behavior operation types (here there is only consumption C and no other absorbing state behavior operation types, so this step is only for self-reinforcement in this example). Obtain the adjusted state transition probability matrix. For example, the original transition probability from consumption (C) to repayment (R) is 0.3, and after adjustment, it becomes 0.2 (decreased by 0.1), and the self-transition probability from consumption (C) to itself (C) changes from 0.5 to 0.6 (increased by 0.1). Perform column normalization on the adjusted state transition probability matrix so that the sum of the elements in each column is 1 to obtain the re-normalized state transition probability matrix. Use the re-normalized state transition probability matrix as the updated state transition probability matrix and input it into the portrait label generation process for the next time period. Record the update history of the set of absorbing state behavior operation types for subsequent dynamic correction of the portrait label set.

[0033] Step S140: Based on the behavior operation types in the absorbing state behavior operation type set and the corresponding context attributes, generate a set of portrait tags for the target user. Each portrait tag in the set of portrait tags is used to describe the stable characteristics of the target user in a specified behavior dimension.

[0034] Specifically, the context attribute feature vectors corresponding to each behavior operation type can be extracted from the absorbing state behavior operation type set (here it is consumption C). Suppose the operation frequency of consumption (C) is 10 times per month, the operation duration (which can be understood as the average processing time for each consumption here) is 2 minutes, and the associated resource identifier is bank account A to which the credit card belongs.

[0035] Match the context attribute feature vectors with the tag definition rules in the preset portrait tag knowledge base. In the field of financial service default risk, there may be tag definition rules for consumption behavior in the portrait tag knowledge base, such as high-frequency consumption (number of monthly consumptions greater than 8 times), short-duration consumption (each consumption duration less than 3 minutes), associated with a specific account (the associated account is bank account A), etc. Determine the candidate portrait tag list corresponding to each behavior operation type. Here, the candidate portrait tag list may include tags such as high-frequency consumption, short-duration consumption, and associated with a specific account.

[0036] Perform semantic aggregation on the tags in the candidate portrait tag list, and merge multiple sub-tags describing the same behavior dimension. For example, high-frequency consumption and short-duration consumption both describe the characteristics related to the time dimension of consumption, and can be merged into an aggregated tag of high-frequency short-duration consumption. Generate an aggregated tag set.

[0037] Based on the occurrence frequency of each portrait tag in the aggregated tag set (here it is assumed that the high-frequency short-duration consumption tag has appeared 5 times in past analyses) and the time distribution density (for example, evenly distributed in the past year), calculate the confidence weight of each portrait tag. Suppose according to a specific calculation rule (such as a weighted calculation based on the occurrence frequency and time distribution density), the confidence weight of the high-frequency short-duration consumption tag is 0.8.

[0038] Sort and filter the aggregated tag set according to the confidence weights, and retain the tags with weight values greater than the preset threshold (assumed to be 0.6) to generate a set of portrait tags. Here, the weight of the high-frequency short-duration consumption tag, 0.8, is greater than 0.6, so this tag enters the set of portrait tags.

[0039] Step S150: After matching and verifying the set of portrait tags with the preset portrait tag knowledge base, output the dynamically updated portrait tag sequence corresponding to the target user.

[0040] Specifically, the portrait label set (here it is high-frequency short-duration consumption) can be matched and verified with a preset portrait label knowledge base. In the field of financial service default risk, the portrait label knowledge base may have risk assessment information corresponding to different label combinations. For example, the high-frequency short-duration consumption label may be associated with a certain default risk, which may indicate that the user's consumption habit is relatively impulsive or there are potential capital turnover problems. After the matching and verification, a dynamically updated portrait label sequence corresponding to the target user is output. This portrait label sequence may be dynamically updated as the user's subsequent behavior changes. For example, if the user changes their consumption habit in the future, reducing the consumption frequency or increasing the consumption duration, then the portrait label sequence will be regenerated and updated based on the new behavior data to accurately reflect the user's characteristics in terms of financial service default risk.

[0041] Based on the above steps, the embodiment of the present application can comprehensively capture the temporal correlation between user behaviors by obtaining the detailed behavior trajectory data of the target user within a preset time period, including the behavior operation type and its context attributes, and constructing a state transition probability matrix based on the homogeneous transition model. On this basis, by determining the absorbing state behavior operation type set according to the state transition probability matrix and then generating a portrait label set, it can accurately describe the stable characteristics of the user in the specified behavior dimension, effectively avoiding misjudgment caused by short-term behavior fluctuations of the user. Next, using the homogeneous transition model, the user behavior data is transformed into a state transition probability matrix, and the absorbing state behavior operation type is determined by analyzing the steady-state convergence of the state nodes in the state transition probability matrix, greatly improving the scientificity and reliability of the portrait label generation process, making the generated portrait label more accurately reflect the true behavior characteristics of the user. Finally, after generating the portrait label set, it is matched and verified with the preset portrait label knowledge base, and then a dynamically updated portrait label sequence is output, ensuring that the portrait label can not only reflect the latest behavior characteristics of the user at present, but also conform to the existing knowledge base standards and knowledge systems. On the one hand, the dynamic update mechanism enables the portrait label to be adjusted in a timely manner as the user's behavior changes, always maintaining real-time tracking and accurate portrayal of the user's behavior; on the other hand, the matching and verification process ensures the consistency and accuracy of the generated portrait label within the entire system knowledge framework, improving the generality and usability of the portrait label in different application scenarios.

[0042] In a possible implementation manner, step S110 includes:

[0043] Step S111, extracting an original behavior event sequence from the interaction log of the target user, where the original behavior event sequence includes an event trigger timestamp, an event operation type identifier, and event context parameters.

[0044] Specifically, the event trigger timestamp accurately records the time of each credit card operation of the user. For example, a transaction was made at 13:15:00 on May 10, 2022, and a repayment operation was made at 15:30:00 on May 15, 2022, etc.; the event operation type identifier clarifies the type of operation, such as consumption, repayment, cash withdrawal, etc.; the event context parameters contain rich information, such as the transaction amount during consumption is 500 yuan, the transaction location is a large shopping mall, the address code of the shopping mall in the credit card system is 1234, the industry to which the transaction merchant belongs is retail, the repayment amount during the repayment operation is 800 yuan, etc.

[0045] Step S112: Perform data cleaning on the original behavior event sequence, remove duplicate event records and event records with abnormal timestamps, and obtain the cleaned behavior event sequence.

[0046] Specifically, for the timestamp aspect, check whether there are event records earlier than the system start recording time (such as the system starts recording on January 1, 2020) or later than the current system time (assuming the current time is July 2023). If so, mark them as abnormal timestamp events. For example, due to system failures, there may be a transaction record with a timestamp of December 30, 2019, which obviously does not meet the requirements. For the event operation type identifier, if it is empty (possibly due to partial loss during data transmission resulting in no operation type identifier for a certain transaction) or does not conform to the preset coding specification (such as the coding being maliciously tampered with or entered in the wrong format), mark the event record as an invalid operation type event. Then traverse every two consecutive event records. If the timestamp interval between the two is less than the preset minimum operation interval (the normal credit card transaction interval is at least 1 minute. If it is less than 1 minute, it may be a duplicate record), for example, if there are two transaction records with the same amount and the same merchant at 13:15:00 and 13:15:10, mark them as duplicate trigger events. Remove these abnormal timestamp events, invalid operation type events, and duplicate trigger events from the original behavior event sequence to obtain the intermediate cleaned event sequence. Then perform integrity verification on the remaining event records in the intermediate cleaned event sequence. If the transaction amount of a certain transaction record is missing, fill in the default value by querying the average amount of the user's same type of transactions or the common amount of transactions of the same type of merchants, thereby generating the cleaned behavior event sequence.

[0047] Step S113: Segment and cut the cleaned behavior event sequence according to the preset time window granularity to generate behavior subsequences corresponding to multiple time windows, where the event trigger timestamps within each behavior subsequence satisfy the time range constraint of the time window.

[0048] Specifically, for example, during the credit card promotion season (such as November - December 2022), transactions are frequent. If the number of event triggers exceeds the density threshold (assuming 60 transactions per day), the time window granularity is reduced to the first sub - granularity, changing from a weekly - divided window to a daily - divided window; while during the off - peak hours on weekdays (such as some weekdays in March 2023), transactions are scarce. If the number of event triggers is below the sparsity threshold (assuming 3 transactions per day), the time window granularity is expanded to the second sub - granularity, changing from a daily - divided window to a bi - weekly - divided window. Based on the adjusted time window granularity, the cleaned behavioral event sequence is divided into multiple consecutive or partially overlapping time window intervals. The frequency of the event operation type identifiers within each time window interval is counted to determine the main behavioral operation type. For example, within a certain time window, there are 20 transactions in total, including 12 consumption transactions, 5 repayment transactions, and 3 cash withdrawal transactions. Then consumption is the main behavioral operation type of this time window.

[0049] Step S114: Perform clustering analysis on the event operation type identifiers in each of the behavioral subsequences to obtain the main behavioral operation type and the set of auxiliary behavioral operation types corresponding to the time window.

[0050] Specifically, first convert the behavioral subsequence corresponding to each time window into a sequence vector of event operation type identifiers. For example, the operation sequence within a time window is [consumption, repayment, consumption, cash withdrawal, repayment], which is converted into the corresponding vector representation. Use a preset clustering model (such as the hierarchical clustering algorithm) to perform unsupervised clustering on the sequence vectors to obtain multiple candidate behavior clusters. Calculate the similarity between the center vector of each candidate behavior cluster and the preset benchmark behavior pattern vector (such as the vector representing the proportion relationship of consumption, repayment, and cash withdrawal in the normal credit card usage pattern). Assume there are three candidate behavior clusters. Through calculation, it is found that the similarity between the first candidate behavior cluster and the benchmark behavior pattern vector is the highest. Determine it as the main behavior cluster, and the corresponding operation type, consumption, is the main behavioral operation type of this time window. For other candidate behavior clusters, if the frequency of the operation type in them is greater than the auxiliary frequency threshold (assuming 3 times), then it is added to the set of auxiliary behavioral operation types. For example, if the frequency of cash withdrawal in a certain candidate behavior cluster is greater than 3 times, then cash withdrawal is added to the set of auxiliary behavioral operation types.

[0051] Step S115: Based on the event context parameters, the main behavioral operation type, and the corresponding set of auxiliary behavioral operation types, generate the behavioral operation type and context attributes corresponding to each discrete time point in the user behavior trajectory data.

[0052] For example, the main behavior operation type corresponding to a certain discrete time point is consumption, the transaction amount is 300 yuan, the transaction location is a certain supermarket, its address code is 5678, the industry belongs to retail, and the set of auxiliary behavior operation types includes cash withdrawal, and the average amount of cash withdrawal is 200 yuan (obtained by counting the cash withdrawal operations within this time window), etc. Thus, the behavior operation type and context attributes corresponding to each discrete time point are generated.

[0053] In a possible implementation manner, step S112 includes:

[0054] Step S1121, detecting whether there are event records in the original behavior event sequence whose timestamps are earlier than the system record start time or later than the current system time, and marking them as abnormal timestamp events.

[0055] In the credit card system, the original behavior event sequence is extracted from the interaction logs of the target user. This original behavior event sequence contains a lot of key information. For example, the event trigger timestamp accurately records the moment when each credit card operation occurs, the event operation type identifier clearly defines the type of operation, and the event context parameters contain additional information related to the operation.

[0056] Specifically, the credit card system has clear regulations on the time range of credit card operation records. The system record start time of the credit card system is set to January 1, 2020, and the current system time is assumed to be July 2023. In the original behavior event sequence, the timestamps of the events are checked one by one. When a transaction record with a timestamp of December 30, 2019 is found, since this time is earlier than the credit card system record start time, according to the rules, this event record is marked as an abnormal timestamp event. Similarly, if there is a record with a timestamp of August 1, 2023 (assumed), because it is later than the current system time, it is also marked as an abnormal timestamp event.

[0057] Step S1122, identifying event records in the original behavior event sequence whose event operation type identifiers are empty or do not conform to the preset coding specification, and marking them as invalid operation type events.

[0058] For example, it may be found that the operation type identifier in a certain transaction record is empty, which may be due to partial loss or writing errors during data transmission. In addition, the credit card operation type identifier has a preset coding specification. For example, the consumption operation code is 1, the repayment is 2, the cash withdrawal is 3, etc. If it is found that the operation type identifier in a certain record is 5 (not conforming to the preset coding specification), then this event record is marked as an invalid operation type event.

[0059] Step S1123: Traverse every two consecutive event records in the original behavior event sequence. If the time stamp interval between the two consecutive event records is less than the preset minimum operation interval, mark it as a repeatedly triggered event.

[0060] Specifically, there are certain time interval requirements for normal credit card operations. Assume that the preset minimum operation interval is 1 minute. For example, there are two consecutive transaction records. The first transaction occurred at 13:15:00 on May 10, 2022, with an amount of 500 yuan, and the second transaction occurred at 13:15:10 on May 10, 2022, with the same amount of 500 yuan, and the context parameters such as the transaction location and merchant are exactly the same. This indicates that these two events are very likely to be repeatedly triggered, so mark them as repeatedly triggered events.

[0061] Step S1124: Remove the abnormal time stamp events, invalid operation type events, and repeatedly triggered events from the original behavior event sequence to generate an intermediate cleaned event sequence.

[0062] For example, if there are 100 records in the original behavior event sequence, and after inspection, there are 5 abnormal time stamp events, 3 invalid operation type events, and 2 repeatedly triggered events. Then, after removing these 10 records, the remaining 90 records constitute the intermediate cleaned event sequence.

[0063] Step S1125: Perform integrity verification on the remaining event records in the intermediate cleaned event sequence, and complete the default values of the missing context parameters to generate the cleaned behavior event sequence.

[0064] For example, check the event context parameters of each transaction record. If a certain consumption transaction record lacks an important context parameter such as the transaction amount. At this time, it is necessary to complete the missing default value. It can be completed by querying the average amount of the same type of transactions of this user (such as consumption in the same merchant or within the same time period). Assume that the average amount of the past 10 consumptions in the same merchant is 300 yuan, then the amount of this transaction is completed to 300 yuan. After the above integrity verification, all event records have complete information, thus generating the cleaned behavior event sequence.

[0065] In a possible implementation manner, step S113 includes:

[0066] Step S1131: Dynamically adjust the length of the time window granularity according to the distribution density of the event trigger time stamps in the cleaned behavior event sequence.

[0067] Step S1132: If the number of event triggers exceeds the density threshold within a preset first time period, reduce the time window granularity to a first sub-granularity.

[0068] Step S1133, if the number of event triggers within the first time period is lower than the sparse threshold, expand the time window granularity to the second sub-granularity.

[0069] Taking credit card transactions as an example, the transaction frequencies vary greatly in different time periods. Suppose a first time period is preset, such as from November 1, 2022 to December 31, 2022 (this is a credit card promotion season), and the density threshold is set at 60 transactions per day. If within this time period, it is statistically found that the number of credit card transaction event triggers of a certain user exceeds this density threshold, then the time window granularity needs to be reduced to the first sub-granularity. Originally, the time window might be divided by week, and now it becomes divided by day. On the contrary, if the sparse threshold is set at 3 transactions per day from March 1, 2023 to March 31, 2023 (non-peak working days), when it is found that the number of credit card transaction event triggers of the user within this time period is lower than this sparse threshold, the time window granularity is expanded to the second sub-granularity. For example, the time window originally divided by day now becomes divided by two weeks.

[0070] Step S1134, based on the adjusted time window granularity, divide the sequence of cleaned behavior events into multiple consecutive or partially overlapping time window intervals.

[0071] For example, after adjustment, the time window granularity is divided by day. Starting from November 1, 2022, then time window intervals such as November 1, 2022 and November 2, 2022 are divided in sequence.

[0072] Step S1135, conduct a frequency count on the event operation type identifiers within each time window interval, and take the operation type with the highest frequency as the main behavior operation type.

[0073] For example, within the time window interval of November 1, 2022, there are a total of 20 transaction records, among which there are 12 consumption operations, 5 repayment operations, and 3 cash withdrawal operations. Since the consumption operation has the highest frequency, within this time window, the consumption operation is determined as the main behavior operation type corresponding to this time window.

[0074] In a possible implementation manner, step S114 includes:

[0075] Step S1141, convert the behavior subsequence corresponding to each time window into a sequence vector of event operation type identifiers.

[0076] For example, for an operation sequence within a time window as [consumption, repayment, consumption, cash withdrawal, repayment], if the consumption operation is encoded as 1, the repayment as 2, and the cash withdrawal as 3, then this behavior subsequence is converted into a sequence vector of [1, 2, 1, 3, 2].

[0077] Step S1142, perform unsupervised clustering on the sequence vectors using a preset clustering model to obtain multiple candidate behavior clusters.

[0078] Suppose there are sequence vectors corresponding to multiple time windows. Through the K-means clustering algorithm, these vectors are divided into multiple candidate behavior clusters according to their feature similarity. For example, three candidate behavior clusters may be obtained. Cluster 1 contains some sequence vectors with mainly consumption, relatively high consumption frequency, and relatively less repayment; Cluster 2 contains sequence vectors with relatively balanced consumption and repayment; Cluster 3 contains sequence vectors with relatively more cash withdrawal operations.

[0079] Step S1143, calculate the similarity between the central vector of each candidate behavior cluster and a preset benchmark behavior pattern vector.

[0080] Suppose the benchmark behavior pattern vector is set according to the operation type ratio in the normal credit card usage pattern. For example, consumption accounts for 60%, repayment accounts for 30%, and cash withdrawal accounts for 10%. For Cluster 1, calculate the similarity between its central vector and the benchmark behavior pattern vector. The calculation process is as follows: First, determine the central vector of Cluster 1. Suppose the average ratio of consumption operations in Cluster 1 is 70%, the repayment ratio is 20%, and the cash withdrawal ratio is 10%. Calculate the similarity between the two. By calculating the sum of the squares of the differences in the ratio of each operation type and then taking the square root to measure the distance, and then subtracting this distance from 1 to get the similarity. For the consumption operation, the difference is 70% - 60% = 10%; for the repayment operation, the difference is 30% - 20% = 10%; for the cash withdrawal operation, the difference is 10% - 10% = 0%. The sum of the squares of the differences is (10%)²+(10%)²+0² = 0.02, and taking the square root is approximately 0.1414. Then the similarity is 1 - 0.1414 = 0.8586. Calculate the similarities of Cluster 2 and Cluster 3 with the benchmark behavior pattern vector in the same way.

[0081] Step S1144, determine the candidate behavior cluster with the highest similarity as the main behavior cluster, and use the operation type corresponding to the main behavior cluster as the main behavior operation type.

[0082] For example, by comparing the calculated similarities, it is found that the similarity of Cluster 1 is the highest. Then Cluster 1 is determined as the main behavior cluster. Since the operation type corresponding to Cluster 1 is mainly consumption, the main behavior operation type corresponding to this time window is consumption.

[0083] Step S1145, add the types of operation types that appear more frequently than the auxiliary frequency threshold in other candidate behavior clusters to the set of auxiliary behavior operation types.

[0084] For other candidate behavior clusters (cluster 2 and cluster 3), for example, check the frequency of occurrence of operation types therein. Assume that the auxiliary frequency threshold is set to 3 times. In cluster 2, although consumption and repayment are relatively balanced, the cash withdrawal operation appears 4 times, which is greater than the auxiliary frequency threshold. In cluster 3, the cash withdrawal operation appears 5 times, which is also greater than the auxiliary frequency threshold. Then, add the cash withdrawal operation to the set of auxiliary behavior operation types.

[0085] In a possible implementation manner, step S120 includes:

[0086] Step S121, define the state space dimension of the state transition probability matrix according to the set of unique identifiers of all behavior operation types in the user behavior track data.

[0087] In the credit card usage scenario, assume that there are three behavior operation types: consumption, repayment, and cash withdrawal, which are represented by identifiers C, R, and W respectively. Then, the state transition probability matrix is a 3×3 matrix, and the rows and columns correspond to these three behavior operation types respectively.

[0088] Step S122, align the discrete time points in the user behavior track data in chronological order to generate the behavior state migration chain of the target user.

[0089] For example, sorting out the user's operation sequence from the credit card interaction log in chronological order as [C, R, C, W, R], this constitutes the behavior state migration chain, indicating that the user first performs a consumption operation, then a repayment operation, then another consumption operation, then a cash withdrawal operation, and finally a repayment operation again.

[0090] Step S123, count the number of migrations between every two adjacent behavior operation types in the behavior state migration chain to generate the initial state transition number matrix.

[0091] In the above behavior state migration chain [C, R, C, W, R], the migration from C to R occurs 1 time, the migration from R to C occurs 1 time, the migration from C to W occurs 1 time, and the migration from W to R occurs 1 time. Since there is no migration from R to W and from W to C, the initial state transition number matrix is:

[0092]

[0093] Among them, the first row corresponds to the migrations starting from C, and the first column represents the situation of migrating to C, and so on.

[0094] Step S124, perform normalization processing on the initial state transition number matrix to obtain the transition probability values between each pair of behavior operation types, and fill them into the corresponding positions of the state transition probability matrix.

[0095] In a possible implementation, step S124 includes:

[0096] Step S1241: Traverse each row element in the initial state transition count matrix and calculate the sum value of all elements in that row.

[0097] Step S1242: Compare the sum value with a preset minimum row sum threshold. If the sum value is less than the preset minimum row sum threshold, add a pseudo-count compensation value to that row to obtain the compensated row elements.

[0098] For the first row, the elements are 0, 1, 1, and the sum value is 2; for the second row, the elements are 1, 0, 0, and the sum value is 1; for the third row, the elements are 0, 1, 0, and the sum value is 1. Assume the preset minimum row sum threshold is 2. For the second row and the third row, their sum values are less than the preset minimum row sum threshold, so pseudo-count compensation values need to be added to these two rows. Assume the compensation value is 1, then the second row becomes 2, 0, 0, and the third row becomes 1, 1, 0.

[0099] Step S1243: Normalize the compensated row elements so that the sum of all elements in that row is 1 to obtain the normalized probability values.

[0100] Step S1244: Round the normalized probability values to a preset decimal precision to generate a standardized transition probability matrix.

[0101] For the first row, the original elements are 0, 1, 1, and the sum value is 2. After normalization, it becomes 0.0, 0.5, 0.5; for the second row, the original elements are 2, 0, 0, and the sum value is 2. After normalization, it becomes 1.0, 0.0, 0.0; for the third row, the original elements are 1, 1, 0, and the sum value is 2. After normalization, it becomes 0.5, 0.5, 0.0. Round the normalized probability values to the preset decimal precision. Assume the preset decimal precision is two decimal places, then the obtained standardized transition probability matrix is:

[0102]

[0103] Step S1245: Replace the elements in the standardized transition probability matrix that are lower than the preset probability lower limit value with the lower limit value to generate a final state transition probability matrix.

[0104] Assume the preset probability lower limit value is 0.05. Since 0.00 in the matrix is lower than the lower limit value, replace it with 0.05. The obtained final state transition probability matrix is: [0.05 0.50 0.45 0.95 0.05 0.00 0.45 0.50 0.05 ]

[0108] Step S125: Replace the elements with zero transition probability values in the state transition probability matrix with a preset minimum probability value to generate a smoothed target state transition probability matrix.

[0109] Assume that the preset minimum probability value is 0.01. If there are elements with probability values of zero in the matrix, replace them with 0.01 to obtain a smoothed target state transition probability matrix, so as to avoid calculation problems caused by zero probability in subsequent calculations. Thus, the process of constructing a state transition probability matrix based on credit card user behavior trajectory data is completed.

[0110] In a possible implementation, step S130 includes:

[0111] Step S131: Perform power iteration calculation on the state transition probability matrix until the maximum difference between the elements in each row of the state transition probability matrix is less than a preset convergence threshold to obtain a steady-state distribution vector.

[0112] Assume that the initial state transition probability matrix is constructed based on three types of behavioral operation types of credit card users: consumption, repayment, and cash withdrawal (labeled as C, R, and W respectively), as shown below: [0.3 0.5 0.2 0.4 0.3 0.3 0.2 0.4 0.4]

[0116] Perform power iteration calculation, and each iteration multiplies the matrix by itself. For example, the new matrix obtained from the first iteration calculation is:

[0117]

[0118]

[0119]

[0120] After detailed calculation (the detailed calculation process of multiple intermediate iterations is omitted here), until the maximum difference between the elements in each row of the state transition probability matrix is less than a preset convergence threshold (assume the convergence threshold is 0.001), a steady-state distribution vector is obtained. For example, after multiple iterations, the steady-state distribution vector is [0.33, 0.33, 0.34].

[0121] Step S132: Screen out the state nodes with probability values greater than a preset absorption state threshold from the steady-state distribution vector to generate a set of candidate absorption state behavioral operation types.

[0122] Specifically, state nodes with probability values greater than a preset absorbing state threshold (assuming the absorbing state threshold is 0.3) can be screened out from the steady-state distribution vector to generate a set of candidate absorbing state behavior operation types. In this example, the probability values corresponding to consumption (C) and cash withdrawal (W) meet the conditions and enter the set of candidate absorbing state behavior operation types.

[0123] Step S133: For each behavior operation type in the set of candidate absorbing state behavior operation types, verify whether the state self-transition probability of this behavior operation type in a continuous preset number of time periods is continuously greater than a preset stability threshold.

[0124] Step S134: Merge the behavior operation types that are continuously greater than the preset stability threshold into the set of absorbing state behavior operation types, and delete the behavior operation types that fail to pass the verification in the set of candidate absorbing state behavior operation types.

[0125] Specifically, for each behavior operation type in the set of candidate absorbing state behavior operation types, verify whether the state self-transition probability of this behavior operation type in a continuous preset number of time periods (assuming 3 consecutive months) is continuously greater than a preset stability threshold (assuming the stability threshold is 0.6). Taking consumption (C) as an example, check the value of the self-transition probability from consumption to consumption (i.e., the self-transition probability) in the state transition probability matrix for each month. Assume that the self-transition probability from consumption to consumption in the first month is 0.7, 0.65 in the second month, and 0.62 in the third month, all of which are greater than 0.6 and meet the conditions; for cash withdrawal (W), assume that its self-transition probability in the first month is 0.5, 0.55 in the second month, and 0.58 in the third month, which does not meet the conditions. Merge the behavior operation types that are continuously greater than the preset stability threshold (here only consumption C) into the set of absorbing state behavior operation types, and delete the behavior operation types that fail to pass the verification in the set of candidate absorbing state behavior operation types (here is cash withdrawal W).

[0126] Step S135: Increase the self-transition probability of the state nodes corresponding to each behavior operation type in the set of absorbing state behavior operation types by a preset reinforcement coefficient, reduce the transition probability value from the absorbing state behavior operation type to the non-absorbing state behavior operation type, and evenly distribute the difference to the migration paths of other absorbing state behavior operation types to obtain an adjusted state transition probability matrix.

[0127] Among them, step S135 includes:

[0128] Step S1351: Traverse each behavior operation type in the set of absorbing state behavior operation types, and identify the row index of the state node corresponding to this behavior operation type in the state transition probability matrix.

[0129] Step S1352, in the state transition probability matrix, for the row vector corresponding to the row index, extract the self-transition probability value of the state node corresponding to the absorbing state behavior operation type in this row vector.

[0130] Step S1353, multiply the self-transition probability value by a preset reinforcement coefficient to obtain a reinforced self-transition probability value, and replace the original self-transition probability value in the row vector with the reinforced self-transition probability value.

[0131] Step S1354, in the row vector, identify the column indices corresponding to all non-absorbing state behavior operation types, extract the transition probability values corresponding to these column indices, and reduce the transition probability values corresponding to these column indices to the product of the original transition probability value and a preset attenuation coefficient to obtain the attenuated non-absorbing state transition probability values.

[0132] Step S1355, calculate the difference between the original transition probability value corresponding to the column index and the attenuated non-absorbing state transition probability value, generate a difference component for each non-absorbing state transition path, and sum up the difference components of all non-absorbing state transition paths to obtain a total difference amount.

[0133] Step S1356, identify the column indices corresponding to the other absorbing state behavior operation types in the set of absorbing state behavior operation types except the current behavior operation type, count the number of the other absorbing state behavior operation types, and divide the total difference amount by the number of the other absorbing state behavior operation types to obtain an evenly distributed difference amount.

[0134] Step S1357, in the row vector, increase the transition probability value corresponding to each of the other absorbing state behavior operation types by the evenly distributed difference amount.

[0135] Step S1358, repeat the above steps until the transition probability values of the row vectors corresponding to all behavior operation types in the set of absorbing state behavior operation types are all adjusted. Then recombine the adjusted row vectors into an intermediate state transition probability matrix, and perform column normalization on each column of the intermediate state transition probability matrix so that the sum of all elements in each column is equal to 1, generating an adjusted state transition probability matrix.

[0136] Step S1359, manually calibrate the elements in the adjusted state transition probability matrix whose floating-point errors generated by column normalization exceed a preset error threshold to ensure that the error value of each column is less than the preset error threshold, and output the calibrated state transition probability matrix as the finally adjusted state transition probability matrix to the next processing node.

[0137] In this embodiment, the self-transition probability of the state node corresponding to each behavior operation type in the set of absorption-state behavior operation types (here it is consumption C) is increased by a preset reinforcement coefficient (assuming the reinforcement coefficient is 1.2), the transition probability value from the absorption-state behavior operation type to the non-absorption-state behavior operation type is reduced, and the difference is evenly distributed to the migration paths of other absorption-state behavior operation types (here there is only consumption C and no other absorption-state behavior operation types, so this step is only self-reinforcement in this example, but is described according to the complete process). In the state transition probability matrix, first traverse each behavior operation type in the set of absorption-state behavior operation types (here there is only consumption C), and identify the row index of the state node corresponding to consumption C in the state transition probability matrix (assuming it is the first row). In this row vector, extract the self-transition probability value of the state node corresponding to consumption C (assuming the initial self-transition probability value is 0.5). Multiply the self-transition probability value by the preset reinforcement coefficient 1.2 to obtain the reinforced self-transition probability value 0.6 (0.5×1.2 = 0.6), and replace the original self-transition probability value in the row vector with the reinforced self-transition probability value.

[0138] In this row vector, identify the column indices corresponding to all non-absorption-state behavior operation types (here they are repayment R and cash withdrawal W) (the second column and the third column respectively), extract the transition probability values corresponding to the column indices (0.3 and 0.2 respectively), and reduce the transition probability values corresponding to the column indices to the product of the original transition probability value and the preset attenuation coefficient (assuming the attenuation coefficient is 0.8) to obtain the attenuated non-absorption-state transition probability values. For repayment R, the attenuated transition probability value is 0.24 (0.3×0.8 = 0.24); for cash withdrawal W, the attenuated transition probability value is 0.16 (0.2×0.8 = 0.16).

[0139] Calculate the difference between the original transition probability value corresponding to the column index and the attenuated non-absorption-state transition probability value to generate the difference component of each non-absorption-state migration path. For repayment R, the difference component is 0.06 (0.3 - 0.24 = 0.06); for cash withdrawal W, the difference component is 0.04 (0.2 - 0.16 = 0.04). Sum the difference components of all non-absorption-state migration paths to obtain the total difference amount 0.1 (0.06 + 0.04 = 0.1). Since there is only one absorption-state behavior operation type, consumption C, and there are no other absorption-state behavior operation types to allocate the difference, this step is calculated in this example but not actually allocated.

[0140] Repeat the above steps until all row vectors corresponding to the behavior operation types in the absorbing state behavior operation type set have completed the adjustment of the transition probability values. Then recombine all the adjusted row vectors into an intermediate state transition probability matrix. Here, since only one row is adjusted, the intermediate state transition probability matrix is the matrix composed of this adjusted row (assumed to be [0.6, 0.24, 0.16]). Then perform column normalization on each column of the intermediate state transition probability matrix so that the sum of all elements in each column is equal to 1. Calculate the sum of the elements in each column. For the first column, there is only one element 0.6, no adjustment is needed; for the second column, the element is 0.24 and the sum is 0.24. To make its sum equal to 1, divide 0.24 by 0.24 to get 1; for the third column, the element is 0.16 and the sum is 0.16. Divide 0.16 by 0.16 to get 1. The renormalized state transition probability matrix obtained is [0.6, 1, 1] (here, for the sake of simplicity in explanation, the actual situation will be calculated according to the complete matrix).

[0141] During this process, manually calibrate the elements in the renormalized state transition probability matrix whose floating-point errors generated by column normalization exceed the preset error threshold (assuming the error threshold is 0.0001). For example, if an element should be 0.99995 after calculation but is displayed as 0.9998 due to floating-point error, and the error from the theoretical value exceeds 0.0001, it needs to be manually calibrated to 0.9999 to ensure that the error value in each column is less than the preset error threshold. Output the calibrated state transition probability matrix as the finally adjusted state transition probability matrix to the next processing node.

[0142] Step S136, perform column normalization on the adjusted state transition probability matrix again so that the sum of the elements in each column is 1, and obtain the renormalized state transition probability matrix.

[0143] Step S137, input the renormalized state transition probability matrix as the updated state transition probability matrix into the portrait label generation process of the next time period.

[0144] Step S138, record the update history of the absorbing state behavior operation type set for subsequent dynamic correction of the portrait label set.

[0145] For example, record information such as the time when consumption C is determined as the absorbing state behavior operation type and the changes in related probability values, so as to be used as a reference basis when dynamically correcting the portrait label set according to new behavior data in the future.

[0146] In a possible implementation manner, step S140 includes:

[0147] Step S141: Extract the context attribute feature vectors corresponding to each behavior operation type from the set of absorption state behavior operation types. The context attribute feature vectors include operation frequency, operation duration, and associated resource identifier.

[0148] In this embodiment, it is assumed that there is only one behavior operation type, namely consumption, in the set of absorption state behavior operation types. For the consumption operation, in terms of operation frequency, by statistically analyzing the credit card interaction logs, it is found that this user conducts consumption operations 15 times on average per month in the past year; the operation duration can be understood as the average time from the initiation to the completion of each consumption transaction, which is calculated to be 2.5 minutes; the associated resource identifier is the specific bank account number 123456 to which this credit card belongs. Thus, the context attribute feature vector corresponding to the behavior operation type of consumption is formed, that is, the operation frequency is 15 times per month, the operation duration is 2.5 minutes, and the associated resource identifier is the bank account number 123456.

[0149] Step S142: Match the context attribute feature vectors with the label definition rules in the preset portrait label knowledge base to determine the candidate portrait label list corresponding to each behavior operation type.

[0150] Specifically, in the portrait label knowledge base, there are the following definition rules for the consumption operation frequency: more than 10 consumption times per month is high-frequency consumption, 5 - 10 times is medium-frequency consumption, and less than 5 times is low-frequency consumption. Since this user consumes 15 times on average per month, it meets the label definition of high-frequency consumption. For the operation duration, if the duration of each consumption is less than 3 minutes, it is short-duration consumption. The average consumption duration of this user is 2.5 minutes, which meets the label definition of short-duration consumption. For the associated resource identifier, if the bank account number 123456 belongs to a specific high-risk account association type (this is set based on the bank's internal classification of account risks, for example, this account has had some suspected risk transaction histories), then it meets the label definition of high-risk account associated consumption. Therefore, the candidate portrait label list corresponding to the behavior operation type of consumption is high-frequency consumption, short-duration consumption, and high-risk account associated consumption.

[0151] Step S143: Perform semantic aggregation on the labels in the candidate portrait label list, merge multiple sub-labels describing the same behavior dimension, and generate an aggregated label set.

[0152] Specifically, in this example, high-frequency consumption and short-duration consumption both describe the time and frequency characteristics of consumption to a certain extent and belong to the same behavior dimension. They are merged into the aggregated label of high-frequency short-duration consumption. Thus, the aggregated label set contains two labels: high-frequency short-duration consumption and high-risk account associated consumption.

[0153] Step S144: Calculate the confidence weight of each portrait label based on the occurrence frequency and time distribution density of each portrait label in the aggregated label set.

[0154] For example, for high-frequency and short-duration consumption labels, check their occurrence frequencies in past data analyses. Suppose in the past 10 similar analyses, this label appeared 7 times. In terms of its time distribution density, these 7 occurrences were relatively evenly distributed in the past 10 analyses. For high-risk account-associated consumption labels, assume they appeared 3 times in the past 10 analyses and were relatively concentrated in the most recent 3 analyses. When calculating the confidence weight, a calculation method that comprehensively considers the occurrence frequency and time distribution density is adopted. For high-frequency and short-duration consumption labels, first calculate the occurrence frequency weight. Since it appeared 7 times and the total number of analyses is 10 times, the frequency weight is 7÷10 = 0.7. Because its distribution is relatively even, the time distribution density weight is set to 0.8 (a relatively high weight is given to uniform distribution here, and the specific weight setting can be adjusted according to actual business needs). Multiply the frequency weight and the time distribution weight to get the confidence weight of 0.7×0.8 = 0.56. For high-risk account-associated consumption labels, the occurrence frequency weight is 3÷10 = 0.3. Since its distribution is concentrated in the recent period, the time distribution density weight is set to 0.6 (a set weight is given to recent concentrated occurrences, but it is lower than the weight of uniform distribution), and the confidence weight is 0.3×0.6 = 0.18.

[0155] Step S145: Sort and filter the aggregated label set according to the confidence weight, retain the labels with weight values greater than the preset threshold, and generate the portrait label set.

[0156] Suppose the preset threshold is 0.3. Since the confidence weight of the high-frequency and short-duration consumption label is 0.56 which is greater than 0.3, while the confidence weight of the high-risk account-associated consumption label is 0.18 which is less than 0.3, only the high-frequency and short-duration consumption label is retained in the portrait label set. This portrait label set can reflect that the credit card user has stable characteristics of high frequency and short duration in consumption behavior, and this characteristic is of great significance in constructing portraits related to the default risk of financial services.

[0157] Figure 2 FIG. shows a schematic diagram of exemplary hardware and software components of a user portrait label generation system 100 based on a stationary transition model that can implement the ideas of the present application provided by some embodiments of the present application. For example, the processor 120 can be used on the user portrait label generation system 100 based on the stationary transition model and is used to execute the functions in the present application.

[0158] The user portrait label generation system 100 based on the stationary transition model can be a general-purpose server or a special-purpose server, both of which can be used to implement the user portrait label generation method based on the stationary transition model of the present application. Although only one server is shown in the present application, for convenience, the functions described in the present application can be implemented in a distributed manner on multiple similar platforms to balance the processing load.

[0159] For example, the user portrait label generation system 100 based on the stationary transition model can include a network port 110 connected to a network, one or more processors 120 for executing program instructions, a communication bus 130, and different forms of storage media 140, such as disks, ROM, or RAM, or any combination thereof. Exemplarily, the user portrait label generation system 100 based on the stationary transition model can also include program instructions stored in ROM, RAM, or other types of non-transitory storage media, or any combination thereof. The method of the present application can be implemented according to these program instructions. The user portrait label generation system 100 based on the stationary transition model also includes an input / output (I / O) interface 150 between the computer and other input / output devices.

[0160] For ease of explanation, only one processor is described in the user portrait label generation system 100 based on the stationary transition model. However, it should be noted that the user portrait label generation system 100 based on the stationary transition model in the present application can also include multiple processors. Therefore, the steps executed by one processor described in the present application can also be jointly executed or separately executed by multiple processors. For example, if the processor of the user portrait label generation system 100 based on the stationary transition model executes step A and step B, it should be understood that step A and step B can also be jointly executed by two different processors or separately executed in one processor. For example, the first processor executes step A, the second processor executes step B, or the first processor and the second processor jointly execute steps A and B.

[0161] In addition, an embodiment of the present invention also provides a readable storage medium, in which computer-executable instructions are preset. When the processor executes the computer-executable instructions, the user portrait label generation method based on the stationary transition model as described above is implemented.

[0162] It should be noted that, in order to simplify the expression of the disclosure of the present invention and thus help the understanding of one or more embodiments of the invention, in the foregoing description of the embodiments of the present invention, sometimes multiple features are merged into one embodiment, drawing, or description thereof.

Claims

1. A method for generating user portrait labels based on a time-aligned transfer model, characterized in that: The method comprises: Acquire user behavior trajectory data of a target user within a preset time period, wherein the user behavior trajectory data includes behavior operation types and corresponding context attributes at multiple discrete time points; Based on the temporal correlation of each behavior operation type in the user behavior trajectory data, a state transition probability matrix corresponding to the target user is constructed, wherein each element in the state transition probability matrix represents a transition probability of the target user migrating from a first behavior operation type to a second behavior operation type; Determine, according to the steady-state convergence of each state node in the state transition probability matrix, a set of absorbing state behavior operation types corresponding to the target user, each behavior operation type in the set of absorbing state behavior operation types satisfies a state transition probability threshold condition within a continuous time period; Based on the behavior operation types in the absorbing state behavior operation type set and the corresponding context attributes, a portrait tag set of the target user is generated, wherein each portrait tag in the portrait tag set is used to describe the stable characteristics of the target user in a specified behavior dimension; After matching and verifying the portrait tag set with a preset portrait tag knowledge base, outputting a dynamically updated portrait tag sequence corresponding to the target user; The generating the target user's portrait tag set based on the behavior operation types in the absorbing state behavior operation type set and the corresponding context attributes includes: Extracting a context attribute feature vector corresponding to each behavior operation type from the absorbing state behavior operation type set, wherein the context attribute feature vector includes an operation frequency, an operation duration, and an associated resource identifier; Matching the context attribute feature vector with the label definition rules in a preset portrait label knowledge base to determine a candidate portrait label list corresponding to each behavior operation type; Performing semantic aggregation on the tags in the candidate portrait tag list, merging multiple sub-tags describing the same behavior dimension, and generating an aggregated tag set; Calculating the confidence weight of each portrait tag based on the occurrence frequency and time distribution density of each portrait tag in the aggregated tag set; The aggregated tag set is sorted and screened according to the confidence weight, and tags with weight values ​​greater than a preset threshold are retained to generate the portrait tag set.

2. The user portrait label generation method based on the time-aligned transfer model according to claim 1 is characterized in that: The step of obtaining the user behavior trajectory data of the target user within a preset time period includes: Extracting an original behavior event sequence from the target user's interaction log, wherein the original behavior event sequence includes an event triggering timestamp, an event operation type identifier, and an event context parameter; Performing data cleaning on the original behavior event sequence, removing duplicate event records and event records with abnormal timestamps, and obtaining a cleaned behavior event sequence; Segment the cleaned behavior event sequence according to a preset time window granularity to generate behavior subsequences corresponding to multiple time windows, wherein the event triggering timestamp in each behavior subsequence satisfies the time range constraint of the time window; Performing cluster analysis on the event operation type identifiers in each of the behavior subsequences to obtain a set of main behavior operation types and auxiliary behavior operation types corresponding to the time window; Based on the event context parameters and the main behavior operation type and the corresponding auxiliary behavior operation type set, the behavior operation type and context attributes corresponding to each discrete time point in the user behavior trajectory data are generated.

3. The user portrait label generation method based on the time-aligned transfer model according to claim 2 is characterized in that: The data cleaning of the original behavior event sequence, removing duplicate event records and event records with abnormal timestamps, and obtaining a cleaned behavior event sequence includes: Detect whether there is an event record with a timestamp earlier than the system record start time or later than the current system time in the original behavior event sequence, and mark it as an abnormal timestamp event; Identify event records in the original behavior event sequence whose event operation type identifier is empty or does not conform to a preset coding specification, and mark them as invalid operation type events; Traversing every two consecutive event records in the original behavior event sequence, if the time stamp interval between the two consecutive event records is less than a preset minimum operation interval, marking them as repeated triggering events; Remove the abnormal timestamp event, invalid operation type event and repeated trigger event from the original behavior event sequence to generate an intermediate cleaning event sequence; An integrity check is performed on the remaining event records in the intermediate cleaning event sequence, missing default values ​​of context parameters are supplemented, and the post-cleaning behavior event sequence is generated.

4. The method for generating user portrait labels based on a time-aligned transfer model according to claim 2, characterized in that: The step of segmenting the post-cleaning behavior event sequence according to a preset time window granularity to generate behavior subsequences corresponding to multiple time windows includes: Dynamically adjusting the length of the time window granularity according to the distribution density of event triggering timestamps in the post-cleaning behavior event sequence; If the number of event triggering times exceeds a density threshold within a preset first time period, the time window granularity is reduced to a first sub-granularity; If the number of event triggering times in the first time period is lower than a sparse threshold, expanding the time window granularity to a second sub-granularity; Based on the adjusted time window granularity, the post-cleaning behavior event sequence is divided into a plurality of continuous or partially overlapping time window intervals; The frequency of the event operation type identifier within each of the time window intervals is counted, and the operation type with the highest frequency is used as the main behavior operation type.

5. The method for generating user portrait labels based on a time-aligned transfer model according to claim 2, characterized in that: The cluster analysis is performed on the event operation type identifier in each of the behavior subsequences to obtain a set of main behavior operation types and auxiliary behavior operation types corresponding to the time window, including: Converting the behavior subsequence corresponding to each of the time windows into a sequence vector of event operation type identifiers; Using a preset clustering model to perform unsupervised clustering on the sequence vector to obtain multiple candidate behavior clusters; Calculating the similarity between the center vector of each candidate behavior cluster and a preset benchmark behavior pattern vector; Determine the candidate behavior cluster with the highest similarity as the main behavior cluster, and use the operation type corresponding to the main behavior cluster as the main behavior operation type; The operation types in other candidate behavior clusters whose occurrence frequencies are greater than the auxiliary frequency threshold are added to the auxiliary behavior operation type set.

6. The method for generating user portrait labels based on a time-aligned transfer model according to claim 1, characterized in that: The constructing a state transition probability matrix corresponding to the target user based on the temporal correlation of each behavior operation type in the user behavior trajectory data includes: Defining the state space dimension of the state transition probability matrix according to a unique identification set of all behavior operation types in the user behavior trajectory data; Aligning discrete time points in the user behavior trajectory data in chronological order to generate a behavior state transition chain of the target user; Counting the number of migrations between every two adjacent behavior operation types in the behavior state migration chain to generate an initial state transition number matrix; Normalizing the initial state transition frequency matrix to obtain the transition probability value between each pair of behavior operation types, and filling the corresponding position of the state transition probability matrix; The elements with transition probability values ​​of zero in the state transition probability matrix are replaced with a preset minimum probability value to generate a target state transition probability matrix after smoothing.

7. The method for generating user portrait labels based on a time-aligned transfer model according to claim 6, characterized in that: The normalizing process of the initial state transition times matrix to obtain the transition probability value between each pair of behavior operation types and fill the corresponding position of the state transition probability matrix includes: Traverse each row of elements in the initial state transition count matrix and calculate the sum of all elements in the row; Compare the sum value with a preset minimum row sum threshold, and if the sum value is less than the preset minimum row sum threshold, add a pseudo count compensation value to the row to obtain compensated elements in each row; Normalize the probability of each row of elements after compensation so that the sum of all elements in the row is 1, and obtain the normalized probability value; The normalized probability values ​​are rounded to a preset decimal precision to generate a standardized transition probability matrix; The elements in the standardized transition probability matrix that are lower than a preset probability lower limit are replaced with the lower limit value to generate a final state transition probability matrix.

8. The method for generating user portrait labels based on a time-aligned transfer model according to claim 1, characterized in that: The step of determining a set of absorbing state behavior operation types corresponding to the target user according to the steady-state convergence of each state node in the state transition probability matrix includes: Performing power iteration calculation on the state transition probability matrix until the maximum difference of each row element in the state transition probability matrix is ​​less than a preset convergence threshold, thereby obtaining a steady-state distribution vector; Filter out state nodes whose probability values ​​are greater than a preset absorbing state threshold from the steady-state distribution vector, and generate a set of candidate absorbing state behavior operation types; For each behavior operation type in the candidate absorbing state behavior operation type set, verify whether the state self-transition probability of the behavior operation type within a preset number of consecutive time periods is continuously greater than a preset stability threshold; Merge the behavior operation types that are continuously greater than the preset stability threshold into the absorbing state behavior operation type set, and delete the behavior operation types that have not passed the verification in the candidate absorbing state behavior operation type set; The self-transition probability of the state node corresponding to each behavior operation type in the absorbing state behavior operation type set is increased by a preset reinforcement coefficient, the transition probability value from the absorbing state behavior operation type to the non-absorbing state behavior operation type is reduced, and the difference is evenly distributed to the migration paths of other absorbing state behavior operation types, to obtain an adjusted state transition probability matrix; The adjusted state transition probability matrix is ​​re-normalized so that the sum of the elements in each column is 1, thereby obtaining a re-normalized state transition probability matrix; The renormalized state transition probability matrix is ​​used as the updated state transition probability matrix and input into the portrait label generation process of the next time period; Recording the update history of the absorbing state behavior operation type set for dynamic modification of subsequent portrait label sets; The step of increasing the self-transition probability of the state node corresponding to each behavior operation type in the absorbing state behavior operation type set by a preset reinforcement coefficient, reducing the transition probability value from the absorbing state behavior operation type to the non-absorbing state behavior operation type, and evenly distributing the difference to the migration paths of other absorbing state behavior operation types to obtain an adjusted state transition probability matrix includes: Traversing each behavior operation type in the absorbing state behavior operation type set, and identifying the row index of the state node corresponding to the behavior operation type in the state transition probability matrix; In the state transition probability matrix, for the row vector corresponding to the row index, extract the self-transition probability value of the state node corresponding to the absorbing state behavior operation type in the row vector; Multiplying the self-transition probability value by a preset enhancement coefficient to obtain an enhanced self-transition probability value, and replacing the original self-transition probability value in the row vector with the enhanced self-transition probability value; In the row vector, identifying the column indexes corresponding to all non-absorbing state behavior operation types, extracting the transition probability values ​​corresponding to the column indexes, and reducing the transition probability values ​​corresponding to the column indexes to the product of the original transition probability values ​​and a preset attenuation coefficient to obtain the attenuated non-absorbing state transition probability values; Calculating the difference between the original transition probability value corresponding to the column index and the attenuated non-absorbing state transition probability value, generating a difference component for each non-absorbing state transition path, and summing the difference components of all non-absorbing state transfer paths to obtain a total difference; Identify the column indexes corresponding to other absorbing state behavior operation types in the absorbing state behavior operation type set except the current behavior operation type, count the number of the other absorbing state behavior operation types, and divide the total difference by the number of the other absorbing state behavior operation types to obtain an average distribution difference; In the row vector, increasing the transition probability value corresponding to each of the other absorbing state behavior operation types by the average distribution difference amount; Repeat the above steps until the row vectors corresponding to all behavior operation types in the absorbing state behavior operation type set have completed the adjustment of the transfer probability values, recombining all the adjusted row vectors into an intermediate state transfer probability matrix, and performing column normalization processing on each column of the intermediate state transfer probability matrix so that the sum of all elements in each column is equal to 1, thereby generating an adjusted state transfer probability matrix; Manually calibrate the elements in the adjusted state transition probability matrix whose floating-point errors due to column normalization processing exceed the preset error threshold to ensure that the error value of each column is less than the preset error threshold, and output the calibrated state transition probability matrix as the final adjusted state transition probability matrix to the next processing node.

9. A user portrait label generation system based on a time-aligned transfer model, characterized in that: The user portrait label generation system based on the time-aligned transfer model includes a processor and a memory, the memory is connected to the processor, the memory is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the memory to implement the user portrait label generation method based on the time-aligned transfer model described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Precise pushing method and system based on user portrait

    CN118886980A